【发布时间】:2021-10-21 14:11:31
【问题描述】:
假设我有一些数据框,其中一列的某些值多次出现形成组(sn-p 中的列A)。现在我想创建一个新列,例如1 用于每个组的第一个 x(列 C)条目,0 在其他条目中。
我设法完成了第一部分,但我没有找到在xes 中包含条件的好方法,有没有好的方法?
import pandas as pd
df = pd.DataFrame(
{
"A": ["0", "0", "1", "2", "2", "2"], # data to group by
"B": ["a", "b", "c", "d", "e", "f"], # some other irrelevant data to be preserved
"C": ["y", "x", "y", "x", "y", "x"], # only consider the 'x'
}
)
target = pd.DataFrame(
{
"A": ["0", "0", "1", "2", "2", "2"],
"B": ["a", "b", "c", "d", "e", "f"],
"C": ["y", "x", "y", "x", "y", "x"],
"D": [ 0, 1, 0, 1, 0, 0] # first entry per group of 'A' that has an 'C' == 'x'
}
)
# following partial solution doesn't account for filtering by 'x' in 'C'
df['D'] = df.groupby('A')['C'].transform(lambda x: [1 if i == 0 else 0 for i in range(len(x))])
【问题讨论】:
标签: python pandas dataframe pandas-groupby