【发布时间】:2021-05-16 16:55:40
【问题描述】:
我想向 DataFrame 添加一个新的布尔列,以指示给定列的值是否第一次出现在 groupby 组中。
我的 DataFrame 是这样的:
UserID Value
0 1955 30
1 1955 40
2 1955 30
3 1956 30
4 1957 30
5 1957 50
6 1958 30
7 1958 50
8 1958 30
9 1958 30
我想得到这个:
UserID Value IsNewValue
0 1955 30 True
1 1955 40 True
2 1955 30 False
3 1956 30 True
4 1957 30 True
5 1957 50 True
6 1958 30 True
7 1958 30 False
8 1958 30 False
9 1958 30 False
请务必注意,数据集已按用户 ID 和时间戳(此处未显示)排序,我无法更改此排序。
我想出了这段代码,虽然效率极低:
def is_new(group, col):
seen = []
ret = []
for i in range(len(group)):
ret.append(group[col].iloc[i] not in seen)
seen.append(group[col].iloc[i])
group[f'IsNew{col}'] = ret
return group
for col in ['ValueA', 'ValueB', 'ValueC']:
dataset = dataset.groupby('UserID').apply(lambda x: is_new(x, col))
我想知道如何重写代码并使其更高效,可能使用 Pandas 的窗口函数或一些 numpy 功能。
【问题讨论】:
标签: python pandas numpy pandas-groupby