【发布时间】:2018-02-16 17:09:52
【问题描述】:
我有一个熊猫数据框。
import pandas as pd
data = pd.DataFrame({
'a': [0,1,0,0,1,1,0,1],
'b': [0,0,1,0,1,0,1,1],
'c': [0,0,0,1,0,1,1,1],
'rate': [0,0.1,0.11,0.12,0.24,0.27,0.3,0.4]})
a,b,c 是我的频道,我正在添加另一列,显示这些频道的按行总计的总和:
data['total'] = data.a + data.b + data.c
data
a b c rate total
1 1 0 0 0.10 1
2 0 1 0 0.11 1
3 0 0 1 0.12 1
4 1 1 0 0.24 2
5 1 0 1 0.27 2
6 0 1 1 0.30 2
7 1 1 1 0.40 3
我想处理总计 = 1 和总计 = 2 的数据
reduced = data[(data.a == 1) & (data.total == 2)]
print(reduced)
a b c rate total
4 1 1 0 0.24 2
5 1 0 1 0.27 2
我想向这个简化的数据框添加列,如下所示:
a b c rate total prob_a prob_b prob_c
4 1 1 0 0.24 2 0.1 0.11 0
5 1 0 1 0.27 2 0.1 0 0.12
在缩减数据帧的第一行中,prob_c 为 0,因为 C 不存在(ABC => 110)。在缩减数据帧的第二行中,prob_b 为 0,因为 B 不存在(ABC => 101)
在哪里,
# Channel a alone occurs (ABC => 100)
prob_a = data['rate'][(data.a == 1) & (data.total == 1)]
# Channel b alone occurs (ABC => 010)
prob_b = data['rate'][(data.b == 1) & (data.total == 1)]
# Channel c alone occurs (ABC => 001)
prob_c = data['rate'][(data.c == 1) & (data.total == 1)]
我试过了:
reduced['prob_a'] = data['rate'][(data.a == 1) & (data.total == 1)]
reduced['prob_b'] = data['rate'][(data.b == 1) & (data.total == 1)]
reduced['prob_c'] = data['rate'][(data.c == 1) & (data.total == 1)]
print(reduced)
导致此输出:
a b c rate total prob_a prob_b prob_c
4 1 1 0 0.24 2 NaN NaN NaN
5 1 0 1 0.27 2 NaN NaN NaN
【问题讨论】: