【发布时间】:2020-08-09 23:30:54
【问题描述】:
我有一个如下所示的数据框。这是医生预约数据。
B_ID No_Show Session slot_num Cumulative_no_show
1 0.4 S1 1 0.4
2 0.3 S1 2 0.7
3 0.8 S1 3 1.5
4 0.3 S1 4 1.8
5 0.6 S1 5 2.4
6 0.8 S1 6 3.2
7 0.9 S1 7 4.1
8 0.4 S1 8 4.5
9 0.6 S1 9 5.1
12 0.9 S2 1 0.9
13 0.5 S2 2 1.4
14 0.3 S2 3 1.7
15 0.7 S2 4 2.4
20 0.7 S2 5 3.1
16 0.6 S2 6 3.7
17 0.8 S2 7 4.5
19 0.3 S2 8 4.8
当 u_cumulative > 0.8 时,在 No_Show = 0.0 的下方创建一个新行,其 Session 和 slot_num 应与前一个相同,并通过从前一个减去 1 创建一个名为 u_cumulative 的新列。
预期输出:
B_ID No_Show Session slot_num Cumulative_no_show u_cumulative
1 0.4 S1 1 0.4 0.4
2 0.3 S1 2 0.7 0.7
3 0.8 S1 3 1.5 1.5
walkin1 0.0 S1 3 1.5 0.5
4 0.3 S1 4 1.8 0.8
5 0.6 S1 5 2.4 1.4
walkin2 0.0 S1 5 2.4 0.4
6 0.8 S1 6 3.2 1.2
walkin3 0.0 S1 6 3.2 0.2
7 0.9 S1 7 4.1 1.1
walkin4 0.0 S1 7 4.1 0.1
8 0.4 S1 8 4.5 0.5
9 0.6 S1 9 5.1 1.1
walkin5 0.0 S1 7 5.1 0.1
12 0.9 S2 1 0.9 0.9
walkin1 0.0 S2 1 0.9 -0.1
13 0.5 S2 2 1.4 0.4
14 0.3 S2 3 1.7 0.7
15 0.7 S2 4 2.4 1.4
walkin2 0.0 S2 4 2.4 0.4
20 0.7 S2 5 3.1 1.1
walkin3 0.0 S2 5 3.1 0.1
16 0.6 S2 6 3.7 0.7
17 0.8 S2 7 4.5 1.5
walkin4 0.0 S2 7 4.5 0.5
19 0.3 S2 8 4.8 0.8
我尝试在下面计算 u_cumulative
def create_u_columns (ser):
arr_ns = ser.to_numpy()
arr_sn = np.ones(len(ser))
for i in range(len(arr_ns)-1):
if arr_ns[i]>0.6:
# remove 1 to u_no_show
arr_ns[i+1:] -= 1
else:
# increment u_slot_num
arr_sn[i+1:] += 1
#return a dataframe with both columns
return pd.DataFrame({'U_slot_num':arr_sn, 'U_No_show': arr_ns}, index=ser.index)
df[['U_slot_num', 'u_cumulative']] = df.groupby(['Session'])['Cumulative_No_show'].apply(create_u_columns)
但我无法根据上述逻辑创建新行。
【问题讨论】:
-
我想你的意思是如果 u_cumulative > 0.8 而不是 Cumulative_no_show > 0.8
-
@NYCCoder 是的,你是对的,感谢您指出.. 已编辑
-
简单来说,如何计算u_cumulative?
-
@wwnde 何时 Cumulative_no_show 超过 0.8 添加新行从累积无显示中减去 1 使其成为 u_cumulative
标签: pandas numpy loops pandas-groupby