【发布时间】:2019-06-16 04:49:51
【问题描述】:
我有一个 df(下面的一小部分)。我正在尝试为当前行及其所有子行添加额外的 tw0 列 rolled_doc_cnt 和 rolled_doc_cnt_all ,其中包含 doc_cnt/doc_cnt_all 的总和。在下面非常有限的行中,df.at[0,'rolled_doc_cnt'] = 1317 & df.at[0,'rolled_doc_cnt_all'] = 3540
SYMBOL level not-allocatable additional-only doc_cnt doc_cnt_all parent
0 A 2 True False 0 0
1 A01 4 True False 0 0 A
2 A01B 5 True False 0 0 A01
3 A01B 1/00 7 False False 198 244 A01B
4 A01B 1/02 8 False False 230 538 A01B 1/00
5 A01B 1/022 9 False False 83 238 A01B 1/02
6 A01B 1/024 9 False False 28 63 A01B 1/02
7 A01B 1/026 9 False False 100 120 A01B 1/02
8 A01B 1/028 9 False False 27 82 A01B 1/02
9 A01B 1/04 9 False False 29 54 A01B 1/02
10 A01B 1/06 8 False False 78 508 A01B 1/00
11 A01B 1/065 9 False False 118 150 A01B 1/06
12 A01B 1/08 9 False False 71 326 A01B 1/06
13 A01B 1/10 9 False False 14 30 A01B 1/06
14 A01B 1/12 9 False False 24 86 A01B 1/06
15 A01B 1/14 9 False False 44 131 A01B 1/06
16 A01B 1/16 8 False False 159 518 A01B 1/00
17 A01B 1/165 9 False False 50 114 A01B 1/16
18 A01B 1/18 9 False False 64 338 A01B 1/16
我在创建parent 列here 时得到了一些帮助。
def GetParent():
# level 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19, 20
hierarchy = [0, 0, 0, 0, 2, 4, 0, 5, 7, 8, 9,
10, 11, 12, 13, 14, 15, 16, 17, 18, 19]
parent = ['']*len(hierarchy)
def func(row):
# print(row)
symbol, level = row[['SYMBOL', 'level']]
parent_level = hierarchy[level]
parent_symbol = parent[parent_level]
parent[level] = symbol
return pd.Series([parent_symbol], index=['parent'])
return func
# create a column with the parents
st = time()
parents = dfa.apply(GetParent(), axis=1)
dfa = pd.concat([dfa, parents], axis=1)
print((time()-st)/60, 'minutes elapsed')
我尝试在 spyder 中调试此代码,以便在推进 df 行时看到列表 parent 的变化,但我无法弄清楚如何在不跳转到 pandas 的情况下跳转到函数 GetParent()函数apply()。跳入apply() 最终导致我出现递归错误。
我尝试对GetParents() 进行一些修改,以跟踪每个级别的每个符号的文档计数,但后来我意识到我正在跟踪父节点的文档计数,但不是孩子们。那么,使用上面的 df,我如何能够创建类似于以下 df 的内容?
SYMBOL level not-allocatable additional-only doc_cnt doc_cnt_all parent rolled_doc_cnt rolled_doc_cnt_all
0 A 2 TRUE FALSE 0 0 1317 3540
1 A01 4 TRUE FALSE 0 0 A 1317 3540
2 A01B 5 TRUE FALSE 0 0 A01 1317 3540
3 A01B 1/00 7 FALSE FALSE 198 244 A01B 1317 3540
4 A01B 1/02 8 FALSE FALSE 230 538 A01B 1/00 497 1095
5 A01B 1/022 9 FALSE FALSE 83 238 A01B 1/02 83 238
6 A01B 1/024 9 FALSE FALSE 28 63 A01B 1/02 28 63
7 A01B 1/026 9 FALSE FALSE 100 120 A01B 1/02 100 120
8 A01B 1/028 9 FALSE FALSE 27 82 A01B 1/02 27 82
9 A01B 1/04 9 FALSE FALSE 29 54 A01B 1/02 29 54
10 A01B 1/06 8 FALSE FALSE 78 508 A01B 1/00 349 1231
11 A01B 1/065 9 FALSE FALSE 118 150 A01B 1/06 118 150
12 A01B 1/08 9 FALSE FALSE 71 326 A01B 1/06 71 326
13 A01B 1/10 9 FALSE FALSE 14 30 A01B 1/06 14 30
14 A01B 1/12 9 FALSE FALSE 24 86 A01B 1/06 24 86
15 A01B 1/14 9 FALSE FALSE 44 131 A01B 1/06 44 131
16 A01B 1/16 8 FALSE FALSE 159 518 A01B 1/00 273 970
17 A01B 1/165 9 FALSE FALSE 50 114 A01B 1/16 50 114
18 A01B 1/18 9 FALSE FALSE 64 338 A01B 1/16 64 338
也请随时告诉我,我尝试这样做的方式不是最佳的,并建议另一种方式
【问题讨论】:
标签: python-3.x pandas