【发布时间】:2018-02-02 21:34:53
【问题描述】:
我确实有这样的数据框:
import pandas as pd
df = pd.DataFrame({"c0": list('ABC'),
"c1": [" ".join(list('ab')), " ".join(list('def')), " ".join(list('s'))],
"c2": list('DEF')})
c0 c1 c2
0 A a b D
1 B d e f E
2 C s F
我想创建一个如下所示的数据透视表:
c2
c0 c1
A a D
b D
B d E
e E
f E
C s F
因此,c1 中的条目被拆分,然后被视为多索引中使用的单个元素。
我这样做如下:
newdf = pd.DataFrame()
for indi, rowi in df.iterrows():
# get all single elements in string
n_elements = rowi['c1'].split()
# only one element so we can just add the entire row
if len(n_elements) == 1:
newdf = newdf.append(rowi)
# more than one element
else:
for eli in n_elements:
# that allows to add new elements using loc, without it we will have identical index values
if not newdf.empty:
newdf = newdf.reset_index(drop=True)
newdf.index = -1 * newdf.index - 1
# add entire row
newdf = newdf.append(rowi)
# replace the entire string by the single element
newdf.loc[indi, 'c1'] = eli
print newdf.reset_index(drop=True)
产生
c0 c1 c2
0 A a D
1 A b D
2 B d E
3 B e E
4 B f E
5 C s F
那我可以打电话了
pd.pivot_table(newdf, index=['c0', 'c1'], aggfunc=lambda x: ' '.join(set(str(v) for v in x)))
这给了我想要的输出(见上文)。
对于可能非常慢的巨大数据帧,所以我想知道是否有更有效的方法来做到这一点。
【问题讨论】:
标签: python performance pandas