【问题标题】:Pandas: parse merged header columns from ExcelPandas:从 Excel 解析合并的标题列
【发布时间】:2015-02-09 18:49:09
【问题描述】:

excel表格中的数据存储方式如下:

   Area     |          Product1     |      Product2        |      Product3
            |      sales|sales.Value|   sales |sales.Value |  sales |sales.Value
  Location1 |    20     | 20000     |      25 |  10000     |   200  | 100
  Location2 |    30     | 30000     |      3  | 12300      |   213  | 10

产品名称是给定月份的 1000 个左右区域中的每一个区域的 2 行“销售额”和“销售额”两行单元格的合并。同样,过去 5 年的每个月都有单独的文件。此外,新产品已在不同月份添加和删除。因此,不同的月份文件可能如下所示:

   Area     |          Product1     |      Product4        |      Product3

论坛能否建议使用 pandas 读取这些数据的最佳方法? 我不能使用索引,因为每个月的产品列都不一样

理想情况下,我想将上面的初始格式转换为:

 Area      | Product1.sales|Product1.sales.Value| Product2.sales |Product2.sales.Value | 
 Location1 | 20            | 20000              | 25             | 10000               |  
 Location2 | 30            | 30000              | 3              | 12300               | 

import pandas as pd
xl_file = read_excel("file path", skiprow=2, sheetname=0)
/* since the first two rows are always blank */


                  0            1        2               3                      4
      0          NaN          NaN      NaN       Auto loan                    NaN
      1  Branch Code  Branch Name   Region  No of accounts  Portfolio Outstanding
      2         3000       Name1  Central               0                      0
      3         3001       Name2  Central               0                      0

我想将其转换为Auto loan.No of accountAuto loan.Portfolio Outstanding 作为标题。

【问题讨论】:

  • 您能否发布一个示例,说明使用df = pd.read_excel(...) 加载文件时DataFrame 的外观? df.indexdf.columns 是什么?
  • 谢谢,我想通了。 5x12=60 文件中只有 4 个元列组合。所以我只是为所有 4 种组合使用字典。
  • @unutbu:我的编辑是否清楚地表达了我的要求?感谢您的帮助,因为我的解决方案并不优雅。
  • 编辑解释了所需的结果(好),但仍然不清楚(至少对我而言)原始 DataFrame 的样子。原始 DataFrame 是否具有 MultiIndex 列?请发布您的代码,即使它不优雅。它可以帮助我们了解您的 DataFrame 的外观。
  • @unutbu 我在 EDIT2 中添加了数据框 head() 视图。如您所见,我想将其转换为“Auto loan.No of accounts”或类似的可唯一查询的内容。

标签: python excel pandas read-data


【解决方案1】:

假设你的 DataFrame 是df:

import numpy as np
import pandas as pd

nan = np.nan
df = pd.DataFrame([
    (nan, nan, nan, 'Auto loan', nan)
    , ('Branch Code', 'Branch Name', 'Region', 'No of accounts'
       , 'Portfolio Outstanding')
    , (3000, 'Name1', 'Central', 0, 0)
    , (3001, 'Name2', 'Central', 0, 0)
])

让它看起来像这样:

             0            1        2               3                      4
0          NaN          NaN      NaN       Auto loan                    NaN
1  Branch Code  Branch Name   Region  No of accounts  Portfolio Outstanding
2         3000       Name1  Central               0                      0
3         3001       Name2  Central               0                      0

然后首先向前填充前两行中的 NaN(因此传播 'Auto 贷款”,例如)。

df.iloc[0:2] = df.iloc[0:2].fillna(method='ffill', axis=1)

接下来用空字符串填充剩余的NaN:

df.iloc[0:2] = df.iloc[0:2].fillna('')

现在将两行与. 连接在一起并将其分配为列级值:

df.columns = df.iloc[0:2].apply(lambda x: '.'.join([y for y in x if y]), axis=0)

最后,删除前两行:

df = df.iloc[2:]

这会产生

  Branch Code Branch Name   Region Auto loan.No of accounts  \
2        3000      Name1  Central                        0   
3        3001      Name2  Central                        0   

  Auto loan.Portfolio Outstanding  
2                               0  
3                               0  

或者,您可以创建 MultiIndex 列而不是创建平面列索引:

import numpy as np
import pandas as pd

nan = np.nan
df = pd.DataFrame([
    (nan, nan, nan, 'Auto loan', nan)
    , ('Branch Code', 'Branch Name', 'Region', 'No of accounts'
       , 'Portfolio Outstanding')
    , (3000, 'Name1', 'Central', 0, 0)
    , (3001, 'Name2', 'Central', 0, 0)
])
df.iloc[0:2] = df.iloc[0:2].fillna(method='ffill', axis=1)
df.iloc[0:2] = df.iloc[0:2].fillna('Area')

df.columns = pd.MultiIndex.from_tuples(
    zip(*df.iloc[0:2].to_records(index=False).tolist()))
df = df.iloc[2:]

现在df 看起来像这样:

         Area                           Auto loan                      
  Branch Code Branch Name   Region No of accounts Portfolio Outstanding
2        3000      Name1  Central              0                     0
3        3001      Name2  Central              0                     0

该列是一个 MultiIndex:

In [275]: df.columns
Out[275]: 
MultiIndex(levels=[[u'Area', u'Auto loan'], [u'Branch Code', u'Branch Name', u'No of accounts', u'Portfolio Outstanding', u'Region']],
           labels=[[0, 0, 0, 1, 1], [0, 1, 4, 2, 3]])

该列有两个级别。第一级的值为[u'Area', u'Auto loan'],第二级的值为[u'Branch Code', u'Branch Name', u'No of accounts', u'Portfolio Outstanding', u'Region']

然后您可以通过指定两个级别的值来访问列:

print(df.loc[:, ('Area', 'Branch Name')])
# 2    Name1
# 3    Name2
# Name: (Area, Branch Name), dtype: object

print(df.loc[:, ('Auto loan', 'No of accounts')])
# 2    0
# 3    0
# Name: (Auto loan, No of accounts), dtype: object

使用 MultiIndex 的一个优点是您可以轻松选择具有特定级别值的所有列。例如,要选择与 Auto loans 相关的子数据帧,您可以使用:

In [279]: df.loc[:, 'Auto loan']
Out[279]: 
  No of accounts Portfolio Outstanding
2              0                     0
3              0                     0

有关从 MultiIndex 中选择行和列的更多信息,请参阅MultiIndexing Using Slicers

【讨论】:

    猜你喜欢
    • 2019-05-11
    • 1970-01-01
    • 2020-04-24
    • 1970-01-01
    • 2017-12-21
    • 2021-12-02
    • 2014-01-29
    • 2019-04-01
    相关资源
    最近更新 更多