【问题标题】:Combining Pandas dataframes of different Period frequencies结合不同周期频率的 Pandas 数据帧
【发布时间】:2019-09-22 14:26:41
【问题描述】:

假设我有以下两个数据框:

np.random.seed(1)
annual = pd.DataFrame(data=np.random.random((2, 4)), index=index, columns=pd.period_range(start="2015", end="2018", freq="Y"))
quarterly = pd.DataFrame(data=np.random.random((2,3)), index=index, columns=pd.period_range('2019', freq='Q', periods=3))

Annual:

    2015        2016        2017        2018
A   0.417022    0.720324    0.000114    0.302333
B   0.146756    0.092339    0.186260    0.345561

Quarterly:

    2019Q1      2019Q2      2019Q3
A   0.396767    0.538817    0.419195
B   0.685220    0.204452    0.878117

我是否有可能将这两个数据帧组合在一起,以使生成的数据帧df 看起来像下面这样?如果没有,是否有允许我合并两个数据框的解决方法,以便我可以执行df['2019Q2'] - df['2018'] 之类的操作?

    2015        2016        2017        2018        2019Q1      2019Q2      2019Q3
A   0.417022    0.720324    0.000114    0.302333    0.396767    0.538817    0.419195   
B   0.146756    0.092339    0.186260    0.345561    0.685220    0.204452    0.878117

【问题讨论】:

    标签: python pandas dataframe


    【解决方案1】:

    首先concataxis=1,如果以后需要处理,则必须将列名转换为字符串:

    df = pd.concat([annual,quarterly], axis=1).rename(columns=str)
    print (df)
           2015      2016      2017      2018    2019Q1    2019Q2    2019Q3
    A  0.417022  0.720324  0.000114  0.302333  0.396767  0.538817  0.419195
    B  0.146756  0.092339  0.186260  0.345561  0.685220  0.204452  0.878117
    
    print (df.columns)
    Index(['2015', '2016', '2017', '2018', '2019Q1', '2019Q2', '2019Q3'], dtype='object')
    
    print (df['2019Q2'] - df['2018'])
    A    0.236484
    B   -0.141108
    dtype: float64
    

    如果想使用句点,这是可能的,但更复杂:

    df = pd.concat([annual,quarterly], axis=1)
    print (df)
           2015      2016      2017      2018    2019Q1    2019Q2    2019Q3
    A  0.417022  0.720324  0.000114  0.302333  0.396767  0.538817  0.419195
    B  0.146756  0.092339  0.186260  0.345561  0.685220  0.204452  0.878117
    
    print (df[pd.Period('2018', freq='A-DEC')])
    A    0.302333
    B    0.345561
    Name: 2018, dtype: float64
    
    print (df[pd.Period('2019Q2', freq='Q-DEC')])
    A    0.538817
    B    0.204452
    Name: 2019Q2, dtype: float64
    

    print (df[pd.Period('2019Q2', freq='Q-DEC')] - 
           df[pd.Period('2018', freq='A-DEC')])
    

    IncompatibleFrequency:输入与 Period(freq=Q-DEC) 具有不同的 freq=A-DEC

    更改Series的名称以防出错:

    print (df[pd.Period('2019Q2', freq='Q-DEC')].rename('a') - 
           df[pd.Period('2018', freq='A-DEC')].rename('a'))
    
    A    0.236484
    B   -0.141108
    Name: a, dtype: float64
    

    在我看来,如果需要使用 Periods 处理值,最好使用相同的频率:

    annual.columns = annual.columns.to_timestamp('Q').to_period('Q')
    df = pd.concat([annual,quarterly], axis=1)
    print (df)
         2015Q1    2016Q1    2017Q1    2018Q1    2019Q1    2019Q2    2019Q3
    A  0.417022  0.720324  0.000114  0.302333  0.396767  0.538817  0.419195
    B  0.146756  0.092339  0.186260  0.345561  0.685220  0.204452  0.878117
    
    print (df[pd.Period('2019Q2', freq='Q-DEC')] - 
           df[pd.Period('2018Q1', freq='Q-DEC')])
    
    A    0.236484
    B   -0.141108
    dtype: float64
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2013-05-26
      • 2016-03-28
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2015-01-20
      相关资源
      最近更新 更多