【问题标题】:pandas: slicing multiindex data frames... simple series desired熊猫:切片多索引数据帧......需要简单的系列
【发布时间】:2017-10-10 05:00:56
【问题描述】:

我通过从 datareader 获取面板并将其转换为多索引数据框来创建股票数据的多索引。有时当我使用.loc 时,我会得到一个有 1 个索引的系列,有时我会得到一个有两个索引的系列。如何按日期切片并获得具有一个索引的系列?代码会有所帮助...

import pandas_datareader.data as web

# Define the securities to download
symbols = ['AAPL', 'MSFT']

# Define which online source one should use
data_source = 'yahoo'

# Define the period of interest
start_date = '2010-01-01'
end_date = '2010-12-31'

# User pandas_reader.data.DataReader to load the desired data. 
panel = web.DataReader(symbols, data_source, start_date, end_date)

# Convert panel to multiindex dataframe
midf = panel.to_frame()

# for slicing multiindex dataframes it must be sorted
midf = midf.sort_index(level=0)

在这里我选择我想要的列:

adj_close = midf['Adj Close']
adj_close.head()

我得到一个包含两个索引的系列(Dateminor):

Date        minor
2010-01-04  AAPL     27.505054
            SPY      96.833946
2010-01-05  AAPL     27.552608
            SPY      97.090271
2010-01-06  AAPL     27.114347
Name: Adj Close, dtype: float64

现在我使用: 选择苹果来选择所有日期。

aapl_adj_close = adj_close.loc[:, 'AAPL']
aapl_adj_close.head()

并获得索引为Date 的系列。这就是我要找的!

Date
2010-01-04    27.505054
2010-01-05    27.552608
2010-01-06    27.114347
2010-01-07    27.064222
2010-01-08    27.244156
Name: Adj Close, dtype: float64

但是当我实际按日期切片时,我没有得到那个系列:

sliced_aapl_adj_close  = adj_close.loc['2010-01-04':'2010-01-06', 'AAPL']
sliced_aapl_adj_close.head()

我得到一个包含两个索引的系列:

Date        minor
2010-01-04  AAPL     27.505054
2010-01-05  AAPL     27.552608
2010-01-06  AAPL     27.114347
Name: Adj Close, dtype: float64

切片是正确的,值是正确的,但我不希望那里有次要索引(因为我想通过这个系列来绘制)。切片的正确方法是什么?

谢谢!

【问题讨论】:

    标签: python pandas


    【解决方案1】:

    你可以使用:

    df = df.reset_index(level=1, drop=True)
    

    或者:

    df.index = df.index.droplevel(1)
    

    另一种解决方案是通过unstack 重塑DataFrame,然后通过[] 选择:

    df = adj_close.unstack()
    
    print (df)
    minor            AAPL        SPY
    Date                            
    2010-01-04  27.505054  96.833946
    2010-01-05  27.552608  97.090271
    2010-01-06  27.114347        NaN
    
    print (df['AAPL'])
    
    Date
    2010-01-04    27.505054
    2010-01-05    27.552608
    2010-01-06    27.114347
    Name: AAPL, dtype: float64
    

    【讨论】:

    • 谢谢!这样可行!但是为什么两个切片返回不同的东西呢? adj_close.loc['2010-01-04':'2010-01-06', 'AAPL']adj_close.loc[:, 'AAPL'] 看起来很相似。
    • 难题。也许是因为第二种解决方案更通用,所以可以使用adj_close.loc['2010-01-04':'2010-01-05', ['AAPL', 'SPY']]
    猜你喜欢
    • 2020-08-30
    • 2017-02-22
    • 2018-07-28
    • 2020-01-24
    • 2019-10-28
    • 2020-06-19
    • 2020-03-17
    • 2022-11-03
    • 2021-08-07
    相关资源
    最近更新 更多