【发布时间】:2017-06-28 03:30:36
【问题描述】:
这是我的数据框的样子:
Timestamp CAT
0 2016-12-02 23:35:28 200
1 2016-12-02 23:37:43 200
2 2016-12-02 23:40:49 300
3 2016-12-02 23:58:53 400
4 2016-12-02 23:59:02 300
...
这就是我在 Pandas 中尝试做的事情(注意时间戳是分组的):
Timestamp BINS 200 300 400 500
2016-12-02 23:30 2 0 0 0
2016-12-02 23:40 0 1 0 0
2016-12-02 23:50 0 1 1 0
...
我正在尝试创建 10 分钟时间间隔的 bin,以便制作条形图。并将列作为 CAT 值,因此我可以计算每个 CAT 在该时间段内出现的次数。
我目前所拥有的可以创建时间箱:
def create_hist(df, timestamp, freq, fontsize, outfile):
""" Create a histogram of the number of CATs per time period."""
df.set_index(timestamp,drop=False,inplace=True)
to_plot = df[timestamp].groupby(pandas.TimeGrouper(freq=freq)).count()
...
但我的问题是我终其一生都无法弄清楚如何按 CAT 和按时间箱进行分组。我最近的尝试是在进行 groupby 之前使用df.pivot(columns="CAT"),但这只会给我错误:
def create_hist(df, timestamp, freq, fontsize, outfile):
""" Create a histogram of the number of CATs per time period."""
df.pivot(columns="CAT")
df.set_index(timestamp,drop=False,inplace=True)
to_plot = df[timestamp].groupby(pandas.TimeGrouper(freq=freq)).count()
...
这给了我:ValueError: Buffer has wrong number of dimensions (expected 1, got 2)
【问题讨论】:
标签: python pandas pivot histogram pandas-groupby