【发布时间】:2020-06-07 15:42:03
【问题描述】:
我有一个大约 5GB 大小的 CSV,数据结构和类型是这样的:
datetime product name serial number
0 2017-06-24 14:30:15 orange 123456
1 2017-07-04 21:33:50 apple 123456
2 2017-07-06 06:38:52 orange 123456
3 2017-07-10 15:52:07 banana 123456
4 2017-07-10 15:52:51 banana 123456
5 2017-07-10 15:53:18 banana 123456
6 2017-07-11 11:50:40 pineapple 123456
7 2017-07-11 00:53:43 apple 54321
8 2017-07-11 06:23:52 apple 54321
9 2017-07-11 06:23:52 apple 12454
10 2017-07-11 06:23:52 apple 12454
11 2017-07-11 06:23:52 apple 12454
12 2017-07-11 06:23:52 apple 15039
13 2017-07-11 06:23:52 apple 15037
14 2017-07-11 06:23:52 apple 15039
15 2017-07-11 06:23:52 apple 15190
16 2017-07-11 06:23:52 apple 15039
17 2017-07-11 06:23:52 apple 15037
18 2017-07-11 06:23:52 apple 15037
19 2017-07-11 06:23:52 apple 15037
....
few millions more lines
df.dtypes
Out[134]:
datetime datetime64[ns]
name object
events int64
dtype: object
问题 1: 如何按产品名称分组,然后仅统计前 10 位产品的序列号出现次数(出现次数最多的产品在顶部)?
# this does the count, but there are over 10,000 rows, and it is not sorted by counts f
df.groupby(['product name', 'serial number']).agg({'serial number':'count'}).compute()
# expected output (in table form):
product name serial number counts
orange 123456 2
orange 54321 12
apple 123456 1
apple 54321 4
pineapple 123456 16
问题 2: 如何在时间域内绘制一个产品名称的每个序列号的出现次数?
问题 3: 我真的很想绘制一个产品名称随时间域出现的每个“序列号”, 到目前为止,我可以使用以下方法从数据框中挑选出“产品名称”:
df_orange = df[df['proudct name'] == 'orange']
# how do I plot it?
【问题讨论】:
-
如果您有问题 1 和 2 所需的输出示例,这将非常有帮助,因此我们可以通过视觉帮助了解您的需求
标签: pandas csv matplotlib anaconda dask