【发布时间】:2020-09-16 20:09:54
【问题描述】:
标题确实只是一个非常粗略的想法,很可能与实际问题不太匹配。
我有一些股票数据,看起来像这样:
"DateTime","Price","Volume","Group"
2020-05-01 13:30:01.354,174.003,750,2020-05-01
2020-05-01 13:30:01.454,174.003,750,2020-05-01
2020-05-01 13:30:01.612,174.592,750,2020-05-01
2020-05-01 13:30:01.663,174.812,750,2020-05-01
2020-05-01 13:30:01.775,174.742,750,2020-05-01
2020-05-01 13:30:02.090,174.742,2000.0001,2020-05-01
2020-05-01 13:30:02.139,174.742,750,2020-05-01
2020-05-01 13:30:02.190,174.743,2000.0001,2020-05-01
2020-05-01 13:30:02.308,174.612,2000.0001,2020-05-01
2020-05-01 13:30:02.428,174.612,750,2020-05-01
2020-05-01 13:30:02.554,174.522,2000.0001,2020-05-01
2020-05-01 13:30:02.656,174.552,750,2020-05-01
2020-05-01 13:30:02.705,174.522,2000.0001,2020-05-01
2020-05-01 13:30:02.913,174.645,750,2020-05-01
2020-05-01 13:30:02.963,175.002,750,2020-05-01
2020-05-01 13:30:03.013,175.002,2000.0001,2020-05-01
2020-05-01 13:30:03.125,175.002,750,2020-05-01
2020-05-01 13:30:03.312,174.803,750,2020-05-01
2020-05-01 13:30:03.362,175.002,2000.0001,2020-05-01
2020-05-01 13:30:03.876,174.772,750,2020-05-01
2020-05-01 13:30:03.927,174.802,2000.0001,2020-05-01
2020-05-01 13:30:04.052,174.802,2000.0001,2020-05-01
2020-05-01 13:30:04.154,174.692,750,2020-05-01
2020-05-01 13:30:04.203,174.802,750,2020-05-01
2020-05-01 13:30:04.255,174.803,2000.0001,2020-05-01
2020-05-01 13:30:04.304,174.803,2000.0001,2020-05-01
2020-05-01 13:30:04.404,174.802,750,2020-05-01
2020-05-01 13:30:04.455,175.003,2000.0001,2020-05-01
2020-05-01 13:30:04.521,174.803,750,2020-05-01
2020-05-01 13:30:04.649,174.802,750,2020-05-01
2020-05-01 13:30:04.771,174.803,2000.0001,2020-05-01
2020-05-01 13:30:04.822,174.803,2000.0001,2020-05-01
2020-05-01 13:30:04.899,174.702,750,2020-05-01
2020-05-01 13:30:04.950,174.802,750,2020-05-01
2020-05-01 13:30:06.498,174.722,750,2020-05-01
2020-05-01 13:30:07.794,174.723,750,2020-05-01
2020-05-01 13:30:07.843,175.003,2000.0001,2020-05-01
2020-05-01 13:30:08.095,175.002,750,2020-05-01
2020-05-01 13:30:08.466,175.002,750,2020-05-01
2020-05-01 13:30:08.567,175.002,750,2020-05-01
2020-05-01 13:30:08.743,174.982,2000.0001,2020-05-01
2020-05-01 13:30:09.123,175.002,750,2020-05-01
2020-05-01 13:30:09.381,174.982,750,2020-05-01
2020-05-01 13:30:09.893,175.002,750,2020-05-01
2020-05-01 13:30:09.942,174.882,750,2020-05-01
2020-05-01 13:30:09.993,174.962,750,2020-05-01
2020-05-01 13:30:11.404,175.002,2000.0001,2020-05-01
2020-05-01 13:30:11.716,174.963,750,2020-05-01
2020-05-01 13:30:11.932,174.963,750,2020-05-01
2020-05-01 13:30:11.983,175.002,750,2020-05-01
2020-05-01 13:30:12.038,174.962,750,2020-05-01
2020-05-01 13:30:12.414,174.963,2000.0001,2020-05-01
2020-05-01 13:30:12.533,174.863,750,2020-05-01
2020-05-01 13:30:12.585,174.962,2000.0001,2020-05-01
2020-05-01 13:30:13.763,175.002,750,2020-05-01
2020-05-01 13:30:14.473,174.962,750,2020-05-01
2020-05-01 13:30:16.157,174.962,750,2020-05-01
2020-05-01 13:30:16.207,175.002,2000.0001,2020-05-01
2020-05-01 13:30:16.268,175.002,750,2020-05-01
2020-05-01 13:30:18.455,175.002,750,2020-05-01
2020-05-01 13:30:18.506,175.322,750,2020-05-01
2020-05-01 13:30:19.289,175.322,750,2020-05-01
2020-05-01 13:30:19.340,175.342,750,2020-05-01
2020-05-01 13:30:19.953,175.343,750,2020-05-01
2020-05-01 13:30:20.761,175.362,2000.0001,2020-05-01
2020-05-01 13:30:21.588,175.363,750,2020-05-01
2020-05-01 13:30:21.638,175.382,750,2020-05-01
2020-05-01 13:30:22.387,175.383,750,2020-05-01
2020-05-01 13:30:22.486,175.442,750,2020-05-01
2020-05-01 13:30:22.580,175.382,750,2020-05-01
2020-05-01 13:30:23.595,175.442,750,2020-05-01
2020-05-01 13:30:23.645,175.383,750,2020-05-01
2020-05-01 13:30:23.762,175.442,750,2020-05-01
2020-05-01 13:30:24.085,175.382,750,2020-05-01
2020-05-01 13:30:24.134,175.273,2000.0001,2020-05-01
2020-05-01 13:30:24.608,175.272,750,2020-05-01
2020-05-01 13:30:24.658,175.272,750,2020-05-01
2020-05-01 13:30:25.019,175.272,750,2020-05-01
2020-05-01 13:30:25.070,175.332,750,2020-05-01
2020-05-01 13:30:25.238,175.283,750,2020-05-01
2020-05-01 13:30:25.289,175.282,2000.0001,2020-05-01
2020-05-01 13:30:25.749,175.273,750,2020-05-01
2020-05-01 13:30:25.799,175.273,2000.0001,2020-05-01
2020-05-01 13:30:25.863,175.273,750,2020-05-01
2020-05-01 13:30:25.914,175.333,2000.0001,2020-05-01
2020-05-01 13:30:26.073,175.283,750,2020-05-01
2020-05-01 13:30:26.124,175.282,2000.0001,2020-05-01
2020-05-01 13:30:26.187,175.203,750,2020-05-01
2020-05-01 13:30:26.237,175.182,2000.0001,2020-05-01
2020-05-01 13:30:26.710,175.282,2000.0001,2020-05-01
2020-05-01 13:30:27.511,175.282,2000.0001,2020-05-01
2020-05-01 13:30:27.763,175.332,2000.0001,2020-05-01
2020-05-01 13:30:28.187,175.233,750,2020-05-01
2020-05-01 13:30:28.236,175.232,750,2020-05-01
2020-05-01 13:30:28.302,175.232,750,2020-05-01
2020-05-01 13:30:28.353,175.232,2000.0001,2020-05-01
2020-05-01 13:30:28.457,175.152,750,2020-05-01
2020-05-01 13:30:28.507,175.152,750,2020-05-01
2020-05-01 13:30:28.601,175.153,2000.0001,2020-05-01
2020-05-01 13:30:28.894,175.093,750,2020-05-01
2020-05-01 13:30:28.945,175.092,750,2020-05-01
2020-05-01 13:30:29.049,175.093,2000.0001,2020-05-01
我想做的是根据Price 列中值的顺序频率计算Volume 的cumsum。以下是上述csv 的示例r 输出为data.frame/data.table。
DateTime Price Volume Group
1: 2020-05-01 13:30:01.354 174.003 750 2020-05-01
2: 2020-05-01 13:30:01.454 174.003 750 2020-05-01
3: 2020-05-01 13:30:01.612 174.592 750 2020-05-01
4: 2020-05-01 13:30:01.663 174.812 750 2020-05-01
5: 2020-05-01 13:30:01.775 174.742 750 2020-05-01
6: 2020-05-01 13:30:02.090 174.742 2000 2020-05-01
7: 2020-05-01 13:30:02.139 174.742 750 2020-05-01
8: 2020-05-01 13:30:02.190 174.743 2000 2020-05-01
9: 2020-05-01 13:30:02.308 174.612 2000 2020-05-01
10: 2020-05-01 13:30:02.428 174.612 750 2020-05-01
8: 2020-05-01 13:30:02.554 174.522 2000 2020-05-01
9: 2020-05-01 13:30:02.656 174.552 2000 2020-05-01
10: 2020-05-01 13:30:02.705 174.522 750 2020-05-01
为了更详细地解释,我将取前 2 行并将它们对应的体积汇总为输出数据对象的一个元素(最好是data.frame);第 3 行和第 4 行的价格频率为 1(在价格序列中唯一出现),因此它们的交易量不需要相加,并且会在输出数据对象中产生 2 行。第 5、6、7 行的价格是相同的 174.742,然后我再次总结 3 卷并在输出数据对象中生成 1 行。相同的逻辑适用于其余数据。
我一直在尝试dplyr,但无济于事;我能得到的最好的结果是每个价格出现的指数组,但不能保留自然顺序。我想不出一种方法来接近我想要的结果(在矢量化的风格中,我试图避免简单的循环和编写过多的逻辑)。
按照@RonakShah 的建议,我将添加这个预期的输出示例。抱歉,我之前没有添加一个,因为我真的不确定它会是什么样子。但正如我再次考虑的那样,我认为我只需要 cumsum 卷序列的输出,作为单列。我真的不需要另一列进行索引(尽管显然任何额外的列都可以作为非常好的拥有)。
示例输出:
Cumsum.Volume
1: 1500 # aggregated from row 1,2 as they have same price
2: 750 # no aggregation as row 3 has occurred only once
3: 750 # same as above
4: 3500 # aggregation from row 4,5,6 as they have the same price
5: 2000 # row 7 no aggregation for unique price occurrence in sequence
6: 2750 # row 8,9 maps to same price so add them up
7: 2000 # row 10 needs no aggregation
8: 2000 # row 11 needs no aggregation
9: 750 # row 12 needs no aggregation
这是价格和数量之间的映射:
`174.003` appeared twice together, so we cumsum their volumes
`174.592` appeared by itself, so we keep its volume
`174.812` appeared by itself, so we keep its volume
`174.742` appeared 3 times consecutively, so we cumsum/aggregate their volumes
`174.743` appeared by itself, so we keep its volume
`174.612` appeared twice together, so we cumsum their volumes
`174.522` appeared once by itself, we keep its volume
`174.552` appeared by itself, we keep its volume
`174,522` appeared again by itself, we keep its volume
我希望以上补充可以给你们一个更好的主意; cumsum 卷的顺序与价格列的自然顺序一致,但我不需要保留价格列。
关于正确方向的任何提示?
【问题讨论】:
-
您需要
df %>% group_by(Price) %>% mutate(Volume = cumsum(Volume))吗?你能显示预期输出的前几行吗? -
@RonakShah 抱歉,我刚回来;我认为我不需要价格分组,因为我想保持价格序列的原始顺序(实际上我认为我可以完全忽略价格序列,我只需要使用价格创建一个新的累积量列序列)我将尝试添加一个预期输出的示例(我没有添加,因为手工创建有点太乏味并且不太确定它应该是什么样子)
标签: r dataframe datatable vectorization tidyverse