【发布时间】:2016-07-31 12:58:27
【问题描述】:
我有一些数据结构如下:
structure(list(subject = c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 3L, 3L, 3L, 3L, 3L, 3L, 3L, 3L, 4L, 4L, 4L, 4L, 4L, 4L, 4L, 4L), group = structure(c(1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 1L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L, 2L), .Label = c("group1", "group2"), class = "factor"), measurement = c("color", "time", "color", "time", "color", "time", "color", "time", "color", "time", "color", "time", "color", "time", "color", "time", "color", "time", "color", "time", "color", "time", "color", "time", "color", "time", "color", "time", "color", "time", "color", "time"), item_pos = c("1", "1", "2", "2", "3", "3", "4", "4", "1", "1", "2", "2", "3", "3", "4", "4", "1", "1", "2", "2", "3", "3", "4", "4", "1", "1", "2", "2", "3", "3", "4", "4"), value = c("blue", "1508", "orange", "752", "black", "585", "red", "842", "red", "879", "white", "1455", "green", "1757", "orange", "2241", "white", "2251", "yellow", "1740", "red", "1962", "yellow", "1854", "green", "1859", "blue", "2156", "yellow", "2494", "green", "1757"), item = c("A", "A", "B", "B", "B", "B", "A", "A", "A", "A", "B", "B", "B", "B", "A", "A", "C", "C", "C", "C", "D", "D", "D", "D", "C", "C", "C", "C", "D", "D", "D", "D")), .Names = c("subject", "group", "measurement", "item_pos", "value", "item"), row.names = c(NA, -32L), class = "data.frame")
按主题按项目有多个观察结果,因此主题 1 的数据如下所示:
> filter(df.tidy, subject==1)
subject group measurement item_pos value item
1 1 group1 color 1 blue A
2 1 group1 time 1 1508 A
3 1 group1 color 2 orange B
4 1 group1 time 2 752 B
5 1 group1 color 3 black B
6 1 group1 time 3 585 B
7 1 group1 color 4 red A
8 1 group1 time 4 842 A
所以在group 内,每个item 出现两次,每次出现都有一个measurement 的颜色和时间。项目出现的顺序是item_pos。
虽然我喜欢这种长格式,但一位同事需要它稍微“宽一些”,在他们自己的列中按项目重复颜色和时间测量。 所需的格式如下:
subject group item color1 color2 time1 time2
1 group1 A blue red 1508 842
1 group1 B orange black 752 585
...
4 group2 D yellow green 2494 1757
我的感觉是,这应该可以使用gather()、spread() 和其他 dplyr 动词的组合来实现,但我不确定这里的 dplyr 等价物用于(在 for-loop 中)循环按组通过项目并收集后续列中的颜色和时间观察。非常感谢您的帮助!
我咨询过的相关问题:
【问题讨论】: