【发布时间】:2021-06-04 01:06:06
【问题描述】:
我正在处理相当复杂的数据。这是作为数据框df 的简化快照。
ID Measures ME1 ME2 X1 X2
53-21 comm - 01 narrate 2 1 NA NA
53-21 comm - overall 1 NA NA NA
53-21 comm - 10 participate NA NA NA NA
43-65 comm - 02 project 2 3 NA NA
43-65 comm - 01 narrate 1 1 NA NA
67-21 comm - 06 action 2 1 NA NA
67-21 comm - 08 plan 1 1 NA 1
43-65 comm - overall 2 NA NA NA
53-21 comm - exhibit 1 1 NA NA
这里:
ID = 唯一用户 ID
Measures = 为给定用户测量或评估的项目名称
对于每个Measure,用户最多可以在四个不同的项目上评分,例如ME1、ME2、X1 和X2。
我想将此数据转换为将项目放置在行中的格式,即每行一个 ID,并将它们的相应度量值放在附加列中。我需要的重塑数据框是这样的:
ID comm-01-narrate-ME1 comm-01-narrate-ME2 comm-01-narrate-X1 comm-01-narrate-X2 comm-overall-ME1 comm-overall-ME2 comm-overall-X1 comm-overall-X2 comm-10-participate-ME1 comm-10-participate-ME2 comm-10-participate-X1 comm-10-participate-X2 comm-exhibit-ME1 comm-exhibit-ME2 comm-exhibit-X1 comm-exhibit-X2 comm-02-project-ME1 comm-02-project-ME2 comm-02-project-X1 comm-02-project-X2 comm-06-action-ME1 comm-06-action-ME2 comm-06-action-X1 comm-06-action-X2 comm-08-plan-ME1 comm-08-plan-ME2 comm-08-plan-X1 comm-08-plan-X2
53-21 2 1 NA NA 1 NA NA NA NA NA NA NA 1 1 NA NA NA NA NA NA NA NA NA NA NA NA NA NA
43-65 1 1 NA NA 2 NA NA NA NA NA NA NA NA NA NA NA 2 3 NA NA NA NA NA NA NA NA NA NA
67-21 NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA 2 1 NA NA 1 1 NA 1
输入文件df 的dput() 是:
dput(df)
structure(list(ID = structure(c(2L, 2L, 2L, 1L, 1L, 3L, 3L, 1L, 2L),
.Label = c("43-65", "53-21", "67-21"), class = "factor"),
Measures = structure(c(1L, 7L, 5L, 2L, 1L, 3L, 4L, 7L, 6L),
.Label = c("comm - 01 narrate", "comm - 02 project", "comm - 06 action", "comm - 08 plan", "comm - 10 participate", "comm - exhibit", "comm - overall"), class = "factor"),
ME1 = c(2L, 1L, NA, 2L, 1L, 2L, 1L, 2L, 1L),
ME2 = c(1L, NA, NA, 3L, 1L, 1L, 1L, NA, 1L),
X1 = c(NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_, NA_integer_),
X2 = c(NA, NA, NA, NA, NA, NA, 1L, NA, NA)), class = "data.frame", row.names = c(NA, -9L))
我正在努力定义问题,甚至开始处理。数据可以认为是long,但是因为多列,所以也是wide。
任何关于如何实现此输出文件的建议将不胜感激。
感谢您花时间阅读这篇文章。
编辑 1
从相关帖子Convert data from long format to wide format with multiple measure columns我尝试了以下解决方案:
library(data.table)
df2 = dcast(setDT(df), ID~Measures,
value.var=c("ME1", "ME2", "X1", "X2"))
但是,我收到警告:
Aggregate function missing, defaulting to 'length'
这意味着我的数据中的条目全部更改为 1 或 NA。我不希望这种情况发生。
编辑 2
当我在原始数据上测试建议的解决方案时,它失败了。为了更好地解释,我提供了一个与我的原始数据非常相似的小样本。现有的解决方案都不起作用。
dput(df)
structure(list(
Id = c("39fca07f-d62e-494a-4a86-8dec54836c08", "39fca8ee-fe3f-4c85-ab0a-acb3c2db1b9c", "39fca8ed-f34c-b7e3-4229-111155aabe35", "39fca8e9-1e08-1809-c7a8-d2c8a4bc9b00", "39fc6ae5-0de8-4820-eede-343e738e7a4a", "39fca8e9-fbf9-a098-cf8c-322810997ce9"), DeliverId = c("39fb74ce-d5e6-69f6-f733-ee5fbc4689e6", "39fb74ce-d5e6-69f6-f733-ee5fbc4689e6", "39fb74ce-d5e6-69f6-f733-ee5fbc4689e6", "39fb74ce-d5e6-69f6-f733-ee5fbc4689e6", "39fb74ce-d5e6-69f6-f733-ee5fbc4689e6", "39fb74ce-d5e6-69f6-f733-ee5fbc4689e6"),
DeliverN = c("1Assess", "1Assess", "1Assess", "1Assess", "1Assess", "1Assess"),
AssessRId = c("39fb74cf-5fb6-4248-6d08-0e36647e190b", "39fb74cf-5fb6-4248-6d08-0e36647e190b", "39fb74cf-5fb6-4248-6d08-0e36647e190b", "39fb74cf-5fb6-4248-6d08-0e36647e190b", "39fb74cf-5fb6-4248-6d08-0e36647e190b", "39fb74cf-5fb6-4248-6d08-0e36647e190b"),
AssessRN = c("P1", "P2", "P3", "P4", "P5", "P6"),
AssessTId = c("1ee2684c99fa", "fd2dbea08b43", "0e0177a33282", "091b8f805553", "6e5b9301116d", "7a307a90de19"),
AssessTN = c("Comm - 09 Narrate", "Comm - Prog Level Judge", "Comm - O Indi Level Judge", "Comm - 02 Int Prj", "Comm - 10 Learn Comm Participate",
"Comm - 05 Exhibit"),
S.Time = c("21/05/2020 19:47", "23/05/2020 11:06", "23/05/2020 11:05", "23/05/2020 10:59", "11/05/2020 9:58", "23/05/2020 11:00"),
F.Time = c("24/05/2020 11:02", "23/05/2020 11:06", "23/05/2020 11:05",
"23/05/2020 11:00", "23/05/2020 11:04", "23/05/2020 11:03"), CompletedIndi = c(8L, 1L, 8L, 8L, 8L, 8L),
TotalIndi = c(8L, 1L, 8L, 8L, 8L, 8L),
Progress = c(100L, 100L, 100L, 100L, 100L, 100L),
Build = c("Monice Island", "Pink Lasy", "", "", "", ""),
Advice = c("Monica", "Chandler", "", "", "", ""),
TechUserId = c(128L, 129L, 130L, 129L, 129L, 129L),
TechName = c("Barba", "Raymond", "Raymond", "Raymond", "Raymond","Raymond"), TechEmail = c("barber@123.com", "raymond@123.com", "raymond@123.com", "raymond@123.com", "raymond@123.com", "raymond@123.com"),
TechLife = c("0 - 2 years", "Over 10 years", "Over 10 years", "Over 10 years", "Over 10 years", "Over 10 years"),
OtherLife = c("0 - 2 years", "5 - 10 years", "5 - 10 years", "5 - 10 years", "5 - 10 years", "5 - 10 years"),
PersonUId = c(470L, 455L, 455L, 455L, 455L, 455L),
PersonDName = c("Tall Tiffany", "Sharp Steff", "Sharp Steff", "Sharp Steff", "Sharp Steff", "Sharp Steff"),
PersonFName = c("Tall", "Sharp", "Sharp", "Sharp", "Sharp", "Sharp"), PersonLName = c("Tiffany", "Steff", "Steff", "Steff", "Steff", "Steff"), PersonUID = c("2783-4409", "4307-4369", "4307-4369", "4307-4369", "4307-4369", "4307-4369"),
Gender = c("Female", "Female", "Female", "Female", "Female", "Female"), PYear = c(2023L, 2024L, 2024L, 2024L, 2024L, 2024L),
Course = c("Undergrad", "Grad", "Grad", "Grad", "Grad", "Grad"),
Special = c("Yes", "No", "No", "No", "No", "No"),
Q1 = c(2L, 1L, 3L, 3L, 2L, 2L),
Q2 = c(1L, NA, 2L, 2L, 1L, 2L),
Q3 = c(1L, NA, 3L, 3L, 2L, 2L),
Q4 = c(1L, NA, 3L, 3L, 2L, 1L),
Q5 = c(1L, NA, 2L, 2L, 1L, 2L),
Q6 = c(1L, NA, 0L, 1L, 1L, 1L),
Q7 = c(1L, NA, 2L, 1L, 2L, 2L),
Q8 = c(2L, NA, 2L, 1L, 2L, 2L),
Q9 = c(NA, NA, NA, NA, NA, NA),
Q10 = c(NA, NA, NA, NA, NA, NA),
X = c(NA, NA, NA, NA, NA, NA),
X.1 = c(NA, NA, NA, NA, NA, NA),
ListDetails = c("Missing", "Complete", "Complete", "Complete", "Complete", "Complete")),
class = "data.frame", row.names = c(NA, -6L))
想要的输出如下:
Id DeliverId DeliverN AssessRId AssessRN AssessTId S-Time F-Time CompletedIndi TotalIndi Progress Build Advice TechUserId TechName TechEmail TechLife OtherLife PersonDName PersonFName PersonLName PersonUID Gender PYear Course Special ListDetails PersonUId Q1_Comm - 02 Int Prj Q1_Comm - 05 Exhibit Q1_Comm - 09 Narrate Q1_Comm - 10 Learn Comm Participate Q1_Comm - O Indi Level Judge Q1_Comm - Prog Level Judge Q2_Comm - 02 Int Prj Q2_Comm - 05 Exhibit Q2_Comm - 09 Narrate Q2_Comm - 10 Learn Comm Participate Q2_Comm - O Indi Level Judge Q2_Comm - Prog Level Judge Q3_Comm - 02 Int Prj Q3_Comm - 05 Exhibit Q3_Comm - 09 Narrate Q3_Comm - 10 Learn Comm Participate Q3_Comm - O Indi Level Judge Q3_Comm - Prog Level Judge Q4_Comm - 02 Int Prj Q4_Comm - 05 Exhibit Q4_Comm - 09 Narrate Q4_Comm - 10 Learn Comm Participate Q4_Comm - O Indi Level Judge Q4_Comm - Prog Level Judge Q5_Comm - 02 Int Prj Q5_Comm - 05 Exhibit Q5_Comm - 09 Narrate Q5_Comm - 10 Learn Comm Participate Q5_Comm - O Indi Level Judge Q5_Comm - Prog Level Judge Q6_Comm - 02 Int Prj Q6_Comm - 05 Exhibit Q6_Comm - 09 Narrate Q6_Comm - 10 Learn Comm Participate Q6_Comm - O Indi Level Judge Q6_Comm - Prog Level Judge Q7_Comm - 02 Int Prj Q7_Comm - 05 Exhibit Q7_Comm - 09 Narrate Q7_Comm - 10 Learn Comm Participate Q7_Comm - O Indi Level Judge Q7_Comm - Prog Level Judge Q8_Comm - 02 Int Prj Q8_Comm - 05 Exhibit Q8_Comm - 09 Narrate Q8_Comm - 10 Learn Comm Participate Q8_Comm - O Indi Level Judge Q8_Comm - Prog Level Judge Q9_Comm - 02 Int Prj Q9_Comm - 05 Exhibit Q9_Comm - 09 Narrate Q9_Comm - 10 Learn Comm Participate Q9_Comm - O Indi Level Judge Q9_Comm - Prog Level Judge Q10_Comm - 02 Int Prj Q10_Comm - 05 Exhibit Q10_Comm - 09 Narrate Q10_Comm - 10 Learn Comm Participate Q10_Comm - O Indi Level Judge Q10_Comm - Prog Level Judge X_Comm - 02 Int Prj X_Comm - 05 Exhibit X_Comm - 09 Narrate X_Comm - 10 Learn Comm Participate X_Comm - O Indi Level Judge X_Comm - Prog Level Judge X.1_Comm - 02 Int Prj X.1_Comm - 05 Exhibit X.1_Comm - 09 Narrate X.1_Comm - 10 Learn Comm Participate X.1_Comm - O Indi Level Judge X.1_Comm - Prog Level Judge
39fca07f-d62e-494a-4a86-8dec54836c08 39fb74ce-d5e6-69f6-f733-ee5fbc4689e6 1Assess 39fb74cf-5fb6-4248-6d08-0e36647e190b P1 1ee2684c99fa 21/05/2020 19:47 24/05/2020 11:02 8 8 100 Monice Island Monica 128 Barba barber@123.com 0 - 2 years 0 - 2 years Tall Tiffany Tall Tiffany 2783-4409 Female 2023 Undergrad Yes Missing 470 NA NA 2 NA NA NA NA NA 1 NA NA NA NA NA 1 NA NA NA NA NA 1 NA NA NA NA NA 1 NA NA NA NA NA 1 NA NA NA NA NA 1 NA NA NA NA NA 2 NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA
39fca8ee-fe3f-4c85-ab0a-acb3c2db1b9c 39fb74ce-d5e6-69f6-f733-ee5fbc4689e6 1Assess 39fb74cf-5fb6-4248-6d08-0e36647e190b P2 fd2dbea08b43 23/05/2020 11:06 23/05/2020 11:06 1 1 100 Pink Lasy Chandler 129 Raymond raymond@123.com Over 10 years 5 - 10 years Sharp Steff Sharp Steff 4307-4369 Female 2024 Grad No Complete 455 3 2 NA 2 3 1 2 2 NA 1 2 NA 3 2 NA 2 3 NA 3 1 NA 2 3 NA 2 2 NA 1 2 NA 1 1 NA 1 0 NA 1 2 NA 2 2 NA 1 2 NA 2 2 NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA NA
注意:我不认为这个问题是重复的,因为以前的解决方案都不适用于我的数据集。我请求您再次打开我的问题并使其可见。
对于解决方案为何不适用于我的数据集的任何帮助或建议,我将不胜感激。
【问题讨论】:
-
在新的期望输出中,
AssessRN值 P3 到 P6 怎么样。你要放弃那些 -
如果我使用
df %>% pivot_wider(names_from = AssessTN, values_from = Q1:X.1),它确实可以正常工作而不会出现任何错误或警告。数据结构涉及到很多列,所以不清楚,而且你没有展示P3到P6,我理解有些困难