【发布时间】:2018-11-20 11:20:03
【问题描述】:
我有这个数据框df
df <- data.frame(stringsAsFactors=FALSE,
id = c(1L, 2L, 3L, 4L, 5L, 6L),
Country = c("ESP", "ESP", "ESP", "ITA", "ITA", "ITA"),
Year = c(1965L, 1965L, 1965L, 1965L, 1965L, 1965L),
Time.step = c("Month", "Month", "Month", "Month", "Month", "Month"),
GSA.numb = c("GSA 5", "GSA 5", "GSA 5", "GSA 17", "GSA 17", "GSA 17"),
Species = c("Mullus", "Mullus", "Mullus", "Eledone", "Eledone", "Eledone"),
Quantity = c(500L, 200L, 200L, 350L, 350L, 125L)
)
df
id Country Year Time.step GSA.numb Species Quantity
1 ESP 1965 Month GSA 5 Mullus 500
2 ESP 1965 Month GSA 5 Mullus 200
3 ESP 1965 Month GSA 5 Mullus 200
4 ITA 1965 Month GSA 17 Eledone 350
5 ITA 1965 Month GSA 17 Eledone 350
6 ITA 1965 Month GSA 17 Eledone 125
我有一些重复的行,如:3 和 5。 当行重复时,我可以为 F 或 T 逻辑值创建一列:
df$dup <- duplicated(df[,2:7]) #No id!
结果:
id Country Year Time.step GSA.numb Species Quantity dup
1 ESP 1965 Month GSA 5 Mullus 500 FALSE
2 ESP 1965 Month GSA 5 Mullus 200 FALSE
3 ESP 1965 Month GSA 5 Mullus 200 TRUE
4 ITA 1965 Month GSA 17 Eledone 350 FALSE
5 ITA 1965 Month GSA 17 Eledone 350 TRUE
6 ITA 1965 Month GSA 17 Eledone 125 FALSE
现在,我想要一个新列(以动态方式,我的真实df 非常大,有很多行、列和变量)当为 TRUE 时可以查看重复行数,像这样:
aspected.df
id Country Year Time.step GSA.numb Species Quantity dup ref
1 ESP 1965 Month GSA 5 Mullus 500 FALSE NA
2 ESP 1965 Month GSA 5 Mullus 200 FALSE NA
3 ESP 1965 Month GSA 5 Mullus 200 TRUE =id2
4 ITA 1965 Month GSA 17 Eledone 350 FALSE NA
5 ITA 1965 Month GSA 17 Eledone 350 TRUE =id4
6 ITA 1965 Month GSA 17 Eledone 125 FALSE NA
我试过了:
with(df, ave(as.character(Species), df[,2:6], FUN = make.unique))
但结果是:
[1] "Mullus" "Mullus.1" "Mullus.2" "Eledone" "Eledone.1" "Eledone.2"
我想我需要更多的代码输入。哪个功能有用? (duplicated,make.unit, row.names 等等……)
【问题讨论】:
-
请注意,您接受的答案与您想要的输出不相符;你可能想更正你的输出表。
标签: r dataframe duplicates rowname