【发布时间】:2018-07-22 19:58:13
【问题描述】:
我有来自数据表df 的四列,我想从中得出第五列。当前的四个列是 - year、month、id 和 conflict。现在conflict 列只有 1 和 0,对于给定的 id 分组,一旦在一年中出现 1,那么在该年剩余的月份中出现 1。我想将conflict 列更改为新列conflict_mutated,如下所示:如果我们在一个给定的年份,其中任何一个月都包含 1 并且前一年在任何月份都包含 1,我想要当年的月份为conflict_mutated 全为 1,同时保留所有旧的 1。
如果我们有这样的数据:
year month id conflict
1989 6 33 0
1989 7 33 0
1989 8 33 1
1989 9 33 1
1989 10 33 1
1989 11 33 1
1989 12 33 1
1990 1 33 0
1990 3 33 0
1990 3 33 0
1990 4 33 0
1990 5 33 1
1990 6 33 1
1990 7 33 1
1990 8 33 1
1990 9 33 1
1990 10 33 1
1990 11 33 1
1990 12 33 1
所以我希望在第 1、2、3 和 4 个月期间 conlfict 中的 0 为 1,因为它们是相同的 id,并且 1989(前一年)和 1990 中都有 1。前面的示例数据将如下所示:
year month id conflict conflict_mutated
1989 6 33 0 0
1989 7 33 0 0
1989 8 33 1 1
1989 9 33 1 1
1989 10 33 1 1
1989 11 33 1 1
1989 12 33 1 1
1990 1 33 0 1
1990 3 33 0 1
1990 3 33 0 1
1990 4 33 0 1
1990 5 33 1 1
1990 6 33 1 1
1990 7 33 1 1
1990 8 33 1 1
1990 9 33 1 1
1990 10 33 1 1
1990 11 33 1 1
1990 12 33 1 1
我有一个解决方案,但需要将近 3 天才能完成。如下:
conflict_mutated = df$conflict
for (i in 1:length(nrow(df)) {
if (df$year[i] != 1989 & any(filter(df, id == df$id[i],
year == (df$year[i] - 1))$conflict == 1) &
any(filter(df, id == df$id[i], year == df$year[i])$conflict == 1))
{conflict_mutated[i] = 1}
有没有什么方法可以利用 group_by 和 mutate 来让这更快或更好?考虑到分组年份必须考虑并在条件逻辑中与不同的 id 相结合,无法考虑如何完成。
【问题讨论】: