【发布时间】:2019-10-08 12:20:24
【问题描述】:
我有一个数据文件,我想量化从字符串/类别到数字的列。我有一个预制文件,其中包含大约 500 个不同的类别以及它需要成为的相应编号。
所以我的第一个文件看起来有点像:
Type_of_fruit
Banana
Apple
Apple
Kiwi
Passionfruit
Banana
Apple
Orange
Etc.
然后我有第二张表,看起来像这样(翻译表):
Banana | 1
Apple | 2
Kiwi | 3
Passionfruit | 4
Orange | 5
Mango | 6
Grape | 7
Etc.
并且想使用这个转换表在我的原始数据框中创建一个新的量化列:
Type_of_fruit_quantified
1
2
2
3
4
1
2
5
起初我想用 mutate 命令来做,例如 Mutate(Type_of_fruit_quantified = if_else(Type_of_fruit == “香蕉”, 1, if_else(Type_of_fruit == “苹果”, 2, 等等。等等。 但是,翻译表中有大约 500 个不同的类别,这将需要很长时间。我怎样才能更快地做到这一点,例如通过参考翻译表?
重新创建我的模拟数据:
Type_of_fruit <- c("Banana", "Apple", "Apple", "Kiwi", "Passionfruit", "Banana", "Apple", "Orange")
Type_of_fruit_df <- data.frame(Type_of_fruit)
Fruit <- c("Banana", "Apple", "Kiwi", "Passionfruit", "Orange", "Mango", "Grape")
Number <- c(1, 2, 3, 4, 5, 6, 7)
Translation_table <- data.frame(Fruit, Number)
【问题讨论】:
-
Type_of_fruit_df$Type_of_fruit_quantified <- Translation_table$Number[match(Translation_table$Fruit, Type_of_fruit_df$Type_of_fruit)] -
注意^总是比
left_join快
标签: r translation