【发布时间】:2020-04-29 03:59:47
【问题描述】:
我的数据(csv 文件)有一列包含无意义的字符(例如特殊字符、随机小写字母),我想删除它们。
df <- data.frame(Affiliation = c(". Biotechnology Centre, Malaysia Agricultural Research and Development Institute (MARDI), Serdang, Malaysia","**Institute for Research in Molecular Medicine (INFORMM), Universiti Sains Malaysia, Pulau Pinang, Malaysia","aas Massachusetts General Hospital and Harvard Medical School, Center for Human Genetic Research and Department of Neurology , Boston , MA , USA","ac Albert Einstein College of Medicine , Department of Pathology , Bronx , NY , USA"))
每行我要删除的字符数(例如“.”、“**”、“aas”、“ac”)是不确定的,如上所示。
预期输出:
df <- data.frame(Affiliation = c("Biotechnology Centre, Malaysia Agricultural Research and Development Institute (MARDI), Serdang, Malaysia","Institute for Research in Molecular Medicine (INFORMM), Universiti Sains Malaysia, Pulau Pinang, Malaysia","Massachusetts General Hospital and Harvard Medical School, Center for Human Genetic Research and Department of Neurology , Boston , MA , USA","Albert Einstein College of Medicine , Department of Pathology , Bronx , NY , USA"))
我正在考虑使用 dplyr 的 mutate 函数,但我不确定如何去做。
【问题讨论】:
-
你想从
aasvogel(南非秃鹰的古老术语)中删除aas吗?如果不是,确定是否删除aas或任何其他命名字符串的规则 是什么?