【问题标题】:Extract year from string and text data从字符串和文本数据中提取年份
【发布时间】:2016-06-13 03:09:42
【问题描述】:

我需要从具有这些性质的值的向量中提取开始年份和结束年份。

 yr<- c("June 2013 – Present (2 years 9 months)", "January 2012 – June 2013 (1 year 6 months)","2006 – Present (10 years)","2002 – 2006 (4 years)")


 yr
 June 2013 – Present (2 years 9 months)
 January 2012 – June 2013 (1 year 6 months)
 2006 – Present (10 years)
 2002 – 2006 (4 years)

我期待这样的输出。有人有建议吗?

 start_yr       end_yr

2013            2016
2012            2013
2006            2016
2002            2006

【问题讨论】:

  • gsub "present" 与 2016 并提取四位数字。试试看

标签: regex r lubridate stringi


【解决方案1】:
x <- gsub("present", "2016", yr, ignore.case = TRUE)
x <- regmatches(x, gregexpr("\\d{4}", x))
start_yr <- sapply(x, "[[", 1)
end_yr <- sapply(x, "[[", 2)

这会将开始年份和结束年份保存在 2 个单独的变量中,如果您希望将它们放在一个变量中,只需编辑代码并设置 y$start_yr y$end_yr

【讨论】:

  • 我有一个叫做“character(0)”的东西,它正在蔓延并得到这个错误“FUN(X[[i]], ...) 中的错误:下标越界”。有关删除该行的任何建议?
【解决方案2】:

另一种解决方案是使用stringr

library(stringr)
x <- str_replace(yr, "Present", 2016)
DF <- as.data.frame(str_extract_all(x, "\\d{4}", simplify = T))
names(DF) <- c("start_yr", "end_yr")
DF

你会得到

      start_yr end_yr
1     2013   2016
2     2012   2013
3     2006   2016
4     2002   2006

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2021-04-14
    • 1970-01-01
    • 2020-09-18
    • 2021-05-10
    • 2023-03-17
    • 1970-01-01
    • 1970-01-01
    • 2012-11-22
    相关资源
    最近更新 更多