【问题标题】:How to convert reference list to data frame?如何将参考列表转换为数据框?
【发布时间】:2019-10-08 06:38:07
【问题描述】:

我有一个参考列表,例如,

references <- c(
  "Dumitru, T.A., Smith, D., Chang, E.Z., and Graham, S.A., 2001, Uplift, exhumation, and deformation in the Japanese Mt Everest, Paleozoic and Mesozoic tectonic evolution of central Africa: from continental assembly to intracontinental deformation: Journal of Neverland, v. 3, no. 192, p. 71-199.",
  "Dumitru, T.A., Smith, D., Chang, E.Z., and Graham, S.A., 2001, Uplift, exhumation, and deformation in the Japanese Mt Everest, Paleozoic and Mesozoic tectonic evolution of central Africa: from continental assembly to intracontinental deformation: Journal of Neverland, no. 3.",
  "Dumitru, T.A., Smith, D., Chang, E.Z., and Graham, S.A., 2001, Uplift, exhumation, and deformation in the Japanese Mt Everest, Paleozoic and Mesozoic tectonic evolution of central Africa: from continental assembly to intracontinental deformation: Journal of Neverland, p. 71-199."
)

我尝试过(?&lt;=:)(?.*)(?=(v\.)|(no\.)|(p\.)),但正则表达式返回“从大陆组装到大陆变形:梦幻岛杂志,第 3 卷,第 3 期”。 192,页。不是我想要提取的。

(?<=:)(?:[^:].*?)(?=(, v\.)|(, no\.)|(, p\.))

我期待的是“梦幻岛日志”,但回归是“从大陆组装到大陆内变形:梦幻岛日志”

【问题讨论】:

  • 等等,你是真的用R,还是其他语言来做这个正则表达式?
  • 是的,我使用的是 R。所以实际上所有的 '\' 都应该是 '\\'

标签: r regex string regex-lookarounds regex-greedy


【解决方案1】:

这里我们只是匹配捕获组中最后一个冒号之前的文本到下一个逗号

stringr::str_match(references, ": ((?!:)[^,:]*),")[,2]
# [1] "Journal of Neverland" "Journal of Neverland" "Journal of Neverland"

【讨论】:

    【解决方案2】:

    你可以使用

    :\s*\K[^:]*?(?=,\s*(?:v|no|p)\.)
    

    regex demo

    详情

    • : - 冒号
    • \s* - 0+ 个空格
    • \K - 匹配重置运算符
    • [^:]*? - 除了: 之外的零个或多个字符,但尽可能少的*? 是非贪婪的
    • (?=,\s*(?:v|no|p)\.) - 正向前瞻,需要 ,,然后是 0+ 个空格,然后是 vnop,紧跟在当前位置右侧的 .

    在 R 中:

    regmatches(references, regexpr(":\\s*\\K[^:]*?(?=,\\s*(?:v|no|p)\\.)", references, perl=TRUE))
    

    R demo online:

    references <- c(
      "Dumitru, T.A., Smith, D., Chang, E.Z., and Graham, S.A., 2001, Uplift, exhumation, and deformation in the Japanese Mt Everest, Paleozoic and Mesozoic tectonic evolution of central Africa: from continental assembly to intracontinental deformation: Journal of Neverland, v. 3, no. 192, p. 71-199.",
      "Dumitru, T.A., Smith, D., Chang, E.Z., and Graham, S.A., 2001, Uplift, exhumation, and deformation in the Japanese Mt Everest, Paleozoic and Mesozoic tectonic evolution of central Africa: from continental assembly to intracontinental deformation: Journal of Neverland, no. 3.",
      "Dumitru, T.A., Smith, D., Chang, E.Z., and Graham, S.A., 2001, Uplift, exhumation, and deformation in the Japanese Mt Everest, Paleozoic and Mesozoic tectonic evolution of central Africa: from continental assembly to intracontinental deformation: Journal of Neverland, p. 71-199."
    )
    regmatches(references, regexpr(":\\s*\\K[^:]*?(?=,\\s*(?:v|no|p)\\.)", references, perl=TRUE))
    ## => [1] "Journal of Neverland" "Journal of Neverland" "Journal of Neverland"
    

    如果您更喜欢基于stringr 的解决方案,请使用任一

    > str_extract(references, "(?<=:\\s)[^:]*?(?=,\\s*(?:v|no|p)\\.)")
    [1] "Journal of Neverland" "Journal of Neverland" "Journal of Neverland"
    

    或者,如果: 之后的空格可以是 0 或多个:

    > str_match(references, ":\\s*([^:]*?)(?:,\\s*(?:v|no|p)\\.)")[,2]
    [1] "Journal of Neverland" "Journal of Neverland" "Journal of Neverland"
    

    【讨论】:

    • 进一步了解正则表达式:'(?=(, v\\.)|(, [(\\d{0,})|(p\\.)])| (, no\\.))' 字面上可以是 '(?=,\\s*(?:v|no|p|\d+p)\\.)' 。对于不包含 ':',使用 [^:]*? (非贪婪)。谢谢你的建议。真的很有帮助!
    • @JiulinGuo 如果我的解决方案对您有用,请考虑接受/支持答案。如果还有什么不清楚的地方请告知。
    【解决方案3】:

    这是gsub 解决方案

    gsub('.*: (.*?), (?=v|no|p).*','\\1', references, perl=TRUE)
    # [1] "Journal of Neverland" "Journal of Neverland" "Journal of Neverland"
    

    或者,也可以使用strsplit

    vapply(strsplit(references, ': *|, *', perl=TRUE),
           function (l) {
             k <- which(startsWith(l, 'p. ') | startsWith(l, 'v. ') | startsWith(l, 'no. '))
             k <- k[1] - 1
             return (l[k]) 
           }, character (1))
    # [1] "Journal of Neverland" "Journal of Neverland" "Journal of Neverland"
    

    【讨论】:

      猜你喜欢
      • 2019-11-16
      • 2019-06-26
      • 2020-07-29
      • 2019-09-13
      • 2020-03-04
      • 1970-01-01
      • 2020-04-15
      • 2020-04-17
      相关资源
      最近更新 更多