【问题标题】:Convert string vector to dataframe in R在R中将字符串向量转换为数据框
【发布时间】:2018-04-08 06:50:41
【问题描述】:

我正在开展一个快速抓取项目,其中涉及抓取历史 NFL 足球数据。下面是我的数据的快速浏览:

allgames_thisweek = c("Chicago Bears 21, Tampa Bay Buccaneers 9 -- Box Score", "Cleveland Browns 28, Cincinnati Bengals 20 -- Box Score", 
"Dallas Cowboys 26, Pittsburgh Steelers 9 -- Box Score", "Detroit Lions 31, Atlanta Falcons 28 (OT)  -- Box Score", 
"Green Bay Packers 16, Minnesota Vikings 10 -- Box Score", "Indianapolis Colts 45, Houston Oilers 21 -- Box Score", 
"Kansas City Chiefs 30, New Orleans Saints 17 -- Box Score", 
"Los Angeles Rams 14, Arizona Cardinals 12 -- Box Score", "Miami Dolphins 39, New England Patriots 35 -- Box Score", 
"New York Giants 28, Philadelphia Eagles 23 -- Box Score", "New York Jets 23, Buffalo Bills 3 -- Box Score", 
"San Diego Chargers 37, Denver Broncos 34 -- Box Score", "San Francisco 49ers 44, Los Angeles Raiders 14 -- Box Score", 
"Seattle Seahawks 28, Washington Redskins 7 -- Box Score")

allgames_thisweek[1]
"Chicago Bears 21, Tampa Bay Buccaneers 9 -- Box Score"

每一行有以下数据[team1, team1score, team2, team2score, --, Box Score]

我的数据格式都完全相同,这意味着第一队的得分后面总是有一个逗号,第二队的得分后面总是有一个--。我想创建一个包含 4 列(team1、team1score、team2、team2score)的数据框,因此输出可能如下所示:

output_df
            team1    team1score                  team2   team2score
1.  Chicago Bears            21  Tampba Bay Buccaneers            9

对我如何实现这一点有任何想法吗?任何帮助表示赞赏!谢谢

【问题讨论】:

  • 类似这样的东西 - unlist(strsplit(allgames_thisweek[1], ',|--')) 把字符串变成3个字符串,这是一个好的开始

标签: r string dataframe data-manipulation


【解决方案1】:

您可以使用dplyr + stringr 做到这一点:

library(dplyr)
library(stringr)

string %>%
  str_replace("(?<=\\d)\\s.*--.+$", "") %>%
  str_replace_all("\\s(?=\\d+\\b)", ",") %>%
  strsplit(",") %>%
  do.call(rbind, .) %>%
  data.frame() %>%
  setNames(c("team1", "team1score", "team2", "team2score"))

结果:

                 team1 team1score                 team2 team2score
1        Chicago Bears         21  Tampa Bay Buccaneers          9
2     Cleveland Browns         28    Cincinnati Bengals         20
3       Dallas Cowboys         26   Pittsburgh Steelers          9
4        Detroit Lions         31       Atlanta Falcons         28
5    Green Bay Packers         16     Minnesota Vikings         10
6   Indianapolis Colts         45        Houston Oilers         21
7   Kansas City Chiefs         30    New Orleans Saints         17
8     Los Angeles Rams         14     Arizona Cardinals         12
9       Miami Dolphins         39  New England Patriots         35
10     New York Giants         28   Philadelphia Eagles         23
11       New York Jets         23         Buffalo Bills          3
12  San Diego Chargers         37        Denver Broncos         34
13 San Francisco 49ers         44   Los Angeles Raiders         14
14    Seattle Seahawks         28   Washington Redskins          7

注意事项:

  1. (?&lt;=\\d)\\s.*--.+$ 匹配空格 (\\s) 后跟任意字符零次或多次 (.*)、文字 --、任意字符一次或多次 (.+),并以字符串结尾 ($)。这个模式有一个额外的条件,它必须跟随一个数字(?&lt;=\\d)
  2. (?&lt;=...) 被称为正向后视,它检查 after 之后的内容是否紧跟 ... 中的模式。
  3. \\s(?=\\d+\\b) 匹配紧跟在 ((?=...)) 数字一次或多次 和单词边界 (\\b) 之后的空格。所以这匹配了团队名称和团队分数之间的空格。
  4. (?=...) 是一个正向前瞻,它检查 之前 中的内容是否立即遵循 ... 中的模式。

数据:

string = c("Chicago Bears 21, Tampa Bay Buccaneers 9 -- Box Score", "Cleveland Browns 28, Cincinnati Bengals 20 -- Box Score", 
       "Dallas Cowboys 26, Pittsburgh Steelers 9 -- Box Score", "Detroit Lions 31, Atlanta Falcons 28 (OT)  -- Box Score", 
       "Green Bay Packers 16, Minnesota Vikings 10 -- Box Score", "Indianapolis Colts 45, Houston Oilers 21 -- Box Score", 
       "Kansas City Chiefs 30, New Orleans Saints 17 -- Box Score", 
       "Los Angeles Rams 14, Arizona Cardinals 12 -- Box Score", "Miami Dolphins 39, New England Patriots 35 -- Box Score", 
       "New York Giants 28, Philadelphia Eagles 23 -- Box Score", "New York Jets 23, Buffalo Bills 3 -- Box Score", 
       "San Diego Chargers 37, Denver Broncos 34 -- Box Score", "San Francisco 49ers 44, Los Angeles Raiders 14 -- Box Score", 
       "Seattle Seahawks 28, Washington Redskins 7 -- Box Score")

【讨论】:

  • 我在哪里可以了解更多关于在 str_replace() 中使用诸如 "(?
  • @Canovic 为正则表达式添加了一些解释。希望这会有所帮助!
  • 它们被称为正则表达式。 R documentation
  • @Canovic 后视和前瞻是“PERL 正则表达式”的一部分。也许用这个词搜索会有所帮助。
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2017-04-25
  • 2013-06-04
  • 1970-01-01
  • 2012-12-12
  • 2016-05-03
相关资源
最近更新 更多