【发布时间】:2018-04-08 06:50:41
【问题描述】:
我正在开展一个快速抓取项目,其中涉及抓取历史 NFL 足球数据。下面是我的数据的快速浏览:
allgames_thisweek = c("Chicago Bears 21, Tampa Bay Buccaneers 9 -- Box Score", "Cleveland Browns 28, Cincinnati Bengals 20 -- Box Score",
"Dallas Cowboys 26, Pittsburgh Steelers 9 -- Box Score", "Detroit Lions 31, Atlanta Falcons 28 (OT) -- Box Score",
"Green Bay Packers 16, Minnesota Vikings 10 -- Box Score", "Indianapolis Colts 45, Houston Oilers 21 -- Box Score",
"Kansas City Chiefs 30, New Orleans Saints 17 -- Box Score",
"Los Angeles Rams 14, Arizona Cardinals 12 -- Box Score", "Miami Dolphins 39, New England Patriots 35 -- Box Score",
"New York Giants 28, Philadelphia Eagles 23 -- Box Score", "New York Jets 23, Buffalo Bills 3 -- Box Score",
"San Diego Chargers 37, Denver Broncos 34 -- Box Score", "San Francisco 49ers 44, Los Angeles Raiders 14 -- Box Score",
"Seattle Seahawks 28, Washington Redskins 7 -- Box Score")
allgames_thisweek[1]
"Chicago Bears 21, Tampa Bay Buccaneers 9 -- Box Score"
每一行有以下数据[team1, team1score, team2, team2score, --, Box Score]
我的数据格式都完全相同,这意味着第一队的得分后面总是有一个逗号,第二队的得分后面总是有一个--。我想创建一个包含 4 列(team1、team1score、team2、team2score)的数据框,因此输出可能如下所示:
output_df
team1 team1score team2 team2score
1. Chicago Bears 21 Tampba Bay Buccaneers 9
对我如何实现这一点有任何想法吗?任何帮助表示赞赏!谢谢
【问题讨论】:
-
类似这样的东西 - unlist(strsplit(allgames_thisweek[1], ',|--')) 把字符串变成3个字符串,这是一个好的开始
标签: r string dataframe data-manipulation