【问题标题】:Scrape a school's top247 college football recruits of all-time刮取一所学校有史以来排名前 247 位的大学橄榄球新兵
【发布时间】:2021-05-28 15:45:16
【问题描述】:

我正在尝试从以下网页抓取 google colab 上的表格:https://247sports.com/college/penn-state/Sport/Football/AllTimeRecruits/

下面是我尝试使用的python脚本...

Team = 'penn-state'

url = "https://247sports.com/college/" + str(Team) + "/Sport/Football/AllTimeRecruits/"

# Add the `user-agent` otherwise we will get blocked when sending the request
headers = {"user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/90.0.4430.93 Safari/537.36"}

response = requests.get(url, headers = headers).content
soup = BeautifulSoup(response, "html.parser")
data = []

for tag in soup.find_all("li", class_="ri-page__list-item"):  # `[1:]` Since the first result is a table header
    rank = tag.find_next("span", class_="all-time-rank").text
    school = tag.find_next("span", class_="meta").text
    year = tag.find_next("span", class_="meta").text
    name = tag.find_next("a", class_="ri-page__name-link").text
    position = tag.find_next("div", class_="position").text
    height_weight = tag.find_next("div", class_="metrics").text
    rating = tag.find_next("span", class_="score").text
    nat_rank = tag.find_next("a", class_="natrank").text
    state_rank = tag.find_next("a", class_="sttrank").text
    pos_rank = tag.find_next("a", class_="posrank").text
#    status = tag.find_next("p", class_="commit-date withDate").text

    data.append(
        {
            "Rank": rank,
            "Name": name,
            "School": school,
            "Class of": year,
            "Position": position,
            "Height & Weight": height_weight,
            "Rating": rating,
            "National Rank": nat_rank,
            "State Rank": state_rank,
            "Position Rank": pos_rank,
#            "Date": status,
        }
    )

df = pd.DataFrame(data)

df

我想获得一个关于该球员所在的招募班级年份的列。例如,如果一名球员来自“2005 年班级”,我希望“2005”作为“年份”列的列值。

    Rank    Name    School  Class of    Position    Height & Weight Rating  National Rank   State Rank  Position Rank
0   1   Derrick Williams    Eleanor Roosevelt (Greenbelt, MD)   Eleanor Roosevelt (Greenbelt, MD)   WR  6-0 / 190   0.9986  4   1   2
1   2   Micah Parsons   Harrisburg (Harrisburg, PA) Harrisburg (Harrisburg, PA) WDE 6-3 / 235   0.9982  5   1   2
2   3   Justin Shorter  South Brunswick (Monmouth Junction, NJ) ... South Brunswick (Monmouth Junction, NJ) ... WR  6-4 / 213   0.9962  8   1   1
3   4   Dan Connor  Strath Haven (Wallingford, PA)  Strath Haven (Wallingford, PA)  ILB 6-3 / 215   0.9944  13  1   2
4   5   Justin King Gateway (Monroeville, PA)   Gateway (Monroeville, PA)   CB  6-0 / 185   0.9942  15  1   2
... ... ... ... ... ... ... ... ... ... ...
242 243 Will Levis  Xavier (Middletown, CT) Xavier (Middletown, CT) PRO 6-4 / 222   0.8689  652 2   28
243 244 Troy Reeder Salesianum (Wilmington, DE) Salesianum (Wilmington, DE) ILB 6-2 / 230   0.8687  500 2   22
244 245 Jake Cooper Archbishop Wood (Warminster, PA)    Archbishop Wood (Warminster, PA)    ILB 6-1 / 220   0.8686  520 11  17
245 246 Jon Ditto   Gateway (Monroeville, PA)   Gateway (Monroeville, PA)   WR  6-3 / 221   0.8684  417 16  52
246 247 Shareef Miller  George Washington (Philadelphia, PA)    George Washington (Philadelphia, PA)    SDE 6-5 / 230   0.8681  525 12  27
247 rows × 10 columns

但是,我在学校得到了重复。那是因为在 html 中,在观察 html 代码时,在“span”下发现了高中和年份。话虽如此,有没有办法根据 html 的设置方式来抓取高中和年级?

对于如何完成这项工作的任何帮助将不胜感激。

【问题讨论】:

    标签: python-3.x pandas beautifulsoup


    【解决方案1】:

    您有两个 spans 和班级 meta - 第一个用于学校,第二个用于年级(始终按此顺序),因此您可以使用 find_all 找到两者,然后从 school 中提取第一个和第二个的year

    for tag in soup.find_all("li", class_="ri-page__list-item"):
        meta = tag.find_all("span", class_="meta")
        school = meta[0].text
        year = meta[1].text.replace('Class of ', '')
    
        # extract other fields...
        # data.append(...)
    

    【讨论】:

      猜你喜欢
      • 2021-09-22
      • 1970-01-01
      • 2023-04-04
      • 1970-01-01
      • 1970-01-01
      • 2011-06-07
      • 2015-12-19
      • 1970-01-01
      • 2015-08-16
      相关资源
      最近更新 更多