【问题标题】:Scraping Wikipedia tables with Python selectively有选择地使用 Python 抓取维基百科表格
【发布时间】:2018-05-15 16:56:34
【问题描述】:

我在整理 wiki 表格时遇到了麻烦,希望以前做过的人能给我建议。 从 List_of_current_heads_of_state_and_government 我需要国家(使用下面的代码),然后只需要第一次提到国家元首 + 他们的名字。我不确定如何隔离第一次提及,因为它们都在一个单元格中。我试图提取他们的名字给了我这个错误:IndexError: list index out of range。感谢您的帮助!

import requests
from bs4 import BeautifulSoup

wiki = "https://en.wikipedia.org/wiki/List_of_current_heads_of_state_and_government"
website_url = requests.get(wiki).text
soup = BeautifulSoup(website_url,'lxml')

my_table = soup.find('table',{'class':'wikitable plainrowheaders'})
#print(my_table)

states = []
titles = []
names = []
for row in my_table.find_all('tr')[1:]:
    state_cell = row.find_all('a')[0]  
    states.append(state_cell.text)
print(states)
for row in my_table.find_all('td'):
    title_cell = row.find_all('a')[0]
    titles.append(title_cell.text)
print(titles)
for row in my_table.find_all('td'):
    name_cell = row.find_all('a')[1]
    names.append(name_cell.text)
print(names)

理想的输出是 pandas df:

State | Title | Name |

【问题讨论】:

    标签: python-3.x web-scraping beautifulsoup wikipedia


    【解决方案1】:

    我找到了一个超级简单快捷的方法,通过导入 wikipedia python 模块,然后使用 pandas 的 read_html 将其放入数据框。

    您可以从那里应用任何数量的分析。

    import pandas as pd
    import wikipedia as wp
    html = wp.page("List_of_video_games_considered_the_best").html().encode("UTF-8")
    try: 
        df = pd.read_html(html)[1]  # Try 2nd table first as most pages contain contents table first
    except IndexError:
        df = pd.read_html(html)[0]
    print(df.to_string())
    

    或者如果你想从命令行调用它:

    只需拨打python yourfile.py -p Wikipedia_Page_Article_Here

    import pandas as pd
    import argparse
    import wikipedia as wp
    parser = argparse.ArgumentParser()
    parser.add_argument("-p", "--wiki_page", help="Give a wiki page to get table", required=True)
    args = parser.parse_args()
    html = wp.page(args.wiki_page).html().encode("UTF-8")
    try: 
        df = pd.read_html(html)[1]  # Try 2nd table first as most pages contain contents table first
    except IndexError:
        df = pd.read_html(html)[0]
    print(df.to_string())
    

    希望这对那里的人有所帮助!

    【讨论】:

      【解决方案2】:

      如果我能理解您的问题,那么以下内容应该可以帮助您:

      import requests
      from bs4 import BeautifulSoup
      
      URL = "https://en.wikipedia.org/wiki/List_of_current_heads_of_state_and_government"
      
      res = requests.get(URL).text
      soup = BeautifulSoup(res,'lxml')
      for items in soup.find('table', class_='wikitable').find_all('tr')[1::1]:
          data = items.find_all(['th','td'])
          try:
              country = data[0].a.text
              title = data[1].a.text
              name = data[1].a.find_next_sibling().text
          except IndexError:pass
          print("{}|{}|{}".format(country,title,name))
      

      输出:

      Afghanistan|President|Ashraf Ghani
      Albania|President|Ilir Meta
      Algeria|President|Abdelaziz Bouteflika
      Andorra|Episcopal Co-Prince|Joan Enric Vives Sicília
      Angola|President|João Lourenço
      Antigua and Barbuda|Queen|Elizabeth II
      Argentina|President|Mauricio Macri
      

      等等----

      【讨论】:

      • 或者你也可以这样尝试print(country,title,name,sep=" | ")。谢谢。
      • 是的,这正是我想要的。谢谢!
      【解决方案3】:

      它并不完美,但几乎可以这样工作。

      import requests
      from bs4 import BeautifulSoup
      
      wiki = "https://en.wikipedia.org/wiki/List_of_current_heads_of_state_and_government"
      website_url = requests.get(wiki).text
      soup = BeautifulSoup(website_url,'lxml')
      
      my_table = soup.find('table',{'class':'wikitable plainrowheaders'})
      #print(my_table)
      
      states = []
      titles = []
      names = []
      """ for row in my_table.find_all('tr')[1:]:
          state_cell = row.find_all('a')[0]  
          states.append(state_cell.text)
      print(states)
      for row in my_table.find_all('td'):
          title_cell = row.find_all('a')[0]
          titles.append(title_cell.text)
      print(titles) """
      for row in my_table.find_all('td'):
          try:
              names.append(row.find_all('a')[1].text)
          except IndexError:
              names.append(row.find_all('a')[0].text)
      
      print(names)
      

      到目前为止,我可以看到此名称列表中只有一个错误。由于您必须编写异常,该表有点困难。例如,有些名称不是链接,然后代码仅捕获它在该行中找到的第一个链接。但是你只需要为这种情况编写更多的 if 子句。至少我会这样做。

      【讨论】:

        猜你喜欢
        • 1970-01-01
        • 1970-01-01
        • 1970-01-01
        • 2020-07-16
        • 2020-07-08
        • 2019-05-24
        • 1970-01-01
        • 2016-09-08
        • 1970-01-01
        相关资源
        最近更新 更多