【发布时间】:2016-08-11 12:50:07
【问题描述】:
Here is the Website I am trying to scrape http://livingwage.mit.edu/
具体网址来自
http://livingwage.mit.edu/states/01
http://livingwage.mit.edu/states/02
http://livingwage.mit.edu/states/04 (For some reason they skipped 03)
...all the way to...
http://livingwage.mit.edu/states/56
在这些 URL 中的每一个上,我都需要第二个表的最后一行:
http://livingwage.mit.edu/states/01 的示例
要求的税前年收入 $20,260 $42,786 $51,642 $64,767 $34,325 $42,305 $47,345 $53,206 $34,325 $47,691 56,934 美元 66,997 美元
期望输出:
阿拉巴马州 $20,260 $42,786 $51,642 $64,767 $34,325 $42,305 $47,345 $53,206 $34,325 $47,691 $56,934 $66,997
阿拉斯加 $24,070 $49,295 $60,933 $79,871 $38,561 $47,136 $52,233 $61,531 $38,561 $54,433 $66,316 $82,403
...
...
怀俄明州 $20,867 $42,689 $52,007 $65,892 $34,988 $41,887 $46,983 $53,549 $34,988 $47,826 $57,391 $68,424
经过2个小时的折腾,这是我目前所拥有的(我是初学者):
import requests, bs4
res = requests.get('http://livingwage.mit.edu/states/01')
res.raise_for_status()
states = bs4.BeautifulSoup(res.text)
state_name=states.select('h1')
table = states.find_all('table')[1]
rows = table.find_all('tr', 'odd')[4:]
result=[]
result.append(state_name)
result.append(rows)
当我在 Python 控制台中查看 state_name 和 rows 时,它给了我 html 元素
[<h1>Living Wag...Alabama</h1>]
和
[<tr class = "odd... </td> </tr>]
问题 1:这些是我想要的输出中的内容,但是我怎样才能让 python 以字符串格式而不是像上面的 HTML 格式给我呢?
问题 2:如何循环通过 request.get(url01 到 url56)?
感谢您的帮助。
如果你能提供一种更有效的方式来获取我的代码中的 rows 变量,我将不胜感激,因为我到达那里的方式不是很 Pythonic。
【问题讨论】:
标签: python web-scraping beautifulsoup