【发布时间】:2021-03-05 07:43:09
【问题描述】:
我正在尝试使用concurrent.futures 在以下脚本中实现多处理。问题是即使我使用concurrent.futures,性能仍然相同。它似乎对执行过程没有任何影响,这意味着它无法提高性能。
我知道如果我创建另一个函数并将从get_titles() 填充的链接传递给该函数以便从它们的内页中刮取标题,我可以使这个concurrent.futures 工作。但是,我希望使用我在下面创建的功能从登录页面获取标题。
我使用迭代方法而不是递归只是因为如果我选择后者,当调用超过 1000 次时,该函数将抛出递归错误。
这是我迄今为止尝试过的方式 (the site link that I've used within the script is a placeholder):
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
import concurrent.futures as futures
base = 'https://stackoverflow.com'
link = 'https://stackoverflow.com/questions/tagged/web-scraping'
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.190 Safari/537.36',
}
def get_titles(link):
while True:
res = requests.get(link,headers=headers)
soup = BeautifulSoup(res.text,"html.parser")
for item in soup.select(".summary > h3"):
post_title = item.select_one("a.question-hyperlink").get("href")
print(urljoin(base,post_title))
next_page = soup.select_one(".pager > a[rel='next']")
if not next_page: return
link = urljoin(base,next_page.get("href"))
if __name__ == '__main__':
with futures.ThreadPoolExecutor(max_workers=5) as executor:
future_to_url = {executor.submit(get_titles,url): url for url in [link]}
futures.as_completed(future_to_url)
问题:
如何在解析着陆页链接时提高性能?
编辑: 我知道我可以按照以下路线实现相同的目标,但 这不是我最初尝试的样子
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
import concurrent.futures as futures
base = 'https://stackoverflow.com'
links = ['https://stackoverflow.com/questions/tagged/web-scraping?tab=newest&page={}&pagesize=30'.format(i) for i in range(1,5)]
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/88.0.4324.190 Safari/537.36',
}
def get_titles(link):
res = requests.get(link,headers=headers)
soup = BeautifulSoup(res.text,"html.parser")
for item in soup.select(".summary > h3"):
post_title = item.select_one("a.question-hyperlink").get("href")
print(urljoin(base,post_title))
if __name__ == '__main__':
with futures.ThreadPoolExecutor(max_workers=5) as executor:
future_to_url = {executor.submit(get_titles,url): url for url in links}
futures.as_completed(future_to_url)
【问题讨论】:
-
请注意,您永远不会跳出
while True:循环,尝试向寻呼机的最后一页发出无限数量的请求。当没有next_page时,您可能需要break。 -
我忘了包括这一行
if not next_page: return。我已经编辑了脚本。感谢@MatsLindh 的指点。 -
@robots.txt this might be of interest
-
将
html.parser更改为lxml:) 你会看到性能提升
标签: python python-3.x web-scraping concurrent.futures