【发布时间】:2017-05-30 22:16:53
【问题描述】:
我已经通过脚本实现了我的蜘蛛,就像主要示例一样:
import scrapy
class BlogSpider(scrapy.Spider):
name = 'blogspider'
start_urls = ['https://blog.scrapinghub.com']
def parse(self, response):
for title in response.css('h2.entry-title'):
yield {'title': title.css('a ::text').extract_first()}
next_page = response.css('div.prev-post > a ::attr(href)').extract_first()
if next_page:
yield scrapy.Request(response.urljoin(next_page), callback=self.parse)
我运行:
scrapy runspider myspider.py
如果我没有设置或从 startproject 创建,如何更改用户代理?在这里指定:
【问题讨论】: