【问题标题】:Scrapy x path: only get first item in for loopScrapy x路径:仅获取for循环中的第一项
【发布时间】:2019-07-03 15:25:34
【问题描述】:

我正在尝试获取此页面每个元素的详细信息:https://www.mrlodge.de/wohnungen/

我经常使用 for 循环来执行此操作。然而这一次它只返回第一个元素。循环中一定有问题,因为当我使用 getall() 而不是 get() 时,我得到了我需要但没有排序的所有细节。

请帮忙

import scrapy

class MrlodgeSpiderSpider(scrapy.Spider):
    name = 'mrlodge_spider'

    payload = '''
        {mrl_ft%5Bfd%5D%5Bdate_from%5D=&mrl_ft%5Bfd%5D%5Brent_from%5D=1000&mrl_ft%5Bfd%5D%5Brent_to%5D=8500&mrl_ft%5Bfd%5D%5Bpersons%5D=1&mrl_ft%5Bfd%5D%5Bkids%5D=0&mrl_ft%5Bfd%5D%5Brooms_from%5D=1&mrl_ft%5Bfd%5D%5Brooms_to%5D=9&mrl_ft%5Bfd%5D%5Barea_from%5D=20&mrl_ft%5Bfd%5D%5Barea_to%5D=480&mrl_ft%5Bfd%5D%5Bsterm%5D=&mrl_ft%5Bfd%5D%5Bradius%5D=50&mrl_ft%5Bfd%5D%5Bmvv%5D=&mrl_ft%5Bfd%5D%5Bobjecttype_cb%5D%5B%5D=w&mrl_ft%5Bfd%5D%5Bobjecttype_cb%5D%5B%5D=h&mrl_ft%5Bpage%5D=1}
    '''
    def start_requests(self):
        yield scrapy.Request(url='https://www.mrlodge.de/wohnungen/', method='POST',
                    body = self.payload, headers={"content-type": "application/json"})

    def parse(self, response):
        for apartment in response.xpath("//div[@class='mrl-ft-results mrlobject-list']"):
            yield {
                'info': apartment.xpath(".//div[@class='obj-smallinfo']/text()").get()
            }

【问题讨论】:

    标签: python-3.x xpath web-scraping scrapy


    【解决方案1】:

    您需要更改第一个 xpath 查询

    class MrlodgeSpiderSpider(scrapy.Spider):
        name = 'mrlodge_spider'
    
        payload = '''
        {mrl_ft%5Bfd%5D%5Bdate_from%5D=&mrl_ft%5Bfd%5D%5Brent_from%5D=1000&mrl
        _ft%5Bfd%5D%5Brent_to%5D=8500&mrl_ft%5Bfd%5D%5Bpersons%5D=1&mrl_ft%5Bfd
        %5D%5Bkids%5D=0&mrl_ft%5Bfd%5D%5Brooms_from%5D=1&mrl_ft%5Bfd%5D%5Brooms
        _to%5D=9&mrl_ft%5Bfd%5D%5Barea_from%5D=20&mrl_ft%5Bfd%5D%5Barea_to%5D=4
        80&mrl_ft%5Bfd%5D%5Bsterm%5D=&mrl_ft%5Bfd%5D%5Bradius%5D=50&mrl_ft%5Bfd
        %5D%5Bmvv%5D=&mrl_ft%5Bfd%5D%5Bobjecttype_cb%5D%5B%5D=w&mrl_ft%5Bfd%5D%
        5Bobjecttype_cb%5D%5B%5D=h&mrl_ft%5Bpage%5D=1}
    '''
    
        def start_requests(self):
            yield scrapy.Request(
                url='https://www.mrlodge.de/wohnungen/',
                method='POST',
                body=self.payload,
                headers={"content-type": "application/json"},
            )
    
        def parse(self, response):
            for apartment in response.xpath('//div[@class="mrlobject-list__item mrlobject-row"]'):
                yield {
                    'info': apartment.xpath(".//div[@class='obj-smallinfo']/text()").get()
                }
    

    【讨论】:

    • 谢谢,你是救命稻草。//是我的问题线索
    【解决方案2】:

    尝试使用

    //div[contains(@class,'mrlobject-row')] 
    

    而不是

    //div[@class='mrl-ft-results mrlobject-list']
    

    得到想要的结果。

    【讨论】:

      猜你喜欢
      • 2010-11-02
      • 2010-09-19
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2018-05-19
      • 1970-01-01
      相关资源
      最近更新 更多