【问题标题】:How do I scrape /html/head/script fields?如何抓取 /html/head/script 字段?
【发布时间】:2018-12-21 15:22:51
【问题描述】:

我是编程和抓取的新手。有没有办法刮掉它们而不是仅仅加载页面并将其拆开?

例子:

> <script> window.initialState =
> {"ACCOUNT":{"type":"PRODUCTUNIQUE","universe":"Woman","sku":"M1286ZTDT_M884_TU","code":"M1286ZTDT_M884","price":{"value":2950,"currency":"USD"},"status":"NOTFORSALE","eReservation":false,"hasSizeGuide":false,"tracking":[{"events":["addToCart"],"addToCartType":"regular","pageType":"CDC_ProductPage","ecommerce":{"currencyCode":"USD","add":{"products":{"id":"M1286ZTDT_M884_TU","name":"dior
> book tote toile de jouy bag","price":2950,"brand":"Dior Book
> Tote","category":"women/handbags/shopping bags/dior book
> tote","variant":"Multi-coloured","quantity":1,"dimension16":"M1286ZTDT_M884","dimension32":"not
> engraved"}}}}]},{"type":"PRODUCTSECTIONDESCRIPTION","sections":[{"title":"THE
> DESCRIPTION","content":"Dior Book Tote bag in canvas embroidered with
> a multi-coloured Toile de Jouy motif.<br /><br />Reference :
> M1286ZTDT_M884","type":"TEXT"},{"title":"THE
> CHARACTERISTICS","content":"Carried in the hand or on the shoulder <br
> />\nDimensions: 41.5 x 32 x 5
> cm","type":"TEXT"}]},{"type":"PRODUCTDECLINATIONS","declinations":[{"title":"Dior
> Book Tote Toile de Jouy
> bag","color":"Blue","colorCode":"33","uri":"/couture/en_us/horizon/products/couture-M1286ZTDT_M928_TU-dior-book-tote-toile-de-jouy-bag","image":{"target":"DESKTOP","uri":"https://wwws.dior.com/couture/ecommerce/media/catalog/product/cache/1/grid_image_1/460x497/17f82f742ffe127f42dca9de82fb58b1/M/1/1540309423_M1286ZTDT_M928_E01_GH.jpg","width":460,"height":497,"alt":"Click
> here to enlarge the product picture Dior Book Tote Toile de Jouy
> bag"}},{"title":"Dior Book Tote Toile de Jouy
> bag","color":"Burgundy","colorCode":"44","uri":"/couture/en_us/horizon/products/couture-M1286ZTDT_M974_TU-dior-book-tote-toile-de-jouy-bag","image":
> <a...... </script>

================================================ ==========================

【问题讨论】:

    标签: css xpath scrapy


    【解决方案1】:

    您可以像定位任何其他元素一样定位这些 script 元素 - 例如使用 xpath 和 css 选择器:

    script_text = response.xpath("//script[contains(., 'window.initialState')]").extract_first()
    

    然后,为了从脚本文本中提取有用的数据,您可以采用不同的方法 - 通常的方法是使用正则表达式从脚本中提取所需的对象(数组或对象/字典)文本,然后通过json.loads() 将其加载到 Python 数据结构中。

    另一种方法是使用 JS 解析器,例如 slimit,它会在 JavaScript 代码上为您提供类似 ast 的接口。这是working example of using slimit

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2019-11-26
      • 1970-01-01
      • 2016-05-14
      • 2015-06-11
      • 2021-11-18
      • 2015-10-25
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多