【问题标题】:Scrapy: customize Image pipeline with renaming defualt image nameScrapy:具有重命名默认图像名称的自定义图像管道
【发布时间】:2013-08-07 13:36:33
【问题描述】:

我正在使用图像管道从不同网站下载所有图像。

所有图像都成功下载到我定义的文件夹中,但在保存到硬盘之前我无法命名我选择的下载图像。

这是我的代码

pipelines.py

class jellyImagesPipeline(ImagesPipeline):


def image_key(self, url, item):
    name = item['image_name']
    return 'full/%s.jpg' % (name)


def get_media_requests(self, item, info):
    print'Entered get_media_request'
    for image_url in item['image_urls']:
        yield Request(image_url)

Image_spider.py

 def getImage(self, response):
 item = JellyfishItem()
 item['image_urls']= [response.url]
 item['image_name']= response.meta['image_name']
 return item

我需要在我的代码中做哪些更改??

更新 1


管道.py

class jellyImagesPipeline(ImagesPipeline):

    def image_custom_key(self, response):
        print '\n\n image_custom_key \n\n'
        name = response.meta['image_name'][0]
        img_key = 'full/%s.jpg' % (name)
        print "custom image key:", img_key
        return img_key
        
    def get_images(self, response, request, info):
        print "\n\n get_images \n\n"
        for key, image, buf, in super(jellyImagesPipeline, self).get_images(response, request, info):
            yield key, image, buf

        
        key = self.image_custom_key(response)
        orig_image = Image.open(StringIO(response.body))
        image, buf = self.convert_image(orig_image)
        yield key, image, buf
   
    def get_media_requests(self, item, info):
        print "\n\nget_media_requests\n"
        return [Request(x, meta={'image_name': item["image_name"]})
                for x in item.get('image_urls', [])]

更新 2


def image_key(self, image_name):
print 'entered into image_key'
    name = 'homeshop/%s.jpg' %(image_name)
    print name
    return name
    
def get_images(self,request):
    print '\nEntered into get_images'
    key = self.image_key(request.url)
yield key 

def get_media_requests(self, item, info):
print '\n\nEntered media_request'
print item['image_name']
    yield Request(item['image_urls'][0], meta=dict(image_name=item['image_name']))

def item_completed(self, results, item, info):
    print '\n\nentered into item_completed\n'
print 'Name : ', item['image_urls']
print item['image_name']
for tuple in results:
    print tuple

                 
        

【问题讨论】:

  • response.meta['image_name'] 中有什么内容?它仅依赖于 URL 吗?或者<img>@alt 或@title?
  • response.meta['image_name'] 是从 Mysql 表中检索的,它不依赖于 url。完全独立于url
  • 可以使用scrapy进化更简单的解决方案,请参阅Scrapy image download how to use custom filename

标签: image python-imaging-library scrapy


【解决方案1】:

pipelines.py

from scrapy.contrib.pipeline.images import ImagesPipeline
from scrapy.http import Request
from PIL import Image
from cStringIO import StringIO
import re

class jellyImagesPipeline(ImagesPipeline):

    CONVERTED_ORIGINAL = re.compile('^full/[0-9,a-f]+.jpg$')

    # name information coming from the spider, in each item
    # add this information to Requests() for individual images downloads
    # through "meta" dictionary
    def get_media_requests(self, item, info):
        print "get_media_requests"
        return [Request(x, meta={'image_name': item["image_name"]})
                for x in item.get('image_urls', [])]

    # this is where the image is extracted from the HTTP response
    def get_images(self, response, request, info):
        print "get_images"

        for key, image, buf, in super(jellyImagesPipeline, self).get_images(response, request, info):
            if self.CONVERTED_ORIGINAL.match(key):
                key = self.change_filename(key, response)
            yield key, image, buf

    def change_filename(self, key, response):
        return "full/%s.jpg" % response.meta['image_name'][0]

settings.py,确保你有

ITEM_PIPELINES = ['jelly.pipelines.jellyImagesPipeline']
IMAGES_STORE = '/path/to/where/you/want/to/store/images'

示例蜘蛛: 从 Python.org 的主页获取图像,保存图像的名称(和路径)将遵循站点结构,即在名为 www.python.org 的文件夹中

from scrapy.spider import BaseSpider
from scrapy.selector import HtmlXPathSelector
from scrapy.item import Item, Field
import urlparse

class CustomItem(Item):
    image_urls = Field()
    image_names = Field()
    images = Field()

class ImageSpider(BaseSpider):
    name = "customimg"
    allowed_domains = ["www.python.org"]
    start_urls = ['http://www.python.org']

    def parse(self, response):
        hxs = HtmlXPathSelector(response)
        sites = hxs.select('//img')
        items = []
        for site in sites:
            item = CustomItem()
            item['image_urls'] = [urlparse.urljoin(response.url, u) for u in site.select('@src').extract()]
            # the name information for your image
            item['image_name'] = ['whatever_you_want']
            items.append(item)
        return items

【讨论】:

  • 感谢您回答我的问题,但它并没有解决我的问题。图像名称仍然没有改变。请帮我挖掘一下
  • 我用经过测试的代码编辑了我的答案。您有 2 个选择:使用新名称创建另一个图像,或者从内置 ImagesPipeline 更改原始 JPEG 转换图像的名称@
  • 感谢@paul 的回答。我真的很感谢你的努力。我想选择您建议的第二个选项,即从内置 ImagePipeline 更改原始 JPEG 转换图像的名称。
  • 我已根据您的建议更新了有问题的代码,但程序根本没有进入get_imageimage_coustom_key。因此,图像被下载,名称不变。
  • 您是否在 settings.py 文件中设置了ITEM_PIPELINES = ['yourprojectname.pipelines.jellyImagesPipeline']?此外,请确保管道名称一致:在您编辑的问题代码中,JellyImagesPipeline 以“j”开头,然后是“J”(在 super() 调用中)
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 2020-05-29
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
相关资源
最近更新 更多