【问题标题】:Filter images from tweets从推文中过滤图像
【发布时间】:2014-05-12 09:24:00
【问题描述】:

我对 tweepy 很陌生,我想知道如何追踪和存储用户在他/她的推文中发布的图像。我在教程中找到了几种获取用户推文的方法,但我找不到只过滤图像的方法。

我正在使用以下代码来获取用户推文。怎么可能只获取用户图片??

编辑:我像上面一样编辑我的代码:

auth = tweepy.OAuthHandler(CONSUMER_KEY, CONSUMER_SECRET)
auth.set_access_token(OAUTH_TOKEN, OAUTH_SECRET)
api = tweepy.API(auth)
timeline = api.user_timeline(count=10, screen_name = "zenitiss") 
for tweet in timeline: 
   for media in tweet.entities.get("media",[{}]):
      print media
      #checks if there is any media-entity
      if media.get("type",None) == "photo":
          # checks if the entity is of the type "photo"
          image_content=requests.get(media["media_url"])
          print image_content

但是,for 循环似乎不起作用。打印媒体行打印一个空对象。基本上,当我尝试打印用户的网址时,例如 karyperry,我得到了:

{u'url': u'http://t.co/TaP2JZrpxu', u'indices': [42, 64], u'expanded_url':  
u'http://youtu.be/7bDLIV96LD4', u'display_url': u'youtu.be/7bDLIV96LD4'}
{u'url': u'https://t.co/t3hv7VQiPG', u'indices': [42, 65], u'expanded_url': 
u'https://vine.co/v/MgvxZA2qKbV', u'display_url': u'vine.co/v/MgvxZA2qKbV'}
{u'url': u'http://t.co/vnJAAU7KN6', u'indices': [50, 72], u'expanded_url':
u'http://instagram.com/p/n01XZjv-fp/', u'display_url': u'instagram.com/p/n01XZjv-fp/'}
{u'url': u'http://t.co/NycqAwtcgo', u'indices': [78, 100], u'expanded_url':
u'http://bit.ly/1o7xQRj', u'display_url': u'bit.ly/1o7xQRj'}
{u'url': u'http://t.co/BG6ozuRD6D', u'indices': [111, 133], u'expanded_url':
u'http://www.johnnywujek.com/sos', u'display_url': u'johnnywujek.com/sos'}
{u'url': u'http://t.co/nWIQ9ruJ3f', u'indices': [88, 110], u'expanded_url':
u'http://uncf.us/1kSXIwF', u'display_url': u'uncf.us/1kSXIwF'}
{u'url': u'http://t.co/yTbOgqt9fw', u'indices': [101, 123], u'expanded_url':
u'http://instagram.com/p/nvxD8eP-SZ/', u'display_url': u'instagram.com/p/nvxD8eP-SZ/'}

大多数 url 是图像,但是当我在 tweet.entities.get("url",[{}]) 中将 'url' 而不是 'media' 放入循环中时。其中大部分是图片网址。

【问题讨论】:

  • 你能贴出带有图片的推文的源代码吗?
  • 推文的源代码一词是什么意思?
  • 我的意思是一条带有 pic .data 的推文,你可以从中提取图像
  • 如果我打印 tweet.text.encode("utf-8") 它将包含带有 url 的文本。
  • 你分享那个文本..作为输入

标签: python tweepy


【解决方案1】:

推文(它们的 JSON 表示)包含一个“媒体”实体,如 here 所述。假设推文中包含图像,Tweepy 应该按以下方式公开该类型的实体:

tweet.entities["media"]["media_url"]

因此,如果你想存储图像,你只需要下载它,f.e.通过python的请求库。尝试在您的代码中添加类似以下语句的内容(或根据您的需要进行修改):

for media in tweet.entities.get("media",[{}]):
    #checks if there is any media-entity
    if media.get("type",None) == "photo":
        # checks if the entity is of the type "photo"
        image_content=requests.get(media["media_url"])
        # save to file etc.

【讨论】:

  • 我有 alltweets 对象,如何单独访问每条推文?
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2013-05-12
  • 1970-01-01
  • 2011-12-14
  • 1970-01-01
  • 2013-02-21
相关资源
最近更新 更多