【问题标题】:Iteration over the dictionary and extracting values迭代字典并提取值
【发布时间】:2018-05-19 00:01:59
【问题描述】:

我有一个字典(result_dict)如下。

{'11333216@N05': {'person': {'can_buy_pro': 0,
   'description': {'_content': ''},
   'has_stats': '1',
   'iconfarm': 3,
   'iconserver': '2214',
   'id': '11333216@N05',
   'ispro': 0,
   'location': {'_content': ''},
   'mbox_sha1sum': {'_content': '8eb2e248cbad94e2b4a5aae75eb653c7e061a90c'},
   'mobileurl': {'_content': 'https://m.flickr.com/photostream.gne?id=11327876'},
   'nsid': '11333216@N05',
   'path_alias': 'kishansamarasinghe',
   'photos': {'count': {'_content': 442},
    'firstdate': {'_content': '1193073180'},
    'firstdatetaken': {'_content': '2000-01-01 00:49:17'}},
   'photosurl': {'_content': 'https://www.flickr.com/photos/kishansamarasinghe/'},
   'profileurl': {'_content': 'https://www.flickr.com/people/kishansamarasinghe/'},
   'realname': {'_content': 'Kishan Samarasinghe'},
   'timezone': {'label': 'Sri Jayawardenepura',
    'offset': '+06:00',
    'timezone_id': 'Asia/Colombo'},
   'username': {'_content': 'Three Sixty Five Degrees'}},
  'stat': 'ok'},
 '117692977@N08': {'person': {'can_buy_pro': 0,
   'description': {'_content': ''},
   'has_stats': '0',
   'iconfarm': 1,
   'iconserver': '404',
   'id': '117692977@N08',
   'ispro': 0,
   'location': {'_content': 'Almere, The Nederlands'},
   'mobileurl': {'_content': 'https://m.flickr.com/photostream.gne?id=117600164'},
   'nsid': '117692977@N08',
   'path_alias': 'meijsvo',
   'photos': {'count': {'_content': 3237},
    'firstdate': {'_content': '1392469161'},
    'firstdatetaken': {'_content': '2013-06-23 14:39:30'}},
   'photosurl': {'_content': 'https://www.flickr.com/photos/meijsvo/'},
   'profileurl': {'_content': 'https://www.flickr.com/people/meijsvo/'},
   'realname': {'_content': 'Markéta Eijsvogelová'},
   'timezone': {'label': 'Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna',
    'offset': '+01:00',
    'timezone_id': 'Europe/Amsterdam'},
   'username': {'_content': 'meijsvo'}},
  'stat': 'ok'},
 '21539776@N02': {'person': {'can_buy_pro': 0,
   'description': {'_content': ''},
   'has_stats': '1',
   'iconfarm': 0,
   'iconserver': '0',

这包含超过 150 个用户名 (e.g. 11333216@N05) 。我想为每个用户提取'mobileurl' 并创建一个包含usernamemobileurl 列的数据框。我找不到迭代每个用户并提取他的mobileurl 的方法,因为索引是不可能的。但是,我已经为其中一位用户提取了 mobileurl,如下所示。

result_dict['76617062@N08']["person"]["mobileurl"]['_content']

'https://m.flickr.com/photostream.gne?id=76524249'

如果有人可以提供帮助将不胜感激,因为我对 python 有点陌生。

【问题讨论】:

  • "for user in result_dict.keys()" 将遍历所有用户
  • 您确实应该发布有效的 json,我们可以直接将其加载到设施中。即你的情况下的熊猫。

标签: python json pandas dictionary flickr


【解决方案1】:

我认为您也可以尝试更多地使用熊猫方式而不是纯粹的字典迭代。它不一定是最快的,但鉴于您是 python 和 pandas 的新手,我认为知道 pandas 可以很好地处理这个是一件好事。

我假设您使用的是 pandas DataFrame,而不仅仅是字典。您无需将 json 转换为 pandas DataFrame 即可轻松实现相同的目的。即即使您不是熊猫DataFrame,其他答案也可以使用。它们也是有效的 Python 字典语法。

urls = result_dict[result_dict.index=='person'].apply(lambda x: x['mobileurl']['_content'])

在这里,我们选择了索引为person 的所有行,然后我们尝试为每个行apply 一个函数(lambda 是我们将使用的匿名函数) person。在这种情况下,我们使用lambda 函数提取url,然后pandas 将结果转换回pandas DataFrame(或Series)供您使用。

通常我也会关心我的迭代速度有多快。

(以下是在IPython中完成的,一个很好的工具,你可以用它在python中做很多事情。%%timeit是IPython提供的一个神奇的函数供你计算 您的代码可能需要的时间)

%timeit 
urls = result_dict[result_dict.index=='person'].apply(lambda x: x['mobileurl']['_content'])

1000 loops, best of 3: 133 us per loop (us = microsecond, 10e-6)

@SamC 在这里提供了快速解决方案,我可以告诉你。但就像我说的,你不需要DataFrame 来使用他的解决方案。它也适用于普通字典。

【讨论】:

    【解决方案2】:

    遍历键的字典列表,在这种情况下是用户名,然后使用每个键访问每个顶级字典,然后从那里深入所有其他层以找到您需要的确切数据。您示例中的 mobileurl。

    拥有这 2 个变量后,将它们添加到您的数据框中。

    # Iterate through list of users
    for user in result_dict.keys():
    
        # use each username to find the mobileurl you need within
        mobileurl = result_dict[user]["person"]["mobileurl"]["_content"]
    
        # Add the variables 'user' and 'mobileurl' to dataframe as you see fit
    

    【讨论】:

    • @SamC,感谢您的 sn-p,我已经尝试过并将其附加到数据框。正如 stucash 所说,这是最有效的。
    【解决方案3】:
    result_dict = {'11333216@N05': {'person': {'can_buy_pro': 0,
       'description': {'_content': ''},
       'has_stats': '1',
       'iconfarm': 3,
       'iconserver': '2214',
       'id': '11333216@N05',
       'ispro': 0,
       'location': {'_content': ''},
       'mbox_sha1sum': {'_content': '8eb2e248cbad94e2b4a5aae75eb653c7e061a90c'},
       'mobileurl': {'_content': 'https://m.flickr.com/photostream.gne?id=11327876'},
       'nsid': '11333216@N05',
       'path_alias': 'kishansamarasinghe',
       'photos': {'count': {'_content': 442},
        'firstdate': {'_content': '1193073180'},
        'firstdatetaken': {'_content': '2000-01-01 00:49:17'}},
       'photosurl': {'_content': 'https://www.flickr.com/photos/kishansamarasinghe/'},
       'profileurl': {'_content': 'https://www.flickr.com/people/kishansamarasinghe/'},
       'realname': {'_content': 'Kishan Samarasinghe'},
       'timezone': {'label': 'Sri Jayawardenepura',
        'offset': '+06:00',
        'timezone_id': 'Asia/Colombo'},
       'username': {'_content': 'Three Sixty Five Degrees'}},
      'stat': 'ok'},
     '117692977@N08': {'person': {'can_buy_pro': 0,
       'description': {'_content': ''},
       'has_stats': '0',
       'iconfarm': 1,
       'iconserver': '404',
       'id': '117692977@N08',
       'ispro': 0,
       'location': {'_content': 'Almere, The Nederlands'},
       'mobileurl': {'_content': 'https://m.flickr.com/photostream.gne?id=117600164'},
       'nsid': '117692977@N08',
       'path_alias': 'meijsvo',
       'photos': {'count': {'_content': 3237},
        'firstdate': {'_content': '1392469161'},
        'firstdatetaken': {'_content': '2013-06-23 14:39:30'}},
       'photosurl': {'_content': 'https://www.flickr.com/photos/meijsvo/'},
       'profileurl': {'_content': 'https://www.flickr.com/people/meijsvo/'},
       'realname': {'_content': 'Markéta Eijsvogelová'},
       'timezone': {'label': 'Amsterdam, Berlin, Bern, Rome, Stockholm, Vienna',
        'offset': '+01:00',
        'timezone_id': 'Europe/Amsterdam'},
       'username': {'_content': 'meijsvo'}},
      'stat': 'ok'},
     '21539776@N02': {'person': {'can_buy_pro': 0,
       'description': {'_content': ''},
       'has_stats': '1',
       'iconfarm': 0,
       'iconserver': '0'}
    }
    }
    

    对于您的用例,最好使用字典的 iteritems():

    for key, value in result_dict.iteritems():
        print value.get("person", {}).get("mobileurl", {}).get("_content")
    

    输出

    https://m.flickr.com/photostream.gne?id=117600164
    https://m.flickr.com/photostream.gne?id=11327876
    

    【讨论】:

    • 对不起,我看不到你的输出我在代理后面,但是你成功加载了他的 json 示例数据吗?
    • 他的 json 不正确,您必须删除最后一个 ',' 并在最后一个 json 中添加 '}}' 才能工作,谢谢
    • 试过了,它甚至不能让它工作;甚至删除最后一个不完整的用户部分。
    • 等待 @stucash 我将编辑我的答案以获取该 json
    • 对我来说,我使用 j = pd.read_json(json.dumps(invalid_json))
    猜你喜欢
    • 1970-01-01
    • 2021-04-05
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-06-17
    相关资源
    最近更新 更多