【问题标题】:Filter only in one element/array of Twitter JSON file仅过滤 Twitter JSON 文件的一个元素/数组
【发布时间】:2013-09-07 10:53:28
【问题描述】:

我从 Streaming API 爬取了 twitter JSON 文件,得到了一个包含数千行 JSON 数据的文件。但是,该数据包含很多元素,例如“创建日期”“来源”“推文”等。 我实际上想在推文文本中过滤“iphone”这个词。但是,如果我使用 GREP UNIX 进行过滤,它不仅会过滤掉“推文文本”字段,还会过滤掉“源”字段。因此,这意味着一条不包含单词“iphone”但从 Twitter for Iphone 发布的推文(如“来源”字段中所述)也将被过滤。

有没有办法只在一个特定的字段中过滤这个 JSON(在我的例子中是“tweet text”字段)。

这是一个 JSON 行的示例:

{"created_at":"Tue Aug 20 03:48:27 +0000 2013","id":369667218608369666,"id_str":"369667218608369666","text":"@Mattyb_chyeah_ yeah I'm only watching him! :)","source":"\u003ca href=\"http:\/\/twitter.com\/download\/iphone\" rel=\"nofollow\"\u003eTwitter for iPhone\u003c\/a\u003e","truncated":false,"in_reply_to_status_id":369666992334073856,"in_reply_to_status_id_str":"369666992334073856","in_reply_to_user_id":1557571363,"in_reply_to_user_id_str":"1557571363","in_reply_to_screen_name":"Mattyb_chyeah_","user":{"id":1325959333,"id_str":"1325959333","name":"MattyBRapsTexas","screen_name":"MattyBRapsTexas","location":"Atlanta,Georgia","url":"http:\/\/www.instagram.com\/mattybrapstexas","description":"3 RT 6 Mentions He followed me on 4\/15\/13 6\/17\/13 Maddi Jane followed me on 6\/18\/13 @8:25pm! Cimorelli also follows Pizza Hut mentioned me 2 times on 7\/26\/13","protected":false,"followers_count":1095,"friends_count":426,"listed_count":8,"created_at":"Thu Apr 04 02:34:56 +0000 2013","favourites_count":226,"utc_offset":-14400,"time_zone":"Eastern Time (US & Canada)","geo_enabled":false,"verified":false,"statuses_count":3447,"lang":"en","contributors_enabled":false,"is_translator":false,"profile_background_color":"C0DEED","profile_background_image_url":"http:\/\/a0.twimg.com\/images\/themes\/theme1\/bg.png","profile_background_image_url_https":"https:\/\/si0.twimg.com\/images\/themes\/theme1\/bg.png","profile_background_tile":false,"profile_image_url":"http:\/\/a0.twimg.com\/profile_images\/378800000313651225\/afee0cc2286882eeb15f21ed7fae334a_normal.jpeg","profile_image_url_https":"https:\/\/si0.twimg.com\/profile_images\/378800000313651225\/afee0cc2286882eeb15f21ed7fae334a_normal.jpeg","profile_banner_url":"https:\/\/pbs.twimg.com\/profile_banners\/1325959333\/1376759786","profile_link_color":"0084B4","profile_sidebar_border_color":"C0DEED","profile_sidebar_fill_color":"DDEEF6","profile_text_color":"333333","profile_use_background_image":true,"default_profile":true,"default_profile_image":false,"following":null,"follow_request_sent":null,"notifications":null},"geo":null,"coordinates":null,"place":null,"contributors":null,"retweet_count":0,"favorite_count":0,"entities":{"hashtags":[],"symbols":[],"urls":[],"user_mentions":[{"screen_name":"Mattyb_chyeah_","name":"MattyB (\u2661_\u2661\u2740)","id":1557571363,"id_str":"1557571363","indices":[0,15]}]},"favorited":false,"retweeted":false,"filter_level":"medium","lang":"en"

【问题讨论】:

  • 下面的答案有帮助吗?

标签: json twitter filter grep


【解决方案1】:

您在 grep 正则表达式中使用什么?如果您只是将“iphone”用于正则表达式,那么是的,您将获得多次点击。您可以扩展您的正则表达式以仅在源之前的文本部分中匹配 iphone:

grep '"text":".*iphone.*","source":' myfile.txt

将在"text" 之后但在"source" 之前搜索模式iphone。它将忽略该行其余部分的iphone

【讨论】:

    猜你喜欢
    • 2018-03-07
    • 2013-02-26
    • 1970-01-01
    • 2021-12-21
    • 1970-01-01
    • 1970-01-01
    • 2018-12-24
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多