【问题标题】:Nutch Crawler doesn't retrieve news article contentNutch Crawler 不检索新闻文章内容
【发布时间】:2016-08-04 07:28:39
【问题描述】:

我试图从链接中抓取新闻文章:-

Article 1

Article 2

但我没有将页面中的文本放到 index(elasticsearch) 中的内容字段中。

爬取的结果是:-

{
  "took": 2,
  "timed_out": false,
  "_shards": {
    "total": 5,
    "successful": 5,
    "failed": 0
  },
  "hits": {
    "total": 2,
    "max_score": 0.09492774,
    "hits": [
      {
        "_index": "news",
        "_type": "doc",
        "_id": "http://www.bloomberg.com/press-releases/2016-07-08/network-1-announces-settlement-of-patent-litigation-with-apple-inc",
        "_score": 0.09492774,
        "_source": {
          "tstamp": "2016-08-04T07:21:59.614Z",
          "segment": "20160804125156",
          "digest": "d583a81c0c4c7510f5c842ea3b557992",
          "host": "www.bloomberg.com",
          "boost": "1.0",
          "id": "http://www.bloomberg.com/press-releases/2016-07-08/network-1-announces-settlement-of-patent-litigation-with-apple-inc",
          "url": "http://www.bloomberg.com/press-releases/2016-07-08/network-1-announces-settlement-of-patent-litigation-with-apple-inc",
          "content": ""
        }
      },
      {
        "_index": "news",
        "_type": "doc",
        "_id": "http://www.bloomberg.com/press-releases/2016-07-05/apple-donate-life-america-bring-national-organ-donor-registration-to-iphone",
        "_score": 0.009845509,
        "_source": {
          "tstamp": "2016-08-04T07:22:05.708Z",
          "segment": "20160804125156",
          "digest": "2a94a32ffffd0e03647928755e055e30",
          "host": "www.bloomberg.com",
          "boost": "1.0",
          "id": "http://www.bloomberg.com/press-releases/2016-07-05/apple-donate-life-america-bring-national-organ-donor-registration-to-iphone",
          "url": "http://www.bloomberg.com/press-releases/2016-07-05/apple-donate-life-america-bring-national-organ-donor-registration-to-iphone",
          "content": ""
        }
      }
    ]
  }
}

我们可以注意到内容字段是空的。我在 nutch-site.txt 中尝试了不同的选项。但结果还是一样。请帮我解决这个问题。

【问题讨论】:

    标签: web-crawler nutch


    【解决方案1】:

    不知道为什么 nutch 无法提取文章内容。但我找到了使用 Jsoup 的解决方法。我开发了一个自定义解析过滤器插件,它解析整个文档并在解析器过滤器返回的 ParseResult 中设置解析文本。并通过替换 parse-plugins.xml 中的 parse-html 插件来使用我的自定义解析过滤器

    会是这样的:-

       document = Jsoup.parse(new String(content.getContent(),"UTF-8"),content.getUrl());
       parse = parseResult.get(content.getUrl());
       status = parse.getData().getStatus();
       title = document.title();
       parseData = new ParseData(status, title,parse.getData().getOutlinks(), parse.getData().getContentMeta(), parse.getData().getParseMeta());
       parseResult.put(content.getUrl(), new ParseText(document.body().text()), parseData);
    

    【讨论】:

      【解决方案2】:

      只是一个断章取义的答案,但请尝试使用 Apache ManifoldCF 。它为弹性搜索提供了内置连接器,并提供了更好的记录历史来找出数据未编入索引的原因。 ManifoldCF 中的连接器部分允许您指定应在哪个字段中索引您的内容。这是一个很好的开源替代品,可以尝试一下。

      【讨论】:

      • 谢谢 :) 。我去看看。
      • 我想选择特定 div 或任何其他标签内的链接,并获取该链接的内容并将它们编入索引。我们是否可以用流形做这样的事情
      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2014-10-03
      • 1970-01-01
      相关资源
      最近更新 更多