【问题标题】:How to filter stromcrawler data from elasticsearch如何从弹性搜索中过滤 stromcrawler 数据
【发布时间】:2020-06-16 06:46:57
【问题描述】:

我正在使用 apache-storm 1.2.3 和 elasticsearch 7.5.0。我已经成功地从 3k 新闻网站中提取了数据,并在 Grafana 和 kibana 上进行了可视化。我在内容中收到很多垃圾(如广告)。我附上了 CONTENT 的 SS。content 谁能建议我如何过滤它们。我正在考虑将 ES 中的 html 内容提供给一些 python 包。如果不是,我是否在正确的轨道上,请建议我好的解决方案。 提前致谢。

这是 crawler-conf.yaml 文件

config:
  topology.workers: 1
  topology.message.timeout.secs: 300
  topology.max.spout.pending: 100
  topology.debug: false

  fetcher.threads.number: 50

  # override the JVM parameters for the workers
  topology.worker.childopts: "-Xmx2g -Djava.net.preferIPv4Stack=true"

  # mandatory when using Flux
  topology.kryo.register:
    - com.digitalpebble.stormcrawler.Metadata

  # metadata to transfer to the outlinks
  # used by Fetcher for redirections, sitemapparser, etc...
  # these are also persisted for the parent document (see below)
  # metadata.transfer:
  # - customMetadataName

  # lists the metadata to persist to storage
  # these are not transfered to the outlinks
  metadata.persist:
   - _redirTo
    - error.source
   - isSitemap
   - isFeed

  http.agent.name: "Nitesh Singh"
  http.agent.version: "1.0"
  http.agent.description: "built with StormCrawler Elasticsearch Archetype 1.16"
  http.agent.url: "http://someorganization.com/"
  http.agent.email: "nite0sh@gmail.com"

  # The maximum number of bytes for returned HTTP response bodies.
  # The fetched page will be trimmed to 65KB in this case
  # Set -1 to disable the limit.
  http.content.limit: 65536

  # FetcherBolt queue dump => comment out to activate
  # if a file exists on the worker machine with the corresponding port number
  # the FetcherBolt will log the content of its internal queues to the logs
  # fetcherbolt.queue.debug.filepath: "/tmp/fetcher-dump-{port}"

  parsefilters.config.file: "parsefilters.json"
  urlfilters.config.file: "urlfilters.json"

  # revisit a page daily (value in minutes)
  # set it to -1 to never refetch a page
  fetchInterval.default: 1440

  # revisit a page with a fetch error after 2 hours (value in minutes)
  # set it to -1 to never refetch a page
  fetchInterval.fetch.error: 120
fetchInterval.error: -1

  # text extraction for JSoupParserBolt
  textextractor.include.pattern:
   - DIV[id="maincontent"]
   - DIV[itemprop="articleBody"]
   - ARTICLE

  textextractor.exclude.tags:
   - STYLE
   - SCRIPT

  # custom fetch interval to be used when a document has the key/value in its metadata
  # and has been fetched successfully (value in minutes)
  # fetchInterval.FETCH_ERROR.isFeed=true: 30
  # fetchInterval.isFeed=true: 10

  # configuration for the classes extending AbstractIndexerBolt
  # indexer.md.filter: "someKey=aValue"
  indexer.url.fieldname: "url"
  indexer.text.fieldname: "content"
  indexer.canonical.name: "canonical"
  indexer.md.mapping:
  - parse.title=title
  - parse.keywords=keywords
  - parse.description=description
  - domain=domain

  # Metrics consumers:
  topology.metrics.consumer.register:
     - class: "org.apache.storm.metric.LoggingMetricsConsumer"
 parallelism.hint: 1

【问题讨论】:

    标签: elasticsearch web-crawler apache-storm stormcrawler


    【解决方案1】:

    您是否配置了文本提取器?例如

      # text extraction for JSoupParserBolt
      textextractor.include.pattern:
       - DIV[id="maincontent"]
       - DIV[itemprop="articleBody"]
       - ARTICLE
    
      textextractor.exclude.tags:
       - STYLE
       - SCRIPT
    

    如果找到和/或删除排除中指定的元素,这会将文本限制为特定元素。

    大多数新闻网站都会使用某种形式的标签来标记主要内容。

    您作为元素提供的示例,您可以为其添加模式。

    您可以在 ParseFilter 中嵌入各种样板删除库,但它们的准确性差异很大。

    【讨论】:

    • 感谢 Julien 的回复。文本提取器已在 crawler-conf.yaml 文件中配置。
    猜你喜欢
    • 1970-01-01
    • 2012-11-07
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-04-29
    • 2019-01-08
    • 2015-11-20
    • 2016-06-22
    相关资源
    最近更新 更多