【发布时间】:2018-10-23 18:49:00
【问题描述】:
我正在使用 Storm crawler 1.10 和 Elastic Search 6.3.x。例如,我有一个主网站https://www.abce.org,它有像https://abce.org/def 和https://abce.org/ghi 这样的子页面。我想专门抓取https://www.abce.org/ghi下的页面。
我的种子网址是https://www.abce.org/ghi/。
目前我每次都在不同的正则表达式过滤器下应用。
+^https:\/\/www.abce.org\/ghi*+^(?:https?:\/\/)www.abce.org\/ghi(.+)*$+^(?:https?:\/\/)?(?:www\.)?abce\.[a-zA-Z0-9.\S]+$
我测试了我的正则表达式 regexr 它显示有效。但是当我检查 statusindex 时,它的显示只发现了种子 url,没有别的。
【问题讨论】:
标签: regex web-crawler stormcrawler