【发布时间】:2013-12-05 15:52:43
【问题描述】:
我在使用 nokogiri 和 xpath 时遇到了一个奇怪的问题。我想解析一个 HTML 文档并通过 href 值和它们包含的锚文本获取所有链接。
到目前为止,这是我的 xpath:
xpath = "//a[contains(text(), #{link['anchor_text']}) and @href='#{link['target_url']}']"
a = doc.search(xpath)
只要 link['anchor_text'] 是一个没有数字的字符串,它就可以正常工作。
如果我试图获取带有锚文本“11example”的链接,则会引发以下错误:
Invalid expression: //a[contains(text(), 11example) and @href='http://www.example.com/']
也许这只是一个愚蠢的错误,但我不明白为什么会发生此错误。如果我在 xpath 中的 #{link['anchor_text']} 周围加上引号,则没有任何效果。
编辑:这是示例 HTML:
<!DOCTYPE html>
<head>
<title>Example.com</title>
</head>
<body>
<p>
<strong>Here is some text</strong><br />
<a href="example.com" target="_blank">11example</a>Some text here and there
</p>
<p>
<strong>Another text</strong><br />
<a href="example.com/test" target="_blank">example.com</a>Some text here and there
</p>
</body>
Edit2:如果我在 irb 控制台中手动运行这些查询,一切都会按预期工作,但前提是我将文本放在引号中。
提前致谢!
亲切的问候, 疯子
【问题讨论】:
-
把示例 HTML 也给我们..
-
抱歉,我添加了 HTML。
标签: ruby-on-rails ruby xpath nokogiri