【问题标题】:preg_match_all how to find specify only <p> tags that end with <br></p>preg_match_all 如何查找仅指定以 <br></p> 结尾的 <p> 标记
【发布时间】:2013-01-20 17:21:11
【问题描述】:

我一直在玩这个,我只剩下这个......

preg_match_all('/<p>(.*?)<br>(.*?)<p>/s', $offices, $district);

哪个工作正常...但当然有一个记录会导致问题。如何指定和获取 &lt;p&gt; 标签中包含 &lt;br&gt; 的所有文本?最好在 &lt;p&gt; 中指定多个 &lt;br&gt; 并排除 Eddy Lite 标记?

字符串为地址,如:

    <h3>District Offices:</h3>
<p>
317 Dun Avenue<br>Suite 17<br>Port Samson, AK 32675<br>
(XXX) XXX-XXXX<br> 
VOIP: 40800<br> 
FAX (888) xxx-38xx<br> 

</p>

<h4>Staff Assistants:</h4>
<p>Beth Booger and Ly Sweet</p>

<h4>Secretary:</h4>
<p>Eddy Lite </p>

<p>
OK City Hall<br>110 S.E. Five Avenue<br>3rd Floor<br>Corpse, AK 33371<br>
(xxx) 694-xxxx<br> 

</p>

<h4>Staff Assistant:</h4>
<p>Con Sims </p>

    <br />

<h3>Home Office:</h3>

这就是我要返回的内容: 数组 ( [0] => 数组 ( [0] => 317敦大道 套房 17 参孙港,AK 32675 (XXX) XXX-XXXX

Staff Assistants:

[1] =>

Eddy Lite 

OK City Hall
110 S.E. Five Avenue
3rd Floor
Corpse, AK 33371
(xxx) 694-xxxx
Staff Assistant:

) )

任何帮助将不胜感激。 我努力了: preg_match_all('/

(.*?)

/s', $offices, $district); preg_match_all('/

(.*)

/s', $offices, $district); preg_match_all('/

(.?)
(.
?)

/', $offices, $district); preg_match_all('/

(.?)
(.
?)
(.*?)

/s', $offices, $district); preg_match_all('/

(.?)
(.
)
(.*)

/s', $offices, $district);

【问题讨论】:

  • 你具体想做什么?
  • 如果你使用php domdocument会不会更容易?
  • 不要使用正则表达式解析 HTML。您无法使用正则表达式可靠地解析 HTML。一旦 HTML 与您的期望发生变化,您的代码就会被破坏。有关如何使用 PHP 模块正确解析 HTML 的示例,请参阅 htmlparsing.com/php.html

标签: php regex html-parsing


【解决方案1】:

一个简单的解决方法是只允许纯文本和&lt;br&gt; 标签:

preg_match_all('#<p>([^<>]*<br\s*/?>[^<>]*)+</p>#s', $offices, $district);

通常的注意事项:这样的正则表达式仅适用于连贯且众所周知的输入。

【讨论】:

  • 谢谢!奇迹般有效。通常我会使用 DOM,但这是一个特例。
【解决方案2】:

试试这个

'~<p>(?=((?!</p>).)*<br>)((?!Eddy Lite).)*?</p>~s'

【讨论】:

    猜你喜欢
    • 2018-05-19
    • 1970-01-01
    • 2016-10-23
    • 1970-01-01
    • 2017-03-18
    • 1970-01-01
    • 2013-01-31
    • 2011-07-14
    • 2018-11-19
    相关资源
    最近更新 更多