【发布时间】:2021-12-13 19:58:10
【问题描述】:
我正在尝试使用 PHP 的 DOMDocument 和 XPath 将某些短语的所有实例包装在 <span> 中。我的逻辑基于this answer from another post,但这仅允许我在需要选择所有个匹配项时选择节点内的第一个匹配项。
一旦我为第一个匹配项修改了 DOM,我的后续循环会导致错误,在 $after 所在的行中声明 Fatal error: Uncaught Error: Call to a member function splitText() on bool。我很确定这是由修改标记引起的,但我一直无法弄清楚原因。
我在这里做错了什么?
/**
* Automatically wrap various forms of CCJM in a class for branding purposes
*
* @link https://stackoverflow.com/a/6009594/654480
*
* @param string $content
* @return string
*/
function ccjm_branding_filter(string $content): string {
if (! (is_admin() && ! wp_doing_ajax()) && $content) {
$DOM = new DOMDocument();
/**
* Use internal errors to get around HTML5 warnings
*/
libxml_use_internal_errors(true);
/**
* Load in the content, with proper encoding and an `<html>` wrapper required for parsing
*/
$DOM->loadHTML("<?xml encoding='utf-8' ?><html>{$content}</html>", LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD);
/**
* Clear errors to get around HTML5 warnings
*/
libxml_clear_errors();
/**
* Initialize XPath
*/
$XPath = new DOMXPath($DOM);
/**
* Retrieve all text nodes, except those within scripts
*/
$text = $XPath->query("//text()[not(parent::script)]");
foreach ($text as $node) {
/**
* Find all matches, including offset
*/
preg_match_all("/(C\.? ?C\.?(?:JM| Johnson (?:&|&|&|and) Malhotra)(?: Engineers, LTD\.?|, P\.?C\.?)?)/i", $node->textContent, $matches, PREG_OFFSET_CAPTURE);
/**
* Wrap each match in appropriate span
*/
foreach ($matches as $group) {
foreach ($group as $key => $match) {
/**
* Determine the offset and the length of the match
*/
$offset = $match[1];
$length = strlen($match[0]);
/**
* Isolate the match and what comes after it
*/
$word = $node->splitText($offset);
$after = $word->splitText($length);
/**
* Create the wrapping span
*/
$span = $DOM->createElement("span");
$span->setAttribute("class", "__brand");
/**
* Replace the word with the span, and then re-insert the word within it
*/
$word->parentNode->replaceChild($span, $word);
$span->appendChild($word);
break; // it always errors after the first loop
}
}
}
/**
* Save changes, remove unneeded tags
*/
$content = implode(array_map([$DOM->documentElement->ownerDocument, "saveHTML"], iterator_to_array($DOM->documentElement->childNodes)));
}
return $content;
}
add_filter("ccjm_final_output", "ccjm_branding_filter");
示例内容(“C.C. Johnson & Malhotra, P.C.”和“CCJM”的所有实例都匹配,但只有第一个可以成功修改):
C.C. Johnson & Malhotra, P.C. (CCJM) was an integral member of a large Design Team for a 16.5-mile-long Public-Private Partnership (P3) Purple Line Project. The east-west light rail system extends from New Carrollton in PG County, MD to Bethesda in MO County, MD with 21 stations and one short tunnel. CCJM was Engineer of Record (EOR) for the design of eight (8) Bridges and design reviews for 35 transit/highway bridges and over 100 retaining walls of different lengths/types adjacent to bridges and in areas of cut/fill. CCJM designed utility structures for 42,000 LF of relocated water mains and 19,000 LF of relocated sewer mains meeting Washington Suburban Sanitary Commission (WSSC), Md Dept of Transportation (MDOT) MTA, and Local Standards.
编辑 1:做一些测试,当我输出 $node->textContent 时,我看到它在第一个循环之后发生了变化......所以我认为发生的事情是在我做 $node->splitText($offset) 之后,它实际上是在更新整个节点,所以后续的偏移量不起作用。
【问题讨论】:
-
这是您拥有的实际代码吗?我算大括号中的不匹配
-
哎呀,我做了很多修改并试图清理一堆我注释掉的东西。我刚刚用我当前使用的确切代码更新了它。
标签: php regex wordpress xpath domdocument