【发布时间】:2019-08-02 07:48:43
【问题描述】:
我已经使用漂亮的汤包解析了一个 url 以获取它的文本。我想删除条款和条件部分中的所有文本,即“关键条款:...... T&Cs apply”段落中的所有单词。
以下是我尝试过的:
import re
#"text" is part of the text contained in the url
text="Welcome to Company Key.
Key Terms; Single bets only. Any returns from the free bet will be paid
back into your account minus the free bet stake. Free bets can only be
placed at maximum odds of 5.00 (4/1). Bonus will expire midnight, Tuesday
26th February 2019. Bonus T&Cs and General T&Cs apply.
"
rex=re.compile('Key\ (.*?)T&Cs.')"""to remove words between "Key" and
"T&Cs" """
terms_and_cons=rex.findall(text)
text=re.sub("|".join(terms_and_cons)," ",text)
#I also tried: text=re.sub(terms_and_cons[0]," ",text)
print(text)
上面只是保持字符串'text'不变,即使列表“terms_and_cons”是非空的。如何成功删除“Key”和“T&Cs”之间的单词?请帮我。很长一段时间以来,我一直被困在这段据说很简单的代码上,这真的令人沮丧。谢谢。
【问题讨论】:
-
你可以在你的正则表达式的开头添加一个 ^ 符号来否定前瞻,这样 term n cons 变量只会得到没有被正则表达式过滤的东西吗?