【发布时间】:2021-12-01 15:06:48
【问题描述】:
我有一串文本,其中包含从“1”开始的编号段落。到“221.”,但是,有些段落不符合顺序,我想删除它们。以下是数据的外观:
text = """1. Shares of Paras Defence and Space Technologies gained 2.85 times.
2. The company, engaged in manufacturing and testing of defence and space engineering products.
"3. Its stock ended at Rs 499 versus issue price of Rs 175 per share.
42. On July 23, Zomato NSE 0.00 % Ltd. listed on the Indian stock exchanges.
43. That was exactly a week after the food-delivery and restaurant discovery platform's initial public offering went live.
4. Paras Defence’s IPO, which closed on September 23, had generated bids worth Rs 38,021 crore.
5. It surpassed the previous record of Salasar Technologies’ IPO.
14. NBFCs are betting big time on the IPO.
6. Paras Defence is one of the few players having an edge in defence deals."""
从上面的文本中,我想删除不按顺序的段落的内容,即。 “42.”、“43.”和“14”。
所需的输出:
relevant_text = '1. Shares of Paras Defence and Space Technologies gained 2.85 times.
2. The company, engaged in manufacturing and testing of defence and space engineering products.
3. Its stock ended at Rs 499 versus issue price of Rs 175 per share.
4. Paras Defence’s IPO, which closed on September 23, had generated bids worth Rs 38,021 crore.
5. It surpassed the previous record of Salasar Technologies’ IPO.
6. Paras Defence is one of the few players having an edge in defence deals.'
我尝试匹配模式,但不知道如何继续。另外,我不确定正则表达式模式是否正确,因为它匹配 '1.'、'2.'等等,但不是“3.”。这是我想出的:
text_sequence = []
pattern = re.compile('(\s|["])[0-9]{1,3}\.\s')
matches = pattern.finditer(text)
for match in matches:
for r in range(1, 999):
if str(r) in match.group():
text_sequence.append(match.span())
text_sequence.append(match.group())
print(text_sequence)
有没有办法得到想要的输出?
P.S:我从这段代码中得到的匹配有重复的结果。
【问题讨论】: