【发布时间】:2021-01-07 02:52:16
【问题描述】:
我相信我有一个简单的问题。这是我想划分的文学作品的数据样本:
WholeBook = "Random info - at beginning-man. "+ ...
"Random info still continues. "+ ...
"CHAPTER 1 " + ...
"1 This is sentence one of verse one, "+ ...
"This still sentence one of verse one. "+ ...
"2 This is sentence one of verse two. "+ ...
"This is sentence two of verse two. "+ ...
"3 This is sentence one of verse three; "+ ...
"this still sentence one of verse three. "+ ...
"CHAPTER 2 " + ...
"Random info in middle two. "+ ...
"Random info still continues again. "+ ...
"1 This is sentence four? "+ ...
"2 This is sentence five, "+ ...
"3 this still sentence five but verse three!"+ ...
"Random info at end's end.";
我想将以下数据划分为这样的表格(这就是解决方案的外观):
但是,我目前的解决方案是这样的:
因此第 1 行是不正确的,但第 2 行是正确的。否则说,如果“第 # 章”之后确实有信息,我的解决方案有效,但如果没有信息,则无效。这是产生此解决方案的代码:
[tokens, RandomInfoMiddle] = regexp(WholeBook, '(CHAPTER \d)\s*(.*?)1', 'tokens', 'match');
RandomInfoMiddle = RandomInfoMiddle';
RandomInfoMiddle = regexprep(RandomInfoMiddle,'CHAPTER \d+ (.+) \d$','$1'); %Delete "Chapter+Nr" + ...1
% To explain the regular expression (CHAPTER \d)\.\s*(.*?)1:
% (CHAPTER \d) matches CHAPTER with any number, and the () brackets surrounding it will capture the match in the tokens variable.
% \. matches the period
% \s* matches any possible whitespace
% (.*?)1 will capture any text till the next 1 in the text. Note the question mark to make it match lazy, otherwise it will match all the text till the last 1 in str.
请帮助我找到第一张图片/表格中描述的解决方案。 (我怀疑使用 if 语句加上正确的正则表达式。)
感谢所有帮助。
【问题讨论】:
-
可能是这样的
^(CHAPTER \d+)\r?\n((?:(?!\d+\b).*(?:\r?\n|$))+)regex101.com/r/qc4LHr/1
标签: regex matlab if-statement datatable