【发布时间】:2019-05-14 07:04:12
【问题描述】:
我有以下 srt(字幕)文件:
import pysrt
srt = """
01
00:02:14,000 --> 00:02:18,000
I understand how customers do their choice. So
02
00:02:19,000 --> 00:02:24,000
what is the choice of packaging that they prefer when they have to pick up something in a shelf?
03
00:02:24,000 --> 00:02:29,000
What is the choice of the store where they will go shopping? What specific
04
00:02:29,000 --> 00:02:34,000
product they will purchase and also what is the brand that they will
05
00:02:34,000 --> 00:02:39,000
prefer. And of course many of the choices that are relevant in the context of marketing.
"""
如您所见,字幕奇怪地分开了。我希望每个字幕都以完整的句子结尾,如下所示:
srt = """
01
00:02:14,000 --> 00:02:18,000
I understand how customers do their choice.
02
00:02:19,000 --> 00:02:24,000
So what is the choice of packaging that they prefer when they have to pick up something in a shelf?
03
00:02:24,000 --> 00:02:29,000
What is the choice of the store where they will go shopping?
04
00:02:29,000 --> 00:02:34,000
What specific product they will purchase and also what is the brand that they will prefer.
05
00:02:34,000 --> 00:02:39,000
And of course many of the choices that are relevant in the context of marketing.
"""
我想知道如何使用 Python 来实现这一点。字幕文字可以使用pysrt打开:
import pysrt
srt = """
01
00:02:14,000 --> 00:02:18,000
I understand how customers do their choice. So
02
00:02:19,000 --> 00:02:24,000
what is the choice of packaging that they prefer when they have to pick up something in a shelf?
03
00:02:24,000 --> 00:02:29,000
What is the choice of the store where they will go shopping? What specific
04
00:02:29,000 --> 00:02:34,000
product they will purchase and also what is the brand that they will
05
00:02:34,000 --> 00:02:39,000
prefer. And of course many of the choices that are relevant in the context of marketing."""
with open("test.srt", "w") as text_file:
text_file.write(srt)
sub = pysrt.open("test.srt")
text = sub.text
**编辑:**
根据@Chris 的回答,我尝试了:
from operator import itemgetter
srt = """
01
00:02:14,000 --> 00:02:18,000
understand how customers do their choice. So
02
00:02:19,000 --> 00:02:24,000
what is the choice of packaging that they prefer when they have to pick up something in a shelf?
03
00:02:24,000 --> 00:02:29,000
What is the choice of the store where they will go shopping? What specific
04
00:02:29,000 --> 00:02:34,000
product they will purchase and also what is the brand that they will
05
00:02:34,000 --> 00:02:39,000
prefer. And of course many of the choices that are relevant in the context of marketing.
"""
l = [s.split('\n') for s in srt.strip().split('\n\n')]
whole = ' '.join(map(itemgetter(2), l))
for i, sen in enumerate(re.findall(r'([A-Z][^\.!?]*[\.!?])', whole)):
l[i][2] = sen
print('\n\n'.join('\n'.join(s) for s in l))
但我得到的结果与输入完全相同...
01 00:02:14,000 --> 00:02:18,000 understand how customers do their choice. So 02 00:02:19,000 --> 00:02:24,000 what is the choice of packaging that they prefer when they have to pick up something in a shelf? 03 00:02:24,000 --> 00:02:29,000 What is the choice of the store where they will go shopping? What specific 04 00:02:29,000 --> 00:02:34,000 product they will purchase and also what is the brand that they will 05 00:02:34,000 --> 00:02:39,000 prefer. And of course many of the choices that are relevant in the context of marketing.
我做错了什么?
【问题讨论】:
-
@PaulRooney 好点!不太确定,该怎么做,虽然我不知道说一定数量的单词需要多长时间。但是,可以通过将给定时间段(即
00:02:34,000 --> 00:02:39,000)除以该时间段内的字母数量来找到平均值。
标签: python