【问题标题】:Split() in python how to use if there is condition has to skip for some value如果有条件必须跳过某个值,python中的Split()如何使用
【发布时间】:2019-11-18 20:37:44
【问题描述】:

我是python新手,我想将包含电影名称和发行年份的一列中的数据分成多列,所以我找到了拆分功能。

数据按标题(年份)组织。

我在 python 中尝试的是:

movies['title'].str.split('(', 1, expand = True)

以下情况发生异常:

City of Lost Children, The (Cité des enfants perdus, La) (1999)

失落儿童之城,The。 Cité des enfants perdus, La) (1999)

我所期望的只是 1999 年)进入第二列。

我需要你的帮助!

【问题讨论】:

标签: python pandas split


【解决方案1】:

我建议pd.Series.str.rsplit:

给定一个系列s

print(s)
0    City of Lost Children, The (Cité des enfants perdus, La) (1999)
1    'City of Lost Children, The. Cité des enfants perdus, La) (1999)'
dtype: object

使用s.str.rsplit('(', 1, expand=True)

                                                   0      1
0  City of Lost Children, The (Cité des enfants p...  1999)
1  City of Lost Children, The. Cité des enfants p...  1999)

【讨论】:

    【解决方案2】:

    我投票赞成在这里使用re.findall(.*?) \((\d{4})\) 模式:

    input = """City of Lost Children, The (Cité des enfants perdus, La) (1999)
               City of Lost Children, The. Cité des enfants perdus, La) (1999)"""
    
    matches = re.findall(r'\s*(.*?) \((\d{4})\)', input)
    print(matches)
    

    打印出来:

    [('City of Lost Children, The (Cité des enfants perdus, La)', '1999'),
     ('City of Lost Children, The. Cité des enfants perdus, La)', '1999')]
    

    【讨论】:

    • 整洁!也许对于熊猫df.title.str.findall(r'(.*?) \((\d{4})\)')
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-02-15
    • 1970-01-01
    • 1970-01-01
    • 2022-09-24
    • 2020-12-08
    • 2022-08-18
    相关资源
    最近更新 更多