【问题标题】:Parse phone number and string into new columns in pandas dataframe将电话号码和字符串解析为 pandas 数据框中的新列
【发布时间】:2019-12-06 18:37:37
【问题描述】:

我在单个列 address 中有一个地址列表,我将如何将电话号码和餐厅类别解析到新列中?我的数据框看起来像这样

  address
0 Arnie Morton's of Chicago 435 S. La Cienega Blvd. Los Angeles 310-246-1501 Steakhouses                                                                    
1 Art's Deli 12224 Ventura Blvd. Studio City 818-762-1221 Delis                                                                                             
2 Bel-Air Hotel 701 Stone Canyon Rd. Bel Air 310-472-1211 French Bistro 

我想去哪里

  address | phone_number | category
0 Arnie Morton's of Chicago 435 S. La Cienega Blvd. Los Angeles | 310-246-1501 | Steakhouses                                                                    
1 Art's Deli 12224 Ventura Blvd. Studio City | 818-762-1221 | Delis                                                                                             
2 Bel-Air Hotel 701 Stone Canyon Rd. Bel Air | 310-472-1211 | French Bistro 

有人有什么建议吗?

【问题讨论】:

  • 地址是否总是像您在示例中显示的那样位于末尾?
  • 试过this方法? regex 可以例如成为'[0-9]{3}-[0-9]{3}-[0-9]{4}'

标签: python pandas


【解决方案1】:

使用str.extractstr.split

  1. 我们为phone_number 提取模式numbers dash numbers dash numbers
  2. 我们在3 numbers followed by a space 的模式上进行拆分,并在category 中抓取它之后的部分。我们为此使用positive lookbehind,即正则表达式中的?<=
df['phone_number'] = df['address'].str.extract('(\d+-\d+-\d+)')
df['category'] = df['address'].str.split('(?<=\d{3})\s').str[-1]

输出

                                                                                  address  phone_number       category
0  Arnie Morton's of Chicago 435 S. La Cienega Blvd. Los Angeles 310-246-1501 Steakhouses  310-246-1501    Steakhouses
1                           Art's Deli 12224 Ventura Blvd. Studio City 818-762-1221 Delis  818-762-1221          Delis
2                   Bel-Air Hotel 701 Stone Canyon Rd. Bel Air 310-472-1211 French Bistro  310-472-1211  French Bistro

【讨论】:

    【解决方案2】:

    尝试使用带有str.extract 的正则表达式。

    例如:

    df = pd.DataFrame({'address':["Arnie Morton's of Chicago 435 S. La Cienega Blvd. Los Angeles 310-246-1501 Steakhouses", 
                                  "Art's Deli 12224 Ventura Blvd. Studio City 818-762-1221 Delis",
                                  "Bel-Air Hotel 701 Stone Canyon Rd. Bel Air 310-472-1211 French Bistro"]})
    df[["address", "phone_number", "category"]] = df["address"].str.extract(r"(?P<address>.*?)(?P<phone_number>\b\d{3}\-\d{3}\-\d{4}\b)(?P<category>.*$)")
    print(df)
    

    输出:

                                                 address  phone_number  \
    0  Arnie Morton's of Chicago 435 S. La Cienega Bl...  310-246-1501   
    1        Art's Deli 12224 Ventura Blvd. Studio City   818-762-1221   
    2        Bel-Air Hotel 701 Stone Canyon Rd. Bel Air   310-472-1211   
    
             category  
    0     Steakhouses  
    1           Delis  
    2   French Bistro  
    

    注意::假设地址内容始终为address--phone_number--category

    【讨论】:

      猜你喜欢
      • 2023-03-09
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2010-09-18
      • 1970-01-01
      • 2018-05-08
      • 2012-09-14
      相关资源
      最近更新 更多