【问题标题】:Modifying data in DataFrame based on date and string length根据日期和字符串长度修改DataFrame中的数据
【发布时间】:2018-12-18 08:21:16
【问题描述】:

我需要清理 Pandas DataFrame 中的一些数据并为此苦苦挣扎。

样本数据:

Date       | ID     | Name             | Address
-----------------------------------------------------------------------------------------------
1-4-1987   | 124578 | T.Hilpert        | 518 Hessel Plaza Lake Lonzo, AZ 11863
23-6-1990  | 947383 | Birdie Reynolds  | 964 Weissnat Green Suite 568 Rennerbury
12-5-1960  | 746732 | Earline Schulist | 57367 Alfredo Vista East Bertaburgh
9-9-2010   | 947383 | Birdie Reynolds  | 964 Weissnat Green Suite 568 Rennerbury, WV 16241-5205
27-12-2017 | 124578 | Theresia Hilpert | 518 Hessel Plaza Lake Lonzo

我想做的就是这个。按 ID 分组,从最近的日期获取名称并获取最长的地址字符串。将这些用于所有出现的 ID(在两个新列中:Name_newAddress_New)。请在下面找到所需的样本:

Date       | ID     | Name             | Address                                                | Name_New         | Address_New
---------------------------------------------------------------------------------------------------------------------------------------------------------------------------
27-12-2017 | 124578 | Theresia Hilpert | 518 Hessel Plaza Lake Lonzo                            | Theresia Hilpert | 518 Hessel Plaza Lake Lonzo, AZ 11863
1-4-1987   | 124578 | T. Hilpert       | 518 Hessel Plaza Lake Lonzo, AZ 11863                  | Theresia Hilpert | 518 Hessel Plaza Lake Lonzo, AZ 11863
23-6-1990  | 947383 | Birdie Reynolds  | 964 Weissnat Green Suite 568 Rennerbury                | Birdie Reynolds  | 964 Weissnat Green Suite 568 Rennerbury, WV 16241-5205
9-9-2010   | 947383 | Birdie Reynolds  | 964 Weissnat Green Suite 568 Rennerbury, WV 16241-5205 | Birdie Reynolds  | 964 Weissnat Green Suite 568 Rennerbury, WV 16241-5205
12-5-1960  | 746732 | Earline Schulist | 57367 Alfredo Vista East Bertaburgh                    | Earline Schulist | 57367 Alfredo Vista East Bertaburgh

我已经尝试过了,但无法将它组合起来以获得所需的结果。

def f1(s):
    return max(s, key=len)

df_new = df['New_Address'] = df.groupby('ID').agg({'Address': f1})


df_new = df[df.groupby('ID').Date.transform('max') == df['Date']]

我们特别感谢您的帮助。

【问题讨论】:

    标签: python python-3.x pandas dataframe


    【解决方案1】:

    使用transform返回Series,大小与原始DataFrame相同,然后按Name列创建索引并通过Date通过idxmax获取最大值:

    df['Date'] = pd.to_datetime(df['Date'], format='%d-%m-%Y')
    df['Address_New'] = df.groupby('ID')['Address'].transform(lambda s: max(s, key=len))
    df['Name_New'] = df.set_index('Name').groupby('ID')['Date'].transform('idxmax').values
    print (df)
            Date      ID              Name  \
    0 1987-04-01  124578         T.Hilpert   
    1 1990-06-23  947383   Birdie Reynolds   
    2 1960-05-12  746732  Earline Schulist   
    3 2010-09-09  947383   Birdie Reynolds   
    4 2017-12-27  124578  Theresia Hilpert   
    
                                                 Address  \
    0              518 Hessel Plaza Lake Lonzo, AZ 11863   
    1            964 Weissnat Green Suite 568 Rennerbury   
    2                57367 Alfredo Vista East Bertaburgh   
    3  964 Weissnat Green Suite 568 Rennerbury, WV 16...   
    4                        518 Hessel Plaza Lake Lonzo   
    
                                             Address_New          Name_New  
    0              518 Hessel Plaza Lake Lonzo, AZ 11863  Theresia Hilpert  
    1  964 Weissnat Green Suite 568 Rennerbury, WV 16...   Birdie Reynolds  
    2                57367 Alfredo Vista East Bertaburgh  Earline Schulist  
    3  964 Weissnat Green Suite 568 Rennerbury, WV 16...   Birdie Reynolds  
    4              518 Hessel Plaza Lake Lonzo, AZ 11863  Theresia Hilpert  
    

    【讨论】:

    • 我设法将地址更改为新的New_Address 列。我想为 Name 做同样的事情(不是创建一个新的 DataFrame。我只是用它来测试我的代码。我试过df['New_Name'] = df.groupby('ID')['Name'].Date.transform('max') == df['Date']。但是我有一个错误:AttributeError: 'SeriesGroupBy' object has no attribute 'Date'
    • @JohnDoe - 我认为你很亲密,需要df['New_Name'] = df.groupby('ID')['Name'].transform(lambda s: max(s, key=len))
    • 我想获取最近日期的名字。对于地址,它只是最长的字符串。
    • @JohnDoe - 不确定是否理解 - df.groupby('ID').Date.transform('max') 返回最新日期,df.groupby('ID').Name.transform('max') 也在工作,它返回 returns the one that is at the bottom of the alphabetic list(见 this)所以需要最大长度吗?你能解释更多吗?
    • 非常感谢。效果很好!!也感谢您提供解释的链接,
    猜你喜欢
    • 2012-01-19
    • 2020-06-15
    • 1970-01-01
    • 2013-11-25
    • 1970-01-01
    • 2016-07-08
    • 2021-06-27
    • 1970-01-01
    相关资源
    最近更新 更多