【问题标题】:In pandas apply method, duplicate the row based on condition在 pandas apply 方法中,根据条件复制行
【发布时间】:2019-01-05 00:21:49
【问题描述】:

这是我的 df 示例:

pd.DataFrame([["1", "2"], ["1", "2"], ["3", "other_value"]],
                     columns=["a", "b"])
    a   b
0   1   2
1   1   2
2   3   other_value

我想达到这个:

pd.DataFrame([["1", "2"], ["1", "2"], ["3", "other_value"], ["3", "row_duplicated_with_edits_in_this_column"]],
                     columns=["a", "b"])
    a   b
0   1   2
1   1   2
2   3   other_value
3   3   row_duplicated_with_edits_in_this_column

规则是使用 apply 方法,做一些检查(为了保持示例简单,我不包括这些检查),但在某些条件下,对于 apply 函数中的某些行,复制该行,进行编辑到该行并在 df 中插入两行。

比如:

def f(row):
   if condition:
      row["a"] = 3
   elif condition:
      row["a"] = 4
   elif condition:
      row_duplicated = row.copy()
      row_duplicated["a"] = 5 # I need also this row to be included in the df

   return row
df.apply(f, axis=1)

我不想将重复的行存储在我班级的某个地方并在最后添加它们。我想即时进行。

我已经看到了这个pandas: apply function to DataFrame that can return multiple rows,但我不确定 groupby 是否可以在这里帮助我。

谢谢

【问题讨论】:

    标签: python pandas


    【解决方案1】:

    这是在列表理解中使用df.iterrows 的一种方法。您需要将行追加到循环中,然后进行连接。

    def func(row):
       if row['a'] == "3":
            row2 = row.copy()
            # make edits to row2
            return pd.concat([row, row2], axis=1)
       return row
    
    pd.concat([func(row) for _, row in df.iterrows()], ignore_index=True, axis=1).T
    
       a            b
    0  1            2
    1  1            2
    2  3  other_value
    3  3  other_value
    

    我发现在我的情况下,没有ignore_index=True 会更好,因为我稍后会合并 2 个 dfs。

    【讨论】:

    • 谢谢,这行得通,最好使用apply(),因为我使用df.query().apply().combine_first(),但您的解决方案仍然可以进行少量编辑,而无需将数据存储在任何地方。
    【解决方案2】:

    您的逻辑似乎主要是可矢量化的。由于输出中的行顺序似乎很重要,您可以将默认的 RangeIndex 增加 0.5,然后使用 sort_index

    def row_appends(x):
        newrows = x.loc[x['a'].isin(['3', '4', '5'])].copy()
        newrows.loc[x['a'] == '3', 'b'] = 10  # make conditional edit
        newrows.loc[x['a'] == '4', 'b'] = 20  # make conditional edit
        newrows.index = newrows.index + 0.5
        return newrows
    
    res = pd.concat([df, df.pipe(row_appends)])\
            .sort_index().reset_index(drop=True)
    
    print(res)
    
       a            b
    0  1            2
    1  1            2
    2  3  other_value
    3  3           10
    

    【讨论】:

      【解决方案3】:

      我会将它矢量化,逐个类别进行:

      df[df_condition_1]["a"] = 3
      df[df_condition_2]["a"] = 4
      
      duplicates = df[df_condition_3] # somehow we store it ?     
      duplicates["a"] = 5 
      
      #then 
      df.join(duplicates, how='outer')
      

      此解决方案是否适合您的需求?

      【讨论】:

      • 谢谢,会更快,但是我的条件很多,而且涉及多个函数,所以这种解决方案会降低代码的可读性。
      猜你喜欢
      • 2017-08-20
      • 2020-02-20
      • 2021-08-05
      • 2023-02-20
      • 2021-04-02
      • 1970-01-01
      • 2020-04-24
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多