【问题标题】:How to convert pandas DataFrame result to user defined json format如何将 pandas DataFrame 结果转换为用户定义的 json 格式
【发布时间】:2023-04-06 22:42:01
【问题描述】:
data_df = pandas.read_csv('details.csv')
data_df = data_df.replace('Null', np.nan)
df = data_df.groupby(['country', 'branch']).count()
df = df.drop('sales', axis=1)  
df = df.reset_index()
print df

我想将 Data Frame(df) 的结果转换为我在下面提到的用户定义的 json 格式。打印结果后(df)我会得到表格中的结果

country     branch      no_of_employee     total_salary    count_DOB   count_email
  x            a            30                 2500000        20            25
  x            b            20                 350000         15            20
  y            c            30                 4500000        30            30
  z            d            40                 5500000        40            40
  z            e            10                 1000000        10            10
  z            f            15                 1500000        15            15

我想把它转换成 Json。我想要的格式是

x
   {
      a
        {
              no.of employees:30
              total salary:2500000
              count_email:25
         }
       b
         {
              no.of employees:20
              total salary:350000
              count_email:25

           }
     }

   y
     {

        c
         {
              no.of employees:30
              total salary:4500000
              count_email:30

           }
      }
   z
     {
       d
         {
              no.of employees:40
              total salary:550000
              count_email:40
         }
       e
         {
              no.of employees:10
              total salary:100000
              count_email:15

         }
        f
         {
              no.of employees:15
              total salary:1500000
              count_email:15

         }
    }

请注意,我不希望 Json 中的数据帧结果中的所有字段(例如:count_DOB)

【问题讨论】:

    标签: python json pandas dataframe


    【解决方案1】:

    您可以将groupbyapply to_dict 和最后一个to_json 一起使用:

      country branch  no_of_employee  total_salary  count_DOB  count_email
    0       x      a              30       2500000         20           25
    1       x      b              20        350000         15           20
    2       y      c              30       4500000         30           30
    3       z      d              40       5500000         40           40
    4       z      e              10       1000000         10           10
    5       z      f              15       1500000         15           15
    
    g = df.groupby('country')[["branch", "no_of_employee","total_salary", "count_email"]]
                                  .apply(lambda x: x.set_index('branch').to_dict(orient='index'))
    print g.to_json()
    
    {
        "x": {
            "a": {
                "total_salary": 2500000,
                "no_of_employee": 30,
                "count_email": 25
            },
            "b": {
                "total_salary": 350000,
                "no_of_employee": 20,
                "count_email": 20
            }
        },
        "y": {
            "c": {
                "total_salary": 4500000,
                "no_of_employee": 30,
                "count_email": 30
            }
        },
        "z": {
            "e": {
                "total_salary": 1000000,
                "no_of_employee": 10,
                "count_email": 10
            },
            "d": {
                "total_salary": 5500000,
                "no_of_employee": 40,
                "count_email": 40
            },
            "f": {
                "total_salary": 1500000,
                "no_of_employee": 15,
                "count_email": 15
            }
        }
    }
    

    我尝试print g.to_dict(),但 JSON 无效(检查它here)。

    【讨论】:

    • 运行此代码时出现 TypeError: to_dict() got an unexpected keyword argument 'orient'
    • 它不适用于您的真实数据样本?
    • 请用这个df - df = pd.DataFrame({'count_email': {0: 25, 1: 20, 2: 30, 3: 40, 4: 10, 5: 15}, 'country': {0: 'x', 1: 'x', 2: 'y', 3: 'z', 4: 'z', 5: 'z'}, 'count_DOB': {0: 20, 1: 15, 2: 30, 3: 40, 4: 10, 5: 15}, 'branch': {0: 'a', 1: 'b', 2: 'c', 3: 'd', 4: 'e', 5: 'f'}, 'total_salary': {0: 2500000, 1: 350000, 2: 4500000, 3: 5500000, 4: 1000000, 5: 1500000}, 'no_of_employee': {0: 30, 1: 20, 2: 30, 3: 40, 4: 10, 5: 15}}) 检查它。还是同样的错误?如果是,pandas 使用什么版本 - print pd.__version__
    • 现在一切正常。我所做的是卸载熊猫并安装另一个版本。现在,一切正常
    • 非常感谢。我目前的熊猫版本是 0.17.1。我尝试使用 json Lint 并且 JSON 有效
    猜你喜欢
    • 2017-01-08
    • 1970-01-01
    • 2021-02-24
    • 2014-01-05
    • 2021-12-29
    • 2016-06-12
    • 2016-09-24
    • 1970-01-01
    • 2017-04-13
    相关资源
    最近更新 更多