【问题标题】:Transforming Raw Data to a Table for reporting Python/Pandas将原始数据转换为表格以报告 Python/Pandas
【发布时间】:2022-06-29 23:52:54
【问题描述】:

请耐心等待,因为我正在自学。

基本上,我有这个原始数据,其中我得到了日期和 SLT 百分比,这是一个计算加上一个状态。

我想要将它们分组为行,计算每个月有多少已完成和未完成的列,并计算第三列的 SLT 百分比的平均值/平均值。

我一直在尝试做一个 grouper 或 groupby 或 unstack 并且也在 groupby 上做平均但我总是得到不正确的数据。我可以在 excel pivot 上轻松做到这一点,但我很难在 Python Dataframe 上重新创建它

原始数据:

ID SLT Date SLT Percent SLT State
1 5/28/2018 1 Made
2 11/13/2018 0 Mised
11 3/6/2019 0 Missed
12 5/20/2019 1 Made
13 10/25/2021 1 Made
14 11/12/2019 1 Made
18 6/4/2020 1 Made
19 6/11/2020 1 Made
20 8/6/2020 1 Made
21 12/9/2021 0 Missed
22 5/16/2022 1 Made
23 3/22/2018 0 Missed
24 3/20/2018 0 Missed
25 5/11/2018 1 Made
26 12/20/2018 0 Missed
27 5/12/2022 1 Made
28 10/7/2021 1 Made
29 3/21/2019 1 Made
30 4/24/2019 0 Missed

输出表:

Date Made Missed Percent
2020-5 10 2 80%
2020-6 25 15 60%
2020-7 50 23 23%

【问题讨论】:

标签: python pandas dataframe numpy data-analysis


【解决方案1】:

IIUC,你可以试试

df['SLT Date (Target)'] = pd.to_datetime(df['SLT Date (Target)']).dt.strftime('%Y-%b')
out = df.pivot_table(index='SLT Date (Target)', columns='SLT State', values='sltpercent', aggfunc='sum')
out.index = pd.MultiIndex.from_arrays(out.index.str.split('-'))
out['Percent'] = out['Made']/(out['Made']+out['Missed'])

【讨论】:

    猜你喜欢
    • 2020-03-21
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-10-26
    • 1970-01-01
    • 2021-07-21
    • 2014-08-19
    • 2021-09-27
    相关资源
    最近更新 更多