【问题标题】:How to create a co-occurence matrix of product orders in python?如何创建在python中订购的产品的共现矩阵?
【发布时间】:2021-02-26 10:04:23
【问题描述】:

假设我们有以下数据框,其中包括客户订单 (order_id) 和单个订单包含的产品 (product_id):

import pandas as pd

df = pd.DataFrame({'order_id' : [1, 1, 1, 1, 2, 2, 2, 3, 3, 3, 3],
                   'product_id' : [365, 48750, 3333, 9877, 48750, 32001, 3333, 3333, 365, 11202, 365]})
print(df)

   order_id product_id
0         1        365
1         1      48750
2         1       3333
3         1       9877
4         2      48750
5         2      32001
6         2       3333
7         3       3333
8         3        365
9         3      11202
10        3        365

了解产品对在同一个购物篮中出现的频率会很有趣。

如何在python中创建一个看起来像这样的共现矩阵:

       365  48750  3333  9877  32001  11202
365      1      1     2     1      0      1
48750    1      0     2     1      1      0
3333     2      2     0     1      1      1
9877     1      1     1     0      0      0
32001    0      1     1     0      0      0
11202    1      0     1     0      0      0

非常感谢您的帮助!

【问题讨论】:

  • 如果您列出到目前为止您尝试过的内容将会很有帮助。
  • @clarked 在 R 中我可以做到这一点,但我是 Python 新手,甚至不知道从哪里开始。很抱歉给您带来不便。

标签: python pandas product


【解决方案1】:

我们首先按 order_id 对 df 进行分组,然后在每个组中计算所有可能的对。请注意,我们首先按 product_id 排序,因此不同组中的相同对始终按相同顺序排列

import itertools
all_pairs = []
for _, group in df.sort_values('product_id').groupby('order_id'):
    all_pairs += list(itertools.combinations(group['product_id'],2))

all_pairs

我们从所有订单中获得所有对的列表

[('3333', '365'),
 ('3333', '48750'),
 ('3333', '9877'),
 ('365', '48750'),
 ('365', '9877'),
 ('48750', '9877'),
 ('32001', '3333'),
 ('32001', '48750'),
 ('3333', '48750'),
 ('11202', '3333'),
 ('11202', '365'),
 ('11202', '365'),
 ('3333', '365'),
 ('3333', '365'),
 ('365', '365')]

现在我们计算重复项

from collections import Counter

count_dict = dict(Counter(all_pairs))
count_dict

所以我们得到每对的计数,基本上是你想要的

{('3333', '365'): 3,
 ('3333', '48750'): 2,
 ('3333', '9877'): 1,
 ('365', '48750'): 1,
 ('365', '9877'): 1,
 ('48750', '9877'): 1,
 ('32001', '3333'): 1,
 ('32001', '48750'): 1,
 ('11202', '3333'): 1,
 ('11202', '365'): 2,
 ('365', '365'): 1}

将其放回叉积表中需要一些工作,关键是通过调用.apply(pd.Series) 将元组拆分为列,并最终通过unstack 将其中一列移动到列名:

(pd.DataFrame.from_dict(count_dict, orient='index')
    .reset_index(0)
    .set_index(0)['index']
    .apply(pd.Series)
    .rename(columns = {0:'pid1',1:'pid2'})
    .reset_index()
    .rename(columns = {0:'count'})
    .set_index(['pid1', 'pid2'] )
    .unstack()
    .fillna(0))

这会生成一个“紧凑”形式的表格,其中仅包含至少出现一对的产品


count
pid2    3333 365    48750  9877
pid1                
11202   1.0  2.0    0.0    0.0
32001   1.0  0.0    1.0    0.0
3333    0.0  3.0    2.0    1.0
365     0.0  1.0    1.0    1.0
48750   0.0  0.0    0.0    1.0

更新 这是上面的一个相当简化的版本,在 cmets 中进行了各种讨论

import numpy as np
import pandas as pd
from collections import Counter

# we start as in the original solution but use permutations not combinations
all_pairs = []
for _, group in df.sort_values('product_id').groupby('order_id'):
    all_pairs += list(itertools.permutations(group['product_id'],2))
count_dict = dict(Counter(all_pairs))

# We create permutations for _all_ product_ids ... note we use unique() but also product(..) to allow for (365,265) combinations
total_pairs = list(itertools.product(df['product_id'].unique(),repeat = 2))

# pull out first and second elements separately
pid1 = [p[0] for p in total_pairs]
pid2 = [p[1] for p in total_pairs]

# and get the count for those permutations that exist from count_dict. Use 0
# for those that do not
count = [count_dict.get(p,0) for p in total_pairs]

# Now a bit of dataFrame magic
df_cross = pd.DataFrame({'pid1':pid1, 'pid2':pid2, 'count':count})
df_cross.set_index(['pid1','pid2']).unstack()

我们完成了。 df_cross下方


count
pid2    11202   32001   3333    365 48750   9877
pid1                        
11202   0       0       1       2   0       0
32001   0       0       1       0   1       0
3333    1       1       0       3   2       1
365     2       0       3       2   1       1
48750   0       1       2       1   0       1
9877    0       0       1       1   1       0

【讨论】:

  • 聪明的解决方案!
  • @piterbarg 非常感谢您的回答!我能让你的方法奏效。 count_dict 表已经非常有用了。不幸的是,决赛桌有点混乱。例如值 (365, 3333) 是 0,而 (3333, 365) 是正确的 3。有没有办法创建一个 nxn 矩阵,其中 n 是产品的数量,在对角线上镜像?跨度>
  • 很高兴它有一些帮助!抱歉,我没有得到最终形状的答案。我再考虑考虑
  • @piterbarg 我能够通过使用 itertools.permutations 而不是 itertools.combinations 到达那里,如果您考虑一下,这是有道理的。它确实计算了 365 两次,与我的表不同,但这实际上是一件好事,您不需要按 product_id 排序。再次非常感谢您,您帮了我很多忙!
  • 太棒了,谢谢你让我知道,非常感谢
【解决方案2】:

这应该是一个很好的起点,也许会有用

pd.crosstab(df['order_id '], df['product_id'])

product_id  365    3333   9877   11202  32001  48750
order_id 
1            1      1      1      0      0      1
2            0      1      0      0      1      1
3            2      1      0      1      0      0

【讨论】:

    【解决方案3】:

    旋转以使每一行对应一个产品,然后将每一行映射到(df * row > 0).sum(1),这表示该产品与其他每个产品同时出现的订单数。

    >>> df = df.pivot_table(index='product_id', columns='order_id', aggfunc='size')
    >>> co_occ = df.apply(lambda row: (df * row > 0).sum(1), axis=1)
    >>> co_occ
    product_id  365    3333   9877   11202  32001  48750
    product_id                                          
    365             2      2      1      1      0      1
    3333            2      3      1      1      1      2
    9877            1      1      1      0      0      1
    11202           1      1      0      1      0      0
    32001           0      1      0      0      1      1
    48750           1      2      1      0      1      2
    

    可以使用np.fill_diagonal(co_occ.values, (df - 1).sum(1))将对角线修改为示例输出所暗示的约定(如果产品至少有两个以相同的顺序出现,则该产品与自身同时出现)。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2012-10-28
      • 2011-08-03
      • 1970-01-01
      • 2023-03-10
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2011-07-18
      相关资源
      最近更新 更多