【问题标题】:Find all players with the longest winning streak查找所有连胜最长的玩家
【发布时间】:2022-01-26 06:13:21
【问题描述】:

我正在尝试解决使用 Python 的3.2. 找到最长连胜纪录的玩家的问题.谁能帮我一个更好的解决方案?另外,我很好奇如何调整这样的解决方案来输出每月或每年最长连胜的球员?

下面是示例数据框players_results

Player_id     Match_Date             Match_Result
401           2021-05-04 00:00:00    W
401           2021-05-09 00:00:00    L
401           2021-05-16 00:00:00    W
401           2021-05-18 00:00:00    W
401           2021-05-22 00:00:00    L
401           2021-06-15 00:00:00    L
401           2021-06-16 00:00:00    W
401           2021-06-18 00:00:00    W
402           2021-05-14 00:00:00    L
402           2021-05-23 00:00:00    L
402           2021-05-24 00:00:00    W
402           2021-06-01 00:00:00    W
402           2021-06-02 00:00:00    W
402           2021-07-01 00:00:00    W
403           2021-05-03 00:00:00    L
403           2021-05-11 00:00:00    W
403           2021-05-12 00:00:00    W
403           2021-05-13 00:00:00    W
403           2021-05-20 00:00:00    W
403           2021-05-25 00:00:00    W
403           2021-07-06 00:00:00    L
404           2021-05-10 00:00:00    W
404           2021-05-16 00:00:00    W
404           2021-05-20 00:00:00    W
404           2021-05-22 00:00:00    W
404           2021-05-28 00:00:00    L
405           2021-05-07 00:00:00    L
405           2021-05-25 00:00:00    W
405           2021-06-06 00:00:00    L
405           2021-06-07 00:00:00    W
405           2021-06-14 00:00:00    W
405           2021-07-01 00:00:00    W

预期输出

player_id     longest_winningstreak
403           5

我的不可扩展代码

df = players_results.groupby('player_id').count()
df = df['match_date'].reset_index()
#record result of player_id = 401 into vector b
a = [] # record the number of consecutive "W" of player_id = 401
x = 0
for i in range(df['match_date'][0]):
    if players_results['match_result'][i] == 'W':
        x = x + 1
        if (i == df['match_date'][0]-1):
            a.append(x)
    else:
        a.append(x)
        x = 0
print(a)
b = []
b.append([df['player_id'][0],max(a)])
print(b)

c = []
y=0
for i in range(df['match_date'][0], df['match_date'][0]+df['match_date'][1]):
    if players_results['match_result'][i] == 'W':
        y = y + 1
        if (i == df['match_date'][0]+df['match_date'][1]-1):
            c.append(y)
    else:
        c.append(y)
        y = 0
#record result of player_id = 402 into the vector b
b.append([df['player_id'][1],max(c)])

d = []
z=0
for i in range(df['match_date'][0]+df['match_date'][1], df['match_date'] [0]+df['match_date'][1]+df['match_date'][2]):
    if players_results['match_result'][i] == 'W':
        z = z + 1
        if (i == df['match_date'][0]+df['match_date'][1]+df['match_date'][2]-1):
            d.append(z)
    else:
        d.append(z)
        z = 0
b.append([df['player_id'][2],max(d)])

e = []
z2=0
for i in range(df['match_date'][0]+df['match_date'][1]+df['match_date'][2], df['match_date'][0]+df['match_date'][1]+df['match_date'][2]+df['match_date'][3]):
    if players_results['match_result'][i] == 'W':
        z2 = z2 + 1
        if (i == df['match_date'][0]+df['match_date'][1]+df['match_date'][2]+df['match_date'][3]-1):
            e.append(z2)
    else:
        e.append(z2)
        z2 = 0
#print(e)
b.append([df['player_id'][3],max(e)])

f = []
z3=0
for i in range(df['match_date'][0]+df['match_date'][1]+df['match_date'][2]+df['match_date'][3], df['match_date'][0]+df['match_date'][1]+df['match_date'][2]+df['match_date'][3]+df['match_date'][4]):
    if players_results['match_result'][i] == 'W':
        z3 = z3 + 1
        if (i == df['match_date'][0]+df['match_date'][1]+df['match_date'][2]+df['match_date'][3]+df['match_date'][4]-1):
            f.append(z3)
    else:
        f.append(z3)
        z3 = 0
#print(e)
b.append([df['player_id'][4],max(f)])   

【问题讨论】:

    标签: python python-3.x pandas


    【解决方案1】:

    过滤由累积和创建的组,比较不等于WSeriesGroupBy.value_counts,然后通过Series.aggSeries.idxmaxmax 使用player_id 获得最大值:

    m = df['Match_Result'].str.upper().ne('W')
    
    s = m.cumsum()[~m].groupby(df['Player_id']).value_counts().reset_index(level=1, drop=True)
    
    df = s.agg({'player_id': 'idxmax', 'longest_winningstreak':'max'}).to_frame(0).T
    print (df)
       player_id  longest_winningstreak
    0        403                      5
    

    每月解决方案:

    df['Match_Date'] = pd.to_datetime(df['Match_Date'])
    
    m = df['Match_Result'].ne('W')
    
    s = (m.cumsum()[~m].groupby([df['Player_id'], df['Match_Date'].dt.to_period('m')])
           .value_counts()
           .reset_index(level=-1, drop=True))
    
    df1 = (s.groupby(level=0)
            .agg([('period',lambda x: x.idxmax()[1]),('longest_winningstreak','max')]))
    
    print (df1)
                period  longest_winningstreak
    Player_id                                
    401        2021-06                      2
    402        2021-06                      2
    403        2021-05                      5
    404        2021-05                      4
    405        2021-06                      2
    

    【讨论】:

      【解决方案2】:

      一种使用itertools.groupby的方式:

      from itertools import groupby
      
      s = df["Match_Result"].str.lower().eq("w")
      
      def longest_pattern(ser):
          lens = [len(list(g)) for k, g in groupby(ser) if k]
          return max(lens)
      
      new_df = s.groupby(df["Player_id"]).apply(longest_pattern)
      

      或者有点诡计,但使用re.findall的另一种方式:

      import re
      
      def longest(string):
          return max(len(i) for i in re.findall("w+", string, flags=re.I))
      
      new_df = df.groupby("Player_id")["Match_Result"].sum().apply(longest)
      

      然后您可以使用pandas.Series.nlargest 获得所需的输出:

      new_df.nlargest(1)
      

      输出:

      Player_id
      403    5
      Name: Match_Result, dtype: int64
      

      【讨论】:

      • 非常感谢您的帮助,克里斯!一个简单的问题:您的df 是原始数据框players_results,对吗?另外,很抱歉我的错字,但Match_Result 列下的所有条目都应该是WL,所以我认为您的代码的第二行需要是s = df["Match_Result"].str.upper().eq("W")
      • 您的第一个代码有效!唯一的小问题是输出不在数据框中。我们是否有任何捷径,或者我们必须自己创建一个?
      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-07-28
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多