【问题标题】:Create new column with incremental values efficiently有效地创建具有增量值的新列
【发布时间】:2018-09-03 07:19:03
【问题描述】:

我正在创建一个具有增量值的列,然后在列的开头附加一个字符串。当用于大数据时,这非常慢。请建议一种更快更有效的方法。

df['New_Column'] = np.arange(df[0])+1
df['New_Column'] = 'str' + df['New_Column'].astype(str)

输入

id  Field   Value
1     A       1
2     B       0     
3     D       1

输出

id  Field   Value   New_Column
1     A       1     str_1
2     B       0     str_2
3     D       1     str_3

【问题讨论】:

标签: python performance pandas numpy


【解决方案1】:

一种可能的解决方案是通过map 将值转换为strings:

df['New_Column'] = np.arange(len(df['a']))+1
df['New_Column'] = 'str_' + df['New_Column'].map(str)

【讨论】:

    【解决方案2】:

    我会再添加两个

    麻木

    from numpy.core.defchararray import add
    
    df.assign(new=add('str_', np.arange(1, len(df) + 1).astype(str)))
    
       id Field  Value    new
    0   1     A      1  str_1
    1   2     B      0  str_2
    2   3     D      1  str_3
    

    f-string 理解中

    Python 3.6+
    df.assign(new=[f'str_{i}' for i in range(1, len(df) + 1)])
    
       id Field  Value    new
    0   1     A      1  str_1
    1   2     B      0  str_2
    2   3     D      1  str_3
    

    时间测试

    结论

    相对于简单的性能,理解力赢得了胜利。请注意,这是 cᴏʟᴅsᴘᴇᴇᴅ 提出的方法。感谢您的支持(谢谢),但让我们在应得的地方给予赞扬。

    Cythonizing 理解似乎没有帮助。 f弦也没有。
    Divakar 的 numexp 在处理更大数据时的性能名列前茅。

    功能

    %load_ext Cython
    

    %%cython
    def gen_list(l, h):
        return ['str_%s' % i for i in range(l, h)]
    

    pir1 = lambda d: d.assign(new=[f'str_{i}' for i in range(1, len(d) + 1)])
    pir2 = lambda d: d.assign(new=add('str_', np.arange(1, len(d) + 1).astype(str)))
    cld1 = lambda d: d.assign(new=['str_%s' % i for i in range(1, len(d) + 1)])
    cld2 = lambda d: d.assign(new=gen_list(1, len(d) + 1))
    jez1 = lambda d: d.assign(new='str_' + pd.Series(np.arange(1, len(d) + 1), d.index).astype(str))
    div1 = lambda d: d.assign(new=create_inc_pattern(prefix_str='str_', start=1, stop=len(d) + 1))
    div2 = lambda d: d.assign(new=create_inc_pattern_numexpr(prefix_str='str_', start=1, stop=len(d) + 1))
    

    测试

    res = pd.DataFrame(
        np.nan, [10, 30, 100, 300, 1000, 3000, 10000, 30000],
        'pir1 pir2 cld1 cld2 jez1 div1 div2'.split()
    )
    
    for i in res.index:
        d = pd.concat([df] * i)
        for j in res.columns:
            stmt = f'{j}(d)'
            setp = f'from __main__ import {j}, d'
            res.at[i, j] = timeit(stmt, setp, number=200)
    

    结果

    res.plot(loglog=True)
    

    res.div(res.min(1), 0)
    
               pir1      pir2      cld1      cld2       jez1      div1      div2
    10     1.243998  1.137877  1.006501  1.000000   1.798684  1.277133  1.427025
    30     1.009771  1.144892  1.012283  1.000000   2.144972  1.210803  1.283230
    100    1.090170  1.567300  1.039085  1.000000   3.134154  1.281968  1.356706
    300    1.061804  2.260091  1.072633  1.000000   4.792343  1.051886  1.305122
    1000   1.135483  3.401408  1.120250  1.033484   7.678876  1.077430  1.000000
    3000   1.310274  5.179131  1.359795  1.362273  13.006764  1.317411  1.000000
    10000  2.110001  7.861251  1.942805  1.696498  17.905551  1.974627  1.000000
    30000  2.188024  8.236724  2.100529  1.872661  18.416222  1.875299  1.000000
    

    更多功能

    def create_inc_pattern(prefix_str, start, stop):
        N = stop - start # count of numbers
        W = int(np.ceil(np.log10(N+1))) # width of numeral part in string
        dl = len(prefix_str)+W # datatype length
        dt = np.uint8 # int datatype for string to-from conversion 
    
        padv = np.full(W,48,dtype=np.uint8)
        a0 = np.r_[np.fromstring(prefix_str,dtype='uint8'), padv]
    
        r = np.arange(start, stop)
    
        addn = (r[:,None] // 10**np.arange(W-1,-1,-1))%10
        a1 = np.repeat(a0[None],N,axis=0)
        a1[:,len(prefix_str):] += addn.astype(dt)
        a1.shape = (-1)
    
        a2 = np.zeros((len(a1),4),dtype=dt)
        a2[:,0] = a1
        return np.frombuffer(a2.ravel(), dtype='U'+str(dl))
    
    import numexpr as ne
    
    def create_inc_pattern_numexpr(prefix_str, start, stop):
        N = stop - start # count of numbers
        W = int(np.ceil(np.log10(N+1))) # width of numeral part in string
        dl = len(prefix_str)+W # datatype length
        dt = np.uint8 # int datatype for string to-from conversion 
    
        padv = np.full(W,48,dtype=np.uint8)
        a0 = np.r_[np.fromstring(prefix_str,dtype='uint8'), padv]
    
        r = np.arange(start, stop)
    
        r2D = r[:,None]
        s = 10**np.arange(W-1,-1,-1)
        addn = ne.evaluate('(r2D/s)%10')
        a1 = np.repeat(a0[None],N,axis=0)
        a1[:,len(prefix_str):] += addn.astype(dt)
        a1.shape = (-1)
    
        a2 = np.zeros((len(a1),4),dtype=dt)
        a2[:,0] = a1
        return np.frombuffer(a2.ravel(), dtype='U'+str(dl))
    

    【讨论】:

    • 也加我的吗? :)
    • @Divakar 已更新。看看astype(str) 是否花费太多
    • @piRSquared 是的,.astype(str) 看起来是一个很好的解决方案,无需太多开销。所以,继续吧,我想说的是。
    • 从我刚才跑的,你失去了所有的优势。 Cython 重回榜首。我怀疑一定有其他方法。
    • 更新时间
    【解决方案3】:

    建议的方法

    在对字符串和数字 dtype 进行了相当多的修改并利用它们之间的简单互操作性之后,我最终得到了用零填充的字符串,因为 NumPy 做得很好并且允许以这种方式进行矢量化操作 -

    def create_inc_pattern(prefix_str, start, stop):
        N = stop - start # count of numbers
        W = int(np.ceil(np.log10(stop+1))) # width of numeral part in string
    
        padv = np.full(W,48,dtype=np.uint8)
        a0 = np.r_[np.fromstring(prefix_str,dtype='uint8'), padv]
        a1 = np.repeat(a0[None],N,axis=0)
    
        r = np.arange(start, stop)
        addn = (r[:,None] // 10**np.arange(W-1,-1,-1))%10
        a1[:,len(prefix_str):] += addn.astype(a1.dtype)
        return a1.view('S'+str(a1.shape[1])).ravel()
    

    引入 numexpr 以实现更快的广播 + 模运算 -

    import numexpr as ne
    
    def create_inc_pattern_numexpr(prefix_str, start, stop):
        N = stop - start # count of numbers
        W = int(np.ceil(np.log10(stop+1))) # width of numeral part in string
    
        padv = np.full(W,48,dtype=np.uint8)
        a0 = np.r_[np.fromstring(prefix_str,dtype='uint8'), padv]
        a1 = np.repeat(a0[None],N,axis=0)
    
        r = np.arange(start, stop)
        r2D = r[:,None]
        s = 10**np.arange(W-1,-1,-1)
        addn = ne.evaluate('(r2D/s)%10')
        a1[:,len(prefix_str):] += addn.astype(a1.dtype)
        return a1.view('S'+str(a1.shape[1])).ravel()
    

    所以,作为新列使用:

    df['New_Column'] = create_inc_pattern(prefix_str='str_', start=1, stop=len(df)+1)
    

    示例运行 -

    In [334]: create_inc_pattern_numexpr(prefix_str='str_', start=1, stop=14)
    Out[334]: 
    array(['str_01', 'str_02', 'str_03', 'str_04', 'str_05', 'str_06',
           'str_07', 'str_08', 'str_09', 'str_10', 'str_11', 'str_12', 'str_13'], 
          dtype='|S6')
    
    In [338]: create_inc_pattern(prefix_str='str_', start=1, stop=124)
    Out[338]: 
    array(['str_001', 'str_002', 'str_003', 'str_004', 'str_005', 'str_006',
           'str_007', 'str_008', 'str_009', 'str_010', 'str_011', 'str_012',..
           'str_115', 'str_116', 'str_117', 'str_118', 'str_119', 'str_120',
           'str_121', 'str_122', 'str_123'], 
          dtype='|S7')
    

    说明

    基本概念和逐步示例运行说明

    基本思想是创建 ASCII 等效数字数组,可以通过 dtype 转换查看或转换为字符串。更具体地说,我们将创建 uint8 类型的数字。因此,每个字符串将由一维数字数组表示。对于将转换为二维数字数组的字符串列表,每行(一维数组)代表一个字符串。

    1) 输入:

    In [22]: prefix_str='str_'
        ...: start=15
        ...: stop=24
    

    2) 参数:

    In [23]: N = stop - start # count of numbers
        ...: W = int(np.ceil(np.log10(stop+1))) # width of numeral part in string
    
    In [24]: N,W
    Out[24]: (9, 2)
    

    3) 创建表示起始字符串的一维数字数组:

    In [25]: padv = np.full(W,48,dtype=np.uint8)
        ...: a0 = np.r_[np.fromstring(prefix_str,dtype='uint8'), padv]
    
    In [27]: a0
    Out[27]: array([115, 116, 114,  95,  48,  48], dtype=uint8)
    

    4) 将字符串范围扩展到二维数组:

    In [33]: a1 = np.repeat(a0[None],N,axis=0)
        ...: r = np.arange(start, stop)
        ...: addn = (r[:,None] // 10**np.arange(W-1,-1,-1))%10
        ...: a1[:,len(prefix_str):] += addn.astype(a1.dtype)
    
    In [34]: a1
    Out[34]: 
    array([[115, 116, 114,  95,  49,  53],
           [115, 116, 114,  95,  49,  54],
           [115, 116, 114,  95,  49,  55],
           [115, 116, 114,  95,  49,  56],
           [115, 116, 114,  95,  49,  57],
           [115, 116, 114,  95,  50,  48],
           [115, 116, 114,  95,  50,  49],
           [115, 116, 114,  95,  50,  50],
           [115, 116, 114,  95,  50,  51]], dtype=uint8)
    

    5) 因此,每一行代表一个字符串的 ascii 等价物,每个字符串都脱离了所需的输出。让我们完成最后一步:

    In [35]: a1.view('S'+str(a1.shape[1])).ravel()
    Out[35]: 
    array(['str_15', 'str_16', 'str_17', 'str_18', 'str_19', 'str_20',
           'str_21', 'str_22', 'str_23'], 
          dtype='|S6')
    

    时间

    这是一个针对列表理解版本的快速时间测试,从其他帖子的时间来看似乎效果最好 -

    In [339]: N = 10000
    
    In [340]: %timeit ['str_%s'%i for i in range(N)]
    1000 loops, best of 3: 1.12 ms per loop
    
    In [341]: %timeit create_inc_pattern_numexpr(prefix_str='str_', start=1, stop=N)
    1000 loops, best of 3: 490 µs per loop
    
    In [342]: N = 100000
    
    In [343]: %timeit ['str_%s'%i for i in range(N)]
    100 loops, best of 3: 14 ms per loop
    
    In [344]: %timeit create_inc_pattern_numexpr(prefix_str='str_', start=1, stop=N)
    100 loops, best of 3: 4 ms per loop
    

    Python-3 代码

    在 Python-3 上,要获取字符串 dtype 数组,我们需要在中间的 int dtype 数组上再填充几个零。因此,Python-3 的不带和带 numexpr 版本的等价物最终变成了类似的东西 -

    方法 #1(无 numexpr):

    def create_inc_pattern(prefix_str, start, stop):
        N = stop - start # count of numbers
        W = int(np.ceil(np.log10(stop+1))) # width of numeral part in string
        dl = len(prefix_str)+W # datatype length
        dt = np.uint8 # int datatype for string to-from conversion 
    
        padv = np.full(W,48,dtype=np.uint8)
        a0 = np.r_[np.fromstring(prefix_str,dtype='uint8'), padv]
    
        r = np.arange(start, stop)
    
        addn = (r[:,None] // 10**np.arange(W-1,-1,-1))%10
        a1 = np.repeat(a0[None],N,axis=0)
        a1[:,len(prefix_str):] += addn.astype(dt)
        a1.shape = (-1)
    
        a2 = np.zeros((len(a1),4),dtype=dt)
        a2[:,0] = a1
        return np.frombuffer(a2.ravel(), dtype='U'+str(dl))
    

    方法 #2(使用 numexpr):

    import numexpr as ne
    
    def create_inc_pattern_numexpr(prefix_str, start, stop):
        N = stop - start # count of numbers
        W = int(np.ceil(np.log10(stop+1))) # width of numeral part in string
        dl = len(prefix_str)+W # datatype length
        dt = np.uint8 # int datatype for string to-from conversion 
    
        padv = np.full(W,48,dtype=np.uint8)
        a0 = np.r_[np.fromstring(prefix_str,dtype='uint8'), padv]
    
        r = np.arange(start, stop)
    
        r2D = r[:,None]
        s = 10**np.arange(W-1,-1,-1)
        addn = ne.evaluate('(r2D/s)%10')
        a1 = np.repeat(a0[None],N,axis=0)
        a1[:,len(prefix_str):] += addn.astype(dt)
        a1.shape = (-1)
    
        a2 = np.zeros((len(a1),4),dtype=dt)
        a2[:,0] = a1
        return np.frombuffer(a2.ravel(), dtype='U'+str(dl))
    

    时间安排 -

    In [8]: N = 100000
    
    In [9]: %timeit ['str_%s'%i for i in range(N)]
    100 loops, best of 3: 18.5 ms per loop
    
    In [10]: %timeit create_inc_pattern_numexpr(prefix_str='str_', start=1, stop=N)
    100 loops, best of 3: 6.06 ms per loop
    

    【讨论】:

    • ??? 这值得赏金。如果不是太麻烦的话,能不能加个cmet,让小凡人能看懂是怎么回事? ;)
    • 我正在获取字节字符串。我必须编码。所以现在我必须先真正理解你的代码(-:仍在研究它
    • @piRSquared 啊,一定是 Python3 的东西。该死!我在 Python-2 上。
    • 一个简单的astype(str) 修复了它,但我确信这很笨拙。暂时不使用它。
    • 我会建议切换到 python3,因为有很多性能改进(包括列表 comp,我敢打赌)
    【解决方案4】:

    当所有其他方法都失败时,使用 列表推导

    df['NewColumn'] = ['str_%s' %i for i in range(1, len(df) + 1)]
    

    如果您对函数进行 cythonize,则可能会进一步加速:

    %load_ext Cython
    
    %%cython
    def gen_list(l, h):
        return ['str_%s' %i for i in range(l, h)]
    

    注意,此代码在 Python3.6.0 (IPython6.2.1) 上运行。感谢 cmets 中的@hpaulj 改进了解决方案。


    # @jezrael's fastest solution
    
    %%timeit
    df['NewColumn'] = np.arange(len(df['a'])) + 1
    df['NewColumn'] = 'str_' + df['New_Column'].map(str)
    
    547 ms ± 13.6 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
    

    # in this post - no cython
    
    %timeit df['NewColumn'] = ['str_%s'%i for i in range(n)]
    409 ms ± 9.36 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
    

    # cythonized list comp 
    
    %timeit df['NewColumn'] = gen_list(1, len(df) + 1)
    370 ms ± 9.23 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)
    

    【讨论】:

    • 嗯,有趣,对我来说列表理解比较慢,可能是因为 windows?不确定
    • @jezrael 非常有趣,我想这取决于很多因素,很可能是操作系统,因为我在 Unix 机器上运行它。
    • 是的,我已经尝试过了,但没有添加到解决方案中,因为速度较慢;)
    • @DJK - 我再次运行计时,更好的时间和更好的标准
    • @DJK 感谢您指出这一点。我当然没有注意到:/
    猜你喜欢
    • 2017-12-17
    • 2014-05-24
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多