【问题标题】:how to copy values of one column of a dataframe to another column of other dataframe in pandas?如何将数据框的一列的值复制到熊猫中其他数据框的另一列?
【发布时间】:2020-12-30 04:18:17
【问题描述】:

在将字符串列的值逐个复制到另一个数据框的列时,我得到了这个包含方括号的输出:

chk.at[index,'StartLocation1'] = chkn['StartLocation1'].values
chk.at[index,'EndLocation1'] = chkn['EndLocation1'].values

0        [Petrol Pump-Ramji Ambedkar Nagar]
1                           [V Enterprises]
2                                   [Baola]
3                         [Dharmajyot-Vapi]
4    [KINGSTON TOWER VASAI Dominos-(THANE)]
Name: StartLocation1, dtype: object

所以我想进一步删除这个 [] 括号: 我已经应用了这个:

chk['EndLocation1'].str.strip('[]').astype(str)

0      nan
1      nan
2      nan
3      nan
4      nan

但是,我有 nan 值。请支持!

看看这是我的全部代码:

chk['StartLocation1'] = ''
chk['EndLocation1'] = ''

for index, row in chk.iterrows():
    start = row.StartTime
    end = row.EndTime
    reg = row.RegistrationNo

    query = "SELECT TOP 1 RegistrationNo, GPSDateTime, Location  FROM GPSEventsDataCurrentWeek where GPSDateTime Between 'start_date' and 'end_date' and RegistrationNo = 'reg' and GroundSpeed > 0 ORDER BY GPSDateTime ASC"

    query = query.replace('start_date', start.strftime('%m/%d/%Y %H:%M:%S'))
    query = query.replace('end_date', end.strftime('%m/%d/%Y %H:%M:%S'))
    query = query.replace('reg', str(reg))

    chk1 = pd.read_sql(query, con=engine)
    

    chk1 = chk1.rename({'Location': 'StartLocation1','GPSDateTime': 'StartTime'}, axis=1)
   

    query2 = "SELECT TOP 1 RegistrationNo, GPSDateTime, Location  FROM GPSEventsDataCurrentWeek where GPSDateTime Between 'start_date' and 'end_date' and RegistrationNo = 'reg' and GroundSpeed > 0 ORDER BY GPSDateTime DESC"

    query2 = query2.replace('start_date', start.strftime('%m/%d/%Y %H:%M:%S'))
    query2 = query2.replace('end_date', end.strftime('%m/%d/%Y %H:%M:%S'))
    query2 = query2.replace('reg', str(reg))

    chk2 = pd.read_sql(query2, con=engine)
    

    chk2 = chk2.rename({'Location': 'EndLocation1','GPSDateTime': 'EndTime'}, axis=1)
  
    chkn = pd.merge(chk1,chk2, on = ['RegistrationNo'], how = 'outer')
    print(chkn[['StartLocation1','EndLocation1']])
    

    chk.at['StartLocation1'] = chkn['StartLocation1'].values
    chk.at[index,'EndLocation1'] = chkn['EndLocation1'].values

这是我的数据框 chk:

Companyid   RegistrationNo  Date    Hour    Value   RunningDuration StartTime   EndTime
0   236.0   MH-01-CJ-3571   2020-09-01  0.0 True    00:42:00    2020-09-01 00:08:00 2020-09-01 00:59:00
1   236.0   MH-01-CV-7460   2020-09-01  0.0 True    00:49:00    2020-09-01 00:09:00 2020-09-01 00:58:00
2   654.0   MH-04-JK-4102   2020-09-01  0.0 True    00:03:00    2020-09-01 00:11:00 2020-09-01 00:24:00
3   654.0   DN-09-R-9421    2020-09-01  0.0 True    00:02:00    2020-09-01 00:24:00 2020-09-01 00:54:00
4   236.0   MH-01-CV-7456   2020-09-01  0.0 True    00:04:00    2020-09-01 00:38:00 2020-09-01 00:42:00

【问题讨论】:

  • 你原来的DataFrame是什么样的?为什么不直接分配列? chk['StartLocation1'] = chkn['StartLocation1']
  • 为什么要一一做这个?
  • @NYCCoder 我已经编辑了我的代码:你会明白为什么?
  • @ScottBoston 请检查编辑后的代码!

标签: python python-3.x pandas python-2.7 python-requests


【解决方案1】:

也许一个可能的解决方案是通过列表元素的连接选择列表中的唯一值 (join list of lists in python)。 df.values 返回一个 n 数组对象,因此返回的值是一个长度为 1 的列表,但建议使用 df.to_numpy 代替 (https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.values.html)

无论如何,我认为以这种方式分配列并不是最好的方法。

【讨论】:

    猜你喜欢
    • 2018-08-03
    • 1970-01-01
    • 1970-01-01
    • 2019-04-22
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2017-04-14
    • 2018-02-27
    相关资源
    最近更新 更多