【问题标题】:Insert delimiters using rule lists使用规则列表插入分隔符
【发布时间】:2021-09-02 17:13:40
【问题描述】:

我有没有分隔符的数据。共有 10 列,其中每列的起始数字和每列值的长度指定如下:

starting_digit = [1,5,7,9,12,14,15,16,17,19]
col_length = [4,2,2,3,2,1,1,1,2,8]
df.columns = ["year","st","stfips","county","registry","race","hispanic","sex","age","pop"]

我可以知道如何应用“规则”(starting_digit 和 col_length)将数据分成不同的列吗?

data:
1969AL01001991910000000159
1969AL01001991910100000657
1969AL01001991910200001137
1969AL01001991910300000956
1969AL01001991910400000721
1969AL01001991910500000424
....
2019WY56035992921000000001
2019WY56035992921100000001
2019WY56035992921200000003
2019WY56035992921300000002
2019WY56035992921400000003

【问题讨论】:

  • 我得到col_lengthstarting_digit 是做什么的?
  • 是的,任何一个都是必需的......

标签: python pandas dataframe split


【解决方案1】:

试试pandas.read_fwf():

df = pd.read_fwf("data.txt", 
                  widths=[4,2,2,3,2,1,1,1,2,8], 
                  names=["year","st","stfips","county","registry","race","hispanic","sex","age","pop"])

【讨论】:

  • 谢谢!我现在正在尝试这个..由于数据很大,创建 df 需要更多时间
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2016-10-11
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2018-02-22
  • 2011-01-16
相关资源
最近更新 更多