【问题标题】:unpack_from() des not work with big filesunpack_from() 不适用于大文件
【发布时间】:2015-07-04 18:43:49
【问题描述】:

我尝试使用 python3 从一些固定宽度格式的文件(来自here)中读取数据。

如果我只预选几行,它可以正常工作,但如果我想去 通过洞文件(大约 1000 行,每行 611 个块,4 个字符 = 2444 个字符)python 告诉我,struct.Struct(bytes).unpackFrom(bytes) 需要a buffer of at least 2444 bytes,目前我不知道为什么它没有这么大的缓冲区。

我在 64 位 Linux 上运行,4 Gig RAM 和 20 Gig Swap 可能会有所帮助。

sn-p的代码是这样的:

#edit
"""rowMask is 611 times 4s, just to prevent you from counting it... """
rowMask="4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s4s"
def readUsableFields(filename,stdPath):
    usableFields=[]
    with open(stdPath+filename,"r") as f:
            count_line=0
            for line in f:
                    count_col=0
                    fields=struct.Struct(bytes(rowMask,"UTF-8")).unpack_from(bytes(line,"UTF-8"))
                    for field in fields:
                            if(field!=-999):
                                    usableFields.append([count_line,count_col])
                            count_col+=1
                    count_line+=1
            return usableFields

我还查看了 thisthis ,但它们都不是我问题的答案。

一些帮助会很好,如果我的问题是重复的(我没有找到)请告诉我。

【问题讨论】:

  • 什么是rowMask?你的线的长度是多少?您文件中的所有行实际上都具有相同的长度吗?也许在unpack_from 调用之前添加一个print(len(line)) 语句以进行调试。
  • @Blckknght 我添加了信息。问题是我不知道为什么它没有得到 2444 字节的缓冲区
  • 一些可能相关或不相关的 cmets:您应该使用 "4s"*611 而不是显式编写长字符串。您应该以二进制模式打开文件,而不是使用(默认)文本模式,然后在 unpack_from 调用中重新编码为 UTF-8。如果这些事情都没有帮助,我怀疑您的数据在某处有一条坏线(可能是页眉或页脚?)。也许捕捉到异常并在那时打印出line 的值?
  • 哦,有时我觉得自己很愚蠢!这是页脚!谢谢。

标签: python-3.x struct fixed-width


【解决方案1】:

由于许多固定宽度的文件都有页脚(或页眉)代码

在页脚上会失败,因为它可能没有正确的长度。

所以你必须检查正确的线长:

rowMask="4s"*611
def readUsableFields(filename,stdPath):
    usableFields=[]

    with open(stdPath+filename,"r") as f:
            count_line=0
            for line in f:
                    count_col=0
                    # len(line) = 611 * 4 +1
                    # as there is a trailing '\0'
                    if(len(line)!=2445):
                            continue
                    fields=struct.Struct(bytes(rowMask,"UTF-8")).unpack_from(bytes(line,"UTF-8"))
                    for field in fields:
                            if(field!=-999):
                                    usableFields.append([count_line,count_col])
                            count_col+=1
                    count_line+=1

            f.close()
    return usableFields

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 2013-11-06
    • 2012-11-06
    • 1970-01-01
    • 2013-03-07
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2019-06-01
    相关资源
    最近更新 更多