【问题标题】:Python record level comparision 2 large delimited filesPython记录级别比较2大分隔文件
【发布时间】:2014-10-26 08:22:46
【问题描述】:

我有 2 个大的分隔文件。

需要帮助:

a) 我需要根据两个文件的 键列 获取 行数

b) 根据两个文件中的键列查找重复项

c) 从两个文件中获取重复计数

d) 应将副本创建为单独的文件

e) 获取两个文件中的通用记录

f) 排序两个文件(共同记录)

g) 排序后比较两个文件并获得不匹配计数

h) 不匹配记录应创建为单独的文件。

任何帮助将不胜感激。

【问题讨论】:

    标签: python csv python-3.x


    【解决方案1】:

    由于您必须对文件进行排序,因此您必须将它们加载到内存中,您可能会执行以下操作:

    #see @ http://www.grantjenks.com/docs/sortedcontainers/ for information about sorted containers.
    #they are efficient for huge data.
    from sortedcontainers import SortedList, SortedDict
    
    file1=SortedList()
    file2=SortedList()
    
    delimiter = ";"
    commentSign="#"
    path1="./data1"
    path2="./data2"
    
    def get_values_column(delimited_lines, column_number, delimiter, commentSign):
        values = set()
        for line in delimited_lines:
            if line[0] != commentSign:
                fields = line.split(delimiter)
                values.add(fields[column_number])
        return values
    
    def count_not_in_other(collection1, collection2):
        uniq1 = []
        uniq2 = []
        for elem in collection1:
            if elem not in collection2:
                uniq1.append(elem)
    
        for elem in collection2:
            if elem not in collection1:
                uniq2.append(elem)
    
        return (uniq1,uniq2)
    
    def from_line_list_to_line_count(line_list):
        lines = SortedDict()
    
        for line in line_list:
            if line not in lines.keys():
                lines[line] = 0
            lines[line] += 1
    
        return lines
    
    def duplicated_lines(line_list):
        lines_count = from_line_list_to_line_count(line_list)
        return list(filter( lambda x: lines_count[x]>1, lines_count.keys()))
    
    if __name__ == "__main__":
    
        with open(path1, "r") as io1, open(path2,"r") as io2 :
            #copy the file in memory and sort them.
            for line in io1:
                file1.add(line)
    
            for line in io2:
                file2.add(line)
    
        with open(path1, "w") as io1, open(path2,"w") as io2 :
            #rewrite sorted files
            for line in file1:
                io1.write(line)
    
            for line in file2:
                io2.write(line)
    
            print("There is {0} different key value in {1}".format(len(get_values_column(file1, 0, delimiter, commentSign)), path1))
            print("There is {0} different key value in {1}".format(len(get_values_column(file2, 0, delimiter, commentSign)), path2))
    
            uniques = count_not_in_other(file1, file2)
            print("There is {0} lines present in file1 that are not present in file2".format(len(uniques[0])))
            print("There is {0} lines present in file2 that are not present in file1".format(len(uniques[1])))
    
    
            print("file1 duplicated lines are : {0}".format(duplicated_lines(file1)))
            print("file2 duplicated lines are : {0}".format(duplicated_lines(file2)))       
    

    我将它与这些数据一起使用: 数据1

    #id; name; value
    10; foo; 100
    10; foo; 100
    10; foo; 101
    11; foo; 50
    13; bar; 500
    

    数据2

    #id; name; value
    10; foo; 100
    11; foo; 50
    13; bar; 500
    18; bar foo; 46
    10; foo; 100
    10; foo; 101
    18; bar foo; 46
    

    当你要求一份非常完整的工作时却没有提供任何关于你之前尝试过什么的线索,我只会让你这样做。尝试理解代码并完成它。现在,您对文件进行排序并获取每个文件的(数量)键。

    注意:我与 sortedcontainers 库没有任何关系。

    【讨论】:

    • 谢谢......我明白了......现在我留下了重复计数和不匹配计数。重复计数应该基于#id,对于不匹配,它应该像元素到元素一样。 .something it should look like Mismatch on column NAME, row number 5 , Source value is : 13 and Target value is : 10
    猜你喜欢
    • 1970-01-01
    • 2010-09-18
    • 1970-01-01
    • 2020-07-21
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2011-07-19
    相关资源
    最近更新 更多