【问题标题】:Generate unique IDs for a list of strings with duplicates为具有重复项的字符串列表生成唯一 ID
【发布时间】:2018-07-15 18:09:12
【问题描述】:

我想为从文本文件中读取的字符串生成 ID。如果字符串是重复的,我希望字符串的第一个实例有一个包含 6 个字符的 ID。对于该字符串的重复项,我希望 ID 与原始 ID 相同,但多了两个字符。我的逻辑有问题。这是我到目前为止所做的:

from itertools import groupby
import uuid
f = open('test.txt', 'r')
addresses = f.readlines()

list_of_addresses = ['Address']
list_of_ids = ['ID']


for x in addresses:
    list_of_addresses.append(x)


def find_duplicates():

    for x, y in groupby(sorted(list_of_addresses)):
        id = str(uuid.uuid4().get_hex().upper()[0:6])
        j = len(list(y))
        if j > 1:
            print str(j) + " instances of " + x
            list_of_ids.append(id)
        print list_of_ids

find_duplicates()

我应该如何处理这个问题?

编辑:这是test.txt的内容:

123 Test
123 Test
123 Test
321 Test
567 Test
567 Test

还有输出:

3 occurences of 123 Test

['ID', 'C10DD8']
['ID', 'C10DD8']
2 occurences of 567 Test

['ID', 'C10DD8', '595C5E']
['ID', 'C10DD8', '595C5E']

【问题讨论】:

  • 请给出示例输入和预期输出
  • 而重复的“字符串”是指重复的行重复一行中的单词吗?
  • @pylang 抱歉,添加了输入/输出。我的意思是重复的文本条目。
  • 再次查看您的输出。您缺少 321 并且您的 id 与您的重复项相同。您提到了再添加两个字符。

标签: python list uniqueidentifier


【解决方案1】:

如果字符串是重复的,我希望字符串的第一个实例有一个包含 6 个字符的 ID。对于该字符串的重复项,我希望 ID 与原始 ID 相同,但多了两个字符。

尝试使用collections.defaultdict

给定

import ctypes
import collections as ct


filename = "test.txt"


def read_file(fname):
    """Read lines from a file."""
    with open(fname, "r") as f:
        for line in f:
            yield line.strip()

代码

dd = ct.defaultdict(list)
for x in read_file(filename):
    key = str(ctypes.c_size_t(hash(x)).value)      # make positive hashes
    if key[:6] not in dd:
        dd[key[:6]].append(x)
    else:
        dd[key[:8]].append(x)

dd

输出

defaultdict(list,
            {'133259': ['123 Test'],
             '13325942': ['123 Test', '123 Test'],
             '210763': ['567 Test'],
             '21076377': ['567 Test'],
             '240895': ['321 Test']})

生成的字典对于每个第一次出现的唯一行都有键(长度为 6)。对于每个连续的复制行,两个额外的字符被分割为键。

您可以随心所欲地实现这些键。在这种情况下,我们使用hash() 将密钥与每个唯一行关联起来。然后我们从键中切出所需的序列。另请参阅有关制作positive hash values from ctypes 的帖子。


要检查您的结果,请从 defaultdict 创建相应的查找字典。

# Lookups 
occurrences = ct.defaultdict(int)
ids = ct.defaultdict(list)

for k, v in dd.items():
    key = v[0]
    occurrences[key] += len(v)
    ids[key].append(k)

# View data
for k, v in occurrences.items():
    print("{} instances of {}".format(v, k))
    print("IDs:", ids[k])
    print()

输出

1 instances of 321 Test
IDs: ['240895']

2 instances of 567 Test
IDs: ['21076377', '210763']

3 instances of 123 Test
IDs: ['13325942', '133259']

【讨论】:

  • 这太棒了。感谢您的详尽回复。
【解决方案2】:

您的问题有点令人困惑,我不明白生成 id 的标准是什么,这里我向您展示的只是逻辑而不是精确的解决方案,您可以从逻辑中获得帮助

track={}
with open('file.txt') as f:
    for line_no,line in enumerate(f):
        if line.split()[0] not in track:
            track[line.split()[0]]=[['ID','your_unique_id']]
        else:
            #here put your logic what you want to append if id is dublicate
            track[line.split()[0]].append(['ID','dublicate_id'+str(line_no)])

print(track)

输出:

{'123': [['ID', 'your_unique_id'], ['ID', 'dublicate_id1'], ['ID', 'dublicate_id2']], '321': [['ID', 'your_unique_id']], '567': [['ID', 'your_unique_id'], ['ID', 'dublicate_id5']]}

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 2012-06-17
    • 2013-02-09
    • 2021-07-08
    • 2012-09-04
    • 1970-01-01
    • 2011-01-12
    • 2018-12-10
    相关资源
    最近更新 更多