【发布时间】:2020-10-20 11:25:32
【问题描述】:
我有两个字典 h 和 c。这里 1,2,3 是文件夹名称,IMG_0001... 是每个特定文件夹中包含的所有图像文件。
这是我的真相
h = {'1': ['IMG_0001.png', 'IMG_0002.png', 'IMG_0003.png', 'IMG_0004.png'],
'2': ['IMG_0020.png', 'IMG_0021.png', 'IMG_0022.png', 'IMG_0023.png'],
'3': ['IMG_0051.png', 'IMG_0052.png', 'IMG_0053.png', 'IMG_0054.png']}
这是我的聚类输出图像
c = {'1': ['IMG_0001.png', 'IMG_0002.png', 'IMG_0053.png', 'IMG_0054.png'],
'2': ['IMG_0020.png', 'IMG_0021.png', 'IMG_0022.png', 'IMG_0023.png'],
'3': ['IMG_0003.png', 'IMG_0004.png', 'IMG_0051.png', 'IMG_0052.png']}
现在,我必须检查和比较两个字典并为每个文件夹生成一个 accuracy_score。 如何用python编写代码。有一个集群评估指标 - Adjusted Rand Index (ARI),但不知道我应该如何在这里使用它来比较 groundtruth 和集群字典。感谢你的帮助。非常感谢您的参与。我是python初学者。
import os, pprint
pp = pprint.PrettyPrinter()
h={}
for subdir, dirs, files in os.walk(r"folder_paths"):
for file in files:
key, value = os.path.basename(subdir), file #Get basefolder name & file name
h.setdefault(key, []).append(value) #Form DICT
pp.pprint(h)
#####################################
import os, pprint
pp = pprint.PrettyPrinter()
c={}
for subdir, dirs, files in os.walk(r"folder_paths"):
for file in files:
key, value = os.path.basename(subdir), file #Get basefolder name & file name
c.setdefault(key, []).append(value) #Form DICT
pp.pprint(c)
#####################################
# diff = {}
# #value = set(h.values()).intersection(set(c.values()))
# value = { k : second_dict[k] for k in set(second_dict) - set(first_dict) }
# print(value)
print("Changes in Ground Truth and Clustering")
import dictdiffer
for diff in list(dictdiffer.diff(h, c)):
print(diff)
【问题讨论】:
-
嗨,你能描述一下你想要达到的目标吗?即你想成为比较代码的输出是什么? (假设我不知道您在这种情况下所说的“基本事实”或“聚类”是什么意思)
标签: python python-3.x dictionary cluster-analysis