【问题标题】:comparison Algorithm for two hashes两个哈希的比较算法
【发布时间】:2019-09-17 08:21:55
【问题描述】:

作为初始情况,我有一个 sha1 哈希值。我想将此与一个充满哈希值的文件进行比较,以查看 sha1 哈希值是否包含在具有哈希值的文件中。

更准确地说:

f1=sha1 #value read in
fobj = open("Hashvalues.txt", "r") #open file with hash values
for f1 in fobj:
  print ("hash value found")
else:
  print("HashValue not found")
fobj.close()

文件非常大(11.1GB)

是否有一种有用的算法可以尽可能快地执行搜索?散列文件中的散列值按散列排序。 我认为逐行比较不会是最快的方法,对吗?

编辑: 我将代码更改如下:

f1="9bc34549d565d9505b287de0cd20ac77be1d3f2c" #value read in
with open("pwned-passwords-sha1-ordered-by-hash-v5.txt") as f:
lineList = [line.rstrip('\n\r') for line in open("pwned-passwords-sha1- 
ordered-by-hash-v5.txt")]

def binarySearch(arr, l, r, x):

while l <= r:

    mid = l + (r - l)/2;

    # Check if x is present at mid
    if arr[mid] == x:
        return mid

    # If x is greater, ignore left half
    elif arr[mid] < x:
        l = mid + 1

    # If x is smaller, ignore right half
    else:
        r = mid - 1

# If we reach here, then the element
# was not present
return -1


# Test array
arr = lineList
x = "9bc34549d565d9505b287de0cd20ac77be1d3f2c" #value read in

# Function call
result = binarySearch(arr, 0, len(arr)-1, x)

if result != -1:
   print "Element is present at index % d" % result
else:
   print "Element is not present in array"

但它并没有我想象的那么快。我的实现是否正确?

编辑2:

def binarySearch (l, r, x):

# Check base case
if r >= l:

    mid = l + (r - l)/2

    # If element is present at the middle itself
    if getLineFromFile(mid) == x:
        return mid

    # If element is smaller than mid, then it
    # can only be present in left subarray
    elif getLineFromFile(mid) > x:
        return binarySearch(l, mid-1, x)

    # Else the element can only be present
    # in right subarray
    else:
        return binarySearch(mid + 1, r, x)

else:
    # Element is not present in the array
    return -1

x = '0000000A0E3B9F25FF41DE4B5AC238C2D545C7A8:15'

def getLineFromFile(lineNumber):
 with open('testfile.txt') as f:
  for i, line in enumerate(f):
    if i == lineNumber:
     return line
else:
 print('Not 7 lines in file')
 line = None

# get last Element of List
 def tail():
  for line in open('pwned.txt', 'r'):
    pass
  else:
    print line

 ausgabetail = tail()
 #print ausgabetail
 result = binarySearch( 0, ausgabetail, x)
 if result != -1:
    print "Element is present at index % d" % result
 else:
    print "Element is not present in array"

我现在的问题是为二进制搜索的右侧获取正确的索引。我传递了函数 (l, r, x)。左边从 0 开始。右边应该是文件的结尾,所以最后一行。我试图得到它,但它不起作用。我试图用 Funktion tail() 来解决这个问题。但是如果我在测试时打印 r,我会得到值“无”。 你还有别的想法吗?

【问题讨论】:

  • 一种合理的方法是二进制搜索,在这篇文章中进行了解释:stackoverflow.com/a/5219275/3768871
  • 您可以使用file.seek在O(1)中转到文件中的某个位置,回溯到之前的最后一个换行符,读取该行,并以这种方式进行二进制搜索文件。参见例如here.
  • @OmG 仅当文件中的哈希按字母顺序排序时才为真。
  • @nikoksr 它们按哈希排序
  • 我的问题也是,文件太大而无法加载到内存中。直到现在我的进程在我得到答案之前就被杀死了

标签: python algorithm search binary text-files


【解决方案1】:

查看代码我发现您仍在读取文件中的所有行,这确实是瓶颈。 这不是二分查找。 假设哈希是排序的 您可以只读取文件中的行数。 然后只需执行二进制搜索。您可以使用 Seek 来访问文件中的特定行,而不是读取整个文件,这样您只会读取 log(n) 行数。这应该会提高速度。

例子

def binarySearch(l, r, x): 
....
#change arr[mid] with getLineFromFile(mid)


....

def getLineFromFile(lineNumber):
with open('xxx.txt') as f:
for i, line in enumerate(f):
    if i == lineNumber:
        return line
else:
    print('Not 7 lines in file')
    line = None

【讨论】:

  • 你能举个小例子吗?
  • 感谢您对我有很大帮助! :)。但我遇到了另一个问题:如何设置正确的索引?在此之前,我将数组长度设为 -1。我现在该怎么做?
  • 这不应该是一个问题,有一些方法可以有效地计算文件中的行数stackoverflow.com/questions/9629179/…
  • 再次感谢。我现在已经尝试了很多,但它不起作用。你有没有给我一个示例代码它是如何工作的?即使有您的链接,我也无法正确处理。这让我很沮丧,对不起:(
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2017-09-05
  • 2016-09-05
相关资源
最近更新 更多