【问题标题】:Slow file trawler -- python慢速文件拖网渔船——python
【发布时间】:2020-12-13 00:30:59
【问题描述】:

我编写了一个简短的脚本,在目录树中搜索与"Data*.txt" 匹配的最新文件,但速度非常慢。这是因为我不得不嵌套 for 循环(我怀疑)。

示例目录树:

ROOT
   |-- <directoryNameFoo1>
   |     |-- from  # This stays the same in each subdir...
   |            |-- <directoryNameBar1>
   |                  |-- Data*.txt
   |
   |-- <directoryNameFoo2>
   |     |-- from  # This stays the same in each subdir...
   |            |-- <directoryNameBar2>
   |                  |-- Data*.txt
   |
   |-- <directoryNameFoo3>
   |     |-- from  # This stays the same in each subdir...
   |            |-- <directoryNameBar3>
   |                  |-- Data*.txt

我的问题是:是否有更好/更快的方法来搜索目录结构,以便在每个子目录中找到与 "Data*.txt" 匹配的最新文件?

代码:

#!/usr/bin/env python
# -*- coding: utf-8 -*-

import os
import fnmatch
__basedir = os.path.abspath(os.path.dirname(__file__))

last_ctime = None
vehicle_root = None
file_list = []

for root, dirnames, filenames in os.walk(__basedir):
    vehdata = []
    for filename in fnmatch.filter(filenames, 'Data*.txt'):
        _file = os.path.join(root, filename)
        if vehicle_root == root:
            if os.path.getctime > last_ctime[1]:
                last_ctime = [_file, os.path.getctime(_file)]
            else:
                continue
        else:
            file_list.append(last_ctime)
            vehicle_root = root
            last_ctime = [_file, os.path.getctime(_file)]

        
print(file_list)

【问题讨论】:

    标签: python os.path fnmatch nested-for-loop


    【解决方案1】:

    您可以使用glob 来搜索特定模式数据而无需任何循环。 喜欢,

    import glob
    glob.glob('yourdir/Data*.txt')
    

    当你想在你定义的目录的所有子目录中搜索时,使用glob.glob('yourdir/Data*.txt,recursive=True)

    【讨论】:

    • 我仍然需要遍历 glob 输出的列表几乎需要完全相同的时间。
    猜你喜欢
    • 1970-01-01
    • 2018-11-03
    • 1970-01-01
    • 2016-07-29
    • 2013-01-13
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多