【问题标题】:Using Google Colab how to drive.files().list more than 1000 files from google drive使用 Google Colab 如何 drive.files().list 来自 google drive 的 1000 多个文件
【发布时间】:2022-12-18 07:28:07
【问题描述】:

大约一个月一次,我得到一个包含大量视频(通常大约 700-800 个)的谷歌驱动器文件夹和一个电子表格,其中 A 列按照视频文件中的时间戳顺序填充了所有视频文件的名称姓名。现在我已经有了执行此操作的代码(我将在下面发布)但是这次我在文件夹中有大约 8,400 个视频文件并且该算法的 pageSize 限制为 1,000(最初是 100,我更改了它到 1,000,但这是它可以接受的最高值)如何更改此代码以接受超过 1000

这是初始化一切的部分

!pip install gspread_formatting

import time
import gspread
from gspread import urls
from google.colab import auth
from datetime import datetime
from datetime import timedelta
from gspread_formatting import *
from googleapiclient.discovery import build
from oauth2client.client import GoogleCredentials
from google.auth import default


folder_id = '************************' # change to whatever folder the required videos are in

base_dir = '/Example/drive/videofolder' # change this to whatever folder path you want to grab videos from same as above

file_name_qry_filter = "name contains 'mp4' and name contains 'cam'"

file_pattern="cam*.mp4"

spreadSheetUrl = 'https://docs.google.com/spreadsheets/d/SpreadsheetIDExample/edit#gid=0'
data_drive_id = '***********' # This is the ID of the shared Drive


auth.authenticate_user()
creds, _ = default()
gc = gspread.authorize(creds)
#gc = gspread.authorize(GoogleCredentials.get_application_default())
wb = gc.open_by_url(spreadSheetUrl)
sheet = wb.worksheet('Sheet1')

这是代码的主要部分

prevTimeStamp = None
prevHour = None

def dateChecker(fileName, prevHour):
  strippedFileName = fileName.strip(".mp4")             # get rid of the .mp4 from the end of the file name
  parsedFileName = strippedFileName.split("_")          # split the file name into an array of (0 = Cam#, 1 = yyyy-mm-dd, 2 = hh-mm-ss)
  timeStamp = parsedFileName[2]                         # Grabbed specifically the hh-mm-ss time section from the original file name
  parsedTimeStamp = timeStamp.split("-")                # split the time stamp into an array of (0 = hour, 1 = minute, 2 = second)
  hour = int(parsedTimeStamp[0])                              
  minute = int(parsedTimeStamp[1])
  second = int(parsedTimeStamp[2])                           # set hour, minute, and seccond to it's own variable
  commentCell = "Reset"

  if prevHour == None:
    commentCell = " "

    prevHour = hour

  else:
  
    if 0 <= hour < 24:

      if hour == 0:
        if prevHour == 23:
          commentCell = " "
        else:
          commentCell = "Missing Video1"

      else:
        if hour - prevHour == 1:
          commentCell = " "
        else:
          commentCell = "Missing Video2"

    else:
      commentCell = "Error hour is not between 0 and 23"

    if minute != 0 or 1 < second <60:
      commentCell = "Check Length"

  prevHour = hour

  return commentCell, prevHour




# Drive query variables
parent_folder_qry_filter = "'" + folder_id + "' in parents"  #you shouldn't ever need to change this
query = file_name_qry_filter + " and " + parent_folder_qry_filter
drive_service = build('drive', 'v3')

# Build request and call Drive API
page_token = None
response = drive_service.files().list(q=query,
                                      corpora='drive',
                                      supportsAllDrives='true',
                                      includeItemsFromAllDrives='true',
                                      driveId=data_drive_id,
                                      pageSize=1000,
                                      fields='nextPageToken, files(id, name, webViewLink)',  # you can add extra fields in the files() if you need more information about the files you're grabbing
                                      pageToken=page_token).execute()
i = 1
array = [[],[]]  
# Parse/print results
for file in response.get('files', []):
    array.insert(i-1, [file.get('name'), file.get('webViewLink')]) # If you add extra fields above, this is where you will have to start changing the code to make it accomadate the extra fields
    i = i + 1  


array.sort()
array_sorted = [x for x in array if x]  #Idk man this is some alien shit I just copied it from the internet and it worked, it somehow removes any extra blank objects in the array that aren't supposed to be there 
arrayLength = len(array_sorted)
print(arrayLength)

commentCell = 'Error'

# for file_name in array_sorted:
#   date_gap, start_date, end_date = date_checker(file_name[0])
#   if prev_end_date == None:
#     print('hello')
#   elif start_date != prev_end_date:
#     date_gap = 'Missing Video'

for file_name in array_sorted:
  commentCell, prevHour = dateChecker(file_name[0], prevHour)

  time.sleep(0.3)
  #insertRow = [file_name[0], "Not Processed",  " ", date_gap, " ", " ", " ", " ", base_dir + '/' + file_name[0], " ", file_name[1], " ", " ", " "]
  insertRow = [file_name[0], "Not Processed",  " ", commentCell, " ", " ", " ", " ", " ", " ", " ", " ", " ", " ", " ", " ", " ", " "]
  sheet.append_row(insertRow, value_input_option='USER_ENTERED')

现在我知道问题与

page_token = None
response = drive_service.files().list(q=query,
                                      corpora='drive',
                                      supportsAllDrives='true',
                                      includeItemsFromAllDrives='true',
                                      driveId=data_drive_id,
                                      pageSize=1000,
                                      fields='nextPageToken, files(id, name, webViewLink)',  # you can add extra fields in the files() if you need more information about the files you're grabbing
                                      pageToken=page_token).execute()

在代码主体部分的中间。我显然已经尝试过将 pageSize 限制更改为 10,000,但我知道那是行不通的,我是对的,它回来了

HttpError:<HttpError 400 请求 https://www.googleapis.com/drive/v3/files?q=name+contains+%27mp4%27+and+name+contains+%27cam%27+and+%271ANmLGlNr-Cu0BvH2aRrAh_GXEDk1nWvf%27+in+parents&corpora=drive&supportsAllDrives=true&includeItemsFromAllDrives=true&driveId=0AF92uuRq-00KUk9PVA&pageSize=10000&fields=nextPageToken%2C+files%28id%2C+name%2C+webViewLink%29&alt=json 返回“无效值‘10000’。值必须在范围内:[1, 1000]”。详细信息:“无效值‘10000’。值必须在范围内:[1, 1000]”>

我的一个想法是拥有多个页面,每个页面有 1000 个页面,而不是遍历它们,但我几乎不理解这部分代码在一年前设置时是如何工作的,从那以后我除了运行之外没有接触过 google colab这个算法每次我尝试用谷歌搜索如何执行此操作或查找谷歌驱动器 API 或其他任何东西时,一切都会返回如何下载和上传几个文件,我需要的只是获取名称列表所有文件。

【问题讨论】:

    标签: google-sheets jupyter-notebook google-drive-api google-colaboratory


    【解决方案1】:

    documentation 解释了如何使用 pageToken 进行分页(该页面用于 Calendar API,但它在 Drive 中的工作原理相同):

    为了检索下一页,执行与之前完全相同的请求,并附加一个 pageToken 字段,其中包含上一页的 nextPageToken 值。在检索到所有结果之前,将在后续页面上提供新的 nextPageToken。

    本质上,您需要一个循环,在该循环中运行files.list(),检索pageToken,然后再次运行它,同时向其提供前一个令牌,直到您停止获取令牌为止。

    对于您的具体情况,您可以尝试用以下内容替换“问题”sn-p:

        page_token = ""
        filelist = {}
        while True:
            response = drive_service.files().list(q=query,
                                        corpora='drive',
                                        supportsAllDrives='true',
                                        includeItemsFromAllDrives='true',
                                        driveId=data_drive_id,
                                        pageSize=1000,
                                        fields='nextPageToken, files(id, name, webViewLink)',
                                        pageToken=page_token).execute()
    
            page_token = response.get('nextPageToken', None)
            filelist.setdefault("files",[]).extend(response.get('files'))
    
            if (not page_token):
                break
    
        response = filelist
    

    正如我所描述的那样,循环 files.list() 并将结果添加到 filelist 变量,然后在 API 停止返回页面标记时中断循环。最后,我只是将 filelist 的值分配给了 response 变量,因为这就是您在其余代码中使用的值。它应该以相同的方式解析,但这次带有完整的结果列表。

    资料来源:

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2020-03-28
      • 2021-08-14
      • 2019-12-16
      • 1970-01-01
      • 2021-02-24
      • 2020-04-08
      • 1970-01-01
      相关资源
      最近更新 更多