【发布时间】:2014-07-01 22:49:49
【问题描述】:
您好,我一直在使用此代码 sn-p 从网站下载文件,目前小于 1GB 的文件都很好。但我注意到一个 1.5GB 的文件不完整
# s is requests session object
r = s.get(fileUrl, headers=headers, stream=True)
start_time = time.time()
with open(local_filename, 'wb') as f:
count = 1
block_size = 512
try:
total_size = int(r.headers.get('content-length'))
print 'file total size :',total_size
except TypeError:
print 'using dummy length !!!'
total_size = 10000000
for chunk in r.iter_content(chunk_size=block_size):
if chunk: # filter out keep-alive new chunks
duration = time.time() - start_time
progress_size = int(count * block_size)
if duration == 0:
duration = 0.1
speed = int(progress_size / (1024 * duration))
percent = int(count * block_size * 100 / total_size)
sys.stdout.write("\r...%d%%, %d MB, %d KB/s, %d seconds passed" %
(percent, progress_size / (1024 * 1024), speed, duration))
f.write(chunk)
f.flush()
count += 1
使用最新请求 2.2.1 python 2.6.6, centos 6.4 文件下载总是停止在 66.7% 1024MB,我错过了什么? 输出:
file total size : 1581244542
...67%, 1024 MB, 5687 KB/s, 184 seconds passed
iter_content() 返回的生成器似乎认为所有块都已检索并且没有错误。顺便说一句,异常部分没有运行,因为服务器确实在响应头中返回了内容长度。
【问题讨论】:
-
注意“b” = 位,而“B” = 字节(这可能是你的意思)
-
@Jonathon 好的 ... orz,我更新了帖子
-
s.get(...)中的s是什么? -
@Lego
s是请求会话对象...。我从中下载的站点需要身份验证,我省略了这些代码 -
@Shuman,你解决问题了吗?这里也一样....
标签: python web-scraping urllib python-requests