【发布时间】:2013-04-29 14:50:38
【问题描述】:
我用 Python 写了一个简单的脚本。
它解析网页中的超链接,然后检索这些链接以解析一些信息。
我有类似的脚本在运行并重新使用 writefunction 没有任何问题,但由于某种原因它失败了,我不知道为什么。
一般卷曲初始化:
storage = StringIO.StringIO()
c = pycurl.Curl()
c.setopt(pycurl.USERAGENT, USER_AGENT)
c.setopt(pycurl.COOKIEFILE, "")
c.setopt(pycurl.POST, 0)
c.setopt(pycurl.FOLLOWLOCATION, 1)
#Similar scripts are working this way, why this script not?
c.setopt(c.WRITEFUNCTION, storage.write)
第一次调用检索链接:
URL = "http://whatever"
REFERER = URL
c.setopt(pycurl.URL, URL)
c.setopt(pycurl.REFERER, REFERER)
c.perform()
#Write page to file
content = storage.getvalue()
f = open("updates.html", "w")
f.writelines(content)
f.close()
... Here the magic happens and links are extracted ...
现在循环这些链接:
for i, member in enumerate(urls):
URL = urls[i]
print "url:", URL
c.setopt(pycurl.URL, URL)
c.perform()
#Write page to file
#Still the data from previous!
content = storage.getvalue()
f = open("update.html", "w")
f.writelines(content)
f.close()
#print content
... Gather some information ...
... Close objects etc ...
【问题讨论】:
-
您可以在循环中尝试
c.setopt(c.WRITEFUNCTION, f.write)以避免将数据附加到同一个对象。如果Curl()是可重用的,这可能就足够了。 -
不,这不起作用,我以前尝试过,我认为它只是传递一个引用。第一页的字符串长度是否可能太大(网页相当大,与我用 Curl 和 Python 检索的其他内容相比。)
标签: python curl pycurl stringio