【问题标题】:html5lib: TypeError: __init__() got an unexpected keyword argument 'encoding'html5lib: TypeError: __init__() 得到了一个意外的关键字参数“编码”
【发布时间】:2016-12-29 09:40:26
【问题描述】:

我正在尝试安装html5lib。起初我尝试安装最新版本(8 或 9 个 9),但它与我的 BeautifulSoup 冲突,所以我决定尝试旧版本(0.9999999,seven nines)。我安装了它,但是当我尝试使用它时:

>>> with urlopen("http://example.com/") as f:
    document = html5lib.parse(f, encoding=f.info().get_content_charset())

我收到一个错误:

Traceback (most recent call last):
  File "<pyshell#11>", line 2, in <module>
    document = html5lib.parse(f, encoding=f.info().get_content_charset())
  File "C:\Python\Python35-32\lib\site-packages\html5lib\html5parser.py", line 35, in parse
    return p.parse(doc, **kwargs)
  File "C:\Python\Python35-32\lib\site-packages\html5lib\html5parser.py", line 235, in parse
    self._parse(stream, False, None, *args, **kwargs)
  File "C:\Python\Python35-32\lib\site-packages\html5lib\html5parser.py", line 85, in _parse
    self.tokenizer = _tokenizer.HTMLTokenizer(stream, parser=self, **kwargs)
  File "C:\Python\Python35-32\lib\site-packages\html5lib\_tokenizer.py", line 36, in __init__
    self.stream = HTMLInputStream(stream, **kwargs)
  File "C:\Python\Python35-32\lib\site-packages\html5lib\_inputstream.py", line 151, in HTMLInputStream
    return HTMLBinaryInputStream(source, **kwargs)
TypeError: __init__() got an unexpected keyword argument 'encoding'

出了什么问题,我该怎么办?

【问题讨论】:

    标签: web-scraping beautifulsoup html5lib


    【解决方案1】:

    我看到关于 bs4 的最新版本的 html5lib 中出现了问题,html5lib.treebuilders._base 不再存在,usng bs4 4.4.1 似乎是最新的兼容版本有 7 个 9,一旦按如下方式安装它就可以正常工作:

     pip3 install -U html5lib=="0.9999999"
    

    使用 bs4 4.4.1 测试:

    In [1]: import bs4
    
    In [2]: bs4.__version__
    Out[2]: '4.4.1'
    
    In [3]: import html5lib
    
    In [4]: html5lib.__version__
    Out[4]: '0.9999999'
    
    In [5]: from urllib.request import  urlopen
    
    In [6]: with urlopen("http://example.com/") as f:
       ...:         document = html5lib.parse(f, encoding=f.info().get_content_charset())
       ...:     
    
    In [7]: 
    

    在这个commit中可以看到Rename treebuilders._base to .base to reflect public status改名的变化:

    你看到的错误是因为你还在使用最新版本,在html5lib/_inputstream.py中,HTMLBinaryInputStream没有编码参数:

    class HTMLBinaryInputStream(HTMLUnicodeInputStream):
        """Provides a unicode stream of characters to the HTMLTokenizer.
    
        This class takes care of character encoding and removing or replacing
        incorrect byte-sequences and also provides column and line tracking.
    
        """
    
        def __init__(self, source, override_encoding=None, transport_encoding=None,
                     same_origin_parent_encoding=None, likely_encoding=None,
                     default_encoding="windows-1252", useChardet=True):
    

    设置 override_encoding=f.info().get_content_charset() 应该可以解决问题。

    升级到最新版本的 bs4 也可以使用最新版本的 html5lib:

    In [16]: bs4.__version__
    Out[16]: '4.5.1'
    
    In [17]: html5lib.__version__
    Out[17]: '0.999999999'
    
    In [18]: with urlopen("http://example.com/") as f:
                 document = html5lib.parse(f, override_encoding=f.info().get_content_charset())
       ....:     
    
    In [19]: 
    

    【讨论】:

      猜你喜欢
      • 2018-06-08
      • 2016-04-18
      • 2012-12-11
      • 1970-01-01
      • 2017-10-16
      • 2020-10-22
      • 2021-11-17
      • 2016-11-01
      • 2021-10-21
      相关资源
      最近更新 更多