【问题标题】:Python BeautifulSoup - How to extract this textPython BeautifulSoup - 如何提取此文本
【发布时间】:2015-07-18 09:13:19
【问题描述】:

当前 Python 脚本:

import win_unicode_console
win_unicode_console.enable()

import requests
from bs4 import BeautifulSoup

data = '''
<div class="info">
    <h1>Company Title</h1>
    <p class="type">Company type</p>
    <p class="address"><strong>ZIP, City</strong></p>
    <p class="address"><strong>Street 123</strong></p>
    <p style="margin-top:10px;"> Phone: <strong>(111) 123-456-78</strong><br />
        Fax: <strong>(222) 321-654-87</strong><br />
        Phone: <strong>(333) 87-654-321</strong><br />
        Fax: <strong>(444) 000-1111-2222</strong><br />
    </p>
    <p style="margin-top:10px;"> E-mail: <a href="mailto:mail@domain.com">mail@domain.com</a><br />
    E-mail: <a href="mailto:mail2@domain.com">mail2@domain.com</a><br />
    </p>
    <p> Web: <a href="http://www.domain.com" target="_blank">www.domain.com</a><br />
    </p>
    <p style="margin-top:10px;"> ID: <strong>123456789</strong><br />
        VAT: <strong>987654321</strong> </p>
    <p class="del" style="margin-top:10px;">Some info:</p>
    <ul>
        <li><a href="#category">&raquo; Category</a></li>
    </ul>
</div>
'''

html = BeautifulSoup(data, "html.parser")

p = html.find_all('p', attrs={'class': None})

for pp in p:
    print(pp.contents)

它返回以下内容:

[' Phone: ', <strong>123-456-78</strong>, <br/>, '\n\t\tFax: ', <strong>321-654-87</strong>, <br/>, '\n\t\tPhone: ', <strong>87-654-321</strong>, <br/>, '\n\t\tFax: ', <strong>000-1111-2222</strong>, <br/>, '\n']
[' E-mail: ', <a href="mailto:mail@domain.com">mail@domain.com</a>, <br/>, '\n\tE-mail: ', <a href="mailto:mail2@domain.com">mail2@domain.com</a>, <br/>, '\n']
[' Web: ', <a href="http://www.domain.com" target="_blank">www.domain.com</a>, <br/>, '\n']
[' ID: ', <strong>123456789</strong>, <br/>, '\n\t\tVAT: ', <strong>987654321</strong>, ' ']

问题: 我不知道如何提取电话、传真和电子邮件、id、增值税的文本并从中创建数组,例如:

phones = [123-456-78, 87-654-321]
faxes = [321-654-87, 000-1111-2222]
emails = [mail@domain.com, mail2@domain.com]
id = [123456789]
vat = [987654321]

【问题讨论】:

    标签: python python-3.x beautifulsoup extraction data-extraction


    【解决方案1】:

    您可以在拆分后使用 defaultdict 对数据进行分组:

    html = BeautifulSoup(data, "html.parser")
    
    p = html.find_all('p', attrs={'class': None})
    from collections import defaultdict
    
    d = defaultdict(list)
    for pp in p:
        spl = iter(pp.text.split(None,1))
        for ele in spl:
            d[ele.rstrip(":")].append(next(spl).rstrip())
    
    print(d)
    defaultdict(<class 'list'>, {'Phone': ['123-456-78', '87-654-321'],
    'Fax': ['321-654-87', '000-1111-2222'], 'E-mail': ['mail@domain.com',
    'mail2@domain.com'], 'VAT': ['987654321'], 'Web': ['www.domain.com'], 
    'ID': ['123456789']})
    

    拆分文本为您提供数据列表:

    ['Phone:', '123-456-78', 'Fax:', '321-654-87', 'Phone:', '87-654-321', 'Fax:', '000-1111-2222']
    ['E-mail:', 'mail@domain.com', 'E-mail:', 'mail2@domain.com']
    ['Web:', 'www.domain.com']
    ['ID:', '123456789', 'VAT:', '987654321']
    

    所以我们使用每两个元素作为键/值对。附加重复键。

    为了您在传真和电话号码中捕捉空格的编辑,只需将其拆分为带有拆分线的行并在空白处拆分一次: 从集合导入默认字典

    d = defaultdict(list)
    for pp in p:
        spl = pp.text.splitlines()
        for ele in spl:
            k, v = ele.strip().split(None, 1)
            d[k.rstrip(":")].append(v.rstrip())
    

    输出:

    defaultdict(<class 'list'>, {'Fax': ['(222) 321-654-87', '(444) 000-1111-2222'],
     'Web': ['www.domain.com'], 'ID': ['123456789'], 'E-mail': ['mail@domain.com', 'mail2@domain.com'],
     'VAT': ['987654321'], 'Phone': ['(111) 123-456-78', '(333) 87-654-321']})
    

    【讨论】:

    • 抱歉,当电话号码类似于(111) 222-333-4444时出现错误
    • 感谢您的更新!我现在遇到另一个问题。电话包含:'Phone': ['(111) 123-456-78\n\t\tFax: (222) 321-654-87\n\t\tPhone: (333) 87-654-321\n\t\tFax: (444) 000-1111-2222'] 我的意思是它没有拆分为PhoneFax
    • 我解决了这个问题:for pp in p: spl = iter(pp.text.splitlines()) for ele in spl: for spl2 in ele.splitlines(): spl3 = iter(spl2.split(None, 1)) print(spl3) for ele2 in spl3: d[ele2.rstrip(":")].append(next(spl3).rstrip()) 它并不优雅,但它可以工作:)
    • @RhymeGuy,编辑应该用更少的代码完成你所需要的;)
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2016-12-27
    • 1970-01-01
    • 2020-07-25
    • 2014-05-22
    相关资源
    最近更新 更多