【问题标题】:Error extracting text from website: AttributeError 'NoneType' object has no attribute 'get_text'从网站提取文本时出错:AttributeError 'NoneType' 对象没有属性 'get_text'
【发布时间】:2018-05-12 17:17:09
【问题描述】:

我正在抓取这个网站并使用 .get_text().strip() 获取“标题”和“类别”作为文本。

我在使用相同的方法将“作者”提取为文本时遇到问题。

data2 = {
    'url' : [],
    'title' : [],
    'category': [],
    'author': [],
} 

url_pattern = "https://www.nature.com/nature/articles?searchType=journalSearch&sort=PubDate&year=2018&page={}"
count_min = 1
count_max = 3

while count_min <= count_max: 
    print (count_min)
    url = url_pattern.format(count_min)
    r = requests.get(url)
    try: 
        soup = BeautifulSoup(r.content, 'lxml')
        for links in soup.find_all('article'):
            data2['url'].append(links.a.attrs['href']) 
            data2['title'].append(links.h3.get_text().strip())
            data2["category"].append(links.span.get_text().strip()) 
            data2["author"].append(links.find('span', {"itemprop": "name"}).get_text().strip()) #??????

    except Exception as exc:
        print(exc.__class__.__name__, exc)

    time.sleep(0.1)
    count_min = count_min + 1

print ("Fertig.")
df = pd.DataFrame( data2 )
df

df 应该打印一个带有“author”、“category”、“title”、“url”的表格。打印异常给了我以下提示:AttributeError 'NoneType' object has no attribute 'get_text'。但我收到以下消息,而不是表格。

ValueError                                Traceback (most recent call last)
<ipython-input-34-9bfb92af1135> in <module>()
     29 
     30 print ("Fertig.")
---> 31 df = pd.DataFrame( data2 )
     32 df

~/anaconda3/lib/python3.6/site-packages/pandas/core/frame.py in __init__(self, data, index, columns, dtype, copy)
    328                                  dtype=dtype, copy=copy)
    329         elif isinstance(data, dict):
--> 330             mgr = self._init_dict(data, index, columns, dtype=dtype)
    331         elif isinstance(data, ma.MaskedArray):
    332             import numpy.ma.mrecords as mrecords

~/anaconda3/lib/python3.6/site-packages/pandas/core/frame.py in _init_dict(self, data, index, columns, dtype)
    459             arrays = [data[k] for k in keys]
    460 
--> 461         return _arrays_to_mgr(arrays, data_names, index, columns, dtype=dtype)
    462 
    463     def _init_ndarray(self, values, index, columns, dtype=None, copy=False):

~/anaconda3/lib/python3.6/site-packages/pandas/core/frame.py in _arrays_to_mgr(arrays, arr_names, index, columns, dtype)
   6161     # figure out the index, if necessary
   6162     if index is None:
-> 6163         index = extract_index(arrays)
   6164     else:
   6165         index = _ensure_index(index)

~/anaconda3/lib/python3.6/site-packages/pandas/core/frame.py in extract_index(data)
   6209             lengths = list(set(raw_lengths))
   6210             if len(lengths) > 1:
-> 6211                 raise ValueError('arrays must all be same length')
   6212 
   6213             if have_dicts:

ValueError: arrays must all be same length 

如何改进我的代码以提取“作者”姓名?

【问题讨论】:

  • 我们需要比“它不起作用”更多的细节。实际发生了什么,应该发生什么?如果您收到错误消息,请发布准确、完整的错误消息,包括堆栈跟踪。
  • try: ... except Exception: pass 隐藏了有用的异常信息。

标签: python tags


【解决方案1】:

您非常接近 - 我建议您做几件事。首先,我建议仔细查看 HTML——在这种情况下,作者姓名实际上位于 ul 中,其中每个 li 包含一个 span,其中 itemprop'name'。但是,并非所有文章都具有任何作者姓名。在这种情况下,使用您当前的代码,对links.find('div', {'itemprop': 'name'}) 的调用将返回NoneNone 当然没有get_text 的属性。这意味着该行将引发错误,在这种情况下,这只会导致没有值被附加到 data2 'author' 列表中。我建议将作者存储在这样的列表中:

authors = []
ul = links.find('ul', itemprop='creator')
for author in ul.find_all('span', itemprop='name'):
    authors.append(author.text.strip())
data2['authors'].append(authors)

这处理了没有作者的情况,正如我们所期望的那样,“作者”是一个空列表。

作为旁注,将您的代码放在一个

try:
    ...
except:
    pass

construct 通常被认为是不好的做法,这正是您现在看到的原因。默默地忽略错误可以使您的程序看起来运行正常,而实际上任何数量的事情都可能出错。至少,将错误信息打印到stdout 并不是一个坏主意。即使只是做这样的事情也比什么都不做要好:

try:
    ...
except Exception as exc:
    print(exc.__class__.__name__, exc)

但是,对于调试,通常也需要完整的回溯。为此,您可以使用traceback 模块。

import traceback
try:
    ...
except:
    traceback.print_exc()

【讨论】:

    【解决方案2】:

    而不是使用strip方法。创建一个包含所有项目的变量,然后使用 for 循环并利用 .text

    author = links.findAll('span', {"itemprop": "name"})
    for i in author:
        data2["author"].append(i.text) #??????
    

    打印

    'author': ['Mark Zastrow', 'Barbara Mühlemann', 'Terry C. Jones', 'Peter de Barros Damgaard', 'Morten E. Allentoft', 'Irina Shevnina', 'Andrey Logvin', 'Emma Usmanova', ......

    【讨论】:

      猜你喜欢
      • 2015-04-07
      • 1970-01-01
      • 2017-12-28
      • 1970-01-01
      • 2020-05-25
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2021-07-25
      相关资源
      最近更新 更多