【问题标题】:Web scraping with Newspaper3k, got only 50 articles用 Newspaper3k 抓取网页,只得到 50 篇文章
【发布时间】:2020-09-07 18:47:38
【问题描述】:

我想抓取一个法国网站上的数据,结果只有 50 篇文章。这个网站有50多篇文章。我哪里错了?

我的目标是爬取本站所有文章。

我试过了:

import newspaper

legorafi_paper = newspaper.build('http://www.legorafi.fr/', memoize_articles=False)

# Empty list to put all urls
papers = []

for article in legorafi_paper.articles:
    papers.append(article.url)

print(legorafi_paper.size())

这次打印的结果是 50 篇文章。

我不明白为什么newspaper3k 只会抓取 50 篇文章而不是更多。

我尝试过的更新:

def Foo(firstTime = []):
    if firstTime == []:
        WebDriverWait(driver, 30).until(EC.frame_to_be_available_and_switch_to_it((By.CSS_SELECTOR,"div#appconsent>iframe")))
        firstTime.append('Not Empty')
    else:
        print('Cookies already accepted')


%%time


categories = ['societe', 'politique']


import time
from selenium import webdriver
from selenium.common.exceptions import NoSuchElementException
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.action_chains import ActionChains

import newspaper
import requests
from newspaper.utils import BeautifulSoup
from newspaper import Article

categories = ['people', 'sports']
papers = []


driver = webdriver.Chrome(executable_path="/Users/name/Downloads/chromedriver 4")
driver.get('http://www.legorafi.fr/')


for category in categories:
    url = 'http://www.legorafi.fr/category/' + category
    #WebDriverWait(self.driver, 10)
    driver.get(url)
    Foo()
    WebDriverWait(driver, 30).until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button.button--filled>span.baseText"))).click()

    pagesToGet = 2

    title = []
    content = []
    for page in range(1, pagesToGet+1):
        print('Processing page :', page)
        #url = 'http://www.legorafi.fr/category/france/politique/page/'+str(page)
        print(driver.current_url)
        #print(url)

        time.sleep(3)

        raw_html = requests.get(url)
        soup = BeautifulSoup(raw_html.text, 'html.parser')
        for articles_tags in soup.findAll('div', {'class': 'articles'}):
            for article_href in articles_tags.find_all('a', href=True):
                if not str(article_href['href']).endswith('#commentaires'):
                    urls_set.add(article_href['href'])
                    papers.append(article_href['href'])


        for url in papers:
            article = Article(url)
            article.download()
            article.parse()
            if article.title not in title:
                title.append(article.title)
            if article.text not in content:
                content.append(article.text)
            #print(article.title,article.text)

        time.sleep(5)
        driver.execute_script("window.scrollTo(0,document.body.scrollHeight)")
        driver.find_element_by_xpath("//a[contains(text(),'Suivant')]").click()
        time.sleep(10)

【问题讨论】:

  • Stackoverflow 有特殊的方法(和快捷键)来格式化多行代码。
  • 网站可能会在一些计数后阻止您,您可以联系他们的网络管理员,他们可以提供比您可以抓取的更好的集合,如果您正在做一些科学或学习他们可能是免费的可以从中受益!
  • 谢谢@ti7 你知道我能不能用python代码绕过它吗?
  • @LJRB 您可以使用proxy or proxies,这将允许您的程序充当多个独立客户端而不是单个客户端。但是,直接要求更好地访问数据(可能就像没有 50 页限制的帐户一样简单)并在您正在制作的作品中引用它们可能就是他们要求您获得更高质量访问的全部内容(他们知道任何人都可以编写程序来阅读他们的网站,如果你有一个网站,你会看到许多机器人正在积极地阅读你的网站)。

标签: python newspaper3k


【解决方案1】:

2020 年 9 月 21 日更新

我重新检查了您的代码,它工作正常,因为它正在提取Le Gorafi 主页上的所有文章。此页面上的文章是分类页面的亮点,例如社会、体育等。

以下示例来自主页的源代码。这些文章中的每一篇也都列在体育类别页面上。

<div class="cat sports">
    <a href="http://www.legorafi.fr/category/sports/">
       <h4>Sports</h4>
          <ul>
              <li>
                 <a href="http://www.legorafi.fr/2020/07/24/chaque-annee-25-des-lutteurs-doivent-etre-operes-pour-defaire-les-noeuds-avec-leur-bras/" title="Voir l'article 'Chaque année, 25% des lutteurs doivent être opérés pour défaire les nœuds avec leur bras'">
                  Chaque année, 25% des lutteurs doivent être opérés pour défaire les nœuds avec leur bras</a>
              </li>
               <li>
                <a href="http://www.legorafi.fr/2020/07/09/frank-mccourt-lom-nest-pas-a-vendre-sauf-contre-beaucoup-dargent/" title="Voir l'article 'Frank McCourt « L'OM n'est pas à vendre sauf contre beaucoup d'argent »'">
                  Frank McCourt « L'OM n'est pas à vendre sauf contre beaucoup d'argent </a>
              </li>
              <li>
                <a href="http://www.legorafi.fr/2020/06/10/euphorique-un-parieur-appelle-son-fils-betclic/" title="Voir l'article 'Euphorique, un parieur appelle son fils Betclic'">
                  Euphorique, un parieur appelle son fils Betclic                 </a>
              </li>
           </ul>
               <img src="http://www.legorafi.fr/wp-content/uploads/2015/08/rubrique_sport1-300x165.jpg"></a>
        </div>
              </div>

主页上似乎有 35 个独特的文章条目。

import newspaper

legorafi_paper = newspaper.build('http://www.legorafi.fr', memoize_articles=False)

papers = []
urls_set = set()
for article in legorafi_paper.articles:
   # check to see if the article url is not within the urls_set
   if article.url not in urls_set:
     # add the unique article url to the set
     urls_set.add(article.url)
     # remove all links for article commentaires
     if not str(article.url).endswith('#commentaires'):
        papers.append(article.url)

 print(len(papers)) 
 # output
 35

如果我将上面代码中的 URL 更改为:http://www.legorafi.fr/category/sports,它将返回与http://www.legorafi.fr 相同数量的文章。在GitHub上查看Newspaper的源代码后,似乎该模块正在使用urlparse,它似乎正在使用netloc段网址解析netlocwww.legorafi.fr。我注意到这是 Newspaper 的一个已知问题,基于此打开的 issue.

要获取所有文章变得更加复杂,因为您必须使用一些额外的模块,包括 requestsBeautifulSoup。这 后者可以从 Newspaper 中调用。下面的代码可以通过requestsBeautifulSoup精炼得到主页面和分类页面源代码内的所有文章。

import newspaper
import requests
from newspaper.utils import BeautifulSoup

papers = []
urls_set = set()

legorafi_paper = newspaper.build('http://www.legorafi.fr', 
fetch_images=False, memoize_articles=False)
for article in legorafi_paper.articles:
   if article.url not in urls_set:
     urls_set.add(article.url)
     if not str(article.url).endswith('#commentaires'):
       papers.append(article.url)

 
legorafi_urls = {'monde-libre': 'http://www.legorafi.fr/category/monde-libre',
             'politique': 'http://www.legorafi.fr/category/france/politique',
             'societe': 'http://www.legorafi.fr/category/france/societe',
             'economie': 'http://www.legorafi.fr/category/france/economie',
             'culture': 'http://www.legorafi.fr/category/culture',
             'people': 'http://www.legorafi.fr/category/people',
             'sports': 'http://www.legorafi.fr/category/sports',
             'hi-tech': 'http://www.legorafi.fr/category/hi-tech',
             'sciences': 'http://www.legorafi.fr/category/sciences',
             'ledito': 'http://www.legorafi.fr/category/ledito/'
             }


for category, url in legorafi_urls.items():
   raw_html = requests.get(url)
   soup = BeautifulSoup(raw_html.text, 'html.parser')
   for articles_tags in soup.findAll('div', {'class': 'articles'}):
      for article_href in articles_tags.find_all('a', href=True):
         if not str(article_href['href']).endswith('#commentaires'):
           urls_set.add(article_href['href'])
           papers.append(article_href['href'])

   print(len(papers))
   # output
   155

如果您需要获取某个分类页面的子页面中列出的文章(politique 目前有 120 个子页面),那么您必须使用 Selenium 之类的东西来点击链接。

希望此代码可以帮助您更接近实现目标。

【讨论】:

  • 感谢您的帮助 @Life is complex ,我尝试了您的代码,但每个类别都有相同的 53 个网址。我能做什么?
  • 哼...LMK 看看这个。
  • @LJRB 请在您的问题中提供一些详细信息,说明您想从网站legorafi.fr 或其子页面上抓取哪些文章。一旦你这样做了,我就可以解决我的答案。
  • 你好,如果可能的话,我想把所有类别和所有页面上的所有文章都刮掉。
  • @LJRB 看看这些问题是否有用 - stackoverflow.com/search?q=Selenium++newspaper
猜你喜欢
  • 1970-01-01
  • 1970-01-01
  • 1970-01-01
  • 2021-01-22
  • 1970-01-01
  • 1970-01-01
  • 2020-10-11
  • 2021-06-14
  • 2019-04-29
相关资源
最近更新 更多