【问题标题】:Dealing with run inconsistencies with web scraping使用网络抓取处理运行不一致
【发布时间】:2021-11-20 22:22:36
【问题描述】:

我正在从 vanguard 网站抓取有关共同基金的数据,而我的代码在两次运行之间让我的数据出现了一些不一致的情况。我怎样才能使我的抓取代码更加健壮以避免这些?

我正在从这个page 中抓取数据并尝试在特征表中获取平均持续时间。

有时所有的代码都会顺利通过,而有时它会错过页面上的一些数据。我认为这与数据完全加载之前发生的抓取有关,但它只是有时会发生。

这是 2 次背靠背运行的输出,显示它成功抓取了一个代码,然后在下一次运行中丢失了数据。

VBIRX
Fund total net assets $74.7 billion
Number of bonds 2654
Average effective maturity 2.9 years
Average duration 2.8 years
Yield to maturity 0.5%
VSGBX
Fund total net assets $8.5 billion
Number of bonds 195
Average effective maturity 3.4 years
Average duration 1.7 years
Yield to maturity 0.4%
VFSTX # Here the data for VFSTX is successfully scraped
Fund total net assets $79.3 billion
Number of bonds 2519
Average effective maturity 2.8 years
Average duration 2.7 years
Yield to maturity 1.0%
VFISX
Fund total net assets $7.8 billion
Number of bonds 75
 2.2 years
 2.2 years
Yield to maturity 0.3%

# here data is missing for VFISX, Second run:

VBIRX
Fund total net assets $74.7 billion
Number of bonds 2654
Average effective maturity 2.9 years
Average duration 2.8 years
Yield to maturity 0.5%
VSGBX
Fund total net assets $8.5 billion
Number of bonds 195
Average effective maturity 3.4 years
Average duration 1.7 years
Yield to maturity 0.4%
VFSTX
Fund total net assets $79.3 billion
Number of bonds 2519
 2.8 years
 2.7 years
Yield to maturity 1.0%
# Here data is missing for VFSTX even though it worked in the previous run

主要问题是对于某些股票代码,表格的长度不同,所以我使用字典来存储数据,使用相关标签作为键。对于某些运行,“平均有效成熟度”和“平均持续时间”标签丢失,搞砸了我访问数据的方式。

正如您从我的输出中看到的那样,代码有时会起作用,我不确定选择等待页面上加载不同的元素是否会解决它。我应该如何确定我的问题?

这是我正在使用的相关代码:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from selenium.common.exceptions import TimeoutException
import os
import csv


def extractOverviewTable(htmlTable):
    table = htmlTable.find('tbody')
    rows = table.findAll('tr')
    returnDict = {}
    for row in rows:
        cols = row.findAll('td')
        key = cols[0].find('span').text.replace('\n', '')
        value = cols[1].text.replace('\n', '')
        if 'Layer' in key:
            key = key[:key.index('Layer')]
        print(key, value)    
        returnDict[key] = value
        
    return returnDict
    


def main():

    dirname = os.path.dirname(__file__)
    symbols = []
    with open(os.path.join(dirname, 'symbols.csv')) as csvfile:
        reader = csv.reader(csvfile)
        for row in reader:
            if row:
                symbols.append(row[0])
    symbols = [s.strip() for s in symbols if s.startswith('V')]    
    
    options = webdriver.ChromeOptions()
    options.page_load_strategy = 'normal'
    options.add_argument('--headless')
    browser = webdriver.Chrome(options=options, executable_path=os.path.join(dirname, 'chromedriver'))
    url_vanguard = 'https://investor.vanguard.com/mutual-funds/profile/overview/{}'
    
    for symbol in symbols:   
        browser.get(url_vanguard.format(symbol))
        print(symbol)
        WebDriverWait(browser, 20).until(EC.presence_of_element_located((By.XPATH,'/html/body/div[1]/div[3]/div[3]/div[1]/div/div[1]/div/div/div/div[2]/div/div[2]/div[4]/div[2]/div[2]/div[1]/div/table/tbody/tr[4]')))
        html = browser.page_source
        mySoup = BeautifulSoup(html, 'html.parser')
        htmlData = mySoup.findAll('table',{'role':'presentation'})
        overviewDataList = extractOverviewTable(htmlData[2])

这是我正在使用的 symbols.csv 文件的一个子集:

VBIRX
VSGBX
VFSTX
VFISX
VMLTX
VWSTX
VFIIX
VWEHX
VBILX
VFICX
VFITX

【问题讨论】:

  • 请编辑问题以将其限制为具有足够详细信息的特定问题,以确定适当的答案。

标签: python selenium-webdriver web-scraping beautifulsoup


【解决方案1】:

尝试EC.visibility_of_element_located 而不是EC.presence_of_element_located,如果这不起作用,请尝试在WebDriverWait 语句后添加time.sleep() 1-2 秒。

【讨论】:

    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2021-06-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多