【问题标题】:Is it possible to use regular expressions with pdfquery?是否可以在 pdfquery 中使用正则表达式?
【发布时间】:2015-10-13 19:57:32
【问题描述】:

我们可以使用正则表达式检测 pdf 中的文本(使用 pdfquery 或其他工具)吗?

我知道我们可以做到:

pdf = pdfquery.PDFQuery("tests/samples/IRS_1040A.pdf")
pdf.load()
label = pdf.pq('LTTextLineHorizontal:contains("Cash")')
left_corner = float(label.attr('x0'))
bottom_corner = float(label.attr('y0'))
cash = pdf.pq('LTTextLineHorizontal:in_bbox("%s, %s, %s, %s")' % \
        (left_corner, bottom_corner-30, \
        left_corner+150, bottom_corner)).text()
print cash
'179,000.00'

但我们需要这样的东西:

pdf = pdfquery.PDFQuery("tests/samples/IRS_1040A.pdf")
pdf.load()
label = pdf.pq('LTTextLineHorizontal:regex("\d{1,3}(?:,\d{3})*(?:\.\d{2})?")')
cash = str(label.attr('x0'))
print cash
'179,000.00'

【问题讨论】:

    标签: python regex pdfminer


    【解决方案1】:

    这并不完全是对正则表达式的查找,但它可以格式化/过滤可能的提取:

    def regex_function(pattern, match):
        re_obj = re.search(pattern, match)
        if re_obj != None and len(re_obj.groups()) > 0:
            return re_obj.group(1)
        return None
    
    pdf = pdfquery.PDFQuery("tests/samples/IRS_1040A.pdf")
    
    pattern = ''
    pdf.extract( [
    ('with_parent','LTPage[pageid=1]'),
    ('with_formatter', 'text'),
    ('year', 'LTTextLineHorizontal:contains("Form 1040A (")', 
            lambda match: regex_function(SOME_PATTERN_HERE, match)))
     ])
    

    我没有测试下一个,但它可能也可以:

    def some_regex_function_feature():
        # here you could use some regex.
        return float(this.get('width',0)) * float(this.get('height',0)) > 40000
    
    pdf.pq('LTPage[page_index="1"] *').filter(regex_function_filter_here)
    [<LTTextBoxHorizontal>, <LTRect>, <LTRect>]
    

    【讨论】:

      猜你喜欢
      • 2021-03-09
      • 2018-01-16
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多