【问题标题】:Extracting image from PDF with /CCITTFaxDecode filter使用 /CCITTFaxDecode 过滤器从 PDF 中提取图像
【发布时间】:2011-02-08 03:57:15
【问题描述】:

我有一个由扫描软件生成的 pdf。 pdf 每页有 1 个 TIFF 图像。我想从每一页中提取 TIFF 图像。

我正在使用 iTextSharp,并且我已成功找到图像并且可以从 PdfReader.GetStreamBytesRaw 方法中取回原始字节。问题是,正如我之前的许多人所发现的那样,iTextSharp 不包含 PdfReader.CCITTFaxDecode 方法。

我还知道什么?即使没有 iTextSharp,我也可以在记事本中打开 pdf 并找到带有 /Filter /CCITTFaxDecode 的流,我从 /DecodeParams 知道它正在使用 CCITTFaxDecode 组 4。

有人知道如何从我的 pdf 中获取 CCITTFaxDecode 过滤器图像吗?

干杯, 卡胡

【问题讨论】:

  • 如果有人对单独的 iTextSharp 解决方案感兴趣,还有另一个链接 (kuujinbo.info/iTextSharp/CCITTFaxDecodeExtract.aspx) 建议使用 iTextSharp 5xx 中引入的称为 Parser 的新功能 - 它确实有效。归功于@kuunjinbo,只是在我的情况下,我不得不在结果上使用 ImageConverter 来创建一个位图(不知道为什么)
  • 如果您从解析器中跟踪代码,它最终会调用 TIFFFaxDecompressor 类......这似乎非常有缺陷。它忽略了它给出的一些标志,例如,当 Height=0 时,它甚至不会尝试实现

标签: image pdf itextsharp extract


【解决方案1】:

已经为此写了extension (c#)。

PdfDictionary item;
if (item.IsImage()) {
  Image image = item.ToImage();
}

【讨论】:

  • 实际上这个扩展运行良好且正确。这里已经做了很多工作。所以我建议使用这个而不是自己编写代码。
【解决方案2】:

这里是python实现:

import PyPDF2
import struct

"""
Links:
PDF format: http://www.adobe.com/content/dam/Adobe/en/devnet/acrobat/pdfs/pdf_reference_1-7.pdf
CCITT Group 4: https://www.itu.int/rec/dologin_pub.asp?lang=e&id=T-REC-T.6-198811-I!!PDF-E&type=items
Extract images from pdf: http://stackoverflow.com/questions/2693820/extract-images-from-pdf-without-resampling-in-python
Extract images coded with CCITTFaxDecode in .net: http://stackoverflow.com/questions/2641770/extracting-image-from-pdf-with-ccittfaxdecode-filter
TIFF format and tags: http://www.awaresystems.be/imaging/tiff/faq.html
"""


def tiff_header_for_CCITT(width, height, img_size, CCITT_group=4):
    tiff_header_struct = '<' + '2s' + 'h' + 'l' + 'h' + 'hhll' * 8 + 'h'
    return struct.pack(tiff_header_struct,
                       b'II',  # Byte order indication: Little indian
                       42,  # Version number (always 42)
                       8,  # Offset to first IFD
                       8,  # Number of tags in IFD
                       256, 4, 1, width,  # ImageWidth, LONG, 1, width
                       257, 4, 1, height,  # ImageLength, LONG, 1, lenght
                       258, 3, 1, 1,  # BitsPerSample, SHORT, 1, 1
                       259, 3, 1, CCITT_group,  # Compression, SHORT, 1, 4 = CCITT Group 4 fax encoding
                       262, 3, 1, 0,  # Threshholding, SHORT, 1, 0 = WhiteIsZero
                       273, 4, 1, struct.calcsize(tiff_header_struct),  # StripOffsets, LONG, 1, len of header
                       278, 4, 1, height,  # RowsPerStrip, LONG, 1, lenght
                       279, 4, 1, img_size,  # StripByteCounts, LONG, 1, size of image
                       0  # last IFD
                       )

pdf_filename = 'scan.pdf'
pdf_file = open(pdf_filename, 'rb')
cond_scan_reader = PyPDF2.PdfFileReader(pdf_file)
for i in range(0, cond_scan_reader.getNumPages()):
    page = cond_scan_reader.getPage(i)
    xObject = page['/Resources']['/XObject'].getObject()
    for obj in xObject:
        if xObject[obj]['/Subtype'] == '/Image':
            """
            The  CCITTFaxDecode filter decodes image data that has been encoded using
            either Group 3 or Group 4 CCITT facsimile (fax) encoding. CCITT encoding is
            designed to achieve efficient compression of monochrome (1 bit per pixel) image
            data at relatively low resolutions, and so is useful only for bitmap image data, not
            for color images, grayscale images, or general data.

            K < 0 --- Pure two-dimensional encoding (Group 4)
            K = 0 --- Pure one-dimensional encoding (Group 3, 1-D)
            K > 0 --- Mixed one- and two-dimensional encoding (Group 3, 2-D)
            """
            if xObject[obj]['/Filter'] == '/CCITTFaxDecode':
                if xObject[obj]['/DecodeParms']['/K'] == -1:
                    CCITT_group = 4
                else:
                    CCITT_group = 3
                width = xObject[obj]['/Width']
                height = xObject[obj]['/Height']
                data = xObject[obj]._data  # sorry, getData() does not work for CCITTFaxDecode
                img_size = len(data)
                tiff_header = tiff_header_for_CCITT(width, height, img_size, CCITT_group)
                img_name = obj[1:] + '.tiff'
                with open(img_name, 'wb') as img_file:
                    img_file.write(tiff_header + data)
                #
                # import io
                # from PIL import Image
                # im = Image.open(io.BytesIO(tiff_header + data))
pdf_file.close()

【讨论】:

  • 根据标签和问题正文,操作使用iTextSharp。因此,python 实现并不能回答这个问题。
  • 根据 TIFF 规范 (link),我认为您的变量 tiff_header_struct 应为 '&lt;' + '2s' + 'H' + 'L' + 'H' + 'HHLL' * 8 + 'L'。特别注意末尾的'L'
【解决方案3】:

实际上,vbcrlfuser 的回答确实对我有所帮助,但是对于当前版本的 BitMiracle.LibTiff.NET,代码并不完全正确,因为我可以下载它。在当前版本中,等效代码如下所示:

using iTextSharp.text.pdf;
using BitMiracle.LibTiff.Classic;

...
      Tiff tiff = Tiff.Open("C:\\test.tif", "w");
      tiff.SetField(TiffTag.IMAGEWIDTH, UInt32.Parse(pd.Get(PdfName.WIDTH).ToString()));
      tiff.SetField(TiffTag.IMAGELENGTH, UInt32.Parse(pd.Get(PdfName.HEIGHT).ToString()));
      tiff.SetField(TiffTag.COMPRESSION, Compression.CCITTFAX4);
      tiff.SetField(TiffTag.BITSPERSAMPLE, UInt32.Parse(pd.Get(PdfName.BITSPERCOMPONENT).ToString()));
      tiff.SetField(TiffTag.SAMPLESPERPIXEL, 1);
      tiff.WriteRawStrip(0, raw, raw.Length);
      tiff.Close();

使用上面的代码,我终于在 C:\test.tif 中得到了一个有效的 Tiff 文件。谢谢你,vbcrlfuser!

【讨论】:

  • raw 是什么?你能展示完整的例子吗?
  • B.K.这可能有点晚了,但我猜是这样。见 vbcrlfuser 的回答... byte[] data = PdfReader.GetStreamBytesRaw((PRStream)pdfStream); (使用方式与 raw 相同)
【解决方案4】:

这个库...http://www.bitmiracle.com/libtiff/ 和下面的这个例子应该能让你完成 99% 的路

string filter = pd.Get(PdfName.FILTER).ToString();
string width = pd.Get(PdfName.WIDTH).ToString();
string height = pd.Get(PdfName.HEIGHT).ToString();
string bpp = pd.Get(PdfName.BITSPERCOMPONENT).ToString();

switch (filter)
{
   case "/CCITTFaxDecode":

      byte[] data = PdfReader.GetStreamBytesRaw((PRStream)pdfStream);
      int tiff = TIFFOpen("example.tif", "w");
      TIFFSetField(tiff, (uint)BitMiracle.LibTiff.Classic.TiffTag.IMAGEWIDTH,(uint)Int32.Parse(width));
      TIFFSetField(tiff, (uint)BitMiracle.LibTiff.Classic.TiffTag.IMAGEHEIGHT, (uint)Int32.Parse(height));
      TIFFSetField(tiff, (uint)BitMiarcle.LibTiff.Classic.TiffTag.COMPRESSION, (uint)BitMiracle.Libtiff.Classic.Compression.CCITTFAX4);
      TIFFSetField(tiff, (uint)BitMiracle.LibTiff.Classic.TiffTag.BITSPERSAMPLE, (uint)Int32.Parse(bpp));
      TIFFSetField(tiff, (uint)BitMiarcle.Libtiff.Classic.TiffTag.SAMPLESPERPIXEL,1 );

      IntPtr pointer = Marshal.AllocHGlobal(data.length);
      Marshal.copy(data, 0, pointer, data.length);
      TIFFWriteRawStrip(tiff, 0, pointer, data.length);
      TIFFClose(tiff);

      break;




      break;

}

【讨论】:

猜你喜欢
  • 2019-08-02
  • 1970-01-01
  • 2012-02-01
  • 2013-12-22
  • 2015-07-01
  • 2011-08-22
  • 2019-10-15
  • 1970-01-01
相关资源
最近更新 更多