【问题标题】:Reading a IDX file type in Java在 Java 中读取 IDX 文件类型
【发布时间】:2013-06-24 15:19:00
【问题描述】:

我已经用 Java 构建了一个图像分类器,我想针对此处提供的图像进行测试:http://yann.lecun.com/exdb/mnist/

很遗憾,如果您下载 train-images-idx3-ubyte.gz 或其他 3 个文件中的任何一个,它们都是文件类型:.idx1-ubyte

第一个问题: 我想知道是否有人可以指导我如何将 .idx1-ubyte 制作为位图 (.bmp) 文件?

第二个问题: 或者我一般如何阅读这些文件?

有关 IDX 文件格式的信息: IDX 文件格式是各种数值类型的向量和多维矩阵的简单格式。 基本格式是:

magic number 
size in dimension 0 
size in dimension 1 
size in dimension 2 
..... 
size in dimension N 
data

幻数是一个整数(MSB 在前)。前 2 个字节始终为 0。

第三个字节编码数据的类型:

0x08: unsigned byte 
0x09: signed byte 
0x0B: short (2 bytes) 
0x0C: int (4 bytes) 
0x0D: float (4 bytes) 
0x0E: double (8 bytes)

第 4 字节编码向量/矩阵的维数:1 表示向量,2 表示矩阵....

每个维度的大小都是 4 字节整数(MSB 优先,高端,就像在大多数非英特尔处理器中一样)。

数据像 C 数组一样存储,即最后一维的索引变化最快。

【问题讨论】:

  • 我不敢相信“直截了当”这个词在原始字节和高端编码的上下文中出现了两次。不想成为史蒂夫·沃兹尼亚克,布拉 - 只是想要我的数据。说真的,知道他们为什么把它弄得这么复杂吗?
  • 这个善良的灵魂 (Joseph Redmon) 在他的网站上提供了 MNIST 数据的 csv 下载:pjreddie.com/projects/mnist-in-csv

标签: java file-io bitmap


【解决方案1】:

非常简单,正如 WPrecht 所说:“URL 描述了您必须解码的格式”。这是我的 idx 文件的 ImageSet 导出器,不是很干净,但可以完成它必须做的事情。

public class IdxReader {

    public static void main(String[] args) {
        // TODO Auto-generated method stub
        FileInputStream inImage = null;
        FileInputStream inLabel = null;

        String inputImagePath = "CBIR_Project/imagesRaw/MNIST/train-images-idx3-ubyte";
        String inputLabelPath = "CBIR_Project/imagesRaw/MNIST/train-labels-idx1-ubyte";

        String outputPath = "CBIR_Project/images/MNIST_Database_ARGB/";

        int[] hashMap = new int[10]; 

        try {
            inImage = new FileInputStream(inputImagePath);
            inLabel = new FileInputStream(inputLabelPath);

            int magicNumberImages = (inImage.read() << 24) | (inImage.read() << 16) | (inImage.read() << 8) | (inImage.read());
            int numberOfImages = (inImage.read() << 24) | (inImage.read() << 16) | (inImage.read() << 8) | (inImage.read());
            int numberOfRows  = (inImage.read() << 24) | (inImage.read() << 16) | (inImage.read() << 8) | (inImage.read());
            int numberOfColumns = (inImage.read() << 24) | (inImage.read() << 16) | (inImage.read() << 8) | (inImage.read());

            int magicNumberLabels = (inLabel.read() << 24) | (inLabel.read() << 16) | (inLabel.read() << 8) | (inLabel.read());
            int numberOfLabels = (inLabel.read() << 24) | (inLabel.read() << 16) | (inLabel.read() << 8) | (inLabel.read());

            BufferedImage image = new BufferedImage(numberOfColumns, numberOfRows, BufferedImage.TYPE_INT_ARGB);
            int numberOfPixels = numberOfRows * numberOfColumns;
            int[] imgPixels = new int[numberOfPixels];

            for(int i = 0; i < numberOfImages; i++) {

                if(i % 100 == 0) {System.out.println("Number of images extracted: " + i);}

                for(int p = 0; p < numberOfPixels; p++) {
                    int gray = 255 - inImage.read();
                    imgPixels[p] = 0xFF000000 | (gray<<16) | (gray<<8) | gray;
                }

                image.setRGB(0, 0, numberOfColumns, numberOfRows, imgPixels, 0, numberOfColumns);

                int label = inLabel.read();

                hashMap[label]++;
                File outputfile = new File(outputPath + label + "_0" + hashMap[label] + ".png");

                ImageIO.write(image, "png", outputfile);
            }

        } catch (FileNotFoundException e) {
            // TODO Auto-generated catch block
            e.printStackTrace();
        } catch (IOException e) {
            // TODO Auto-generated catch block
            e.printStackTrace();
        } finally {
            if (inImage != null) {
                try {
                    inImage.close();
                } catch (IOException e) {
                    // TODO Auto-generated catch block
                    e.printStackTrace();
                }
            }
            if (inLabel != null) {
                try {
                    inLabel.close();
                } catch (IOException e) {
                    // TODO Auto-generated catch block
                    e.printStackTrace();
                }
            }
        }
    }

}

【讨论】:

    【解决方案2】:

    我创建了一些类用于使用 Java 读取 the MNIST handwritten digits data set。这些类可以在从下载站点上提供的文件中解压缩(解压缩)后读取这些文件。允许读取原始(压缩)文件的类是小型MnistReader 项目的一部分。

    以下这些类是独立的(意味着它们不依赖于第三方库)并且本质上属于公共领域——意味着它们可以被复制到自己的项目中。 (署名将不胜感激,但不是必需的):

    MnistDecompressedReader 类:

    import java.io.DataInputStream;
    import java.io.FileInputStream;
    import java.io.IOException;
    import java.io.InputStream;
    import java.nio.file.Path;
    import java.util.Objects;
    import java.util.function.Consumer;
    
    /**
     * A class for reading the MNIST data set from the <b>decompressed</b> 
     * (unzipped) files that are published at
     * <a href="http://yann.lecun.com/exdb/mnist/">
     * http://yann.lecun.com/exdb/mnist/</a>. 
     */
    public class MnistDecompressedReader
    {
        /**
         * Default constructor
         */
        public MnistDecompressedReader()
        {
            // Default constructor
        }
    
        /**
         * Read the MNIST training data from the given directory. The data is 
         * assumed to be located in files with their default names,
         * <b>decompressed</b> from the original files: 
         * extension) : 
         * <code>train-images.idx3-ubyte</code> and
         * <code>train-labels.idx1-ubyte</code>.
         * 
         * @param inputDirectoryPath The input directory
         * @param consumer The consumer that will receive the resulting 
         * {@link MnistEntry} instances
         * @throws IOException If an IO error occurs
         */
        public void readDecompressedTraining(Path inputDirectoryPath, 
            Consumer<? super MnistEntry> consumer) throws IOException
        {
            String trainImagesFileName = "train-images.idx3-ubyte";
            String trainLabelsFileName = "train-labels.idx1-ubyte";
            Path imagesFilePath = inputDirectoryPath.resolve(trainImagesFileName);
            Path labelsFilePath = inputDirectoryPath.resolve(trainLabelsFileName);
            readDecompressed(imagesFilePath, labelsFilePath, consumer);
        }
    
        /**
         * Read the MNIST training data from the given directory. The data is 
         * assumed to be located in files with their default names,
         * <b>decompressed</b> from the original files: 
         * extension) : 
         * <code>t10k-images.idx3-ubyte</code> and
         * <code>t10k-labels.idx1-ubyte</code>.
         * 
         * @param inputDirectoryPath The input directory
         * @param consumer The consumer that will receive the resulting 
         * {@link MnistEntry} instances
         * @throws IOException If an IO error occurs
         */
        public void readDecompressedTesting(Path inputDirectoryPath, 
            Consumer<? super MnistEntry> consumer) throws IOException
        {
            String testImagesFileName = "t10k-images.idx3-ubyte";
            String testLabelsFileName = "t10k-labels.idx1-ubyte";
            Path imagesFilePath = inputDirectoryPath.resolve(testImagesFileName);
            Path labelsFilePath = inputDirectoryPath.resolve(testLabelsFileName);
            readDecompressed(imagesFilePath, labelsFilePath, consumer);
        }
    
    
        /**
         * Read the MNIST data from the specified (decompressed) files.
         * 
         * @param imagesFilePath The path of the images file
         * @param labelsFilePath The path of the labels file
         * @param consumer The consumer that will receive the resulting 
         * {@link MnistEntry} instances
         * @throws IOException If an IO error occurs
         */
        public void readDecompressed(Path imagesFilePath, Path labelsFilePath, 
            Consumer<? super MnistEntry> consumer) throws IOException
        {
            try (InputStream decompressedImagesInputStream = 
                new FileInputStream(imagesFilePath.toFile());
                InputStream decompressedLabelsInputStream = 
                    new FileInputStream(labelsFilePath.toFile()))
            {
                readDecompressed(
                    decompressedImagesInputStream, 
                    decompressedLabelsInputStream, 
                    consumer);
            }
        }
    
        /**
         * Read the MNIST data from the given (decompressed) input streams.
         * The caller is responsible for closing the given streams.
         * 
         * @param decompressedImagesInputStream The decompressed input stream
         * containing the image data 
         * @param decompressedLabelsInputStream The decompressed input stream
         * containing the label data
         * @param consumer The consumer that will receive the resulting 
         * {@link MnistEntry} instances
         * @throws IOException If an IO error occurs
         */
        public void readDecompressed(
            InputStream decompressedImagesInputStream, 
            InputStream decompressedLabelsInputStream, 
            Consumer<? super MnistEntry> consumer) throws IOException
        {
            Objects.requireNonNull(consumer, "The consumer may not be null");
    
            DataInputStream imagesDataInputStream = 
                new DataInputStream(decompressedImagesInputStream);
            DataInputStream labelsDataInputStream = 
                new DataInputStream(decompressedLabelsInputStream);
    
            int magicImages = imagesDataInputStream.readInt();
            if (magicImages != 0x803)
            {
                throw new IOException("Expected magic header of 0x803 "
                    + "for images, but found " + magicImages);
            }
    
            int magicLabels = labelsDataInputStream.readInt();
            if (magicLabels != 0x801)
            {
                throw new IOException("Expected magic header of 0x801 "
                    + "for labels, but found " + magicLabels);
            }
    
            int numberOfImages = imagesDataInputStream.readInt();
            int numberOfLabels = labelsDataInputStream.readInt();
    
            if (numberOfImages != numberOfLabels)
            {
                throw new IOException("Found " + numberOfImages 
                    + " images but " + numberOfLabels + " labels");
            }
    
            int numRows = imagesDataInputStream.readInt();
            int numCols = imagesDataInputStream.readInt();
    
            for (int n = 0; n < numberOfImages; n++)
            {
                byte label = labelsDataInputStream.readByte();
                byte imageData[] = new byte[numRows * numCols];
                read(imagesDataInputStream, imageData);
    
                MnistEntry mnistEntry = new MnistEntry(
                    n, label, numRows, numCols, imageData);
                consumer.accept(mnistEntry);
            }
        }
    
        /**
         * Read bytes from the given input stream, filling the given array
         * 
         * @param inputStream The input stream
         * @param data The array to be filled
         * @throws IOException If the input stream does not contain enough bytes
         * to fill the array, or any other IO error occurs
         */
        private static void read(InputStream inputStream, byte data[]) 
            throws IOException
        {
            int offset = 0;
            while (true)
            {
                int read = inputStream.read(
                    data, offset, data.length - offset);
                if (read < 0)
                {
                    break;
                }
                offset += read;
                if (offset == data.length)
                {
                    return;
                }
            }
            throw new IOException("Tried to read " + data.length
                + " bytes, but only found " + offset);
        }
    }
    

    MnistEntry 类:

    import java.awt.image.BufferedImage;
    import java.awt.image.DataBuffer;
    import java.awt.image.DataBufferByte;
    
    /**
     * An entry of the MNIST data set. Instances of this class will be passed
     * to the consumer that is given to the {@link MnistCompressedReader} and
     * {@link MnistDecompressedReader} reading methods.
     */
    public class MnistEntry
    {
        /**
         * The index of the entry
         */
        private final int index;
    
        /**
         * The class label of the entry
         */
        private final byte label;
    
        /**
         * The number of rows of the image data
         */
        private final int numRows;
    
        /**
         * The number of columns of the image data
         */
        private final int numCols;
    
        /**
         * The image data 
         */
        private final byte[] imageData;        
    
        /**
         * Default constructor
         * 
         * @param index The index
         * @param label The label
         * @param numRows The number of rows
         * @param numCols The number of columns
         * @param imageData The image data
         */
        MnistEntry(int index, byte label, int numRows, int numCols,
            byte[] imageData)
        {
            this.index = index;
            this.label = label;
            this.numRows = numRows;
            this.numCols = numCols;
            this.imageData = imageData;
        }
    
        /**
         * Returns the index of the entry
         * 
         * @return The index
         */
        public int getIndex()
        {
            return index;
        }
    
        /**
         * Returns the class label of the entry. This is a value in [0,9], 
         * indicating which digit is shown in the entry
         * 
         * @return The class label
         */
        public byte getLabel()
        {
            return label;
        }
    
        /**
         * Returns the number of rows of the image data. 
         * This will usually be 28.
         * 
         * @return The number of rows
         */
        public int getNumRows()
        {
            return numRows;
        }
    
        /**
         * Returns the number of columns of the image data. 
         * This will usually be 28.
         * 
         * @return The number of columns
         */
        public int getNumCols()
        {
            return numCols;
        }
    
        /**
         * Returns a <i>reference</i> to the image data. This will be an array
         * of length <code>numRows * numCols</code>, containing values 
         * in [0,255] indicating the brightness of the pixels.
         * 
         * @return The image data
         */
        public byte[] getImageData()
        {
            return imageData;
        }
    
        /**
         * Creates a new buffered image from the image data that is stored
         * in this entry.
         * 
         * @return The image
         */
        public BufferedImage createImage()
        {
            BufferedImage image = new BufferedImage(getNumCols(),
                getNumRows(), BufferedImage.TYPE_BYTE_GRAY);
            DataBuffer dataBuffer = image.getRaster().getDataBuffer();
            DataBufferByte dataBufferByte = (DataBufferByte) dataBuffer;
            byte data[] = dataBufferByte.getData();
            System.arraycopy(getImageData(), 0, data, 0, data.length);
            return image;
        }
    
    
        @Override
        public String toString()
        {
            String indexString = String.format("%05d", index);
            return "MnistEntry[" 
            + "index=" + indexString + "," 
            + "label=" + label + "]";
        }
    
    }
    

    阅读器可用于读取未压缩的文件。结果将是MnistEntry 传递给消费者的实例:

    MnistDecompressedReader mnistReader = new MnistDecompressedReader();
    mnistReader.readDecompressedTraining(Paths.get("./data"), mnistEntry -> 
    {
        System.out.println("Read entry " + mnistEntry);
        BufferedImage image = mnistEntry.createImage();
        ...
    });
    

    MnistReader 项目包含几个examples,这些类可用于读取压缩或未压缩数据,或从 MNIST 条目生成 PNG 图像。

    【讨论】:

    • 您正在以字节数组的形式读取图像数据,但在 idx 文件中,像素值是无符号字节。由于 Java 没有无符号字节且字节的最大值为 127,因此您需要将这些字节转换为 short/int 以便从图像数据中获取正确的像素值。
    • @MarkoR 没有。数据被读入数组,数组包含字节,字节只是位,没有内在含义。它们是“0...255 范围内的无符号值”还是“-128...127 范围内的有符号值”只是解释这些值的问题。当它们被发送到 BufferedImage 类型为 TYPE_BYTE_GRAY 时,它们被解释为 unsigned 值。 (只需查看链接项目中 ReadAndSaveImages 示例的输出)。
    • 我同意解释这些值很重要,但是如果有人要复制您的代码并使用 MnistEntry 中的 getImageData() 并在 Java 中访问它,他们将获得从 -128 到 127 的值而不是从 0 到 255。很容易错过这个细节,这就是我写评论的原因,提醒人们这些字节值需要解释为无符号字节。
    【解决方案3】:

    该 URL 描述了您必须解码的格式,并且他们提到它是非标准的,因此显而易见的 Google 搜索不会出现任何使用代码。但是,它非常直接,带有一个标题,后跟一个 0-255 灰度值的 28x28 像素矩阵。

    读取数据后(记住要注意字节序),创建 BMP 文件就很简单了。

    我向你推荐以下文章:

    How to make bmp image from pixel byte array in java

    他们的问题是关于颜色的,但他们的代码已经适用于灰度,这就是你所需要的,你应该能够从那个 sn-p 中得到一些东西。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 2016-03-18
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2022-07-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      相关资源
      最近更新 更多