【问题标题】:Folding in (estimating topics for new documents) in LDA using Mallet in Java使用 Java 中的 Mallet 在 LDA 中折叠(估计新文档的主题)
【发布时间】:2012-12-17 22:48:22
【问题描述】:

我正在通过 Java 使用 Mallet,但我不知道如何根据我训练过的现有主题模型评估新文档。

我生成模型的初始代码与Mallett Developers Guide for Topic Modelling 中的代码非常相似,之后我只是将模型保存为Java 对象。在稍后的过程中,我从文件中重新加载该 Java 对象,通过 .addInstances() 添加新实例,然后希望仅根据原始训练集中找到的主题评估这些新实例。

This stats.SE thread 提供了一些高级建议,但我看不出如何将它们应用到 Mallet 框架中。

非常感谢任何帮助。

【问题讨论】:

    标签: java mallet topic-modeling


    【解决方案1】:

    推理实际上也列在问题中提供的example link 中(最后几行)。

    对于任何对保存/加载训练模型然后使用它来推断新文档的模型分布的整个代码感兴趣的人 - 这里有一些 sn-ps:

    model.estimate() 完成后,您就拥有了实际训练的模型,因此您可以使用标准 Java ObjectOutputStream 对其进行序列化(因为ParallelTopicModel 实现了Serializable):

    try {
        FileOutputStream outFile = new FileOutputStream("model.ser");
        ObjectOutputStream oos = new ObjectOutputStream(outFile);
        oos.writeObject(model);
        oos.close();
    } catch (FileNotFoundException ex) {
        // handle this error
    } catch (IOException ex) {
        // handle this error
    }
    

    但请注意,当您推断时,您还需要通过同一管道传递新句子(如 Instance)以便对其进行预处理(tokenzie 等),因此,您还需要保存管道列表(因为我们使用SerialPipe什么时候可以创建一个实例然后序列化它):

    // initialize the pipelist (using in model training)
    SerialPipes pipes = new SerialPipes(pipeList);
    
    try {
        FileOutputStream outFile = new FileOutputStream("pipes.ser");
        ObjectOutputStream oos = new ObjectOutputStream(outFile);
        oos.writeObject(pipes);
        oos.close();
    } catch (FileNotFoundException ex) {
        // handle error
    } catch (IOException ex) {
        // handle error
    }
    

    为了加载模型/管道并将它们用于推理,我们需要反序列化:

    private static void InferByModel(String sentence) {
        // define model and pipeline
        ParallelTopicModel model = null;
        SerialPipes pipes = null;
    
        // load the model
        try {
            FileInputStream outFile = new FileInputStream("model.ser");
            ObjectInputStream oos = new ObjectInputStream(outFile);
            model = (ParallelTopicModel) oos.readObject();
        } catch (IOException ex) {
            System.out.println("Could not read model from file: " + ex);
        } catch (ClassNotFoundException ex) {
            System.out.println("Could not load the model: " + ex);
        }
    
        // load the pipeline
        try {
            FileInputStream outFile = new FileInputStream("pipes.ser");
            ObjectInputStream oos = new ObjectInputStream(outFile);
            pipes = (SerialPipes) oos.readObject();
        } catch (IOException ex) {
            System.out.println("Could not read pipes from file: " + ex);
        } catch (ClassNotFoundException ex) {
            System.out.println("Could not load the pipes: " + ex);
        }
    
        // if both are properly loaded
        if (model != null && pipes != null){
    
            // Create a new instance named "test instance" with empty target 
            // and source fields note we are using the pipes list here
            InstanceList testing = new InstanceList(pipes);   
            testing.addThruPipe(
                new Instance(sentence, null, "test instance", null));
    
            // here we get an inferencer from our loaded model and use it
            TopicInferencer inferencer = model.getInferencer();
            double[] testProbabilities = inferencer
                       .getSampledDistribution(testing.get(0), 10, 1, 5);
            System.out.println("0\t" + testProbabilities[0]);
        }
    }
    

    由于某种原因,我没有得到与原始模型完全相同的推断 - 但这是另一个问题的问题(如果有人知道,我很乐意听到)

    【讨论】:

    • 为了在连续推理后获得相同的结果,您只需添加 inferencer.setRandomSeed(1)。但是,如果我将推理器用于已在模型中使用的文本文档,则无法获得相同的主题分布。
    • 感谢有用的代码 sn-ps。但在性能方面,最好在训练模型时直接生成推理器,而不是先加载整个模型,然后从中获取推理器。
    • @Phauly - 有时您想训练一个模型,然后稍后再使用它。在这种情况下,如果训练模型有意义,然后将其保存以在必要时用于推理(如果我理解您的评论)
    【解决方案2】:

    我在slide-deck from Mallet's lead developer 中找到了答案:

    TopicInferencer inferencer = model.getInferencer();
    double[] topicProbs = inferencer.getSampledDistribution(newInstance, 100, 10, 10);
    

    【讨论】:

    • 这是要走的路。此外,如果您想在训练后保存模型并在测试前加载它(以保持它们分开),您可以查看这个答案stackoverflow.com/a/44379106/1042409
    猜你喜欢
    • 1970-01-01
    • 2020-06-09
    • 2011-07-04
    • 2015-09-11
    • 2015-10-22
    • 1970-01-01
    • 2020-12-25
    • 2020-07-07
    • 1970-01-01
    相关资源
    最近更新 更多