【问题标题】:TopicModel: How to query documents by topic model "topic"?TopicModel:如何通过主题模型“topic”查询文档?
【发布时间】:2018-07-26 05:09:37
【问题描述】:

下面我创建了一个完整的可重现示例来计算给定 DataFrame 的主题模型。

import numpy as np  
import pandas as pd

data = pd.DataFrame({'Body': ['Here goes one example sentence that is generic',
                  'My car drives really fast and I have no brakes',
                  'Your car is slow and needs no brakes', 
                  'Your and my vehicle are both not as fast as the airplane']})

from sklearn.decomposition import LatentDirichletAllocation
from sklearn.feature_extraction.text import CountVectorizer

vectorizer = CountVectorizer(lowercase = True, analyzer = 'word')

data_vectorized = vectorizer.fit_transform(data.Body)
lda_model = LatentDirichletAllocation(n_components=4, 
                                      learning_method='online', 
                                      random_state=0,
                                      verbose=1)
lda_topic_matrix = lda_model.fit_transform(data_vectorized)

问题:如何按主题过滤文档?如果是这样,文档可以有多个主题标签,还是需要一个阈值?

最后,我喜欢将每个文档标记为“1”,这取决于它是否具有主题 2 和主题 3 的高负载,否则为“0”。

【问题讨论】:

    标签: python scikit-learn lda topic-modeling


    【解决方案1】:

    lda_topic_matrix 包含文档属于特定主题/标签的概率分布。在人类中,这意味着每行总和为 1,而每个索引处的值是该文档属于特定主题的概率。因此,每个文档都有不同程度的所有主题标签。如果您有 4 个主题,则具有所有标签的文档将在 lda_topic_matrix 中具有相应的行,类似于 [0.25, 0.25, 0.25, 0.25]。只有单个主题(“0”)的文档行将变成类似[0.97, 0.01, 0.01, 0.01] 的文档,而具有两个主题(“1”和“2”)的文档的分布将类似于[0.01, 0.54, 0.44, 0.01]

    所以最简单的做法是选择概率最高的话题,检查是2还是3

    main_topic_of_document = np.argmax(lda_topic_matrix, axis=1)
    tagged = ((main_topic_of_document==2) | (main_topic_of_document==3)).astype(np.int64)
    

    This article 很好地解释了 LDA 的内部机制。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 2016-03-24
      • 1970-01-01
      • 2021-07-27
      • 2015-11-15
      • 2012-04-16
      • 2013-01-02
      • 2019-07-11
      相关资源
      最近更新 更多