【发布时间】:2019-04-11 19:47:35
【问题描述】:
我有一个包含 2 列的 pandas 数据框,我想在其中之一中使用 sklearn TfidfVectorizer 进行 文本分类。但是,此列是列表列表,并且 TFIDF 想要将原始输入作为文本。在this question 中,他们提供了一个解决方案,以防我们只有一个列表列表,但我想问一下如何在我的数据框的每一行中应用这个函数,哪一行包含一个列表列表。提前谢谢你。
Input:
0 [[this, is, the], [first, row], [of, dataframe]]
1 [[that, is, the], [second], [row, of, dataframe]]
2 [[etc], [etc, etc]]
想要的输出:
0 ['this is the', 'first row', 'of dataframe']
1 ['that is the', 'second', 'row of dataframe']
2 ['etc', 'etc etc']
【问题讨论】:
-
您能添加一些示例输入吗?
-
我更新了丹尼尔的问题
标签: python list dataframe tfidfvectorizer