【发布时间】:2020-05-06 00:16:25
【问题描述】:
我参考了这篇文章,该文章讨论了如何使用 reverse_map 策略从 keras 中标记器的 text_to_sequences 函数中获取文本。
我想知道是否有一个函数可以为 text_to_matrix 函数取回文本。
例子:
from tensorflow.keras.preprocessing.text import Tokenizer
docs = ['Well done!',
'Good work',
'Great effort',
'nice work',
'Excellent!']
# create the tokenizer
t = Tokenizer()
# fit the tokenizer on the documents
t.fit_on_texts(docs)
print(t)
encoded_docs = t.texts_to_matrix(docs, mode='count')
print(encoded_docs)
print(t.word_index.items())
Output:
<keras_preprocessing.text.Tokenizer object at 0x7f746b6594e0>
[[0. 0. 1. 1. 0. 0. 0. 0. 0.]
[0. 1. 0. 0. 1. 0. 0. 0. 0.]
[0. 0. 0. 0. 0. 1. 1. 0. 0.]
[0. 1. 0. 0. 0. 0. 0. 1. 0.]
[0. 0. 0. 0. 0. 0. 0. 0. 1.]]
dict_items([('work', 1), ('well', 2), ('done', 3), ('good', 4), ('great', 5), ('effort', 6),
('nice', 7), ('excellent', 8)])
如何从 one-hot 矩阵中取回文档?
【问题讨论】:
标签: python-3.x text keras tokenize one-hot-encoding