字符串列表/数组到 numpy 浮点数组答案

【问题标题】：List/Array of strings to numpy float array字符串列表/数组到 numpy 浮点数组
【发布时间】：2017-04-29 01:26:33
【问题描述】：

我是 scikit learn 和 numpy 的新手。如何表示由字符串列表/数组组成的数据集，例如

[["aa bb","a","bbb","à"], [bb cc","c","ddd","à"], ["kkk","a","","a"]]

到一个 dtype 浮点数的 numpy 数组？

【问题讨论】：

什么？？？将字符串转换为浮点数？顺便说一句，它与 sklearn 无关
好吧，也许我没有使用正确的术语，但@datawrestler 理解了我的问题并给出了非常有用的建议。还是谢谢。

标签： python numpy scikit-learn

【解决方案1】：

我认为您正在寻找的是您的话的数字表示。您可以使用 gensim 并将每个单词映射到一个令牌 id 并从中创建您的 numpy 数组，如下所示：

import numpy as np
from gensim import corpora 

toconvert = [["aa bb","a","bbb","à"], ["bb", "cc","c","ddd","à"], ["kkk","a","","a"]]

# convert your list of lists into token id's. For example, 'aa bb' could be represented as a 2, a as a 1, etc.
tdict = corpora.Dictionary(toconvert)

# given nested structure, you can append nested numpy arrays
newlist = []
for l in toconvert:
    tmplist = []
    for word in l:
        # append to intermediate list the id for the given word under observation
        tmplist.append(tdict.token2id[word])
    # convert to numpy array and append to main list
    newlist.append(np.array(tmplist).astype(float)) # type float

print(newlist) # desired output: [array([ 2.,  0.,  1.,  0.]), array([ 5.,  3.,  4.,  6.,  0.]), array([ 7.,  0.,  8.,  0.])]

# and to see what id's represent which strings:
tdict[0] # 'a'

【讨论】：

感谢@datawrestler 您提供的答案。挺好用的。