【问题标题】:Python Scikit-Learn DecisionTreeClassifier.fit() throws KeyError: 'default'Python Scikit-Learn DecisionTreeClassifier.fit() 抛出 KeyError: 'default'
【发布时间】:2020-02-26 15:16:44
【问题描述】:

我有一个小数据集,正在尝试使用 sklearn 创建决策树分类器。我使用 sklearn.tree.DecisionTreeClassifier 作为模型,并使用它的 .fit() 函数来拟合数据。环顾四周,我找不到遇到同样问题的其他人。

将数据加载到一个数组并将标签加载到另一个数组后,打印出两个数组(数据和标签)给出:

[[0. 1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 1. 1. 0.]
 [1. 1. 1. 0. 0. 0. 0. 0. 1. 0. 0. 0. 0. 0. 0. 0.]
 [1. 1. 1. 0. 1. 0. 1. 0. 0. 0. 0. 0. 0. 0. 0. 1.]
 [1. 1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 1. 0.]
 [0. 1. 1. 1. 1. 0. 1. 1. 1. 0. 1. 1. 0. 1. 0. 0.]
 [0. 1. 1. 1. 1. 1. 1. 1. 0. 0. 1. 0. 0. 1. 1. 0.]
 [0. 1. 1. 1. 1. 1. 1. 1. 1. 0. 1. 0. 1. 0. 0. 0.]
 [1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0.]
 [0. 1. 1. 1. 0. 0. 0. 0. 0. 1. 1. 1. 0. 1. 1. 1.]
 [0. 1. 1. 1. 0. 0. 0. 0. 1. 0. 1. 1. 0. 1. 0. 0.]
 [0. 1. 1. 0. 0. 0. 1. 0. 0. 0. 1. 0. 0. 1. 0. 1.]
 [1. 0. 1. 0. 1. 0. 1. 0. 1. 0. 1. 0. 0. 0. 0. 0.]
 [1. 0. 1. 0. 1. 0. 1. 0. 1. 1. 0. 0. 0. 0. 0. 0.]]

['Alcaligenes_faecalis' 'Bacillus_circulans' 'Bacillus_megaterium'
 'Bacillus_sphaericus' 'Citrobacter_freundii' 'Enterobacter_aerogenes'
 'Escherichia_coli' 'Micrococcus_luteus' 'Proteus_mirabilis'
 'Salmonella_arizonae' 'Serratia_marcescens' 'Staphylococcus_epidermidis'
 'Staphylococcus_saprophyticus']

我定义了一个函数来进行拟合,我尝试删除该函数并直接运行 .fit() 函数:

def decisiontree(data, labels, criterion = "gini", splitter = "default", max_depth = None): #expects *2d data and 1d labels

    model = sklearn.tree.DecisionTreeClassifier(criterion = criterion, splitter = splitter, max_depth = max_depth)
    model = model.fit(data,labels)

    return model

然后我调用了函数:

model = decisiontree(data, labels)

此时,引发了 KeyError:

KeyError                                  Traceback (most recent call last)
<ipython-input-21-3574397ccfb6> in <module>
----> 1 model = decisiontree(data, labels)

<ipython-input-18-e85883291477> in decisiontree(data, labels, criterion, splitter, max_depth)
      2 
      3     model = sklearn.tree.DecisionTreeClassifier(criterion = criterion, splitter = splitter, max_depth = max_depth)
----> 4     model = model.fit(data,labels)
      5 
      6     return model

~/anaconda3/lib/python3.7/site-packages/sklearn/tree/_classes.py in fit(self, X, y, sample_weight, check_input, X_idx_sorted)
    875             sample_weight=sample_weight,
    876             check_input=check_input,
--> 877             X_idx_sorted=X_idx_sorted)
    878         return self
    879 

~/anaconda3/lib/python3.7/site-packages/sklearn/tree/_classes.py in fit(self, X, y, sample_weight, check_input, X_idx_sorted)
    333         splitter = self.splitter
    334         if not isinstance(self.splitter, Splitter):
--> 335             splitter = SPLITTERS[self.splitter](criterion,
    336                                                 self.max_features_,
    337                                                 min_samples_leaf,

KeyError: 'default'

数据存储在data.csv中:

Alcaligenes_faecalis,0,1,0,0,0,0,0,0,0,0,0,0,0,1,1,0
Bacillus_circulans,1,1,1,0,0,0,0,0,1,0,0,0,0,0,0,0
Bacillus_megaterium,1,1,1,0,1,0,1,0,0,0,0,0,0,0,0,1
Bacillus_sphaericus,1,1,0,0,0,0,0,0,0,0,0,0,0,0,1,0
Citrobacter_freundii,0,1,1,1,1,0,1,1,1,0,1,1,0,1,0,0
Enterobacter_aerogenes,0,1,1,1,1,1,1,1,0,0,1,0,0,1,1,0
Escherichia_coli,0,1,1,1,1,1,1,1,1,0,1,0,1,0,0,0
Micrococcus_luteus,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0
Proteus_mirabilis,0,1,1,1,0,0,0,0,0,1,1,1,0,1,1,1
Salmonella_arizonae,0,1,1,1,0,0,0,0,1,0,1,1,0,1,0,0
Serratia_marcescens,0,1,1,0,0,0,1,0,0,0,1,0,0,1,0,1
Staphylococcus_epidermidis,1,0,1,0,1,0,1,0,1,0,1,0,0,0,0,0
Staphylococcus_saprophyticus,1,0,1,0,1,0,1,0,1,1,0,0,0,0,0,0

【问题讨论】:

    标签: python python-3.x scikit-learn keyerror


    【解决方案1】:

    sklearn.tree.DecisionTreeClassifier spliter 参数没有default 值,默认值为best,因此您可以使用:

    def decisiontree(data, labels, criterion = "gini", splitter = "best", max_depth = None): #expects *2d data and 1d labels
    
        model = sklearn.tree.DecisionTreeClassifier(criterion = criterion, splitter = splitter, max_depth = max_depth)
        model = model.fit(data,labels)
    
        return model
    

    【讨论】:

    • 现在我觉得自己很傻,应该更仔细地阅读文档。
    【解决方案2】:

    根据the docssplitter 必须是“最佳”或“随机”。

    【讨论】:

      猜你喜欢
      • 2018-03-19
      • 2012-03-26
      • 2018-08-03
      • 2013-01-18
      • 2018-06-27
      • 1970-01-01
      • 2018-06-27
      相关资源
      最近更新 更多