【发布时间】:2020-02-22 12:06:58
【问题描述】:
我一直在尝试做出一个预测,该预测由我使用决策树算法创建的模型中的 DataFrame 组成。
我的模型得分为 0.96。然后,我尝试使用该模型对留下但出错的 DataFrame 人进行预测。目标是根据留下的 DataFrame 预测未来将离开公司的人。
如何实现这个目标?
所以我做的是:
- 从我的 github 读取 DF 并将它们分成离开和未离开的人
df = pd.read_csv('https://raw.githubusercontent.com/bhaskoro-muthohar/DataScienceLearning/master/HR_comma_sep.csv')
leftdf = df[df['left']==1]
notleftdf =df[df['left']==0]
- 为模型生成准备数据
df.salary = df.salary.map({'low':0,'medium':1,'high':2})
df.salary
X = df.drop(['left','sales'],axis=1)
y = df['left']
- 拆分训练集和测试集
import numpy as np
from sklearn.model_selection import train_test_split
#splitting the train and test sets
X_train, X_test, y_train, y_test= train_test_split(X,y,random_state=0, stratify=y)
- 训练它
from sklearn import tree
clftree = tree.DecisionTreeClassifier(max_depth=3)
clftree.fit(X_train,y_train)
- 评估模型
y_pred = clftree.predict(X_test)
print("Test set prediction:\n {}".format(y_pred))
print("Test set score: {:.2f}".format(clftree.score(X_test, y_test)))
结果是
测试集分数:0.96
- 然后我尝试使用尚未离开公司的人的 DataFrame 进行预测
X_new = notleftdf.drop(['left','sales'],axis=1)
#Map salary to 0,1,2
X_new.salary = X_new.salary.map({'low':0,'medium':1,'high':2})
X_new.salary
prediction_will_left = clftree.predict(X_new)
print("Prediction: {}".format(prediction_will_left))
print("Predicted target name: {}".format(
notleftdf['left'][prediction_will_left]
))
我得到的错误是:
KeyError: "None of [Int64Index([0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n ...\n 0, 0, 0, 0, 0, 0, 1, 0, 0, 0],\n dtype='int64', length=11428)] are in the [index]"
如何解决?
PS:完整的脚本链接是here
【问题讨论】:
-
目前尚不清楚您到底想做什么。错误很明显,没有找到 index 的值。但请提供具体细节作为问题的一部分,以便获得快速帮助。
-
@SupratimHaldar 很抱歉不清楚,我尝试使用该模型对留下但出错的 DataFrame 人进行预测。目标是根据留下的 DataFrame 预测未来将离开公司的人。
-
请提供Minimal and Reproducible example(即这里的人们可以重现您遇到的错误的最少代码量)。
-
@Xukrao 已经编辑过先生。
标签: python machine-learning classification decision-tree supervised-learning