【发布时间】:2019-02-23 23:46:07
【问题描述】:
我是机器学习的初学者,我正在尝试通过 Kaggle 的 TItanic 问题来学习。我已经完成了我的代码并获得了 0.78 的准确度分数,但现在我需要生成一个包含 418 个条目 + 标题行的 CSV 文件,但我不知道该怎么做它。
这是我应该制作的示例:
PassengerId,Survived
892,0
893,1
894,0
Etc.
数据来自我的test_predictions
这是我的代码:
import pandas as pd
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
"""Assigning the train & test datasets' adresses to variables"""
train_path = "C:\\Users\\Omar\\Downloads\\Titanic Data\\train.csv"
test_path = "C:\\Users\\Omar\\Downloads\\Titanic Data\\test.csv"
"""Using pandas' read_csv() function to read the datasets
and then assigning them to their own variables"""
train_data = pd.read_csv(train_path)
test_data = pd.read_csv(test_path)
"""Using pandas' factorize() function to represent genders (male/female)
with binary values (0/1)"""
train_data['Sex'] = pd.factorize(train_data.Sex)[0]
test_data['Sex'] = pd.factorize(test_data.Sex)[0]
"""Replacing missing values in the training and test dataset with 0"""
train_data.fillna(0.0, inplace = True)
test_data.fillna(0.0, inplace = True)
"""Selecting features for training"""
columns_of_interest = ['Pclass', 'Sex', 'Age']
"""Dropping missing/NaN values from the training dataset"""
filtered_titanic_data = train_data.dropna(axis=0)
"""Using the predictory features in the data to handle the x axis"""
x = filtered_titanic_data[columns_of_interest]
"""The survival (what we're trying to find) is the y axis"""
y = filtered_titanic_data.Survived
"""Splitting the train data with test"""
train_x, val_x, train_y, val_y = train_test_split(x, y, random_state=0)
"""Assigning the DecisionClassifier model to a variable"""
titanic_model = DecisionTreeClassifier()
"""Fitting the x and y values with the model"""
titanic_model.fit(train_x, train_y)
"""Predicting the x-axis"""
val_predictions = titanic_model.predict(val_x)
"""Assigning the feature columns from the test to a variable"""
test_x = test_data[columns_of_interest]
"""Predicting the test by feeding its x axis into the model"""
test_predictions = titanic_model.predict(test_x)
"""Printing the prediction"""
print(val_predictions)
"""Checking for the accuracy"""
print(accuracy_score(val_y, val_predictions))
"""Printing the test prediction"""
print(test_predictions)
【问题讨论】:
-
问题是什么?您的解决方案有什么不足 - 它做什么或不做什么是不正确的?您是否收到错误/异常?
-
How to produce a CSV file with Python with specific entries? -
请按照您创建此帐户时的建议阅读并遵循帮助文档中的发布指南。 Minimal, complete, verifiable example 适用于此。在您发布 MCVE 代码并准确描述问题之前,我们无法有效地帮助您。我们应该能够将您发布的代码粘贴到文本文件中并重现您描述的问题。
-
我已经编辑了问题。
-
通常会为您提供样本提交文件。如果您将其作为 DataFrame,则只需执行
submission['Survived'] = test_predictions。下一行将从 pandas 的 DataFrame 创建 csv 文件。submission.to_csv('filename.csv', index=False)
标签: python pandas machine-learning scikit-learn kaggle