Can You Survive the Titanic? Build a Machine Learning Predictor
A classic hands-on classification project rebuilt as a complete beginner-friendly handbook from raw CSV to a working prediction app.
What you will build
A web app that takes passenger details and predicts survival probability while comparing multiple classification models.
- • A leakage-safe preprocessing + classification pipeline.
- • Logistic Regression and Random Forest comparison.
- • Accuracy, classification report, confusion matrix and ROC-AUC.
- • A saved model file and working browser prediction app.
- • A GitHub-ready project folder you can show in interviews.
Tools you will use
Before you begin
Use the classic Titanic passenger dataset with a train.csv containing the target column Survived. The familiar Kaggle Titanic dataset uses exactly this structure. Put the file inside the project's data folder before you run the training code.
Install Python and VS Code
Open your browser. Go to the official Python website, download Python 3.10+ and run the installer. On Windows, tick Add Python to PATH before selecting Install.
Install VS Code. Open VS Code → click Extensions on the left → search for Python → install the Microsoft Python extension.
Check before continuing: Opening a terminal and running python --version prints Python 3.10 or newer.
Create the project folder
Open VS Code → File → Open Folder → create a new folder named titanic-survival-predictor.
In the Explorer panel create these folders: data, notebooks, and src. At the project root create empty files app.py and requirements.txt.
├── data/
├── notebooks/
├── src/
├── app.py
└── requirements.txt
Check before continuing: VS Code Explorer shows data, notebooks, src, app.py and requirements.txt.
Create a virtual environment and install the tools
In VS Code click Terminal → New Terminal. Run the commands below. Use the activation command for your operating system.
python -m venv .venv
# Windows PowerShell
.venv\Scripts\Activate.ps1
# macOS / Linux
source .venv/bin/activate
python -m pip install --upgrade pip
pip install pandas numpy scikit-learn matplotlib seaborn joblib gradio jupyterlab
pip freeze > requirements.txtCheck before continuing: pip show scikit-learn and pip show gradio both return package information.
Download the Titanic dataset
Open the Titanic machine-learning competition dataset in your browser. Download the dataset files and extract them.
Copy train.csv into your project's data folder. For this handbook, training uses train.csv because it contains the Survived answer column.
Check before continuing: data/train.csv exists and opens in VS Code with columns such as PassengerId, Survived, Pclass, Sex, Age and Fare.
Start JupyterLab and inspect the data
Return to the VS Code terminal and run jupyter lab. A browser tab should open. In JupyterLab click File → New → Notebook → Python 3. Save it as notebooks/01_titanic.ipynb.
Paste the inspection code below into the first code cell and press Shift + Enter.
import pandas as pd
df = pd.read_csv("data/train.csv")
print(df.shape)
print(df.head())
print(df.info())
print(df.isna().sum().sort_values(ascending=False).head(10))
print(df["Survived"].value_counts(normalize=True))Check before continuing: The notebook prints 891 rows for the standard Titanic training dataset and shows missing values in Age and Cabin.
Decide what the model is allowed to use
Use Pclass, Sex, Age, SibSp, Parch, Fare and Embarked. The target is Survived.
Do not put the target into the features. Do not learn imputation values or encodings from the complete dataset before splitting; that would leak test information into training.
Check before continuing: Your feature list does not contain Survived and contains only information available for a passenger before the outcome.
Create train and test sets
Create a new file src/train_model.py. The training script below performs the split before fitting preprocessing. A fixed random_state makes your teaching result reproducible.
Check before continuing: The training and test sets have similar survival proportions because stratify=y was used.
Build one preprocessing pipeline
The numeric branch fills missing values with the training median and scales values. The categorical branch fills missing categories with the most frequent training value and one-hot encodes categories.
Keeping preprocessing inside the Scikit-learn Pipeline is important: the same transformations will later be reused by the browser app.
Check before continuing: Calling pipeline.fit(...) completes without manually filling missing Age values or manually creating dummy columns.
Train Logistic Regression and Random Forest
Paste the complete script below into src/train_model.py. Save the file. In the VS Code terminal run python src/train_model.py.
import joblib
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.ensemble import RandomForestClassifier
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (
accuracy_score,
classification_report,
confusion_matrix,
roc_auc_score,
)
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
df = pd.read_csv("data/train.csv")
features = [
"Pclass", "Sex", "Age", "SibSp",
"Parch", "Fare", "Embarked"
]
X = df[features]
y = df["Survived"]
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.20,
stratify=y,
random_state=42,
)
numeric_features = ["Age", "SibSp", "Parch", "Fare"]
categorical_features = ["Pclass", "Sex", "Embarked"]
preprocess = ColumnTransformer([
(
"numeric",
Pipeline([
("imputer", SimpleImputer(strategy="median")),
("scaler", StandardScaler()),
]),
numeric_features,
),
(
"categorical",
Pipeline([
("imputer", SimpleImputer(strategy="most_frequent")),
("onehot", OneHotEncoder(handle_unknown="ignore")),
]),
categorical_features,
),
])
models = {
"Logistic Regression": LogisticRegression(max_iter=1000, random_state=42),
"Random Forest": RandomForestClassifier(
n_estimators=300,
min_samples_leaf=2,
random_state=42,
),
}
best_pipeline = None
best_name = None
best_auc = -1
for name, estimator in models.items():
pipeline = Pipeline([
("prepare", preprocess),
("model", estimator),
])
pipeline.fit(X_train, y_train)
prediction = pipeline.predict(X_test)
probability = pipeline.predict_proba(X_test)[:, 1]
accuracy = accuracy_score(y_test, prediction)
auc = roc_auc_score(y_test, probability)
print("\n", name)
print("Accuracy:", round(accuracy, 3))
print("ROC-AUC:", round(auc, 3))
print("Confusion matrix:")
print(confusion_matrix(y_test, prediction))
print(classification_report(y_test, prediction))
if auc > best_auc:
best_auc = auc
best_name = name
best_pipeline = pipeline
joblib.dump(best_pipeline, "titanic_model.joblib")
print("\nSaved:", best_name)
print("Model file: titanic_model.joblib")On a standard 891-row Titanic training dataset, a sensible result is usually around the high-70% to low-80% accuracy range. Your exact value can vary with library versions and modeling choices.
Check before continuing: The terminal prints metrics for both models and creates titanic_model.joblib.
Read the confusion matrix instead of stopping at accuracy
If the confusion matrix is [[98, 12], [23, 46]], the model correctly classified 98 non-survivors and 46 survivors, while making 12 false-positive and 23 false-negative errors.
Use ROC-AUC to judge ranking quality across thresholds, but do not treat one metric as the entire story.
Check before continuing: You can point to false positives and false negatives and explain what each means for the survival prediction task.
Save the entire pipeline
The training script saves the complete preprocessing + classifier pipeline, not only the classifier. This means the app can accept raw values such as female, 3 and S without recreating transformations by hand.
Check before continuing: The project root contains titanic_model.joblib and loading it with joblib.load(...) succeeds.
Build the prediction web app
Open app.py in VS Code. Replace its contents with the code below and save it.
import joblib
import pandas as pd
import gradio as gr
model = joblib.load("titanic_model.joblib")
def predict_survival(
passenger_class,
sex,
age,
siblings_spouses,
parents_children,
fare,
embarked,
):
row = pd.DataFrame([{
"Pclass": int(passenger_class),
"Sex": sex,
"Age": float(age),
"SibSp": int(siblings_spouses),
"Parch": int(parents_children),
"Fare": float(fare),
"Embarked": embarked,
}])
probability = float(model.predict_proba(row)[0, 1])
label = (
"Likely survived"
if probability >= 0.50
else "Likely did not survive"
)
return {
"Survived": probability,
"Did not survive": 1 - probability,
}, f"{label} — survival probability {probability:.1%}"
demo = gr.Interface(
fn=predict_survival,
inputs=[
gr.Dropdown([1, 2, 3], value=3, label="Passenger class"),
gr.Dropdown(["female", "male"], value="female", label="Sex"),
gr.Number(value=29, label="Age"),
gr.Number(value=0, precision=0, label="Siblings / spouses aboard"),
gr.Number(value=0, precision=0, label="Parents / children aboard"),
gr.Number(value=32.2, label="Fare"),
gr.Dropdown(["C", "Q", "S"], value="S", label="Embarked port"),
],
outputs=[
gr.Label(label="Prediction"),
gr.Textbox(label="Result"),
],
title="Titanic Survival Predictor",
description="Enter passenger details and estimate survival probability.",
)
demo.launch()Check before continuing: python app.py starts a local Gradio URL without an import or model-loading error.
Run the app in your browser
In the terminal make sure your virtual environment is active, then run:
python app.pyGradio prints a local address, commonly http://127.0.0.1:7860. Hold Ctrl and click the URL, or copy it into your browser.
Check before continuing: The page shows Titanic Survival Predictor, seven passenger inputs and a Predict button.
Test two very different passengers
Test A: Class 1, female, age 38, SibSp 1, Parch 0, Fare 71.28, Embarked C.
Test B: Class 3, male, age 22, SibSp 1, Parch 0, Fare 7.25, Embarked S.
Click Predict for each case. The point is not to prove causation; it is to verify that the complete raw-input → preprocessing → model → probability path works.
Check before continuing: The two test passengers return visibly different survival probabilities.
Create a Git repository
Create a file named .gitignore and add .venv/, __pycache__/ and notebook checkpoint folders. Do not commit secrets.
Open GitHub → click New repository → name it titanic-survival-predictor → create the empty repository. Then run:
git init
git add .
git commit -m "Build Titanic survival predictor"
# Create an empty GitHub repository first, then copy its URL:
git branch -M main
git remote add origin https://github.com/YOUR-USERNAME/titanic-survival-predictor.git
git push -u origin mainCheck before continuing: Refreshing the GitHub repository shows your source files and requirements.txt but not your .venv folder.
Explain the project in an interview
A strong explanation is: “I used the Titanic classification problem to build an end-to-end Scikit-learn pipeline. I split before preprocessing, imputed numeric and categorical features inside the pipeline, one-hot encoded categories, compared Logistic Regression with Random Forest, evaluated with confusion matrix and ROC-AUC, saved the full pipeline and exposed it through a Gradio app.”
Then discuss one improvement you would make next: feature engineering such as family size/title extraction, cross-validation, calibration or threshold analysis.
Check before continuing: You can explain the problem, leakage prevention, preprocessing, baseline, evaluation and deployment without reading the code.
Common problems and exact fixes
| Problem | Why it happens | What to do |
|---|---|---|
| python is not recognized | Python was not added to PATH | Re-run the installer or use the Python launcher py on Windows. |
| FileNotFoundError: data/train.csv | CSV is in the wrong folder | Put train.csv inside project-root/data and run from the project root. |
| ModuleNotFoundError | Virtual environment is inactive or package not installed | Activate .venv and run pip install -r requirements.txt. |
| Input columns do not match | App fields differ from training feature names | Use exactly Pclass, Sex, Age, SibSp, Parch, Fare and Embarked. |
| App loads but prediction fails | Only the model was saved, not preprocessing | Save and load the complete Scikit-learn Pipeline. |
What this one project revises from the curriculum
Screenshot standard for this handbook
We will only publish screenshots captured from the real tools used while running the project. No AI-generated fake IDE, notebook or application screenshots will be used. The screenshot pass will show the VS Code folder, Jupyter data inspection, training output and the running prediction app.