Skip to main content
All project handbooks
HANDBOOK V1 · SCREENSHOTS NEXTBeginner4-6 hours

Can You Survive the Titanic? Build a Machine Learning Predictor

A classic hands-on classification project rebuilt as a complete beginner-friendly handbook from raw CSV to a working prediction app.

What you will build

A web app that takes passenger details and predicts survival probability while comparing multiple classification models.

  • • A leakage-safe preprocessing + classification pipeline.
  • • Logistic Regression and Random Forest comparison.
  • • Accuracy, classification report, confusion matrix and ROC-AUC.
  • • A saved model file and working browser prediction app.
  • • A GitHub-ready project folder you can show in interviews.

Tools you will use

PythonVS CodeJupyterLabPandasNumPyMatplotlibSeabornScikit-learnLogistic RegressionRandom ForestGradioGitGitHub

Before you begin

Use the classic Titanic passenger dataset with a train.csv containing the target column Survived. The familiar Kaggle Titanic dataset uses exactly this structure. Put the file inside the project's data folder before you run the training code.

1

Install Python and VS Code

Open your browser. Go to the official Python website, download Python 3.10+ and run the installer. On Windows, tick Add Python to PATH before selecting Install.

Install VS Code. Open VS Code → click Extensions on the left → search for Python → install the Microsoft Python extension.

Check before continuing: Opening a terminal and running python --version prints Python 3.10 or newer.

2

Create the project folder

Open VS Code → File → Open Folder → create a new folder named titanic-survival-predictor.

In the Explorer panel create these folders: data, notebooks, and src. At the project root create empty files app.py and requirements.txt.

titanic-survival-predictor/
├── data/
├── notebooks/
├── src/
├── app.py
└── requirements.txt

Check before continuing: VS Code Explorer shows data, notebooks, src, app.py and requirements.txt.

3

Create a virtual environment and install the tools

In VS Code click Terminal → New Terminal. Run the commands below. Use the activation command for your operating system.

Create the environmentRunnable
python -m venv .venv

# Windows PowerShell
.venv\Scripts\Activate.ps1

# macOS / Linux
source .venv/bin/activate

python -m pip install --upgrade pip
pip install pandas numpy scikit-learn matplotlib seaborn joblib gradio jupyterlab

pip freeze > requirements.txt

Check before continuing: pip show scikit-learn and pip show gradio both return package information.

4

Download the Titanic dataset

Open the Titanic machine-learning competition dataset in your browser. Download the dataset files and extract them.

Copy train.csv into your project's data folder. For this handbook, training uses train.csv because it contains the Survived answer column.

Check before continuing: data/train.csv exists and opens in VS Code with columns such as PassengerId, Survived, Pclass, Sex, Age and Fare.

5

Start JupyterLab and inspect the data

Return to the VS Code terminal and run jupyter lab. A browser tab should open. In JupyterLab click File → New → Notebook → Python 3. Save it as notebooks/01_titanic.ipynb.

Paste the inspection code below into the first code cell and press Shift + Enter.

Inspect the CSVRunnable
import pandas as pd

df = pd.read_csv("data/train.csv")

print(df.shape)
print(df.head())
print(df.info())
print(df.isna().sum().sort_values(ascending=False).head(10))
print(df["Survived"].value_counts(normalize=True))

Check before continuing: The notebook prints 891 rows for the standard Titanic training dataset and shows missing values in Age and Cabin.

6

Decide what the model is allowed to use

Use Pclass, Sex, Age, SibSp, Parch, Fare and Embarked. The target is Survived.

Do not put the target into the features. Do not learn imputation values or encodings from the complete dataset before splitting; that would leak test information into training.

Check before continuing: Your feature list does not contain Survived and contains only information available for a passenger before the outcome.

7

Create train and test sets

Create a new file src/train_model.py. The training script below performs the split before fitting preprocessing. A fixed random_state makes your teaching result reproducible.

Check before continuing: The training and test sets have similar survival proportions because stratify=y was used.

8

Build one preprocessing pipeline

The numeric branch fills missing values with the training median and scales values. The categorical branch fills missing categories with the most frequent training value and one-hot encodes categories.

Keeping preprocessing inside the Scikit-learn Pipeline is important: the same transformations will later be reused by the browser app.

Check before continuing: Calling pipeline.fit(...) completes without manually filling missing Age values or manually creating dummy columns.

9

Train Logistic Regression and Random Forest

Paste the complete script below into src/train_model.py. Save the file. In the VS Code terminal run python src/train_model.py.

Complete training scriptRunnable
import joblib
import pandas as pd

from sklearn.compose import ColumnTransformer
from sklearn.ensemble import RandomForestClassifier
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (
    accuracy_score,
    classification_report,
    confusion_matrix,
    roc_auc_score,
)
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

df = pd.read_csv("data/train.csv")

features = [
    "Pclass", "Sex", "Age", "SibSp",
    "Parch", "Fare", "Embarked"
]

X = df[features]
y = df["Survived"]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.20,
    stratify=y,
    random_state=42,
)

numeric_features = ["Age", "SibSp", "Parch", "Fare"]
categorical_features = ["Pclass", "Sex", "Embarked"]

preprocess = ColumnTransformer([
    (
        "numeric",
        Pipeline([
            ("imputer", SimpleImputer(strategy="median")),
            ("scaler", StandardScaler()),
        ]),
        numeric_features,
    ),
    (
        "categorical",
        Pipeline([
            ("imputer", SimpleImputer(strategy="most_frequent")),
            ("onehot", OneHotEncoder(handle_unknown="ignore")),
        ]),
        categorical_features,
    ),
])

models = {
    "Logistic Regression": LogisticRegression(max_iter=1000, random_state=42),
    "Random Forest": RandomForestClassifier(
        n_estimators=300,
        min_samples_leaf=2,
        random_state=42,
    ),
}

best_pipeline = None
best_name = None
best_auc = -1

for name, estimator in models.items():
    pipeline = Pipeline([
        ("prepare", preprocess),
        ("model", estimator),
    ])

    pipeline.fit(X_train, y_train)

    prediction = pipeline.predict(X_test)
    probability = pipeline.predict_proba(X_test)[:, 1]

    accuracy = accuracy_score(y_test, prediction)
    auc = roc_auc_score(y_test, probability)

    print("\n", name)
    print("Accuracy:", round(accuracy, 3))
    print("ROC-AUC:", round(auc, 3))
    print("Confusion matrix:")
    print(confusion_matrix(y_test, prediction))
    print(classification_report(y_test, prediction))

    if auc > best_auc:
        best_auc = auc
        best_name = name
        best_pipeline = pipeline

joblib.dump(best_pipeline, "titanic_model.joblib")

print("\nSaved:", best_name)
print("Model file: titanic_model.joblib")

On a standard 891-row Titanic training dataset, a sensible result is usually around the high-70% to low-80% accuracy range. Your exact value can vary with library versions and modeling choices.

Check before continuing: The terminal prints metrics for both models and creates titanic_model.joblib.

10

Read the confusion matrix instead of stopping at accuracy

If the confusion matrix is [[98, 12], [23, 46]], the model correctly classified 98 non-survivors and 46 survivors, while making 12 false-positive and 23 false-negative errors.

Use ROC-AUC to judge ranking quality across thresholds, but do not treat one metric as the entire story.

Check before continuing: You can point to false positives and false negatives and explain what each means for the survival prediction task.

11

Save the entire pipeline

The training script saves the complete preprocessing + classifier pipeline, not only the classifier. This means the app can accept raw values such as female, 3 and S without recreating transformations by hand.

Check before continuing: The project root contains titanic_model.joblib and loading it with joblib.load(...) succeeds.

12

Build the prediction web app

Open app.py in VS Code. Replace its contents with the code below and save it.

Gradio prediction appRunnable
import joblib
import pandas as pd
import gradio as gr

model = joblib.load("titanic_model.joblib")

def predict_survival(
    passenger_class,
    sex,
    age,
    siblings_spouses,
    parents_children,
    fare,
    embarked,
):
    row = pd.DataFrame([{
        "Pclass": int(passenger_class),
        "Sex": sex,
        "Age": float(age),
        "SibSp": int(siblings_spouses),
        "Parch": int(parents_children),
        "Fare": float(fare),
        "Embarked": embarked,
    }])

    probability = float(model.predict_proba(row)[0, 1])

    label = (
        "Likely survived"
        if probability >= 0.50
        else "Likely did not survive"
    )

    return {
        "Survived": probability,
        "Did not survive": 1 - probability,
    }, f"{label} — survival probability {probability:.1%}"

demo = gr.Interface(
    fn=predict_survival,
    inputs=[
        gr.Dropdown([1, 2, 3], value=3, label="Passenger class"),
        gr.Dropdown(["female", "male"], value="female", label="Sex"),
        gr.Number(value=29, label="Age"),
        gr.Number(value=0, precision=0, label="Siblings / spouses aboard"),
        gr.Number(value=0, precision=0, label="Parents / children aboard"),
        gr.Number(value=32.2, label="Fare"),
        gr.Dropdown(["C", "Q", "S"], value="S", label="Embarked port"),
    ],
    outputs=[
        gr.Label(label="Prediction"),
        gr.Textbox(label="Result"),
    ],
    title="Titanic Survival Predictor",
    description="Enter passenger details and estimate survival probability.",
)

demo.launch()

Check before continuing: python app.py starts a local Gradio URL without an import or model-loading error.

13

Run the app in your browser

In the terminal make sure your virtual environment is active, then run:

Start the appRunnable
python app.py

Gradio prints a local address, commonly http://127.0.0.1:7860. Hold Ctrl and click the URL, or copy it into your browser.

Check before continuing: The page shows Titanic Survival Predictor, seven passenger inputs and a Predict button.

14

Test two very different passengers

Test A: Class 1, female, age 38, SibSp 1, Parch 0, Fare 71.28, Embarked C.

Test B: Class 3, male, age 22, SibSp 1, Parch 0, Fare 7.25, Embarked S.

Click Predict for each case. The point is not to prove causation; it is to verify that the complete raw-input → preprocessing → model → probability path works.

Check before continuing: The two test passengers return visibly different survival probabilities.

15

Create a Git repository

Create a file named .gitignore and add .venv/, __pycache__/ and notebook checkpoint folders. Do not commit secrets.

Open GitHub → click New repository → name it titanic-survival-predictor → create the empty repository. Then run:

Push the project to GitHubRunnable
git init
git add .
git commit -m "Build Titanic survival predictor"

# Create an empty GitHub repository first, then copy its URL:
git branch -M main
git remote add origin https://github.com/YOUR-USERNAME/titanic-survival-predictor.git
git push -u origin main

Check before continuing: Refreshing the GitHub repository shows your source files and requirements.txt but not your .venv folder.

16

Explain the project in an interview

A strong explanation is: “I used the Titanic classification problem to build an end-to-end Scikit-learn pipeline. I split before preprocessing, imputed numeric and categorical features inside the pipeline, one-hot encoded categories, compared Logistic Regression with Random Forest, evaluated with confusion matrix and ROC-AUC, saved the full pipeline and exposed it through a Gradio app.”

Then discuss one improvement you would make next: feature engineering such as family size/title extraction, cross-validation, calibration or threshold analysis.

Check before continuing: You can explain the problem, leakage prevention, preprocessing, baseline, evaluation and deployment without reading the code.

Common problems and exact fixes

ProblemWhy it happensWhat to do
python is not recognizedPython was not added to PATHRe-run the installer or use the Python launcher py on Windows.
FileNotFoundError: data/train.csvCSV is in the wrong folderPut train.csv inside project-root/data and run from the project root.
ModuleNotFoundErrorVirtual environment is inactive or package not installedActivate .venv and run pip install -r requirements.txt.
Input columns do not matchApp fields differ from training feature namesUse exactly Pclass, Sex, Age, SibSp, Parch, Fare and Embarked.
App loads but prediction failsOnly the model was saved, not preprocessingSave and load the complete Scikit-learn Pipeline.

What this one project revises from the curriculum

PythonEDAMissing DataEncodingScalingClassificationModel EvaluationCross-ValidationFeature EngineeringDeployment

Screenshot standard for this handbook

We will only publish screenshots captured from the real tools used while running the project. No AI-generated fake IDE, notebook or application screenshots will be used. The screenshot pass will show the VS Code folder, Jupyter data inspection, training output and the running prediction app.