Skip to main content
Back to projects

Project 3 · Imbalanced classification

Can AI Catch a Stolen Credit Card Transaction?

Build a fraud-scoring system where 99.8% accuracy can still mean failure. You will learn why rare-event classification needs precision, recall, Average Precision, leakage-safe resampling and an explicit decision threshold.

The practical problem: which transactions should a fraud team investigate?

Imagine a payment company receives thousands of card transactions. Most are genuine. A tiny fraction are fraudulent. The system must assign each transaction a fraud score and decide which ones cross a review threshold.

The trap is that a model can predict “legitimate” for everything and still be correct about 99.83% of this historical dataset. That model would catch zero fraud. Our goal is therefore not maximum accuracy. We need useful fraud recall while keeping false alarms low enough to be operationally sensible.

Input

30 numerical signals: Time, Amount and anonymized PCA components V1–V28.

Output

A fraud score plus a decision: flag for review or do not flag.

Constraint

Only 492 of 284,807 transactions are fraud, so the classes are extremely imbalanced.

Deliverable

A saved model + threshold and a Streamlit review app using untouched holdout examples.

OpenML data
→
70/15/15 split
→
4 imbalance strategies
→
3-fold PR-based CV
→
Choose model
→
Tune threshold on validation
→
Final test once
→
Fraud-review app

The 99.83% accuracy trap — calculate it before building anything

There are 284,807 transactions: 284,315 legitimate and only 492 fraud. If we predict every row as legitimate:

accuracy = 284,315 ÷ 284,807 = 99.827%
fraud recall = 0 ÷ 492 = 0%

This one calculation explains the whole project: ordinary accuracy is not enough when the event we care about is rare.

See the imbalance before modeling

284,315 legitimate transactions versus 492 frauds

Real class-imbalance chart showing legitimate and fraud transaction counts on a log scale
Generated from the verified OpenML dataset. The logarithmic y-axis is necessary because the fraud bar would otherwise be almost invisible.
Open original-size screenshot (opens in a new tab)

Step 1: Understand the public fraud dataset and its privacy limits

We use OpenML dataset 1597 — creditcard. It contains transactions made by European cardholders over two days, with 492 fraud cases among 284,807 transactions.

The original sensitive transaction variables are not public. V1 through V28 are PCA-transformed numerical components. Only Time and Amount retain direct meanings. Class=1 means fraud.

That means this project can teach fraud-model engineering honestly, but it must not invent interpretations such as “V7 means merchant risk” or “V12 means cardholder age.” We simply do not know those original meanings.

Check before continuing

You can explain why V1–V28 are not given human-readable meanings.

Step 2: Create the project environment

requirements.txttextConfiguration
pandas==3.0.6
numpy==2.5.3
scipy==1.16.3
scikit-learn==1.9.1
imbalanced-learn==0.14.2
pyarrow==21.0.0
matplotlib==3.11.2
joblib==1.6.0
streamlit==1.65.0
pytest==9.0.2
Windows PowerShell setuppowershellRunnable
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install -r requirements.txt
python -m pip check

Check before continuing

The environment installs without dependency conflicts.

Step 3: Download the exact OpenML dataset reproducibly

The downloader first checks OpenML metadata, then downloads the official parquet file, verifies its SHA256 fingerprint, schema, row count and fraud count before saving it locally.

Open complete download_data.py
download_data.pypythonRunnable
"""Download and validate the public OpenML credit-card fraud dataset."""

from __future__ import annotations

from pathlib import Path
import hashlib
import json
import urllib.request

import pandas as pd

ROOT = Path(__file__).resolve().parent
DATA_DIR = ROOT / "data"
DATA_PATH = DATA_DIR / "creditcard.parquet"

OPENML_ID = 1597
METADATA_URL = f"https://www.openml.org/api/v1/json/data/{OPENML_ID}"
EXPECTED_NAME = "creditcard"
EXPECTED_VERSION = "1"
EXPECTED_LICENSE = "Public"
EXPECTED_ROWS = 284_807
EXPECTED_COLUMNS = 31
EXPECTED_FRAUDS = 492
EXPECTED_SHA256 = "b7efcb35a428bbe22347a05d2437d9177bab07ce61e51214a17bec584ad9496d"
EXPECTED_FEATURES = ["Time", *[f"V{i}" for i in range(1, 29)], "Amount", "Class"]


def download_bytes(url: str) -> bytes:
    request = urllib.request.Request(
        url,
        headers={"User-Agent": "LearnMLAcademy-credit-card-fraud-handbook/1.0"},
    )
    with urllib.request.urlopen(request, timeout=120) as response:
        return response.read()


def main() -> None:
    DATA_DIR.mkdir(parents=True, exist_ok=True)

    metadata = json.loads(download_bytes(METADATA_URL).decode("utf-8"))[
        "data_set_description"
    ]
    if metadata["name"] != EXPECTED_NAME:
        raise RuntimeError(f"Unexpected OpenML dataset name: {metadata['name']!r}")
    if str(metadata["version"]) != EXPECTED_VERSION:
        raise RuntimeError(f"Unexpected OpenML version: {metadata['version']!r}")
    if metadata.get("licence") != EXPECTED_LICENSE:
        raise RuntimeError(f"Unexpected OpenML licence: {metadata.get('licence')!r}")

    parquet_url = metadata.get("parquet_url")
    if not parquet_url:
        raise RuntimeError("OpenML metadata did not provide a parquet_url.")

    raw = download_bytes(parquet_url)
    sha256 = hashlib.sha256(raw).hexdigest()
    if sha256 != EXPECTED_SHA256:
        raise RuntimeError(
            "OpenML parquet fingerprint changed: "
            f"{sha256}; expected {EXPECTED_SHA256}. "
            "Do not continue until the dataset change is reviewed."
        )
    DATA_PATH.write_bytes(raw)

    frame = pd.read_parquet(DATA_PATH)
    if frame.shape != (EXPECTED_ROWS, EXPECTED_COLUMNS):
        raise RuntimeError(
            f"Unexpected shape {frame.shape}; expected "
            f"({EXPECTED_ROWS}, {EXPECTED_COLUMNS})."
        )
    if list(frame.columns) != EXPECTED_FEATURES:
        raise RuntimeError(
            "Unexpected column order/schema. Do not continue until reviewed."
        )

    frame["Class"] = frame["Class"].astype(int)
    frauds = int(frame["Class"].sum())
    if frauds != EXPECTED_FRAUDS:
        raise RuntimeError(f"Unexpected fraud count {frauds}; expected {EXPECTED_FRAUDS}.")
    if set(frame["Class"].unique()) != {0, 1}:
        raise RuntimeError("Class must contain only 0 and 1.")

    print(f"OpenML dataset ID: {OPENML_ID}")
    print(f"Rows: {len(frame):,}")
    print(f"Columns: {frame.shape[1]}")
    print(f"Fraud transactions: {frauds}")
    print(f"Fraud rate: {frauds / len(frame):.6%}")
    print(f"SHA256: {sha256}")
    print(f"Saved: {DATA_PATH}")


if __name__ == "__main__":
    main()
Download and verifypowershellRunnable
python download_data.py
Verified checkpointtextOutput
Rows: 284,807
Columns: 31
Fraud transactions: 492
Fraud rate: 0.172749%
SHA256: b7efcb35a428bbe22347a05d2437d9177bab07ce61e51214a17bec584ad9496d

Check before continuing

The downloader reports 284,807 rows, 492 frauds and the expected SHA256 fingerprint.

Step 4: Create train, validation and final-test sets before experimentation

We need three sets because threshold selection is itself a modeling decision. Training data learns model parameters; validation data chooses the decision threshold; the final test set remains sealed until everything is fixed.

Leakage-safe 70 / 15 / 15 splitpythonRunnable
X_build, X_test, y_build, y_test = train_test_split(
    X, y,
    test_size=0.15,
    stratify=y,
    random_state=42,
)

X_train, X_val, y_train, y_val = train_test_split(
    X_build, y_build,
    test_size=0.15 / 0.85,
    stratify=y_build,
    random_state=42,
)

# 70% train | 15% validation | 15% final test
Verified splittextOutput
Train: 199,364 rows / 344 frauds
Validation: 42,721 rows / 74 frauds
Final test: 42,722 rows / 74 frauds

Check before continuing

The split is 199,364 / 42,721 / 42,722 rows and fraud prevalence stays close across all three parts.

Step 5: Use metrics that can see the rare class

Precision = TP / (TP + FP)
When we flag a transaction, how often is it actually fraud?
Recall = TP / (TP + FN)
Of all actual frauds, how many did we catch?
F1
A harmonic balance between precision and recall at one chosen threshold.
Average Precision
Summarizes the precision-recall trade-off across thresholds and is especially useful for highly imbalanced binary problems.

Check before continuing

You can explain precision, recall and Average Precision without relying on accuracy.

Step 6: Compare four ways to handle the imbalance

StrategyWhat changesWhat it teaches
Logistic RegressionNo special balancingA clean probability baseline
Class-weighted LogisticFraud mistakes receive more training weightCost-sensitive learning
SMOTE + LogisticSynthetic minority examples are generated inside training foldsResampling without leakage
Class-weighted Random ForestWeighted ensemble of decision treesNon-linear ensemble learning

Check before continuing

You can explain what changes between the plain, class-weighted, SMOTE and ensemble strategies.

Step 7: Use SMOTE inside the pipeline — never before the split

SMOTE creates synthetic minority examples by interpolating between existing fraud examples. If you SMOTE before splitting, information derived from a transaction can leak into validation/test data and make evaluation too optimistic.

Correct leakage-safe patternpythonRunnable
smote_logistic = ImbPipeline([
    ("scale", StandardScaler()),
    ("smote", SMOTE(
        sampling_strategy=0.10,
        random_state=42,
        k_neighbors=5,
    )),
    ("model", LogisticRegression(max_iter=1500)),
])

# During cross-validation SMOTE runs only inside each training fold.
# Validation/test rows are never synthetically oversampled.

Check before continuing

You can explain why oversampling the full dataset would contaminate evaluation.

Step 8: Build the training program in logical pieces, then assemble it

The complete file is intentionally shown only after the architecture is clear. Read it as six blocks: load → split → candidate models → cross-validation → validation threshold → final evaluation/save.

Open complete verified src/train_model.py
src/train_model.pypythonRunnable
"""Train the fraud detector from the project root: python src/train_model.py."""

from __future__ import annotations

from pathlib import Path
import json
import platform

import joblib
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import sklearn
from imblearn.over_sampling import SMOTE
from imblearn.pipeline import Pipeline as ImbPipeline
from sklearn.base import clone
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (
    ConfusionMatrixDisplay,
    accuracy_score,
    average_precision_score,
    confusion_matrix,
    f1_score,
    precision_recall_curve,
    precision_score,
    recall_score,
    roc_auc_score,
)
from sklearn.model_selection import StratifiedKFold, cross_validate, train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

ROOT = Path(__file__).resolve().parents[1]
DATA_PATH = ROOT / "data" / "creditcard.parquet"
MODEL_DIR = ROOT / "models"
OUTPUT_DIR = ROOT / "outputs"

FEATURES = ["Time", *[f"V{i}" for i in range(1, 29)], "Amount"]
TARGET = "Class"
SEED = 42
TARGET_RECALL = 0.80


def load_data() -> tuple[pd.DataFrame, pd.Series]:
    if not DATA_PATH.is_file():
        raise FileNotFoundError(
            "Dataset missing. Run python download_data.py from the project root."
        )
    frame = pd.read_parquet(DATA_PATH)
    expected = FEATURES + [TARGET]
    if list(frame.columns) != expected:
        raise ValueError("Dataset schema changed.")
    y = frame[TARGET].astype(int)
    if len(frame) != 284_807 or int(y.sum()) != 492:
        raise ValueError("Dataset counts changed.")
    return frame[FEATURES].astype(float), y


def split_data(X: pd.DataFrame, y: pd.Series):
    X_build, X_test, y_build, y_test = train_test_split(
        X,
        y,
        test_size=0.15,
        stratify=y,
        random_state=SEED,
    )
    validation_fraction_of_build = 0.15 / 0.85
    X_train, X_val, y_train, y_val = train_test_split(
        X_build,
        y_build,
        test_size=validation_fraction_of_build,
        stratify=y_build,
        random_state=SEED,
    )
    return X_train, X_val, X_test, y_train, y_val, y_test


def candidate_models():
    logistic = Pipeline(
        [
            ("scale", StandardScaler()),
            (
                "model",
                LogisticRegression(
                    max_iter=1500,
                    solver="lbfgs",
                    random_state=SEED,
                ),
            ),
        ]
    )
    weighted_logistic = Pipeline(
        [
            ("scale", StandardScaler()),
            (
                "model",
                LogisticRegression(
                    max_iter=1500,
                    solver="lbfgs",
                    class_weight="balanced",
                    random_state=SEED,
                ),
            ),
        ]
    )
    smote_logistic = ImbPipeline(
        [
            ("scale", StandardScaler()),
            (
                "smote",
                SMOTE(
                    sampling_strategy=0.10,
                    random_state=SEED,
                    k_neighbors=5,
                ),
            ),
            (
                "model",
                LogisticRegression(
                    max_iter=1500,
                    solver="lbfgs",
                    random_state=SEED,
                ),
            ),
        ]
    )
    forest = RandomForestClassifier(
        n_estimators=120,
        max_depth=12,
        min_samples_leaf=2,
        class_weight="balanced_subsample",
        n_jobs=-1,
        random_state=SEED,
    )
    return {
        "Logistic Regression": logistic,
        "Class-weighted Logistic Regression": weighted_logistic,
        "SMOTE + Logistic Regression": smote_logistic,
        "Class-weighted Random Forest": forest,
    }


def compare_models(X_train: pd.DataFrame, y_train: pd.Series) -> pd.DataFrame:
    cv = StratifiedKFold(n_splits=3, shuffle=True, random_state=SEED)
    rows = []
    for name, model in candidate_models().items():
        print(f"Cross-validating {name} ...", flush=True)
        scores = cross_validate(
            clone(model),
            X_train,
            y_train,
            cv=cv,
            scoring={
                "ap": "average_precision",
                "roc_auc": "roc_auc",
                "precision": "precision",
                "recall": "recall",
                "f1": "f1",
            },
            n_jobs=1,
            error_score="raise",
        )
        rows.append(
            {
                "model": name,
                "cv_average_precision": float(scores["test_ap"].mean()),
                "cv_roc_auc": float(scores["test_roc_auc"].mean()),
                "cv_precision_at_0_5": float(scores["test_precision"].mean()),
                "cv_recall_at_0_5": float(scores["test_recall"].mean()),
                "cv_f1_at_0_5": float(scores["test_f1"].mean()),
            }
        )
    result = pd.DataFrame(rows).sort_values(
        "cv_average_precision",
        ascending=False,
        kind="stable",
    )
    result.to_csv(OUTPUT_DIR / "model_comparison.csv", index=False)
    return result


def choose_threshold(
    y_val: pd.Series,
    scores: np.ndarray,
    target_recall: float = TARGET_RECALL,
) -> tuple[float, pd.DataFrame]:
    precision, recall, thresholds = precision_recall_curve(y_val, scores)
    table = pd.DataFrame(
        {
            "threshold": thresholds,
            "precision": precision[:-1],
            "recall": recall[:-1],
        }
    )
    table["f1"] = (
        2 * table["precision"] * table["recall"]
        / (table["precision"] + table["recall"] + 1e-12)
    )

    feasible = table[table["recall"] >= target_recall]
    if not feasible.empty:
        chosen = feasible.sort_values(
            ["precision", "threshold"],
            ascending=[False, False],
            kind="stable",
        ).iloc[0]
    else:
        chosen = table.sort_values("f1", ascending=False, kind="stable").iloc[0]

    table.to_csv(OUTPUT_DIR / "validation_thresholds.csv", index=False)
    return float(chosen["threshold"]), table


def metrics_at_threshold(
    y_true: pd.Series,
    scores: np.ndarray,
    threshold: float,
) -> tuple[dict, np.ndarray]:
    predictions = (scores >= threshold).astype(int)
    matrix = confusion_matrix(y_true, predictions, labels=[0, 1])
    metrics = {
        "accuracy": float(accuracy_score(y_true, predictions)),
        "precision": float(precision_score(y_true, predictions, zero_division=0)),
        "recall": float(recall_score(y_true, predictions, zero_division=0)),
        "f1": float(f1_score(y_true, predictions, zero_division=0)),
        "average_precision": float(average_precision_score(y_true, scores)),
        "roc_auc": float(roc_auc_score(y_true, scores)),
    }
    return metrics, matrix


def save_figures(
    y: pd.Series,
    comparison: pd.DataFrame,
    threshold_table: pd.DataFrame,
    chosen_threshold: float,
    y_test: pd.Series,
    test_scores: np.ndarray,
    matrix: np.ndarray,
) -> None:
    counts = y.value_counts().sort_index()
    labels = ["Legitimate (0)", "Fraud (1)"]
    values = [int(counts.get(0, 0)), int(counts.get(1, 0))]
    fig, ax = plt.subplots(figsize=(8, 5))
    bars = ax.bar(labels, values)
    ax.set_yscale("log")
    ax.set_ylabel("Transactions (log scale)")
    ax.set_title("Why fraud detection is an imbalanced classification problem")
    for bar, value in zip(bars, values):
        pct = value / len(y) * 100
        ax.text(
            bar.get_x() + bar.get_width() / 2,
            value,
            f"{value:,}\n{pct:.3f}%",
            ha="center",
            va="bottom",
        )
    fig.tight_layout()
    fig.savefig(OUTPUT_DIR / "class_imbalance.png", dpi=160)
    plt.close(fig)

    fig, ax = plt.subplots(figsize=(9, 5))
    ax.barh(comparison["model"], comparison["cv_average_precision"])
    ax.invert_yaxis()
    ax.set_xlim(0, 1)
    ax.set_xlabel("Mean 3-fold Average Precision")
    ax.set_title("Training-only model comparison")
    fig.tight_layout()
    fig.savefig(OUTPUT_DIR / "model_comparison.png", dpi=160)
    plt.close(fig)

    sampled = threshold_table.iloc[:: max(1, len(threshold_table) // 500)].copy()
    fig, ax = plt.subplots(figsize=(9, 5))
    ax.plot(sampled["threshold"], sampled["precision"], label="Precision")
    ax.plot(sampled["threshold"], sampled["recall"], label="Recall")
    ax.axvline(
        chosen_threshold,
        linestyle="--",
        label=f"Chosen threshold = {chosen_threshold:.3f}",
    )
    ax.set(
        xlabel="Fraud-score threshold",
        ylabel="Metric value",
        ylim=(0, 1.02),
    )
    ax.set_title("Validation-only threshold trade-off")
    ax.legend()
    fig.tight_layout()
    fig.savefig(OUTPUT_DIR / "threshold_tradeoff.png", dpi=160)
    plt.close(fig)

    precision, recall, _ = precision_recall_curve(y_test, test_scores)
    fig, ax = plt.subplots(figsize=(7, 5))
    ax.plot(recall, precision)
    ax.set(
        xlabel="Recall",
        ylabel="Precision",
        xlim=(0, 1),
        ylim=(0, 1.02),
    )
    ax.set_title("Final untouched-test precision-recall curve")
    fig.tight_layout()
    fig.savefig(OUTPUT_DIR / "precision_recall_curve.png", dpi=160)
    plt.close(fig)

    fig, ax = plt.subplots(figsize=(7, 5))
    ConfusionMatrixDisplay(
        matrix,
        display_labels=["Legitimate", "Fraud"],
    ).plot(ax=ax, colorbar=False, values_format="d")
    ax.set_title("Final untouched-test confusion matrix")
    fig.tight_layout()
    fig.savefig(OUTPUT_DIR / "confusion_matrix.png", dpi=160)
    plt.close(fig)


def save_demo_transactions(
    bundle: dict,
    X_test: pd.DataFrame,
    y_test: pd.Series,
) -> None:
    scores = bundle["model"].predict_proba(X_test)[:, 1]
    demo = X_test.copy()
    demo["historical_label"] = y_test.to_numpy()
    demo["fraud_score"] = scores

    fraud_examples = demo[demo["historical_label"] == 1].nlargest(6, "fraud_score")
    legitimate_examples = demo[demo["historical_label"] == 0].nsmallest(6, "fraud_score")
    selected = pd.concat(
        [fraud_examples, legitimate_examples],
        ignore_index=True,
    )
    selected.insert(
        0,
        "example_id",
        [f"TX-{index + 1:02d}" for index in range(len(selected))],
    )
    selected.to_csv(OUTPUT_DIR / "demo_transactions.csv", index=False)


def main() -> None:
    MODEL_DIR.mkdir(exist_ok=True)
    OUTPUT_DIR.mkdir(exist_ok=True)

    X, y = load_data()
    X_train, X_val, X_test, y_train, y_val, y_test = split_data(X, y)

    print(
        f"Rows -> train: {len(X_train):,}, validation: {len(X_val):,}, "
        f"test: {len(X_test):,}"
    )
    print(
        f"Fraud rows -> train: {int(y_train.sum())}, validation: {int(y_val.sum())}, "
        f"test: {int(y_test.sum())}"
    )
    print(f"Always-legitimate accuracy: {(y == 0).mean():.6%}")

    comparison = compare_models(X_train, y_train)
    print("\nTraining-only model comparison:")
    print(comparison.to_string(index=False))

    winner_name = str(comparison.iloc[0]["model"])
    winner = clone(candidate_models()[winner_name])
    winner.fit(X_train, y_train)

    val_scores = winner.predict_proba(X_val)[:, 1]
    threshold, threshold_table = choose_threshold(y_val, val_scores)
    val_metrics, _ = metrics_at_threshold(y_val, val_scores, threshold)

    X_build = pd.concat([X_train, X_val], axis=0)
    y_build = pd.concat([y_train, y_val], axis=0)
    final_model = clone(candidate_models()[winner_name])
    final_model.fit(X_build, y_build)

    test_scores = final_model.predict_proba(X_test)[:, 1]
    test_metrics, matrix = metrics_at_threshold(y_test, test_scores, threshold)

    bundle = {
        "model": final_model,
        "threshold": threshold,
        "features": FEATURES,
        "selected_model": winner_name,
        "target_recall": TARGET_RECALL,
    }
    joblib.dump(bundle, MODEL_DIR / "fraud_detector.joblib")

    reloaded = joblib.load(MODEL_DIR / "fraud_detector.joblib")
    reload_scores = reloaded["model"].predict_proba(X_test.iloc[:25])[:, 1]
    if not np.allclose(reload_scores, test_scores[:25]):
        raise RuntimeError("Reloaded fraud scores differ from in-memory scores.")

    save_demo_transactions(bundle, X_test, y_test)
    save_figures(
        y,
        comparison,
        threshold_table,
        threshold,
        y_test,
        test_scores,
        matrix,
    )

    report = {
        "python": platform.python_version(),
        "scikit_learn": sklearn.__version__,
        "rows": int(len(X)),
        "fraud_rows": int(y.sum()),
        "fraud_rate": float(y.mean()),
        "features": FEATURES,
        "seed": SEED,
        "train_rows": int(len(X_train)),
        "validation_rows": int(len(X_val)),
        "test_rows": int(len(X_test)),
        "selected_model": winner_name,
        "selection_metric": "mean 3-fold average precision on training data",
        "target_validation_recall": TARGET_RECALL,
        "chosen_threshold": threshold,
        "validation_metrics": val_metrics,
        "test_metrics": test_metrics,
        "confusion_matrix": matrix.tolist(),
        "always_legitimate_accuracy": float((y_test == 0).mean()),
    }
    (OUTPUT_DIR / "metrics.json").write_text(
        json.dumps(report, indent=2) + "\n",
        encoding="utf-8",
    )

    print(f"\nSelected model: {winner_name}")
    print(f"Chosen validation threshold: {threshold:.6f}")
    print("Final untouched-test metrics:")
    for name, value in test_metrics.items():
        print(f"  {name}: {value:.6f}")
    print("Confusion matrix [[TN, FP], [FN, TP]]:")
    print(matrix)
    print("Saved: models/fraud_detector.joblib")


if __name__ == "__main__":
    main()

Check before continuing

You can trace data → candidates → CV → threshold → final test → saved bundle.

Step 9: Select the model by training-only Average Precision

ModelAvg PrecisionROC-AUCPrecision @ .5Recall @ .5F1 @ .5
Class-weighted Random Forest0.83600.97180.88630.75570.8152
SMOTE + Logistic Regression0.77000.97780.34790.85460.4944
Logistic Regression0.76980.98000.86950.65970.7482
Class-weighted Logistic Regression0.76930.97990.06640.91270.1239

Training-only model comparison

Real bar chart comparing four fraud-detection models by training-only Average Precision
The class-weighted Random Forest ranked first by the metric we chose before looking at the final test set.
Open original-size screenshot (opens in a new tab)

Notice the trade-off: class-weighted Logistic Regression reaches very high recall at the default 0.5 threshold but precision collapses to about 6.6%. Catching more fraud is useful only if the false-alarm cost remains acceptable.

Check before continuing

Class-weighted Random Forest has the highest mean 3-fold Average Precision: 0.8360.

Step 10: Tune the decision threshold on validation data, not the final test

A probability model gives a score. The threshold converts that score into an operational decision. We ask for at least 80% recall on the validation set, then choose the threshold with the highest precision among those feasible points.

Validation-only threshold rulepythonRunnable
precision, recall, thresholds = precision_recall_curve(
    y_val,
    validation_scores,
)

# Among thresholds that reach at least 80% recall on validation,
# choose the one with the highest precision.
feasible = table[table["recall"] >= 0.80]
chosen = feasible.sort_values(
    ["precision", "threshold"],
    ascending=[False, False],
).iloc[0]

Changing the threshold changes the business trade-off

Real validation-only precision and recall curves across fraud-score thresholds
Lower thresholds generally catch more fraud but also create more false alarms. The dashed line marks the chosen validation threshold.
Open original-size screenshot (opens in a new tab)

Important: meeting an 80% recall target on validation does not guarantee 80% recall on future/test data. The untouched test recall became 74.3%. That difference is exactly why we keep a final test set.

Check before continuing

The chosen threshold is 0.673629 and you understand why 0.5 is not sacred.

Step 11: Open the sealed test set once and interpret every error

Final untouched-test confusion matrix

Real final fraud test confusion matrix with 42641 true negatives, 7 false positives, 19 false negatives and 55 true positives
42,722 transactions were evaluated only after model family and threshold had already been fixed.
Open original-size screenshot (opens in a new tab)

Final precision-recall curve

Real precision-recall curve on the final untouched fraud test set
Average Precision on the untouched test set was 0.7874.
Open original-size screenshot (opens in a new tab)
Precision = 55 / (55 + 7) = 88.71%.
Recall = 55 / (55 + 19) = 74.32%.
F1 = 0.8088.
ROC-AUC = 0.9643, while Average Precision = 0.7874.

Check before continuing

You can derive precision and recall from 42,641 TN, 7 FP, 19 FN and 55 TP.

Step 12: Turn the model + threshold into a fraud-review application

A user does not type V1–V28 from memory. In a real payment system those features would arrive from the transaction pipeline. Our educational app therefore lets you select a verified holdout example, inspect its anonymized inputs, score it and compare the decision with the historical label.

Open complete app.py
app.pypythonRunnable
"""Run with: python -m streamlit run app.py"""

from __future__ import annotations

from pathlib import Path

import joblib
import pandas as pd
import streamlit as st

ROOT = Path(__file__).resolve().parent
MODEL_PATH = ROOT / "models" / "fraud_detector.joblib"
DEMO_PATH = ROOT / "outputs" / "demo_transactions.csv"

st.set_page_config(
    page_title="Credit Card Fraud Detector",
    page_icon="💳",
    layout="wide",
)
st.title("💳 Can AI Catch a Stolen Credit Card Transaction?")
st.caption(
    "Educational fraud-scoring demo using the public anonymized OpenML credit-card dataset"
)
st.info(
    "This is not a banking or payment-security product. The historical dataset is anonymized, "
    "and V1–V28 are PCA-transformed features whose original meanings are not public."
)


@st.cache_resource
def load_bundle():
    return joblib.load(MODEL_PATH)


if not MODEL_PATH.is_file() or not DEMO_PATH.is_file():
    st.error("Training artifacts are missing.")
    st.code(
        "python download_data.py\npython src/train_model.py",
        language="powershell",
    )
    st.stop()

bundle = load_bundle()
demos = pd.read_csv(DEMO_PATH)
features = bundle["features"]
threshold = float(bundle["threshold"])

st.subheader("Score a verified holdout example")
selected_id = st.selectbox(
    "Transaction example",
    demos["example_id"].tolist(),
)
row = demos.loc[demos["example_id"] == selected_id].iloc[0]

left, right, third = st.columns(3)
left.metric("Transaction amount", f"{float(row['Amount']):,.2f}")
right.metric("Seconds from dataset start", f"{float(row['Time']):,.0f}")
third.metric("Decision threshold", f"{threshold:.3f}")

with st.expander("See all anonymized model inputs"):
    st.dataframe(
        pd.DataFrame([row[features].to_dict()]),
        hide_index=True,
        use_container_width=True,
    )

if st.button("Score transaction", type="primary"):
    model_row = pd.DataFrame(
        [[float(row[name]) for name in features]],
        columns=features,
    )
    score = float(bundle["model"].predict_proba(model_row)[0, 1])
    flagged = score >= threshold

    if flagged:
        st.error("Model decision: FLAG FOR FRAUD REVIEW")
    else:
        st.success("Model decision: do not flag at this threshold")

    metric_col, truth_col = st.columns(2)
    metric_col.metric("Fraud score", f"{score:.1%}")
    historical_label = int(row["historical_label"])
    truth_col.metric(
        "Historical label",
        "Fraud" if historical_label == 1 else "Legitimate",
    )

    if flagged and historical_label == 1:
        st.write(
            "This example is a **true positive**: the model flagged a transaction "
            "historically labelled fraud."
        )
    elif flagged and historical_label == 0:
        st.write(
            "This example is a **false positive**: a legitimate historical "
            "transaction was flagged."
        )
    elif not flagged and historical_label == 1:
        st.write(
            "This example is a **false negative**: a historical fraud was missed."
        )
    else:
        st.write(
            "This example is a **true negative**: a legitimate transaction was not flagged."
        )

st.divider()
st.caption(
    "Real fraud systems use richer live signals, cost-sensitive decisions, monitoring, "
    "human review, security controls and continuously changing adversarial patterns."
)
Start the apppowershellRunnable
python -m streamlit run app.py

Real fraud-review app

Real Streamlit credit-card fraud scoring form using a verified holdout transaction
The app displays Amount, Time and the learned decision threshold before scoring.
Open original-size screenshot (opens in a new tab)

Flagged holdout example

Real Streamlit fraud example flagged for review
A real holdout fraud example scored by the saved model and threshold.
Open original-size screenshot (opens in a new tab)

Legitimate holdout example

Real Streamlit legitimate example not flagged at the fraud threshold
A contrasting legitimate holdout example follows the same inference path.
Open original-size screenshot (opens in a new tab)

Check before continuing

The app scores a real holdout example and explains TP, FP, FN or TN.

Step 13: Run automated checks so screenshots are not your only proof

The suite checks the exact dataset contract, split separation, class prevalence, four candidate strategies, threshold logic, saved artifact/report consistency, probability-aware final metrics and both historical classes in the demo set.

Open tests/test_fraud_detector.py
tests/test_fraud_detector.pypythonRunnable
from pathlib import Path
import importlib.util
import json

import joblib
import numpy as np
import pandas as pd
import pytest

ROOT = Path(__file__).resolve().parents[1]
DATA = ROOT / "data" / "creditcard.parquet"
MODEL = ROOT / "models" / "fraud_detector.joblib"
METRICS = ROOT / "outputs" / "metrics.json"
DEMOS = ROOT / "outputs" / "demo_transactions.csv"

spec = importlib.util.spec_from_file_location(
    "fraud_train",
    ROOT / "src" / "train_model.py",
)
train = importlib.util.module_from_spec(spec)
spec.loader.exec_module(train)


def test_dataset_contract():
    frame = pd.read_parquet(DATA)
    assert frame.shape == (284_807, 31)
    assert list(frame.columns) == train.FEATURES + [train.TARGET]
    assert int(frame["Class"].astype(int).sum()) == 492
    assert set(frame["Class"].astype(int).unique()) == {0, 1}


def test_three_way_split_is_disjoint_and_stratified():
    X, y = train.load_data()
    X_train, X_val, X_test, y_train, y_val, y_test = train.split_data(X, y)
    assert len(X_train) + len(X_val) + len(X_test) == len(X)
    assert set(X_train.index).isdisjoint(X_val.index)
    assert set(X_train.index).isdisjoint(X_test.index)
    assert set(X_val.index).isdisjoint(X_test.index)
    for part in [y_train, y_val, y_test]:
        assert 0 < int(part.sum()) < len(part)
        assert abs(float(part.mean()) - float(y.mean())) < 0.0005


def test_candidates_include_imbalance_strategies():
    candidates = train.candidate_models()
    assert set(candidates) == {
        "Logistic Regression",
        "Class-weighted Logistic Regression",
        "SMOTE + Logistic Regression",
        "Class-weighted Random Forest",
    }


def test_threshold_selection_meets_target_when_feasible():
    y = pd.Series([0, 0, 0, 1, 1])
    scores = np.array([0.05, 0.10, 0.20, 0.60, 0.90])
    threshold, table = train.choose_threshold(y, scores, target_recall=1.0)
    row = table.loc[np.isclose(table["threshold"], threshold)].iloc[0]
    assert row["recall"] >= 1.0
    assert 0 <= threshold <= 1


def test_saved_artifact_and_report_exist():
    assert MODEL.is_file()
    assert METRICS.is_file()
    bundle = joblib.load(MODEL)
    report = json.loads(METRICS.read_text(encoding="utf-8"))
    assert 0 < float(bundle["threshold"]) < 1
    assert bundle["selected_model"] == report["selected_model"]
    assert report["selection_metric"].startswith("mean 3-fold average precision")


def test_final_metrics_are_probability_aware():
    report = json.loads(METRICS.read_text(encoding="utf-8"))
    metrics = report["test_metrics"]
    assert 0 <= metrics["average_precision"] <= 1
    assert 0 <= metrics["roc_auc"] <= 1
    assert 0 <= metrics["precision"] <= 1
    assert 0 <= metrics["recall"] <= 1
    assert report["test_rows"] > 40_000


def test_demo_transactions_cover_both_historical_classes():
    demos = pd.read_csv(DEMOS)
    assert len(demos) == 12
    assert set(demos["historical_label"].astype(int)) == {0, 1}
    assert set(train.FEATURES).issubset(demos.columns)


@pytest.mark.parametrize("example_index", [0, 6])
def test_reloaded_model_scores_demo_rows(example_index):
    bundle = joblib.load(MODEL)
    demos = pd.read_csv(DEMOS)
    row = demos.iloc[[example_index]][bundle["features"]]
    score = float(bundle["model"].predict_proba(row)[0, 1])
    assert 0 <= score <= 1
Run testspowershellRunnable
pytest -q

Check before continuing

pytest finishes without failure.

Step 14: Know what SMOTE and supervised classification do not solve

Fraud patterns change because attackers adapt. The dataset is historical and anonymized, has only two days of transactions, and does not expose customer/device/merchant/network context. A production system would need temporal validation, drift monitoring, calibrated decision costs, security controls and human-review workflows.

Anomaly detection is another family of techniques. For example, Isolation Forest looks for observations that are easier to isolate than normal ones. It can help when labels are scarce, but “unusual” is not automatically the same as “fraud,” so it is an extension rather than a magic replacement.

Check before continuing

You can name at least two reasons a production fraud system needs more than this model.

Step 15: Understand the complete project folder

credit-card-fraud/
├── data/creditcard.parquet # downloaded locally
├── models/fraud_detector.joblib # generated model + threshold
├── outputs/ # metrics, plots, demo rows
├── scripts/capture_app_screenshots.py
├── src/train_model.py
├── tests/test_fraud_detector.py
├── app.py
├── download_data.py
├── requirements.txt
└── README.md

Check before continuing

You can point to the file responsible for data acquisition, training, tests and browser inference.

Now change the project yourself

Exercise 1: Change the validation recall target from 80% to 90%. Record what happens to precision and false positives.

Exercise 2: Remove SMOTE and compare class weighting alone. Explain why resampling is not automatically better.

Exercise 3: Try an Isolation Forest as an anomaly-detection extension and compare its ranking with the supervised model.

Exercise 4: Define a simple cost where missing fraud costs 20× more than reviewing a legitimate transaction. Choose a threshold using cost instead of recall.

How to explain this project in an interview

Why is 99.8% accuracy useless here?

Because the majority class is so dominant that predicting every transaction as legitimate achieves that accuracy while fraud recall is zero.

Why Average Precision?

It evaluates the precision-recall ranking across thresholds and focuses attention on performance for the rare positive class.

Why is SMOTE inside a pipeline?

So synthetic minority samples are created only from each training fold; validation and test rows remain natural and uncontaminated.

Why have validation and test separately?

Threshold selection uses validation data. The final test must remain untouched so it measures the completed decision process.

Why did validation target recall not equal test recall?

A threshold chosen on one finite sample will not reproduce identical metrics on unseen data. That generalization gap is real evidence, not an error to hide.

Implementation mastery check

Why does always-legitimate prediction reach about 99.83% accuracy?
What is the difference between precision and recall in fraud review?
Why is PR-oriented evaluation especially useful here?
What does class_weight="balanced" change?
What exactly does SMOTE create?
Why must SMOTE stay inside training folds?
Why do we need a validation set for threshold selection?
What does a threshold of 0.674 mean operationally?
How do 55 TP, 7 FP and 19 FN produce the final precision and recall?
Why can validation recall be 80% while test recall is 74.3%?
What is stored inside fraud_detector.joblib?
Why are V1–V28 not given business meanings?
When might anomaly detection be useful?
What would you add before deploying fraud detection at a real bank?

Complete-project checkpoint

Public OpenML dataset fingerprint verified
Extreme class imbalance visualized
70/15/15 stratified split created
Accuracy trap calculated
Four imbalance strategies compared
SMOTE kept inside training folds
Model selected by training-only Average Precision
Threshold selected on validation only
Final test opened once
Precision/recall derived from confusion counts
Model + threshold saved and reloaded
Streamlit app scores real holdout examples
Automated tests verify contracts
You can explain limitations and extensions