Project 3 · Imbalanced classification
Can AI Catch a Stolen Credit Card Transaction?
Build a fraud-scoring system where 99.8% accuracy can still mean failure. You will learn why rare-event classification needs precision, recall, Average Precision, leakage-safe resampling and an explicit decision threshold.
The practical problem: which transactions should a fraud team investigate?
Imagine a payment company receives thousands of card transactions. Most are genuine. A tiny fraction are fraudulent. The system must assign each transaction a fraud score and decide which ones cross a review threshold.
The trap is that a model can predict “legitimate” for everything and still be correct about 99.83% of this historical dataset. That model would catch zero fraud. Our goal is therefore not maximum accuracy. We need useful fraud recall while keeping false alarms low enough to be operationally sensible.
Input
30 numerical signals: Time, Amount and anonymized PCA components V1–V28.
Output
A fraud score plus a decision: flag for review or do not flag.
Constraint
Only 492 of 284,807 transactions are fraud, so the classes are extremely imbalanced.
Deliverable
A saved model + threshold and a Streamlit review app using untouched holdout examples.
The 99.83% accuracy trap — calculate it before building anything
There are 284,807 transactions: 284,315 legitimate and only 492 fraud. If we predict every row as legitimate:
fraud recall = 0 ÷ 492 = 0%
This one calculation explains the whole project: ordinary accuracy is not enough when the event we care about is rare.
See the imbalance before modeling
284,315 legitimate transactions versus 492 frauds
Step 1: Understand the public fraud dataset and its privacy limits
We use OpenML dataset 1597 — creditcard. It contains transactions made by European cardholders over two days, with 492 fraud cases among 284,807 transactions.
The original sensitive transaction variables are not public. V1 through V28 are PCA-transformed numerical components. Only Time and Amount retain direct meanings. Class=1 means fraud.
That means this project can teach fraud-model engineering honestly, but it must not invent interpretations such as “V7 means merchant risk” or “V12 means cardholder age.” We simply do not know those original meanings.
Check before continuing
Step 2: Create the project environment
pandas==3.0.6
numpy==2.5.3
scipy==1.16.3
scikit-learn==1.9.1
imbalanced-learn==0.14.2
pyarrow==21.0.0
matplotlib==3.11.2
joblib==1.6.0
streamlit==1.65.0
pytest==9.0.2
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install -r requirements.txt
python -m pip checkCheck before continuing
Step 3: Download the exact OpenML dataset reproducibly
The downloader first checks OpenML metadata, then downloads the official parquet file, verifies its SHA256 fingerprint, schema, row count and fraud count before saving it locally.
Open complete download_data.py
"""Download and validate the public OpenML credit-card fraud dataset."""
from __future__ import annotations
from pathlib import Path
import hashlib
import json
import urllib.request
import pandas as pd
ROOT = Path(__file__).resolve().parent
DATA_DIR = ROOT / "data"
DATA_PATH = DATA_DIR / "creditcard.parquet"
OPENML_ID = 1597
METADATA_URL = f"https://www.openml.org/api/v1/json/data/{OPENML_ID}"
EXPECTED_NAME = "creditcard"
EXPECTED_VERSION = "1"
EXPECTED_LICENSE = "Public"
EXPECTED_ROWS = 284_807
EXPECTED_COLUMNS = 31
EXPECTED_FRAUDS = 492
EXPECTED_SHA256 = "b7efcb35a428bbe22347a05d2437d9177bab07ce61e51214a17bec584ad9496d"
EXPECTED_FEATURES = ["Time", *[f"V{i}" for i in range(1, 29)], "Amount", "Class"]
def download_bytes(url: str) -> bytes:
request = urllib.request.Request(
url,
headers={"User-Agent": "LearnMLAcademy-credit-card-fraud-handbook/1.0"},
)
with urllib.request.urlopen(request, timeout=120) as response:
return response.read()
def main() -> None:
DATA_DIR.mkdir(parents=True, exist_ok=True)
metadata = json.loads(download_bytes(METADATA_URL).decode("utf-8"))[
"data_set_description"
]
if metadata["name"] != EXPECTED_NAME:
raise RuntimeError(f"Unexpected OpenML dataset name: {metadata['name']!r}")
if str(metadata["version"]) != EXPECTED_VERSION:
raise RuntimeError(f"Unexpected OpenML version: {metadata['version']!r}")
if metadata.get("licence") != EXPECTED_LICENSE:
raise RuntimeError(f"Unexpected OpenML licence: {metadata.get('licence')!r}")
parquet_url = metadata.get("parquet_url")
if not parquet_url:
raise RuntimeError("OpenML metadata did not provide a parquet_url.")
raw = download_bytes(parquet_url)
sha256 = hashlib.sha256(raw).hexdigest()
if sha256 != EXPECTED_SHA256:
raise RuntimeError(
"OpenML parquet fingerprint changed: "
f"{sha256}; expected {EXPECTED_SHA256}. "
"Do not continue until the dataset change is reviewed."
)
DATA_PATH.write_bytes(raw)
frame = pd.read_parquet(DATA_PATH)
if frame.shape != (EXPECTED_ROWS, EXPECTED_COLUMNS):
raise RuntimeError(
f"Unexpected shape {frame.shape}; expected "
f"({EXPECTED_ROWS}, {EXPECTED_COLUMNS})."
)
if list(frame.columns) != EXPECTED_FEATURES:
raise RuntimeError(
"Unexpected column order/schema. Do not continue until reviewed."
)
frame["Class"] = frame["Class"].astype(int)
frauds = int(frame["Class"].sum())
if frauds != EXPECTED_FRAUDS:
raise RuntimeError(f"Unexpected fraud count {frauds}; expected {EXPECTED_FRAUDS}.")
if set(frame["Class"].unique()) != {0, 1}:
raise RuntimeError("Class must contain only 0 and 1.")
print(f"OpenML dataset ID: {OPENML_ID}")
print(f"Rows: {len(frame):,}")
print(f"Columns: {frame.shape[1]}")
print(f"Fraud transactions: {frauds}")
print(f"Fraud rate: {frauds / len(frame):.6%}")
print(f"SHA256: {sha256}")
print(f"Saved: {DATA_PATH}")
if __name__ == "__main__":
main()
python download_data.pyRows: 284,807
Columns: 31
Fraud transactions: 492
Fraud rate: 0.172749%
SHA256: b7efcb35a428bbe22347a05d2437d9177bab07ce61e51214a17bec584ad9496dCheck before continuing
Step 4: Create train, validation and final-test sets before experimentation
We need three sets because threshold selection is itself a modeling decision. Training data learns model parameters; validation data chooses the decision threshold; the final test set remains sealed until everything is fixed.
X_build, X_test, y_build, y_test = train_test_split(
X, y,
test_size=0.15,
stratify=y,
random_state=42,
)
X_train, X_val, y_train, y_val = train_test_split(
X_build, y_build,
test_size=0.15 / 0.85,
stratify=y_build,
random_state=42,
)
# 70% train | 15% validation | 15% final testTrain: 199,364 rows / 344 frauds
Validation: 42,721 rows / 74 frauds
Final test: 42,722 rows / 74 fraudsCheck before continuing
Step 5: Use metrics that can see the rare class
When we flag a transaction, how often is it actually fraud?
Of all actual frauds, how many did we catch?
A harmonic balance between precision and recall at one chosen threshold.
Summarizes the precision-recall trade-off across thresholds and is especially useful for highly imbalanced binary problems.
Check before continuing
Step 6: Compare four ways to handle the imbalance
Check before continuing
Step 7: Use SMOTE inside the pipeline — never before the split
SMOTE creates synthetic minority examples by interpolating between existing fraud examples. If you SMOTE before splitting, information derived from a transaction can leak into validation/test data and make evaluation too optimistic.
smote_logistic = ImbPipeline([
("scale", StandardScaler()),
("smote", SMOTE(
sampling_strategy=0.10,
random_state=42,
k_neighbors=5,
)),
("model", LogisticRegression(max_iter=1500)),
])
# During cross-validation SMOTE runs only inside each training fold.
# Validation/test rows are never synthetically oversampled.Check before continuing
Step 8: Build the training program in logical pieces, then assemble it
The complete file is intentionally shown only after the architecture is clear. Read it as six blocks: load → split → candidate models → cross-validation → validation threshold → final evaluation/save.
Open complete verified src/train_model.py
"""Train the fraud detector from the project root: python src/train_model.py."""
from __future__ import annotations
from pathlib import Path
import json
import platform
import joblib
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import sklearn
from imblearn.over_sampling import SMOTE
from imblearn.pipeline import Pipeline as ImbPipeline
from sklearn.base import clone
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (
ConfusionMatrixDisplay,
accuracy_score,
average_precision_score,
confusion_matrix,
f1_score,
precision_recall_curve,
precision_score,
recall_score,
roc_auc_score,
)
from sklearn.model_selection import StratifiedKFold, cross_validate, train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
ROOT = Path(__file__).resolve().parents[1]
DATA_PATH = ROOT / "data" / "creditcard.parquet"
MODEL_DIR = ROOT / "models"
OUTPUT_DIR = ROOT / "outputs"
FEATURES = ["Time", *[f"V{i}" for i in range(1, 29)], "Amount"]
TARGET = "Class"
SEED = 42
TARGET_RECALL = 0.80
def load_data() -> tuple[pd.DataFrame, pd.Series]:
if not DATA_PATH.is_file():
raise FileNotFoundError(
"Dataset missing. Run python download_data.py from the project root."
)
frame = pd.read_parquet(DATA_PATH)
expected = FEATURES + [TARGET]
if list(frame.columns) != expected:
raise ValueError("Dataset schema changed.")
y = frame[TARGET].astype(int)
if len(frame) != 284_807 or int(y.sum()) != 492:
raise ValueError("Dataset counts changed.")
return frame[FEATURES].astype(float), y
def split_data(X: pd.DataFrame, y: pd.Series):
X_build, X_test, y_build, y_test = train_test_split(
X,
y,
test_size=0.15,
stratify=y,
random_state=SEED,
)
validation_fraction_of_build = 0.15 / 0.85
X_train, X_val, y_train, y_val = train_test_split(
X_build,
y_build,
test_size=validation_fraction_of_build,
stratify=y_build,
random_state=SEED,
)
return X_train, X_val, X_test, y_train, y_val, y_test
def candidate_models():
logistic = Pipeline(
[
("scale", StandardScaler()),
(
"model",
LogisticRegression(
max_iter=1500,
solver="lbfgs",
random_state=SEED,
),
),
]
)
weighted_logistic = Pipeline(
[
("scale", StandardScaler()),
(
"model",
LogisticRegression(
max_iter=1500,
solver="lbfgs",
class_weight="balanced",
random_state=SEED,
),
),
]
)
smote_logistic = ImbPipeline(
[
("scale", StandardScaler()),
(
"smote",
SMOTE(
sampling_strategy=0.10,
random_state=SEED,
k_neighbors=5,
),
),
(
"model",
LogisticRegression(
max_iter=1500,
solver="lbfgs",
random_state=SEED,
),
),
]
)
forest = RandomForestClassifier(
n_estimators=120,
max_depth=12,
min_samples_leaf=2,
class_weight="balanced_subsample",
n_jobs=-1,
random_state=SEED,
)
return {
"Logistic Regression": logistic,
"Class-weighted Logistic Regression": weighted_logistic,
"SMOTE + Logistic Regression": smote_logistic,
"Class-weighted Random Forest": forest,
}
def compare_models(X_train: pd.DataFrame, y_train: pd.Series) -> pd.DataFrame:
cv = StratifiedKFold(n_splits=3, shuffle=True, random_state=SEED)
rows = []
for name, model in candidate_models().items():
print(f"Cross-validating {name} ...", flush=True)
scores = cross_validate(
clone(model),
X_train,
y_train,
cv=cv,
scoring={
"ap": "average_precision",
"roc_auc": "roc_auc",
"precision": "precision",
"recall": "recall",
"f1": "f1",
},
n_jobs=1,
error_score="raise",
)
rows.append(
{
"model": name,
"cv_average_precision": float(scores["test_ap"].mean()),
"cv_roc_auc": float(scores["test_roc_auc"].mean()),
"cv_precision_at_0_5": float(scores["test_precision"].mean()),
"cv_recall_at_0_5": float(scores["test_recall"].mean()),
"cv_f1_at_0_5": float(scores["test_f1"].mean()),
}
)
result = pd.DataFrame(rows).sort_values(
"cv_average_precision",
ascending=False,
kind="stable",
)
result.to_csv(OUTPUT_DIR / "model_comparison.csv", index=False)
return result
def choose_threshold(
y_val: pd.Series,
scores: np.ndarray,
target_recall: float = TARGET_RECALL,
) -> tuple[float, pd.DataFrame]:
precision, recall, thresholds = precision_recall_curve(y_val, scores)
table = pd.DataFrame(
{
"threshold": thresholds,
"precision": precision[:-1],
"recall": recall[:-1],
}
)
table["f1"] = (
2 * table["precision"] * table["recall"]
/ (table["precision"] + table["recall"] + 1e-12)
)
feasible = table[table["recall"] >= target_recall]
if not feasible.empty:
chosen = feasible.sort_values(
["precision", "threshold"],
ascending=[False, False],
kind="stable",
).iloc[0]
else:
chosen = table.sort_values("f1", ascending=False, kind="stable").iloc[0]
table.to_csv(OUTPUT_DIR / "validation_thresholds.csv", index=False)
return float(chosen["threshold"]), table
def metrics_at_threshold(
y_true: pd.Series,
scores: np.ndarray,
threshold: float,
) -> tuple[dict, np.ndarray]:
predictions = (scores >= threshold).astype(int)
matrix = confusion_matrix(y_true, predictions, labels=[0, 1])
metrics = {
"accuracy": float(accuracy_score(y_true, predictions)),
"precision": float(precision_score(y_true, predictions, zero_division=0)),
"recall": float(recall_score(y_true, predictions, zero_division=0)),
"f1": float(f1_score(y_true, predictions, zero_division=0)),
"average_precision": float(average_precision_score(y_true, scores)),
"roc_auc": float(roc_auc_score(y_true, scores)),
}
return metrics, matrix
def save_figures(
y: pd.Series,
comparison: pd.DataFrame,
threshold_table: pd.DataFrame,
chosen_threshold: float,
y_test: pd.Series,
test_scores: np.ndarray,
matrix: np.ndarray,
) -> None:
counts = y.value_counts().sort_index()
labels = ["Legitimate (0)", "Fraud (1)"]
values = [int(counts.get(0, 0)), int(counts.get(1, 0))]
fig, ax = plt.subplots(figsize=(8, 5))
bars = ax.bar(labels, values)
ax.set_yscale("log")
ax.set_ylabel("Transactions (log scale)")
ax.set_title("Why fraud detection is an imbalanced classification problem")
for bar, value in zip(bars, values):
pct = value / len(y) * 100
ax.text(
bar.get_x() + bar.get_width() / 2,
value,
f"{value:,}\n{pct:.3f}%",
ha="center",
va="bottom",
)
fig.tight_layout()
fig.savefig(OUTPUT_DIR / "class_imbalance.png", dpi=160)
plt.close(fig)
fig, ax = plt.subplots(figsize=(9, 5))
ax.barh(comparison["model"], comparison["cv_average_precision"])
ax.invert_yaxis()
ax.set_xlim(0, 1)
ax.set_xlabel("Mean 3-fold Average Precision")
ax.set_title("Training-only model comparison")
fig.tight_layout()
fig.savefig(OUTPUT_DIR / "model_comparison.png", dpi=160)
plt.close(fig)
sampled = threshold_table.iloc[:: max(1, len(threshold_table) // 500)].copy()
fig, ax = plt.subplots(figsize=(9, 5))
ax.plot(sampled["threshold"], sampled["precision"], label="Precision")
ax.plot(sampled["threshold"], sampled["recall"], label="Recall")
ax.axvline(
chosen_threshold,
linestyle="--",
label=f"Chosen threshold = {chosen_threshold:.3f}",
)
ax.set(
xlabel="Fraud-score threshold",
ylabel="Metric value",
ylim=(0, 1.02),
)
ax.set_title("Validation-only threshold trade-off")
ax.legend()
fig.tight_layout()
fig.savefig(OUTPUT_DIR / "threshold_tradeoff.png", dpi=160)
plt.close(fig)
precision, recall, _ = precision_recall_curve(y_test, test_scores)
fig, ax = plt.subplots(figsize=(7, 5))
ax.plot(recall, precision)
ax.set(
xlabel="Recall",
ylabel="Precision",
xlim=(0, 1),
ylim=(0, 1.02),
)
ax.set_title("Final untouched-test precision-recall curve")
fig.tight_layout()
fig.savefig(OUTPUT_DIR / "precision_recall_curve.png", dpi=160)
plt.close(fig)
fig, ax = plt.subplots(figsize=(7, 5))
ConfusionMatrixDisplay(
matrix,
display_labels=["Legitimate", "Fraud"],
).plot(ax=ax, colorbar=False, values_format="d")
ax.set_title("Final untouched-test confusion matrix")
fig.tight_layout()
fig.savefig(OUTPUT_DIR / "confusion_matrix.png", dpi=160)
plt.close(fig)
def save_demo_transactions(
bundle: dict,
X_test: pd.DataFrame,
y_test: pd.Series,
) -> None:
scores = bundle["model"].predict_proba(X_test)[:, 1]
demo = X_test.copy()
demo["historical_label"] = y_test.to_numpy()
demo["fraud_score"] = scores
fraud_examples = demo[demo["historical_label"] == 1].nlargest(6, "fraud_score")
legitimate_examples = demo[demo["historical_label"] == 0].nsmallest(6, "fraud_score")
selected = pd.concat(
[fraud_examples, legitimate_examples],
ignore_index=True,
)
selected.insert(
0,
"example_id",
[f"TX-{index + 1:02d}" for index in range(len(selected))],
)
selected.to_csv(OUTPUT_DIR / "demo_transactions.csv", index=False)
def main() -> None:
MODEL_DIR.mkdir(exist_ok=True)
OUTPUT_DIR.mkdir(exist_ok=True)
X, y = load_data()
X_train, X_val, X_test, y_train, y_val, y_test = split_data(X, y)
print(
f"Rows -> train: {len(X_train):,}, validation: {len(X_val):,}, "
f"test: {len(X_test):,}"
)
print(
f"Fraud rows -> train: {int(y_train.sum())}, validation: {int(y_val.sum())}, "
f"test: {int(y_test.sum())}"
)
print(f"Always-legitimate accuracy: {(y == 0).mean():.6%}")
comparison = compare_models(X_train, y_train)
print("\nTraining-only model comparison:")
print(comparison.to_string(index=False))
winner_name = str(comparison.iloc[0]["model"])
winner = clone(candidate_models()[winner_name])
winner.fit(X_train, y_train)
val_scores = winner.predict_proba(X_val)[:, 1]
threshold, threshold_table = choose_threshold(y_val, val_scores)
val_metrics, _ = metrics_at_threshold(y_val, val_scores, threshold)
X_build = pd.concat([X_train, X_val], axis=0)
y_build = pd.concat([y_train, y_val], axis=0)
final_model = clone(candidate_models()[winner_name])
final_model.fit(X_build, y_build)
test_scores = final_model.predict_proba(X_test)[:, 1]
test_metrics, matrix = metrics_at_threshold(y_test, test_scores, threshold)
bundle = {
"model": final_model,
"threshold": threshold,
"features": FEATURES,
"selected_model": winner_name,
"target_recall": TARGET_RECALL,
}
joblib.dump(bundle, MODEL_DIR / "fraud_detector.joblib")
reloaded = joblib.load(MODEL_DIR / "fraud_detector.joblib")
reload_scores = reloaded["model"].predict_proba(X_test.iloc[:25])[:, 1]
if not np.allclose(reload_scores, test_scores[:25]):
raise RuntimeError("Reloaded fraud scores differ from in-memory scores.")
save_demo_transactions(bundle, X_test, y_test)
save_figures(
y,
comparison,
threshold_table,
threshold,
y_test,
test_scores,
matrix,
)
report = {
"python": platform.python_version(),
"scikit_learn": sklearn.__version__,
"rows": int(len(X)),
"fraud_rows": int(y.sum()),
"fraud_rate": float(y.mean()),
"features": FEATURES,
"seed": SEED,
"train_rows": int(len(X_train)),
"validation_rows": int(len(X_val)),
"test_rows": int(len(X_test)),
"selected_model": winner_name,
"selection_metric": "mean 3-fold average precision on training data",
"target_validation_recall": TARGET_RECALL,
"chosen_threshold": threshold,
"validation_metrics": val_metrics,
"test_metrics": test_metrics,
"confusion_matrix": matrix.tolist(),
"always_legitimate_accuracy": float((y_test == 0).mean()),
}
(OUTPUT_DIR / "metrics.json").write_text(
json.dumps(report, indent=2) + "\n",
encoding="utf-8",
)
print(f"\nSelected model: {winner_name}")
print(f"Chosen validation threshold: {threshold:.6f}")
print("Final untouched-test metrics:")
for name, value in test_metrics.items():
print(f" {name}: {value:.6f}")
print("Confusion matrix [[TN, FP], [FN, TP]]:")
print(matrix)
print("Saved: models/fraud_detector.joblib")
if __name__ == "__main__":
main()
Check before continuing
Step 9: Select the model by training-only Average Precision
Training-only model comparison
Notice the trade-off: class-weighted Logistic Regression reaches very high recall at the default 0.5 threshold but precision collapses to about 6.6%. Catching more fraud is useful only if the false-alarm cost remains acceptable.
Check before continuing
Step 10: Tune the decision threshold on validation data, not the final test
A probability model gives a score. The threshold converts that score into an operational decision. We ask for at least 80% recall on the validation set, then choose the threshold with the highest precision among those feasible points.
precision, recall, thresholds = precision_recall_curve(
y_val,
validation_scores,
)
# Among thresholds that reach at least 80% recall on validation,
# choose the one with the highest precision.
feasible = table[table["recall"] >= 0.80]
chosen = feasible.sort_values(
["precision", "threshold"],
ascending=[False, False],
).iloc[0]Changing the threshold changes the business trade-off
Important: meeting an 80% recall target on validation does not guarantee 80% recall on future/test data. The untouched test recall became 74.3%. That difference is exactly why we keep a final test set.
Check before continuing
Step 11: Open the sealed test set once and interpret every error
Final untouched-test confusion matrix
Final precision-recall curve
Check before continuing
Step 12: Turn the model + threshold into a fraud-review application
A user does not type V1–V28 from memory. In a real payment system those features would arrive from the transaction pipeline. Our educational app therefore lets you select a verified holdout example, inspect its anonymized inputs, score it and compare the decision with the historical label.
Open complete app.py
"""Run with: python -m streamlit run app.py"""
from __future__ import annotations
from pathlib import Path
import joblib
import pandas as pd
import streamlit as st
ROOT = Path(__file__).resolve().parent
MODEL_PATH = ROOT / "models" / "fraud_detector.joblib"
DEMO_PATH = ROOT / "outputs" / "demo_transactions.csv"
st.set_page_config(
page_title="Credit Card Fraud Detector",
page_icon="💳",
layout="wide",
)
st.title("💳 Can AI Catch a Stolen Credit Card Transaction?")
st.caption(
"Educational fraud-scoring demo using the public anonymized OpenML credit-card dataset"
)
st.info(
"This is not a banking or payment-security product. The historical dataset is anonymized, "
"and V1–V28 are PCA-transformed features whose original meanings are not public."
)
@st.cache_resource
def load_bundle():
return joblib.load(MODEL_PATH)
if not MODEL_PATH.is_file() or not DEMO_PATH.is_file():
st.error("Training artifacts are missing.")
st.code(
"python download_data.py\npython src/train_model.py",
language="powershell",
)
st.stop()
bundle = load_bundle()
demos = pd.read_csv(DEMO_PATH)
features = bundle["features"]
threshold = float(bundle["threshold"])
st.subheader("Score a verified holdout example")
selected_id = st.selectbox(
"Transaction example",
demos["example_id"].tolist(),
)
row = demos.loc[demos["example_id"] == selected_id].iloc[0]
left, right, third = st.columns(3)
left.metric("Transaction amount", f"{float(row['Amount']):,.2f}")
right.metric("Seconds from dataset start", f"{float(row['Time']):,.0f}")
third.metric("Decision threshold", f"{threshold:.3f}")
with st.expander("See all anonymized model inputs"):
st.dataframe(
pd.DataFrame([row[features].to_dict()]),
hide_index=True,
use_container_width=True,
)
if st.button("Score transaction", type="primary"):
model_row = pd.DataFrame(
[[float(row[name]) for name in features]],
columns=features,
)
score = float(bundle["model"].predict_proba(model_row)[0, 1])
flagged = score >= threshold
if flagged:
st.error("Model decision: FLAG FOR FRAUD REVIEW")
else:
st.success("Model decision: do not flag at this threshold")
metric_col, truth_col = st.columns(2)
metric_col.metric("Fraud score", f"{score:.1%}")
historical_label = int(row["historical_label"])
truth_col.metric(
"Historical label",
"Fraud" if historical_label == 1 else "Legitimate",
)
if flagged and historical_label == 1:
st.write(
"This example is a **true positive**: the model flagged a transaction "
"historically labelled fraud."
)
elif flagged and historical_label == 0:
st.write(
"This example is a **false positive**: a legitimate historical "
"transaction was flagged."
)
elif not flagged and historical_label == 1:
st.write(
"This example is a **false negative**: a historical fraud was missed."
)
else:
st.write(
"This example is a **true negative**: a legitimate transaction was not flagged."
)
st.divider()
st.caption(
"Real fraud systems use richer live signals, cost-sensitive decisions, monitoring, "
"human review, security controls and continuously changing adversarial patterns."
)
python -m streamlit run app.pyReal fraud-review app
Flagged holdout example
Legitimate holdout example
Check before continuing
Step 13: Run automated checks so screenshots are not your only proof
The suite checks the exact dataset contract, split separation, class prevalence, four candidate strategies, threshold logic, saved artifact/report consistency, probability-aware final metrics and both historical classes in the demo set.
Open tests/test_fraud_detector.py
from pathlib import Path
import importlib.util
import json
import joblib
import numpy as np
import pandas as pd
import pytest
ROOT = Path(__file__).resolve().parents[1]
DATA = ROOT / "data" / "creditcard.parquet"
MODEL = ROOT / "models" / "fraud_detector.joblib"
METRICS = ROOT / "outputs" / "metrics.json"
DEMOS = ROOT / "outputs" / "demo_transactions.csv"
spec = importlib.util.spec_from_file_location(
"fraud_train",
ROOT / "src" / "train_model.py",
)
train = importlib.util.module_from_spec(spec)
spec.loader.exec_module(train)
def test_dataset_contract():
frame = pd.read_parquet(DATA)
assert frame.shape == (284_807, 31)
assert list(frame.columns) == train.FEATURES + [train.TARGET]
assert int(frame["Class"].astype(int).sum()) == 492
assert set(frame["Class"].astype(int).unique()) == {0, 1}
def test_three_way_split_is_disjoint_and_stratified():
X, y = train.load_data()
X_train, X_val, X_test, y_train, y_val, y_test = train.split_data(X, y)
assert len(X_train) + len(X_val) + len(X_test) == len(X)
assert set(X_train.index).isdisjoint(X_val.index)
assert set(X_train.index).isdisjoint(X_test.index)
assert set(X_val.index).isdisjoint(X_test.index)
for part in [y_train, y_val, y_test]:
assert 0 < int(part.sum()) < len(part)
assert abs(float(part.mean()) - float(y.mean())) < 0.0005
def test_candidates_include_imbalance_strategies():
candidates = train.candidate_models()
assert set(candidates) == {
"Logistic Regression",
"Class-weighted Logistic Regression",
"SMOTE + Logistic Regression",
"Class-weighted Random Forest",
}
def test_threshold_selection_meets_target_when_feasible():
y = pd.Series([0, 0, 0, 1, 1])
scores = np.array([0.05, 0.10, 0.20, 0.60, 0.90])
threshold, table = train.choose_threshold(y, scores, target_recall=1.0)
row = table.loc[np.isclose(table["threshold"], threshold)].iloc[0]
assert row["recall"] >= 1.0
assert 0 <= threshold <= 1
def test_saved_artifact_and_report_exist():
assert MODEL.is_file()
assert METRICS.is_file()
bundle = joblib.load(MODEL)
report = json.loads(METRICS.read_text(encoding="utf-8"))
assert 0 < float(bundle["threshold"]) < 1
assert bundle["selected_model"] == report["selected_model"]
assert report["selection_metric"].startswith("mean 3-fold average precision")
def test_final_metrics_are_probability_aware():
report = json.loads(METRICS.read_text(encoding="utf-8"))
metrics = report["test_metrics"]
assert 0 <= metrics["average_precision"] <= 1
assert 0 <= metrics["roc_auc"] <= 1
assert 0 <= metrics["precision"] <= 1
assert 0 <= metrics["recall"] <= 1
assert report["test_rows"] > 40_000
def test_demo_transactions_cover_both_historical_classes():
demos = pd.read_csv(DEMOS)
assert len(demos) == 12
assert set(demos["historical_label"].astype(int)) == {0, 1}
assert set(train.FEATURES).issubset(demos.columns)
@pytest.mark.parametrize("example_index", [0, 6])
def test_reloaded_model_scores_demo_rows(example_index):
bundle = joblib.load(MODEL)
demos = pd.read_csv(DEMOS)
row = demos.iloc[[example_index]][bundle["features"]]
score = float(bundle["model"].predict_proba(row)[0, 1])
assert 0 <= score <= 1
pytest -qCheck before continuing
Step 14: Know what SMOTE and supervised classification do not solve
Fraud patterns change because attackers adapt. The dataset is historical and anonymized, has only two days of transactions, and does not expose customer/device/merchant/network context. A production system would need temporal validation, drift monitoring, calibrated decision costs, security controls and human-review workflows.
Anomaly detection is another family of techniques. For example, Isolation Forest looks for observations that are easier to isolate than normal ones. It can help when labels are scarce, but “unusual” is not automatically the same as “fraud,” so it is an extension rather than a magic replacement.
Check before continuing
Step 15: Understand the complete project folder
├── data/creditcard.parquet # downloaded locally
├── models/fraud_detector.joblib # generated model + threshold
├── outputs/ # metrics, plots, demo rows
├── scripts/capture_app_screenshots.py
├── src/train_model.py
├── tests/test_fraud_detector.py
├── app.py
├── download_data.py
├── requirements.txt
└── README.md
Check before continuing
Now change the project yourself
Exercise 1: Change the validation recall target from 80% to 90%. Record what happens to precision and false positives.
Exercise 2: Remove SMOTE and compare class weighting alone. Explain why resampling is not automatically better.
Exercise 3: Try an Isolation Forest as an anomaly-detection extension and compare its ranking with the supervised model.
Exercise 4: Define a simple cost where missing fraud costs 20× more than reviewing a legitimate transaction. Choose a threshold using cost instead of recall.
How to explain this project in an interview
Why is 99.8% accuracy useless here?
Because the majority class is so dominant that predicting every transaction as legitimate achieves that accuracy while fraud recall is zero.
Why Average Precision?
It evaluates the precision-recall ranking across thresholds and focuses attention on performance for the rare positive class.
Why is SMOTE inside a pipeline?
So synthetic minority samples are created only from each training fold; validation and test rows remain natural and uncontaminated.
Why have validation and test separately?
Threshold selection uses validation data. The final test must remain untouched so it measures the completed decision process.
Why did validation target recall not equal test recall?
A threshold chosen on one finite sample will not reproduce identical metrics on unseen data. That generalization gap is real evidence, not an error to hide.







