Skip to main content
All project handbooks
FREE PROJECTIntermediateWindows-first instructions

Build Your Own Netflix-Style Movie Recommendation System

Build a complete recommendation engine from an empty folder to a working browser app. You will compare popularity, content-based, collaborative and hybrid recommendation, then test the system and understand exactly where each ranking signal comes from.

What you build
A Streamlit app with four recommendation modes and an explicit cold-start fallback.
Teaching dataset
9,000 synthetic ratings, 1,100 users and 260 movie IDs, released under CC0.
End-to-end workflow
Download → inspect → build → save → test → run → debug → improve.

The practical problem: a user liked one movie—what should we show next?

Imagine you are building the recommendation row for a small streaming app. A user selects Iron Country. Your screen cannot show all 260 movies. It must return a short ranked list of titles that are reasonable next choices, and you should be able to explain where that ranking came from.

There is no single “correct movie” label. Instead we have different kinds of evidence. We can start with movies that are broadly well rated, compare movie attributes such as genre and release decade, compare audience rating patterns, or combine those signals. We also need a safe answer when a brand-new movie has no learned similarity at all.

Baseline

Rank generally strong movies so the product always has a sensible default.

Content signal

Find movies with similar genre + release-decade labels.

Audience signal

Find movies that receive similar rating patterns across users.

Product answer

Blend both signals into a ranked top-N list and fall back safely for cold start.

Ratings + movie data
→
Popularity
→
Content vectors
→
Sparse rating matrix
→
Cosine neighbours
→
45/55 hybrid
→
Top-N list
→
Streamlit app
Verified Streamlit movie recommendation form with Iron Country selected
The finished product. The learner will build this real Streamlit interface, choose a movie and recommendation method, and request a ranked list.
Verified distribution of ratings used by the movie recommendation project
The evidence underneath the app. Recommendation quality depends on interaction data. Before ranking, we inspect what ratings actually exist and later build a sparse movie-by-user matrix from them.

Why this project uses synthetic movie ratings

The recommendation algorithms in this handbook are real, but the movie titles and ratings are synthetic. Datanemics publishes this dataset under CC0 1.0, so a learner can download, modify and reuse it without depending on personal records or a restricted commercial-data licence. The trade-off is realism: these ratings simulate recommendation-system patterns rather than representing real streaming customers.

“Netflix-style” describes the familiar product idea—choosing one movie and receiving ranked suggestions. This project does not use Netflix data and does not claim to reproduce Netflix's production algorithm.

Tools you will use

PythonVS CodeDatanemics CC0 datasetPandasSciPyscikit-learnNearestNeighborsCosine similarityJoblibMatplotlibStreamlitPytestGitGitHub

Topics covered

Recommendation systems · Popularity baseline · Bayesian-style shrinkage · Content-based filtering · Multi-hot encoding · Cosine similarity · Sparse matrices · Collaborative filtering · Item-item similarity · Hybrid ranking · Cold start · Recommendation evaluation · Model persistence · Testing

1

Understand what a recommender returns

Classification asks “which class?” Regression asks “what number?” A recommendation system asks “which items should appear near the top?” The output is therefore an ordered list.

We will build four versions so that every extra layer has a reason to exist: popularity gives a simple baseline, content compares movie attributes, collaborative filtering compares audience behaviour, and the hybrid combines the last two.

Check before continuing: You can explain that the output is a ranked list, not one class label or one numeric prediction.

2

Create the project folder

Create and open the projectpowershellRunnable
mkdir movie-recommender
cd movie-recommender
code .

Create these folders in Explorer: data, models, outputs,scripts, src and tests.

Check before continuing: VS Code opens the movie-recommender folder and you can see an empty Explorer.

3

Create a clean Python environment

Create and activate .venvpowershellRunnable
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
pip install -r requirements.txt

The virtual environment keeps this project's package versions separate from other Python projects. If PowerShell blocks activation, run Set-ExecutionPolicy -Scope Process -ExecutionPolicy Bypass in that terminal and activate again.

Check before continuing: Your terminal prompt starts with (.venv).

4

Create requirements.txt

requirements.txttextConfiguration
pandas==2.3.3
numpy==2.3.3
scikit-learn==1.7.2
scipy==1.16.2
joblib==1.5.2
matplotlib==3.10.6
streamlit==1.50.0
pytest==8.4.2

Pandas reads the ratings, SciPy stores the sparse matrix, scikit-learn performs nearest-neighbour search, Joblib saves the built artifacts, Matplotlib creates evidence, Streamlit serves the app, and Pytest checks behaviour.

Check before continuing: pip install -r requirements.txt finishes without an error.

5

Download the CC0 movie-ratings dataset

Create download_data.py in the project root. The script downloads the CSV from Datanemics, verifies the exact schema and important counts, computes a SHA-256 fingerprint, and saves the file locally.

download_data.py — complete sourcepythonRunnable
"""Download the CC0 synthetic movie-ratings dataset from Datanemics."""

from __future__ import annotations

from pathlib import Path
import hashlib
import urllib.request

import pandas as pd

ROOT = Path(__file__).resolve().parent
DATA_DIR = ROOT / "data"
DATA_PATH = DATA_DIR / "movie-ratings.csv"

DATASET_PAGE = "https://datanemics.com/datasets/movie-ratings/"
CSV_URL = "https://datanemics.com/datasets/data/movie-ratings.csv"
LICENSE = "CC0 1.0 public domain"
EXPECTED_ROWS = 9_000
EXPECTED_USERS = 1_100
EXPECTED_MOVIES = 260
EXPECTED_SHA256 = "b9e41047db97680f0043a8bdcb18fd5cb25d8f4a6209d5ef536f0ba9b457af52"
EXPECTED_COLUMNS = [
    "user_id",
    "movie_id",
    "title",
    "genre",
    "release_year",
    "rating",
    "rated_at",
]


def download_bytes() -> bytes:
    request = urllib.request.Request(
        CSV_URL,
        headers={"User-Agent": "LearnMLAcademy/1.0 educational-project"},
    )
    with urllib.request.urlopen(request, timeout=120) as response:
        return response.read()


def main() -> None:
    DATA_DIR.mkdir(parents=True, exist_ok=True)

    print("Dataset: Datanemics Movie ratings")
    print("Dataset page:", DATASET_PAGE)
    print("License:", LICENSE)
    print("Synthetic dataset: yes")

    raw = download_bytes()
    sha256 = hashlib.sha256(raw).hexdigest()
    if sha256 != EXPECTED_SHA256:
        raise RuntimeError(
            "Downloaded CSV fingerprint changed: "
            f"{sha256}; expected {EXPECTED_SHA256}. "
            "Do not continue until the dataset change is reviewed."
        )
    DATA_PATH.write_bytes(raw)

    frame = pd.read_csv(DATA_PATH)
    if list(frame.columns) != EXPECTED_COLUMNS:
        raise RuntimeError(
            "Unexpected columns: "
            f"{list(frame.columns)}; expected {EXPECTED_COLUMNS}"
        )

    if len(frame) != EXPECTED_ROWS:
        raise RuntimeError(
            f"Unexpected row count: {len(frame):,}; expected {EXPECTED_ROWS:,}"
        )

    users = int(frame["user_id"].nunique())
    movies = int(frame["movie_id"].nunique())
    if users != EXPECTED_USERS or movies != EXPECTED_MOVIES:
        raise RuntimeError(
            "Unexpected entity counts: "
            f"users={users}, movies={movies}; "
            f"expected users={EXPECTED_USERS}, movies={EXPECTED_MOVIES}"
        )

    ratings = pd.to_numeric(frame["rating"], errors="raise")
    if float(ratings.min()) != 0.5 or float(ratings.max()) != 5.0:
        raise RuntimeError(
            f"Unexpected rating range: {ratings.min()} to {ratings.max()}"
        )

    print(f"Verified rows:    {len(frame):,}")
    print(f"Verified users:   {users:,}")
    print(f"Verified movies:  {movies:,}")
    print(f"Rating range:     {ratings.min():.1f} to {ratings.max():.1f}")
    print(f"SHA256:           {sha256}")
    print(f"Saved to:         {DATA_PATH}")


if __name__ == "__main__":
    main()
Download and verifypowershellRunnable
python download_data.py
Expected dataset contracttextOutput
Verified rows:    9,000
Verified users:   1,100
Verified movies:  260
Rating range:     0.5 to 5.0

The dataset has seven columns: user ID, movie ID, title, primary genre, release year, rating and rating timestamp. The timestamp lets us handle the possibility that the same user rated the same movie more than once.

Check before continuing: The script prints 9,000 rows, 1,100 users, 260 movie IDs and a 0.5–5.0 rating range.

6

Keep one latest interaction per user and movie

A matrix cell cannot safely contain two different ratings. The function latest_interactions()sorts by rated_at and keeps the latest rating for each user/movie pair.

This is a developer choice, not a fact of recommendation systems. Another system might average repeated ratings or model them as a time sequence. Here, “latest wins” keeps the teaching matrix easy to interpret.

Check before continuing: You understand why one user/movie pair should occupy one cell in the collaborative matrix.

7

Build a popularity baseline before anything clever

A raw average can be misleading when very few people rated an item. Our popularity score shrinks each movie's mean toward the global mean with a prior equivalent to 25 ratings.

This baseline has two jobs: it gives us something simple to compare against, and it becomes the fallback when a new or unknown movie has no learned similarity information.

Popularity blockpythonRunnable
def build_popularity(interactions, movies):
    summary = (
        interactions.groupby("movie_id")["rating"]
        .agg(["mean", "count"])
        .reset_index()
    )
    global_mean = float(interactions["rating"].mean())
    prior = 25.0

    summary["weighted_score"] = (
        (summary["count"] / (summary["count"] + prior)) * summary["mean"]
        + (prior / (summary["count"] + prior)) * global_mean
    )
    return movies.merge(summary, on="movie_id", how="left")

Read the formula from left to right: movies with many ratings lean toward their own mean; movies with very few ratings are pulled more strongly toward the global mean.

Check before continuing: You can explain why a movie with one 5-star rating should not automatically rank above a movie with many strong ratings.

8

Represent movie content as two labels

Each movie gets two content labels. For example, a romance released in 2012 becomes genre=romance and decade=2010s. A multi-hot encoder converts those labels to a vector of 0s and 1s.

Content ideatextOutput
Iron Country (2012) → [genre=romance, decade=2010s]
Another romance from 2018 → shares both labels
A romance from 1984 → shares genre but not decade

This deliberately simple representation makes it possible to see where a content recommendation came from. A production system could add plot text, actors, languages, embeddings and many other features.

Content-vector blockpythonRunnable
labels = [
    [
        f"genre={row.genre}",
        f"decade={(int(row.release_year) // 10) * 10}s",
    ]
    for row in movies.itertuples(index=False)
]

encoder = MultiLabelBinarizer()
matrix = encoder.fit_transform(labels).astype(float)
matrix = csr_matrix(normalize(matrix, norm="l2"))

content_model = NearestNeighbors(metric="cosine", algorithm="brute")
content_model.fit(matrix)

Check before continuing: You can turn a movie into one genre label and one decade label.

9

Use cosine similarity to find content neighbours

Cosine similarity compares vector direction. Movies sharing content labels point in more similar directions. scikit-learn's NearestNeighbors gives cosine distance, so the project converts it to a score:

Distance to similaritytextOutput
similarity = 1 - cosine_distance

distance 0.00 → similarity 1.00
distance 0.40 → similarity 0.60

Check before continuing: You know that similarity = 1 - cosine distance in this implementation.

10

Build the sparse movie-by-user matrix

Collaborative filtering ignores genre and year. Each row represents one movie, each column represents one user, and an observed cell contains that user's latest rating for that movie.

There are 260 × 1,100 = 286,000 possible movie/user cells. After the latest-interaction rule, the verified build contains 8,291 observed cells—about 2.90% density and 97.10% sparsity. A SciPy CSR sparse matrix stores the observed values without allocating ordinary dense values for every empty cell.

Sparse collaborative matrix blockpythonRunnable
movie_to_row = {
    movie_id: index
    for index, movie_id in enumerate(movie_ids)
}
user_to_col = {
    user_id: index
    for index, user_id in enumerate(user_ids)
}

rows = interactions["movie_id"].map(movie_to_row).to_numpy()
cols = interactions["user_id"].map(user_to_col).to_numpy()
values = interactions["rating"].to_numpy(dtype=float)

matrix = csr_matrix(
    (values, (rows, cols)),
    shape=(len(movie_ids), len(user_ids)),
)

collab_model = NearestNeighbors(metric="cosine", algorithm="brute")
collab_model.fit(matrix)

Check before continuing: You can identify a row, column and observed cell in the collaborative matrix.

11

Let audience behaviour create item-item neighbours

Two movies are collaborative neighbours when their rating vectors point in similar directions across users. This can reveal relationships that metadata does not describe. That is the central difference from content filtering.

This teaching model uses raw 0.5–5 rating vectors. More advanced systems may center ratings, use implicit feedback, factorize the matrix or learn embeddings.

Check before continuing: You can explain why two movies can be collaborative neighbours even if their genres differ.

12

Combine the two signals with a hybrid score

Hybrid ruletextOutput
hybrid_score = 0.45 × content_similarity + 0.55 × collaborative_similarity

The hybrid gives slightly more weight to audience behaviour while preserving a content signal. A real system would tune or learn these weights using offline evaluation and online experiments.

Hybrid ranking blockpythonRunnable
candidates = set(content) | set(collaborative)

scores = {
    index: (
        0.45 * content.get(index, 0.0)
        + 0.55 * collaborative.get(index, 0.0)
    )
    for index in candidates
}

ranked = sorted(
    scores.items(),
    key=lambda pair: (-pair[1], pair[0]),
)[:top_n]

Notice the important design choice: the two similarity scores live on the same 0–1 scale before we blend them. If two components used incompatible scales, the larger numerical scale could dominate the ranking accidentally.

Check before continuing: You know that 45% and 55% are explicit project choices, not universal constants.

13

Assemble the complete recommender engine

By now you have built the important mechanisms separately. One last product rule remains: if the selected movie is unknown, do not invent a similarity. Return the popularity baseline and say why.

Cold-start fallback blockpythonRunnable
if movie_id is None or movie_id not in artifacts["movie_to_row"]:
    popular = popularity_recommendations(
        artifacts,
        top_n=top_n,
    ).copy()

    popular["score"] = (
        popular["weighted_score"] / 5.0
    ).clip(0.0, 1.0)

    popular["reason"] = "popularity fallback for cold start"
    return popular

How the pieces connect

load_data() creates clean ratings and movie tables → latest_interactions() removes repeated user/movie ratings → the popularity/content/collaborative builders create reusable structures →recommend() converts neighbour distances to scores and ranks candidates →recommend_or_fallback() handles unknown items → main() saves the complete Joblib artifact, metrics, chart and deterministic reference output.

Open the complete verified src/build_recommender.py

The complete file is intentionally hidden until you understand its parts. Use the Copy button now to assemble the final executable source.

Complete src/build_recommender.pypythonRunnable
"""Build popularity, content, collaborative and hybrid movie recommenders."""

from __future__ import annotations

from pathlib import Path
import json

import joblib
import matplotlib.pyplot as plt
import pandas as pd
from scipy.sparse import csr_matrix
from sklearn.neighbors import NearestNeighbors
from sklearn.preprocessing import MultiLabelBinarizer, normalize

ROOT = Path(__file__).resolve().parents[1]
DATA_PATH = ROOT / "data" / "movie-ratings.csv"
MODEL_DIR = ROOT / "models"
OUTPUT_DIR = ROOT / "outputs"

REFERENCE_MOVIE_ID = "M1000"
REFERENCE_MOVIE_TITLE = "Iron Country"
MAX_RATING = 5.0


def load_data() -> tuple[pd.DataFrame, pd.DataFrame]:
    if not DATA_PATH.is_file():
        raise FileNotFoundError("Run python download_data.py before building the recommender.")

    frame = pd.read_csv(DATA_PATH)
    required = {
        "user_id",
        "movie_id",
        "title",
        "genre",
        "release_year",
        "rating",
        "rated_at",
    }
    missing = sorted(required.difference(frame.columns))
    if missing:
        raise RuntimeError(f"Dataset is missing required columns: {missing}")

    frame["rating"] = pd.to_numeric(frame["rating"], errors="raise")
    frame["release_year"] = pd.to_numeric(frame["release_year"], errors="raise").astype(int)
    frame["rated_at"] = pd.to_datetime(frame["rated_at"], errors="raise")

    metadata_consistency = (
        frame.groupby("movie_id")[["title", "genre", "release_year"]]
        .nunique(dropna=False)
        .max()
        .max()
    )
    if int(metadata_consistency) != 1:
        raise RuntimeError("At least one movie_id maps to conflicting title/genre/year metadata.")

    movies = (
        frame[["movie_id", "title", "genre", "release_year"]]
        .drop_duplicates(subset=["movie_id"])
        .sort_values("movie_id")
        .reset_index(drop=True)
    )
    ratings = frame[["user_id", "movie_id", "rating", "rated_at"]].copy()
    return ratings, movies


def latest_interactions(ratings: pd.DataFrame) -> pd.DataFrame:
    """Keep one latest rating per user/movie pair for the teaching matrix."""
    return (
        ratings.sort_values("rated_at")
        .drop_duplicates(subset=["user_id", "movie_id"], keep="last")
        .reset_index(drop=True)
    )


def build_popularity(interactions: pd.DataFrame, movies: pd.DataFrame) -> pd.DataFrame:
    """Build a smoothed popularity baseline using a 25-rating prior."""
    summary = (
        interactions.groupby("movie_id")["rating"]
        .agg(["mean", "count"])
        .reset_index()
    )
    global_mean = float(interactions["rating"].mean())
    prior = 25.0
    summary["weighted_score"] = (
        (summary["count"] / (summary["count"] + prior)) * summary["mean"]
        + (prior / (summary["count"] + prior)) * global_mean
    )
    return movies.merge(summary, on="movie_id", how="left").fillna(
        {"mean": global_mean, "count": 0, "weighted_score": global_mean}
    )


def build_content(movies: pd.DataFrame):
    """Encode each movie with two categorical labels: genre and release decade."""
    labels = [
        [
            f"genre={row.genre}",
            f"decade={(int(row.release_year) // 10) * 10}s",
        ]
        for row in movies.itertuples(index=False)
    ]
    encoder = MultiLabelBinarizer()
    matrix = encoder.fit_transform(labels).astype(float)
    matrix = csr_matrix(normalize(matrix, norm="l2"))

    model = NearestNeighbors(metric="cosine", algorithm="brute")
    model.fit(matrix)
    return encoder, matrix, model


def build_collaborative(interactions: pd.DataFrame, movies: pd.DataFrame):
    """Create a sparse movie-by-user matrix from latest observed ratings."""
    movie_ids = movies["movie_id"].tolist()
    user_ids = sorted(interactions["user_id"].unique())
    movie_to_row = {movie_id: index for index, movie_id in enumerate(movie_ids)}
    user_to_col = {user_id: index for index, user_id in enumerate(user_ids)}

    rows = interactions["movie_id"].map(movie_to_row).to_numpy()
    cols = interactions["user_id"].map(user_to_col).to_numpy()
    values = interactions["rating"].to_numpy(dtype=float)

    matrix = csr_matrix(
        (values, (rows, cols)),
        shape=(len(movie_ids), len(user_ids)),
    )
    model = NearestNeighbors(metric="cosine", algorithm="brute")
    model.fit(matrix)
    return matrix, model


def neighbor_scores(model, matrix, row: int, n_candidates: int = 60) -> dict[int, float]:
    count = min(n_candidates + 1, matrix.shape[0])
    distances, indices = model.kneighbors(matrix[row], n_neighbors=count)

    scores: dict[int, float] = {}
    for index, distance in zip(indices[0], distances[0]):
        index = int(index)
        if index == row:
            continue
        scores[index] = max(0.0, min(1.0, 1.0 - float(distance)))
    return scores


def recommend(
    artifacts: dict,
    movie_id: str,
    method: str = "hybrid",
    top_n: int = 10,
) -> pd.DataFrame:
    if top_n < 1:
        raise ValueError("top_n must be at least 1")

    movies = artifacts["movies"]
    movie_to_row = artifacts["movie_to_row"]
    if movie_id not in movie_to_row:
        raise ValueError(f"Unknown movie_id: {movie_id}")

    row = movie_to_row[movie_id]
    content = neighbor_scores(
        artifacts["content_model"],
        artifacts["content_matrix"],
        row,
    )
    collaborative = neighbor_scores(
        artifacts["collab_model"],
        artifacts["collab_matrix"],
        row,
    )

    if method == "content":
        scores = content
        reason = "similar genre + release decade"
    elif method == "collaborative":
        scores = collaborative
        reason = "similar audience rating patterns"
    elif method == "hybrid":
        candidates = set(content) | set(collaborative)
        scores = {
            index: (
                0.45 * content.get(index, 0.0)
                + 0.55 * collaborative.get(index, 0.0)
            )
            for index in candidates
        }
        reason = "45% content + 55% audience similarity"
    else:
        raise ValueError("method must be content, collaborative or hybrid")

    ranked = sorted(scores.items(), key=lambda pair: (-pair[1], pair[0]))[:top_n]
    output = []
    for rank, (index, score) in enumerate(ranked, start=1):
        movie = movies.iloc[index]
        output.append(
            {
                "rank": rank,
                "movie_id": str(movie["movie_id"]),
                "title": str(movie["title"]),
                "genre": str(movie["genre"]),
                "release_year": int(movie["release_year"]),
                "score": round(float(score), 4),
                "reason": reason,
            }
        )
    return pd.DataFrame(output)


def popularity_recommendations(artifacts: dict, top_n: int = 10) -> pd.DataFrame:
    if top_n < 1:
        raise ValueError("top_n must be at least 1")

    popular = (
        artifacts["popularity"]
        .sort_values(
            ["weighted_score", "count", "movie_id"],
            ascending=[False, False, True],
        )
        .head(top_n)
        .copy()
    )
    popular.insert(0, "rank", range(1, len(popular) + 1))
    return popular[
        [
            "rank",
            "movie_id",
            "title",
            "genre",
            "release_year",
            "weighted_score",
            "count",
        ]
    ]


def recommend_or_fallback(
    artifacts: dict,
    movie_id: str | None,
    method: str = "hybrid",
    top_n: int = 10,
) -> pd.DataFrame:
    """Use popularity for an unknown item instead of inventing similarity."""
    if movie_id is None or movie_id not in artifacts["movie_to_row"]:
        popular = popularity_recommendations(artifacts, top_n=top_n).copy()
        popular["score"] = (popular["weighted_score"] / MAX_RATING).clip(0.0, 1.0)
        popular["score"] = popular["score"].round(4)
        popular["reason"] = "popularity fallback for cold start"
        return popular[
            [
                "rank",
                "movie_id",
                "title",
                "genre",
                "release_year",
                "score",
                "reason",
            ]
        ]

    return recommend(
        artifacts,
        movie_id=movie_id,
        method=method,
        top_n=top_n,
    )


def main() -> None:
    MODEL_DIR.mkdir(exist_ok=True)
    OUTPUT_DIR.mkdir(exist_ok=True)

    ratings, movies = load_data()
    interactions = latest_interactions(ratings)

    popularity = build_popularity(interactions, movies)
    encoder, content_matrix, content_model = build_content(movies)
    collab_matrix, collab_model = build_collaborative(interactions, movies)
    movie_to_row = {
        movie_id: index
        for index, movie_id in enumerate(movies["movie_id"])
    }

    artifacts = {
        "movies": movies,
        "popularity": popularity,
        "content_encoder": encoder,
        "content_matrix": content_matrix,
        "content_model": content_model,
        "collab_matrix": collab_matrix,
        "collab_model": collab_model,
        "movie_to_row": movie_to_row,
    }

    model_path = MODEL_DIR / "movie_recommender.joblib"
    joblib.dump(artifacts, model_path)
    reloaded = joblib.load(model_path)

    if REFERENCE_MOVIE_ID not in movie_to_row:
        raise RuntimeError(f"Reference movie {REFERENCE_MOVIE_ID} is missing.")

    reference_title = str(
        movies.loc[movies["movie_id"] == REFERENCE_MOVIE_ID, "title"].iloc[0]
    )
    if reference_title != REFERENCE_MOVIE_TITLE:
        raise RuntimeError(
            "Reference movie metadata changed: "
            f"{REFERENCE_MOVIE_ID} is {reference_title!r}, "
            f"expected {REFERENCE_MOVIE_TITLE!r}"
        )

    reference_frames = []
    for method in ["content", "collaborative", "hybrid"]:
        frame = recommend(
            reloaded,
            REFERENCE_MOVIE_ID,
            method=method,
            top_n=10,
        )
        frame.insert(0, "method", method)
        reference_frames.append(frame)

    pd.concat(reference_frames, ignore_index=True).to_csv(
        OUTPUT_DIR / "iron_country_recommendations.csv",
        index=False,
    )
    popularity_recommendations(reloaded, 10).to_csv(
        OUTPUT_DIR / "popular_movies.csv",
        index=False,
    )

    possible_cells = int(collab_matrix.shape[0] * collab_matrix.shape[1])
    observed_cells = int(collab_matrix.nnz)
    density = observed_cells / possible_cells

    metrics = {
        "raw_rating_rows": int(len(ratings)),
        "unique_user_movie_interactions": int(len(interactions)),
        "users": int(interactions["user_id"].nunique()),
        "movies": int(len(movies)),
        "rating_mean": float(interactions["rating"].mean()),
        "rating_min": float(interactions["rating"].min()),
        "rating_max": float(interactions["rating"].max()),
        "content_features": int(content_matrix.shape[1]),
        "content_feature_labels": list(encoder.classes_),
        "collaborative_shape": list(collab_matrix.shape),
        "observed_rating_cells": observed_cells,
        "possible_rating_cells": possible_cells,
        "rating_density": density,
        "rating_sparsity": 1.0 - density,
        "reference_movie": REFERENCE_MOVIE_TITLE,
        "reference_movie_id": REFERENCE_MOVIE_ID,
        "hybrid_content_weight": 0.45,
        "hybrid_collaborative_weight": 0.55,
    }
    (OUTPUT_DIR / "metrics.json").write_text(
        json.dumps(metrics, indent=2) + "\n",
        encoding="utf-8",
    )

    fig, ax = plt.subplots(figsize=(7, 4))
    interactions["rating"].value_counts().sort_index().plot(kind="bar", ax=ax)
    ax.set(
        title="Synthetic movie-ratings distribution",
        xlabel="Rating",
        ylabel="Count",
    )
    fig.tight_layout()
    fig.savefig(OUTPUT_DIR / "rating_distribution.png", dpi=160)
    plt.close(fig)

    print("Movie recommender built successfully.")
    print(json.dumps(metrics, indent=2))
    print(f"\nHybrid recommendations for {REFERENCE_MOVIE_TITLE}:")
    print(
        recommend(
            reloaded,
            REFERENCE_MOVIE_ID,
            method="hybrid",
            top_n=10,
        ).to_string(index=False)
    )


if __name__ == "__main__":
    main()

Check before continuing: You can connect the popularity, content, collaborative, hybrid and cold-start blocks before copying the complete source.

14

Build the artifacts and inspect Iron Country

Build the recommenderpowershellRunnable
python src/build_recommender.py

M1000 is the fixed reference movie ID and maps to Iron Country. The build writes content, collaborative and hybrid reference lists to outputs/iron_country_recommendations.csv.

Verified hybrid output — first five rowstextOutput
1  Distant Compass   romance  2015  0.5470\n2  Iron Field        romance  2011  0.5342\n3  Quiet Machine     romance  2014  0.4500\n4  Last Lantern      romance  2013  0.4500\n5  Broken Machine    romance  2017  0.4500

The exact order is evidence from this code and dataset—not a universal ranking of movies. If you change the features or weights, you should expect the order to change.

Check before continuing: The command finishes and prints ten hybrid recommendations for Iron Country.

15

Read the real rating-distribution chart

The verified build creates outputs/rating_distribution.png. It shows how frequently each star rating appears after keeping the latest user/movie interaction.

Rating distribution generated by the executable movie recommendation project

Interaction data is not a random sample of every movie a user could have watched. Users choose what to rate, so recommendation data is typically missing in a meaningful, non-random way.

Check before continuing: You can explain what the x-axis, y-axis and bar heights represent.

16

Handle cold start instead of faking a score

Cold start means the system lacks enough history for a new user or item. An unknown movie is absent from both learned matrices, so recommend() rejects it.

recommend_or_fallback() catches that product situation and returns the popularity ranking instead. Its score is normalized to 0–1 so the output schema stays consistent with similarity scores.

Check before continuing: You can explain why an unknown movie cannot have a learned matrix row.

17

Create the Streamlit browser application

app.py — complete sourcepythonRunnable
"""Run from the project root with: python -m streamlit run app.py"""

from pathlib import Path
import importlib.util

import joblib
import streamlit as st

ROOT = Path(__file__).resolve().parent
MODEL_PATH = ROOT / "models" / "movie_recommender.joblib"
REFERENCE_MOVIE_ID = "M1000"

spec = importlib.util.spec_from_file_location(
    "movie_recommender_core",
    ROOT / "src" / "build_recommender.py",
)
core = importlib.util.module_from_spec(spec)
spec.loader.exec_module(core)

st.set_page_config(
    page_title="Movie Recommendation System",
    page_icon="🎬",
    layout="wide",
)
st.title("Build Your Own Netflix-Style Movie Recommendation System")
st.caption(
    "Educational recommender using a CC0 synthetic ratings dataset • "
    "not Netflix's production algorithm"
)

if not MODEL_PATH.is_file():
    st.error("The recommender artifact is missing.")
    st.code(
        "python download_data.py\npython src/build_recommender.py",
        language="powershell",
    )
    st.stop()


@st.cache_resource
def load_artifacts(modified_ns: int):
    return joblib.load(MODEL_PATH)


artifacts = load_artifacts(MODEL_PATH.stat().st_mtime_ns)
movies = artifacts["movies"].sort_values(
    ["title", "release_year", "movie_id"]
).reset_index(drop=True)

method_label = st.radio(
    "Recommendation method",
    ["Hybrid", "Content-based", "Collaborative", "Popular movies"],
    horizontal=True,
)

if method_label == "Popular movies":
    st.subheader("Popular starting points")
    popular = core.popularity_recommendations(artifacts, top_n=10).copy()
    popular["weighted_score"] = popular["weighted_score"].round(3)
    st.dataframe(popular, hide_index=True, use_container_width=True)
    st.info(
        "Popularity is our cold-start fallback. It does not personalize "
        "to a selected movie."
    )
else:
    choices = [
        (
            str(row.movie_id),
            str(row.title),
            int(row.release_year),
        )
        for row in movies.itertuples(index=False)
    ]
    default_choice = next(
        (choice for choice in choices if choice[0] == REFERENCE_MOVIE_ID),
        choices[0],
    )
    selected_choice = st.selectbox(
        "Choose a movie you like",
        choices,
        index=choices.index(default_choice),
        format_func=lambda choice: (
            f"{choice[1]} ({choice[2]}) • ID {choice[0]}"
        ),
    )
    selected_id, selected_title, _selected_year = selected_choice

    method = {
        "Hybrid": "hybrid",
        "Content-based": "content",
        "Collaborative": "collaborative",
    }[method_label]
    top_n = st.slider(
        "How many recommendations?",
        min_value=5,
        max_value=15,
        value=10,
    )

    if st.button("Recommend movies", type="primary"):
        recommendations = core.recommend_or_fallback(
            artifacts,
            movie_id=selected_id,
            method=method,
            top_n=top_n,
        )
        st.subheader(f"Because you chose: {selected_title}")
        st.dataframe(
            recommendations,
            hide_index=True,
            use_container_width=True,
        )
        if method == "content":
            st.caption(
                "Content-based: compare genre + release-decade labels "
                "with cosine similarity."
            )
        elif method == "collaborative":
            st.caption(
                "Collaborative: compare sparse movie-by-user rating patterns."
            )
        else:
            st.caption(
                "Hybrid: 45% content similarity + "
                "55% audience-rating similarity."
            )

st.divider()
st.caption(
    "The ratings and movie titles in this teaching dataset are synthetic. "
    "The recommendation methods are real, but this app is intentionally small. "
    "Production streaming platforms use far more data, ranking stages, "
    "experimentation and infrastructure."
)

The selector displays title, release year and movie ID. The ID is the true lookup key because different movie IDs can share the same synthetic title.

Check before continuing: app.py exists and contains the complete code below.

18

Run the app and generate real recommendations

Start the browser apppowershellRunnable
python -m streamlit run app.py
Real Streamlit form from the verified movie recommendation application

Start with Hybrid, keep Iron Country, leave 10 recommendations and click Recommend movies.

Real Streamlit recommendation results for Iron Country from the verified application

Check before continuing: Iron Country is selected by default and clicking Recommend movies produces a ten-row table.

19

Add tests that check behaviour, not just startup

Create tests/test_recommender.py. The suite checks the dataset contract, latest-interaction rule, three ranking methods, query exclusion, popularity order, cold-start behaviour, normalized fallback scores, sparse storage, content labels and deterministic saved output.

tests/test_recommender.py — complete sourcepythonRunnable
from pathlib import Path
import importlib.util

import joblib
import pandas as pd
import pytest

ROOT = Path(__file__).resolve().parents[1]
MODEL = ROOT / "models" / "movie_recommender.joblib"

spec = importlib.util.spec_from_file_location(
    "recommender",
    ROOT / "src" / "build_recommender.py",
)
core = importlib.util.module_from_spec(spec)
spec.loader.exec_module(core)


@pytest.fixture(scope="session")
def artifacts():
    assert MODEL.is_file(), "Run python src/build_recommender.py first."
    return joblib.load(MODEL)


def test_dataset_contract():
    ratings, movies = core.load_data()
    assert len(ratings) == 9_000
    assert ratings["user_id"].nunique() == 1_100
    assert len(movies) == 260
    assert set(
        ["movie_id", "title", "genre", "release_year"]
    ).issubset(movies.columns)


def test_latest_interactions_have_one_row_per_user_movie():
    ratings, _movies = core.load_data()
    interactions = core.latest_interactions(ratings)
    assert len(interactions) <= len(ratings)
    assert not interactions.duplicated(["user_id", "movie_id"]).any()


@pytest.mark.parametrize("method", ["content", "collaborative", "hybrid"])
def test_recommendations_are_unique_and_exclude_query_movie(
    artifacts,
    method,
):
    result = core.recommend(
        artifacts,
        core.REFERENCE_MOVIE_ID,
        method=method,
        top_n=10,
    )
    assert len(result) == 10
    assert result["movie_id"].is_unique
    assert core.REFERENCE_MOVIE_ID not in set(result["movie_id"])
    assert result["score"].between(0, 1).all()


def test_popularity_fallback_is_ranked(artifacts):
    result = core.popularity_recommendations(artifacts, top_n=10)
    assert len(result) == 10
    assert result["rank"].tolist() == list(range(1, 11))
    assert result["weighted_score"].is_monotonic_decreasing


def test_unknown_movie_is_rejected_by_similarity_api(artifacts):
    with pytest.raises(ValueError, match="Unknown movie_id"):
        core.recommend(
            artifacts,
            "M_DOES_NOT_EXIST",
            method="hybrid",
            top_n=10,
        )


def test_unknown_movie_uses_normalized_cold_start_fallback(artifacts):
    result = core.recommend_or_fallback(
        artifacts,
        movie_id="M_DOES_NOT_EXIST",
        method="hybrid",
        top_n=7,
    )
    assert len(result) == 7
    assert result["rank"].tolist() == list(range(1, 8))
    assert set(result["reason"]) == {
        "popularity fallback for cold start"
    }
    assert result["score"].between(0, 1).all()


def test_top_n_must_be_positive(artifacts):
    with pytest.raises(ValueError, match="top_n"):
        core.recommend(
            artifacts,
            core.REFERENCE_MOVIE_ID,
            method="hybrid",
            top_n=0,
        )
    with pytest.raises(ValueError, match="top_n"):
        core.popularity_recommendations(artifacts, top_n=0)


def test_collaborative_matrix_is_sparse(artifacts):
    ratings, _movies = core.load_data()
    interactions = core.latest_interactions(ratings)

    matrix = artifacts["collab_matrix"]
    assert matrix.shape == (260, 1_100)
    assert matrix.nnz == len(interactions)
    density = matrix.nnz / (matrix.shape[0] * matrix.shape[1])
    assert density < 0.10


def test_content_features_include_genre_and_decade(artifacts):
    labels = set(artifacts["content_encoder"].classes_)
    assert any(label.startswith("genre=") for label in labels)
    assert any(label.startswith("decade=") for label in labels)
    assert artifacts["content_matrix"].shape[0] == 260


def test_reference_output_matches_live_artifact(artifacts):
    reference_path = ROOT / "outputs" / "iron_country_recommendations.csv"
    assert reference_path.is_file()

    saved = pd.read_csv(reference_path)
    hybrid_saved = (
        saved.loc[saved["method"] == "hybrid"]
        .reset_index(drop=True)
    )
    live = core.recommend(
        artifacts,
        core.REFERENCE_MOVIE_ID,
        method="hybrid",
        top_n=10,
    )

    pd.testing.assert_frame_equal(
        hybrid_saved[live.columns].reset_index(drop=True),
        live.reset_index(drop=True),
        check_dtype=False,
    )
Run the automated checkspowershellRunnable
pytest -q

Check before continuing: pytest -q finishes with every test passing.

20

Understand what the Joblib artifact contains

models/movie_recommender.joblib is not a neural-network model. It packages the built movie table, popularity table, content encoder/matrix/model, collaborative matrix/model and the movie-ID-to-row mapping.

The app loads that trusted local artifact instead of rebuilding everything on every browser refresh. Never load arbitrary Joblib files from untrusted sources.

Check before continuing: You can name the tables, sparse matrices, nearest-neighbour models and ID lookup stored in the artifact.

21

Know what this project has not proved

Recommendation evaluation is different from classification accuracy. The dataset records items people rated, not every item they might have liked. A proper offline experiment might hide known interactions and ask whether the system ranks them highly using Hit Rate@K, Precision@K, Recall@K or NDCG.

This project focuses on making the recommendation mechanisms traceable. The exercise section asks you to add an evaluation protocol as the next learning step.

Check before continuing: You do not call the hybrid 'best' merely because its recommendations look reasonable.

22

Understand the complete project folder

movie-recommender/
├── data/ # downloaded CSV, ignored by Git
├── models/ # generated Joblib artifact
├── outputs/ # metrics, reference lists, chart
├── scripts/
│   ├── capture_app_screenshots.py
│   └── verify_handbook_page.py
├── src/
│   └── build_recommender.py
├── tests/
│   └── test_recommender.py
├── app.py
├── download_data.py
├── requirements.txt
├── README.md
└── .gitignore

Check before continuing: You can point to the file responsible for data acquisition, recommendation logic, testing and browser inference.

23

Put the finished source on GitHub

.gitignoretextConfiguration
.venv/
__pycache__/
.pytest_cache/
data/movie-ratings.csv
models/*.joblib
outputs/screenshots/
*.log
Git and GitHub commandspowershellRunnable
git init
git add .
git status
git commit -m "Build movie recommendation system"
git branch -M main
git remote add origin https://github.com/YOUR-USERNAME/movie-recommender.git
git push -u origin main

Read git status before committing. Generated data and model artifacts can be rebuilt; the source files and dependency lock are what another learner needs.

Check before continuing: git status does not show .venv, the downloaded CSV or the generated Joblib file.

Common problems and exact fixes

movie-ratings.csv is missing: run python download_data.py from the project root.

Dataset count/schema verification fails: do not bypass the check. Delete the downloaded file and rerun; the source may have changed.

ModuleNotFoundError: activate .venv, then rerun pip install -r requirements.txt.

Streamlit says the artifact is missing: run python src/build_recommender.py before starting the app.

The selected movie appears in its own recommendations: verify the query-row exclusion inside neighbor_scores().

Two options have the same title: that is why the app shows release year and movie ID; never use title alone as the key.

Collaborative results seem surprising: raw rating-vector cosine is intentionally simple. Try mean-centering or matrix factorization as an extension.

Now change the system yourself

Exercise 1 — Change hybrid weights

Try 70% content and 30% collaborative. Rebuild and explain which recommendations move and which evidence caused the change.

Exercise 2 — Remove the decade feature

Use genre only, rebuild, and compare Iron Country content neighbours. This isolates the contribution of release decade.

Exercise 3 — Break cold start on purpose

Call recommend() with an unknown movie ID, then call recommend_or_fallback(). Explain why the second behaviour is safer in an application.

Exercise 4 — Add held-out evaluation

Hide one known interaction for eligible users and calculate a ranking metric such as Hit Rate@10 or Recall@10 without leaking that held-out interaction into the evidence matrix.

How to explain this project in an interview

Why start with popularity?

It creates a transparent baseline and gives the product a fallback when similarity evidence is unavailable. Shrinkage prevents tiny-rating-count items from dominating the leaderboard.

What is content-based filtering here?

Each movie has a genre label and a release-decade label. Multi-hot vectors plus cosine nearest neighbours retrieve movies with similar metadata.

What is collaborative filtering here?

Each movie is represented by its pattern of ratings across users. Similar rating-vector directions create item-item neighbours independently of genre.

Why use a sparse matrix?

Only a small fraction of the 260 × 1,100 possible movie-user cells have observed ratings, so CSR storage avoids allocating ordinary dense values for empty interactions.

What is cold start?

A new item has no learned matrix row or interaction history. The project explicitly returns the popularity fallback instead of fabricating a similarity.

What is the biggest dataset limitation?

The dataset is synthetic. That makes it safe and reproducible for teaching, but results must not be presented as evidence about real audience preferences.

What would change in a production recommender?

A real streaming product would use real consented interaction data, separate candidate generation from ranking, learn from implicit signals such as views and completion, handle new users and items, add richer or learned features, measure ranking quality offline, run online experiments, control popularity/exposure bias, cache low-latency results, monitor data and model drift, and apply business and safety rules.

Implementation mastery check

Why is the output a ranked list?
Why can a raw average-rating leaderboard be misleading?
What does the 25-rating shrinkage prior do?
Why keep only the latest repeated user/movie rating?
What two labels make up a movie content vector?
How is cosine distance converted to similarity?
What are the rows and columns of the collaborative matrix?
Why is CSR sparse storage useful?
Why can collaborative neighbours cross genre boundaries?
What does the 45/55 hybrid rule mean?
Why must the selected movie be excluded from its own neighbours?
What is cold start and what fallback do we use?
Why are fallback scores normalized to 0–1?
What is stored in movie_recommender.joblib?
Why does a plausible list not prove recommender quality?
What changes because the teaching dataset is synthetic?

Complete-project checkpoint

CC0 synthetic ratings downloaded and verified
Repeated interactions reduced to one latest user/movie rating
Smoothed popularity baseline built
Genre + decade content vectors built
Sparse movie-user collaborative matrix built
Cosine nearest-neighbour retrieval works
Hybrid ranking is deterministic
Cold-start fallback is explicit and score-compatible
Joblib artifact saves and reloads
Automated behavioral tests pass
Streamlit app serves real generated recommendations
Real browser screenshots and chart are visible
Synthetic-data limitation is stated clearly
Production next steps are understood

Run the entire project again from scratch

A reproducible project should not depend on files that happen to exist on your laptop. From a clean clone, these four commands should rebuild the data, recommender, tests and browser app:

Reproduce the projectpowershellRunnable
python download_data.py
python src/build_recommender.py
pytest -q
python -m streamlit run app.py