Skip to main content
All hands-on projects

Project 5 · Time Series · Full beginner handbook

Can We Predict Tomorrow's Sales? Build a Retail Forecaster

An online store owner asks a practical question: how many pounds of orders should we plan for tomorrow? Predicting zero means lost opportunities, but predicting too much could waste stock. Use real sales history, not made-up reported accuracy.

What you will build

An offline, reproducible forecasting pipeline built from official UK retail orders: cleaned daily revenue, lag features, chronological validation, a 28-day held-out test and a Streamlit chart comparing model, actual and baseline.

Exact tools you will use

Python 3.12, VS Code, UCI Online Retail, pandas, NumPy, scikit-learn, Ridge, HistGradientBoosting, Streamlit, pytest, GitHub Actions

Daily orders flow into chronological sales chart, past-day lags, baseline and model then a next-day prediction
The actual order of operations in this learning project, illustrated. Not a screenshot of a trained model.

Real application evidence

Genuine full-page desktop and mobile screenshots captured from a running Streamlit app by GitHub Actions. Retail evidence uses the actual SHA-256-verified UCI workbook and held-out daily evaluation.

Genuine desktop Streamlit Can We Predict Tomorrow's Sales? Build a Retail Forecaster screenshot
Desktop browser evidence
Genuine mobile Streamlit Can We Predict Tomorrow's Sales? Build a Retail Forecaster screenshot
390-pixel mobile browser evidence

Compute a sales feature and a baseline by hand

  1. Assume previous seven days sold £100, £120, £90, £130, £110, £140 and £150.
  2. Prior 7-day mean = (100+120+90+130+110+140+150) / 7 = £120.
  3. Suppose previous Friday was £100. Seasonal naive for Friday predicts £100, not the seven-day mean.
  4. If actual Friday sales are £125, absolute error = |125−100| = £25.
  5. For actual [100,125] and predicted [110,100], MAE = (10+25)/2 = £17.50.
1

Open the project and understand the sales problem

Why: Time series asks what changes next; its ordering matters.

Do this: Install Python 3.12 and VS Code. Open File → Open Folder → projects/retail-forecasting. Open Terminal → New Terminal.

Check: See download_data.py, train.py, app.py and src/forecast.py.

2

Create the virtual environment

Why: Python and library versions should be isolated per project.

Do this: Create and activate the environment; install the pinned-compatible packages.

Type these terminal commandsbashRunnable
python -m venv .venv
# Windows: .venv\Scripts\activate
# macOS/Linux: source .venv/bin/activate
python -m pip install -r requirements.txt

Check: pip finishes without errors.

3

Download the authentic UCI retail data

Why: We need observed dated transactions, not an invented demo series.

Do this: Run the verified official UCI archive downloader. It checks the archive SHA-256 before extracting the original workbook.

Type these terminal commandsbashRunnable
python download_data.py

Check: data/Online Retail.xlsx exists and the archive fingerprint matches.

4

Clean and aggregate daily orders

Why: Cancellations, returns and non-UK orders would change the target definition.

Do this: Read daily_revenue in src/forecast.py. Remove negative/zero quantities and cancelled invoices, retain positive UK orders, multiply quantity by unit price, then aggregate by date.

Check: Daily index includes calendar days. Missing transaction days are zero-filled with an explicit caveat.

5

Build leakage-free features and chronological splits

Why: A model cannot know tomorrow's actual sales today.

Do this: Read make_features. lag_1 and lag_7 are yesterday and the same weekday last week; both precede the target date. Keep last 28 days for final test and preceding 28 days for validation.

Check: The train maximum date precedes validation; validation precedes test. No random time-series shuffle.

6

Compare baseline and models; view actual holdout

Why: A complex model must beat a simple business baseline before it is useful.

Do this: Run train.py. Ridge and gradient boosting compete with the same-weekday baseline using validation MAE. The chosen approach is measured once on the untouched holdout.

Type these terminal commandsbashRunnable
python -m pytest -q
python train.py
python -m streamlit run app.py

Check: artifacts contains metrics.json, test_predictions.csv, next_day.json and forecast.joblib. Dashboard shows actual, baseline and predictions.

Limitations you should understand

Predicts positive-order gross GBP from the UK subset, not recognized net revenue. One-day-ahead holdout evaluation uses observed earlier days and does NOT demonstrate an entire 28-day future forecast. No promotion or weather features.

Complete source code — copy every file

This is the exact code used in the repository, loaded directly as raw source in the website build. Each file below is complete, not abbreviated. Create the named file inside the project folder, paste it, and then run the commands above.

README.md

README.mdmarkdownRunnable
# Project 5 — Can We Predict Tomorrow's Sales?

An online shop owner has to decide how much stock to keep for tomorrow. Can sales from the past month help? We build an actual daily sales forecasting pipeline, compare it against the last-week baseline, and show tomorrow's prediction in a Streamlit dashboard.

## Dataset and tools
Official UCI Online Retail dataset, 541,909 transactions from December 2010 to December 2011. Python 3.12, VS Code, pandas, NumPy, scikit-learn, Streamlit and pytest. Data archive verified by SHA-256. No paid service. We predict positive UK order value rather than net recognized revenue.

## Step-by-step
1. Open VS Code, choose File → Open Folder, select projects/retail-forecasting.
2. Open Terminal → New Terminal. Create a virtual environment: python -m venv .venv.
3. Activate it: Windows .venv\Scripts\activate; macOS/Linux source .venv/bin/activate.
4. Run python -m pip install -r requirements.txt.
5. Run python -m pytest -q. These offline synthetic-data smoke tests do not evaluate the real dataset.
6. Run python download_data.py to download and verify official UCI data.
7. Run python train.py. This loads transactions, removes refunds/cancellations, filters UK, groups revenue per day, creates lagged features and compares models.
8. Open artifacts/metrics.json and artifacts/test_predictions.csv to see chosen model, dates and final errors.
9. Run python -m streamlit run app.py and open localhost:8501 in your browser to view charts and next-day forecast.

## Worked numbers
Suppose the past seven days sold £100, £120, £90, £130, £110, £140, £150. The trailing mean is (100+120+90+130+110+140+150)/7 = £120. If last Friday's amount was £100, seasonal naive predicts £100 for next Friday. Actual £125 means an absolute error of £25. We average absolute errors across 28 days to obtain mean absolute error (MAE). The train / validation / test windows follow calendar order: do not randomly split the time series.

## Important limits
A forecast for each of the last 28 test days uses already observed earlier sales as lagged features: it is a sequence of ONE-day-ahead forecasts, not a 28-day recursive future prediction. Closed days and data outages can resemble zero sales. Promotions, supply shortages, weather and holidays are not modeled. No revenue figures are invented: official dataset results appear only after executing train.py.

## Learner checks
Explain why using the current day's revenue as a predictor causes leakage. Calculate MAE for actual [100, 125], predicted [110, 100]: (10+25)/2 = 17.5. Re-run with different lag features and review validation MAE without repeatedly tuning against the test set.

requirements.txt

requirements.txttextRunnable
numpy>=1.26,<3
pandas>=2.2,<3.1
openpyxl>=3.1,<4
scikit-learn>=1.5,<2
joblib>=1.4,<2
streamlit>=1.37,<2
pytest>=8,<10

download_data.py

download_data.pypythonRunnable
"""Download official UCI Online Retail, checksum-verified (same source as Project 4)."""
from pathlib import Path
from io import BytesIO
import hashlib
import urllib.request
import zipfile

ROOT = Path(__file__).resolve().parent
URL = "https://archive.ics.uci.edu/static/public/352/online+retail.zip"
SHA = "f5385cbb54bbebf7196389109c6b0621faab0c304e3702548165e71c84aede8b"

def main():
    destination = ROOT / "data" / "Online Retail.xlsx"
    destination.parent.mkdir(parents=True, exist_ok=True)
    if destination.exists():
        print("Existing file:", destination)
        return
    request = urllib.request.Request(URL, headers={"User-Agent": "LearnMLAcademy-Forecasting/1.0"})
    with urllib.request.urlopen(request, timeout=180) as response:
        body = response.read()
    digest = hashlib.sha256(body).hexdigest()
    if digest != SHA:
        raise RuntimeError(f"Official data archive hash changed: {digest}")
    with zipfile.ZipFile(BytesIO(body)) as archive:
        candidates = [n for n in archive.namelist() if n.lower().endswith(".xlsx")]
        if len(candidates) != 1:
            raise RuntimeError("Expected a single spreadsheet in the official archive")
        destination.write_bytes(archive.read(candidates[0]))
    print("Downloaded and checksum-verified:", destination)

if __name__ == "__main__":
    main()

src/forecast.py

src/forecast.pypythonRunnable
"""Leakage-aware one-step-ahead forecasting of positive UK ecommerce order value."""
from pathlib import Path
import json

import joblib
import numpy as np
import pandas as pd
from sklearn.base import clone
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.linear_model import Ridge
from sklearn.metrics import mean_absolute_error, mean_squared_error
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

ROOT = Path(__file__).resolve().parents[1]
SEED = 42
LAGS = (1, 7, 14, 28)
FEATURES = [f"lag_{n}" for n in LAGS] + ["prior_7_mean", "prior_28_mean", "weekday", "month"]

def daily_revenue(transactions):
    required = {"InvoiceNo", "InvoiceDate", "Quantity", "UnitPrice", "Country"}
    if not required.issubset(transactions.columns):
        raise ValueError("Missing required UCI transaction columns")
    df = transactions.copy()
    df["InvoiceDate"] = pd.to_datetime(df["InvoiceDate"], errors="coerce")
    for col in ("Quantity", "UnitPrice"):
        df[col] = pd.to_numeric(df[col], errors="coerce")
    mask = df["Country"].eq("United Kingdom") & df["InvoiceDate"].notna()
    mask &= df["Quantity"].gt(0) & df["UnitPrice"].gt(0)
    mask &= ~df["InvoiceNo"].astype(str).str.upper().str.startswith("C")
    df = df.loc[mask].copy()
    if df.empty:
        raise ValueError("No valid positive UK sales")
    df["date"] = df["InvoiceDate"].dt.normalize()
    df["positive_sales"] = df["Quantity"] * df["UnitPrice"]
    observed = df.groupby("date")["positive_sales"].sum().sort_index()
    return observed.reindex(pd.date_range(observed.index.min(), observed.index.max(), freq="D"),
                            fill_value=0.0).rename("positive_gross_gbp")

def make_features(series):
    if not isinstance(series.index, pd.DatetimeIndex) or not series.index.is_monotonic_increasing:
        raise ValueError("Daily sales must have sorted DatetimeIndex")
    if len(series) < 130 or series.index.has_duplicates:
        raise ValueError("At least 130 consecutive distinct dates needed")
    if not pd.date_range(series.index.min(), series.index.max(), freq="D").equals(series.index):
        raise ValueError("Calendar has missing days")
    y = pd.to_numeric(series, errors="raise")
    if y.isna().any() or not np.isfinite(y.to_numpy()).all() or (y < 0).any():
        raise ValueError("Sales must be finite and nonnegative")
    x = pd.DataFrame(index=series.index)
    for lag in LAGS:
        x[f"lag_{lag}"] = y.shift(lag)
    x["prior_7_mean"] = y.shift(1).rolling(7).mean()
    x["prior_28_mean"] = y.shift(1).rolling(28).mean()
    x["weekday"] = series.index.dayofweek
    x["month"] = series.index.month
    x["target"] = y
    return x.dropna()

def split_dates(rows):
    if len(rows) < 84:
        raise ValueError("Need enough dates for 28-day validation and 28-day test")
    train, val, test = rows.iloc[:-56], rows.iloc[-56:-28], rows.iloc[-28:]
    assert len(val) == len(test) == 28 and train.index.max() < val.index.min() < test.index.min()
    return train, val, test

def scores(actual, predicted):
    a, p = np.asarray(actual, float), np.asarray(predicted, float)
    if a.shape != p.shape or not np.isfinite(p).all():
        raise ValueError("Nonfinite or wrong-size predictions")
    return {"mae": float(mean_absolute_error(a, p)),
            "rmse": float(np.sqrt(mean_squared_error(a, p))),
            "wape_pct": float(100 * np.abs(a - p).sum() / max(np.abs(a).sum(), 1e-9))}

def candidates():
    return {
        "Ridge": make_pipeline(StandardScaler(), Ridge(alpha=20.0)),
        "HistGradientBoosting": HistGradientBoostingRegressor(
            max_iter=120, learning_rate=0.05, max_leaf_nodes=8,
            l2_regularization=10.0, random_state=SEED),
    }

def train_and_evaluate(daily, destination, source_label="Original fictional daily-sales test fixture"):
    rows = make_features(daily)
    train, val, test = split_dates(rows)
    chosen_mae = {"same_day_last_week": scores(val.target, val.lag_7)["mae"]}
    for name, candidate in candidates().items():
        fitted = clone(candidate).fit(train[FEATURES], train.target)
        chosen_mae[name] = scores(val.target, np.maximum(0, fitted.predict(val[FEATURES])))["mae"]
    winner = min(chosen_mae, key=chosen_mae.get)
    model = None
    if winner != "same_day_last_week":
        model = clone(candidates()[winner]).fit(
            pd.concat([train, val])[FEATURES], pd.concat([train, val]).target)
        prediction = np.maximum(0, model.predict(test[FEATURES]))
    else:
        prediction = test.lag_7.to_numpy()
    report = {
        "dataset": source_label,
        "target": "gross positive GBP order value per day, not accounting net revenue",
        "forecast_contract": "one-step ahead; observed prior daily actuals are available for each day",
        "train_dates": [str(train.index.min().date()), str(train.index.max().date())],
        "val_dates": [str(val.index.min().date()), str(val.index.max().date())],
        "test_dates": [str(test.index.min().date()), str(test.index.max().date())],
        "n_daily": len(daily),
        "validation_mae": chosen_mae,
        "selected_model": winner,
        "test_selected": scores(test.target, prediction),
        "test_baseline": scores(test.target, test.lag_7),
    }
    destination = Path(destination)
    destination.mkdir(parents=True, exist_ok=True)
    joblib.dump({"model": model, "winner": winner}, destination / "forecast.joblib")
    pd.DataFrame({"actual": test.target, "predicted": prediction,
                  "last_week": test.lag_7}, index=test.index).to_csv(
                      destination / "test_predictions.csv", index_label="date")
    next_date = daily.index[-1] + pd.Timedelta(days=1)
    next_row = {**{f"lag_{i}": float(daily.iloc[-i]) for i in LAGS},
                "prior_7_mean": float(daily.iloc[-7:].mean()),
                "prior_28_mean": float(daily.iloc[-28:].mean()),
                "weekday": int(next_date.dayofweek), "month": int(next_date.month)}
    x = pd.DataFrame([next_row])[FEATURES]
    future = float(x.lag_7.iloc[0] if model is None else max(0, model.predict(x)[0]))
    (destination / "next_day.json").write_text(json.dumps(
        {"date": str(next_date.date()), "forecast_gbp": round(future, 2)}, indent=2) + "\n")
    (destination / "metrics.json").write_text(json.dumps(report, indent=2) + "\n")
    return report

train.py

train.pypythonRunnable
"""Train on real official data: python download_data.py && python train.py."""
import pandas as pd
from src.forecast import ROOT, daily_revenue, train_and_evaluate

source = ROOT / "data" / "Online Retail.xlsx"
if not source.is_file():
    raise SystemExit("Run python download_data.py first.")
df = pd.read_excel(source, engine="openpyxl")
if len(df) != 541_909:
    raise SystemExit(f"Official dataset expected 541909 rows, found {len(df)}")
result = train_and_evaluate(daily_revenue(df), ROOT / "artifacts",
    source_label="Official UCI Online Retail: positive UK orders excluding returns and cancellations")
print("Selected model:", result["selected_model"])
print("Final untouched holdout:", result["test_selected"])

app.py

app.pypythonRunnable
"""Read-only student dashboard; train first to generate real data and artifacts."""
from pathlib import Path
import json
import pandas as pd
import streamlit as st

ROOT = Path(__file__).resolve().parent
st.set_page_config(page_title="Retail Sales Forecast", layout="wide")
st.title("Can We Predict Tomorrow's Sales?")
st.caption("Next-day positive-order gross GBP, not audited net revenue. Historical dataset ends in 2011.")
metric_file = ROOT / "artifacts" / "metrics.json"
if not metric_file.is_file():
    st.warning("First run python download_data.py then python train.py")
    st.stop()
report = json.loads(metric_file.read_text())
st.caption("Data source: " + report["dataset"])
st.subheader("Training and validation")
st.write("Model selected using validation MAE:", report["selected_model"])
st.json(report["validation_mae"])
st.subheader("Untouched chronological 28-day test")
a, b, c = st.columns(3)
a.metric("MAE (£)", f"{report['test_selected']['mae']:,.2f}")
b.metric("RMSE (£)", f"{report['test_selected']['rmse']:,.2f}")
c.metric("WAPE", f"{report['test_selected']['wape_pct']:.1f}%")
data = pd.read_csv(ROOT / "artifacts" / "test_predictions.csv", parse_dates=["date"])
st.line_chart(data.set_index("date")[["actual", "predicted", "last_week"]])
st.dataframe(data, use_container_width=True)
tomorrow = json.loads((ROOT / "artifacts" / "next_day.json").read_text())
st.metric("Next observed date: " + tomorrow["date"], f"£{tomorrow['forecast_gbp']:,.2f}")
st.info("This is one-day-ahead forecasting with observed lags, not a forecast of the entire next month. Zero-filled dates may represent shop closure.")

tests/test_forecast.py

tests/test_forecast.pypythonRunnable
import numpy as np
import pandas as pd
import pytest
from src.forecast import daily_revenue, make_features, split_dates, train_and_evaluate, scores

def sales():
    dates = pd.date_range("2025-01-01", periods=175, freq="D")
    return pd.Series(100 + .35 * np.arange(175) + 10 * np.sin(np.arange(175) * 2*np.pi/7),
                     index=dates)

def test_transaction_cleaning():
    frame = pd.DataFrame({"InvoiceNo": ["100", "C101", "102"],
                          "InvoiceDate": ["2025-01-01"] * 3,
                          "Quantity": [2, -1, 4], "UnitPrice": [5, 5, 5],
                          "Country": ["United Kingdom", "United Kingdom", "France"]})
    assert daily_revenue(frame).iloc[0] == 10

def test_features_only_use_the_past():
    y = sales()
    x = make_features(y)
    assert x.iloc[0].lag_1 == y.iloc[27]
    assert x.iloc[0].lag_7 == y.iloc[21]
    t = y.index[50]
    modified = y.copy()
    modified.loc[t] = 999999
    assert make_features(modified).loc[t, "lag_1"] == x.loc[t, "lag_1"]
    with pytest.raises(ValueError):
        make_features(y.drop(y.index[17]))

def test_split_chronologically():
    train, val, test = split_dates(make_features(sales()))
    assert train.index.max() < val.index.min() < test.index.min()
    assert len(val) == 28 and len(test) == 28

def test_training_and_real_file_outputs(tmp_path):
    report = train_and_evaluate(sales(), tmp_path)
    assert report["selected_model"] in ("Ridge", "HistGradientBoosting", "same_day_last_week")
    assert len(pd.read_csv(tmp_path / "test_predictions.csv")) == 28
    assert (tmp_path / "forecast.joblib").exists()
    assert (tmp_path / "next_day.json").exists()

def test_zero_actuals_wape():
    assert scores([0, 0], [0, 0])["wape_pct"] == 0

Project verification workflow

.github/workflows/three-projects-verify.ymlyamlConfiguration
name: Three Remaining Projects — Engineering Verify
on:
  push:
    branches:
      - feat/complete-three-projects-20261009
  pull_request:
    paths:
      - 'projects/digit-recognizer/**'
      - 'projects/retail-forecasting/**'
      - 'projects/disaster-tweets/**'
      - 'src/pages/*ProjectPage.tsx'
      - 'src/data/projectPortfolio.ts'
      - '.github/workflows/three-projects-verify.yml'
  workflow_dispatch:

jobs:
  retail:
    runs-on: ubuntu-latest
    timeout-minutes: 15
    defaults:
      run:
        working-directory: projects/retail-forecasting
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: '3.12'
          cache: pip
          cache-dependency-path: projects/retail-forecasting/requirements.txt
      - run: python -m pip install -r requirements.txt
      - run: python -m pytest -q
      - run: python -m compileall -q src train.py download_data.py app.py
  disaster:
    runs-on: ubuntu-latest
    timeout-minutes: 15
    defaults:
      run:
        working-directory: projects/disaster-tweets
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: '3.12'
          cache: pip
          cache-dependency-path: projects/disaster-tweets/requirements.txt
      - run: python -m pip install -r requirements.txt
      - run: python -m pytest -q
      - run: python -m compileall -q src train.py app.py
  digits:
    runs-on: ubuntu-latest
    timeout-minutes: 25
    defaults:
      run:
        working-directory: projects/digit-recognizer
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: '3.12'
      - run: python -m pip install -r requirements.txt
      - run: python -m pytest -q
      - run: python -m compileall -q src train.py app.py
  website:
    runs-on: ubuntu-latest
    timeout-minutes: 15
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: '22'
          cache: npm
      - run: npm ci
      - run: npm run lint
      - run: npm run build

Your build checkpoints — Retail Sales Forecaster

Keep the full handbook and all source code visible above. These optional checkpoints help you track what you can actually build and explain. Progress is saved only in this browser.

0 of 5 checkpoints completed

Predict → change → observe → explain

Before: Predict whether last week's sales would beat a simple rolling-average forecast.

Try: Compare the actual chronological validation errors of two lag-based forecasts.

Show your evidence: Compute a mean absolute error from the held-out daily differences.

Environment setup on your computer

Unzip the source first, open its project folder in VS Code, then read its README for dataset/download instructions. Python 3.12 is the documented starting version for this project; follow its README if it specifies a more exact patch release.

Windows PowerShell commands
py -3.12 -m venv .venv
.\.venv\Scripts\python.exe -m pip install -r requirements.txt
.\.venv\Scripts\python.exe -m pip check
macOS / Linux commands
python3.12 -m venv .venv
.venv/bin/python -m pip install -r requirements.txt
.venv/bin/python -m pip check

No global package installation or machine-wide policy changes are necessary. For Windows, explicit environment Python avoids PowerShell activation-policy issues. PyTorch or downloaded datasets may require substantial disk space.

Optional: publish a small demo safely
  1. Finish the local test, save a screenshot and check your actual saved model or index works after restart.
  2. Use a repository you control. Exclude API keys, .env files, personal uploads, unlicensed datasets and generated sensitive artifacts.
  3. Choose a host that supports your actual Python and system dependencies. If a model or data file is generated locally, plan a permitted and reproducible build step before expecting a cloud demo to start.
  4. Test the real hosted application on desktop and mobile, including invalid inputs, empty answers, missing model files and service restarts.
  5. Do not expose a paid AI key or an unrestricted inference endpoint to the public; add user authentication, rate limits and spending limits first. Keep a local-only demonstration if you cannot protect it.

This is an optional safety checklist, not a claim that any project already has a public deployed demo.

Important limitation: One-day-ahead forecasts using observed prior days are not 28-day recursive future forecasts.