Skip to content

About

ChurnCast: an end-to-end data science case study on subscription churn (pandas + scikit-learn, DuckDB SQL or tidymodels): point-in-time features, leakage tests, calibrated models, three-track parity, CI with an evaluation gate, a Dockerised API with a guard-railed AI explainer, monitoring and retraining. main = starter, solution = reference.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

ChurnCast: an end-to-end data science case study on subscription churn, with AI as your pair data scientist

CI CI (solution)

ChurnCast is a guided project from AiCanCode, the capstone of the Data Science track. Over 9 phases (about 40 hours) you take one messy subscription dataset the whole way to a churn model that a retention team can use: requirements, a data contract and a modelling plan, a test-first implementation, data and model tests, CI, a prediction API in Docker, monitoring and retraining after the model goes live, and a retrospective. AI assistants help at every step, and you verify every number they produce.

The situation. TiffinBox is a meal-subscription startup in Pune, Bengaluru, Hyderabad, Nagpur and Indore. About 9 % of paying customers do not renew each month. The retention team can call 10 % of active customers a month with a discount offer, and today they choose whom to call with a rule of thumb (five risk flags). The Head of Growth asks: at the end of each month, who is most likely to cancel next month, and why? Your job is a model that beats the rule on the people the team can actually call, with probabilities honest enough for finance to budget with, and explanations a retention agent can act on.

The data. A 17-month export (January 2025 to May 2026) in four deliberately imperfect CSVs, with the problems real exports have, and one trap: customers.cancellation_reason is filled in after a customer cancels.

Table Raw rows Known problems
customers.csv 5,409 exact duplicates, two date formats, impossible dates, Bangalore / bengaluru / BENGALURU, missing channel and age band, a leaky column
invoices.csv 28,505 exact duplicates, amounts like ₹3,299 / 3,299.00, refunds and failed payments, orphans
usage.csv 27,036 exact duplicates, negative counts, orphans
tickets.csv 6,165 two timestamp formats, resolutions before creation, orphans

What you deliver: a pipeline that turns any export into point-in-time features, compares two baselines with two logistic regressions using grouped, out-of-time validation, and writes metrics.json and a language-neutral model.json; a model card generated from the numbers; a FastAPI prediction API with explanations; and a report (notebook or R Markdown) that runs top to bottom from a clean checkout. It is judged not on the highest AUC, but on whether the evaluation is honest: no leakage, a fair comparison with the rule of thumb, calibrated probabilities, a fairness check, and a clear statement of what the model cannot tell you.

Pick one track. The data, the contract, the golden files and the API are shared; only the pipeline code differs, and the three tracks must produce the same features to 1e-9 and the same metrics and model to 1e-6.

Track Stack Folder
Python Python 3.12 · pandas · scikit-learn · Jupyter · pytest · ruff tracks/python
SQL DuckDB SQL for cleaning, features, labels and every metric · sqlfluff · a thin Python runner (the logistic fit is scikit-learn) · pytest tracks/sql
R R 4.6 · tidymodels (parsnip, rsample, yardstick) · dplyr · R Markdown · testthat · covr · lintr, pinned in a Docker image tracks/r

The prediction API (serving/, from Phase 6) serves any track's model.json. Its optional AI explainer sees only numbered facts (the score and the top contributions, never a customer id or a protected attribute) and is checked by guardrails; offline, it falls back to a deterministic template. Everything runs on your laptop, with no API key and no paid service.

Course repository: https://github.com/AICanCode-org/ds-case-study (main = starter, solution = reference).


Repository layout

.
├── README.md                 ← you are here
├── AI_LOG.md                 ← your log of AI prompts, outputs and how you verified them
├── docs/                     ← the course: phase-0 … phase-8 + GitHub Actions explained
├── spec/                     ← requirements, modelling plan, ADRs (you write them in Phases 1-2)
├── contract/                 ← the shared contract: feature-spec.md (every rule), data-contract.yaml, the metrics
│                               and model schemas, the tiny fixture, metric cases and the golden files
├── data/raw/2026-05/         ← the export (CSV + export.json), generated by scripts/generate_data.py
├── tracks/{python,sql,r}/    ← pipeline code + tests (same Makefile targets in each)
├── serving/                  ← the prediction API + AI explainer (Phase 6)
├── models/                   ← released model artefacts (Phase 6)
├── scripts/                  ← doctor.sh, check-all.sh, verify-tags.sh, build-cms.py, parity.py, eval_gate.py,
│                               model_card.py, validate_outputs.py, generate_data.py (+ smoke-test.sh)
├── cms/                      ← machine-readable export of the course for the AiCanCode site
└── .github/                  ← CI workflows, issue/PR templates

Prerequisites (per operating system)

You need Git, make and Python 3.12+ (the scripts, the SQL runner and the API use it) plus the toolchain of one track. The R track's toolchain is a Docker image (so your R, lintr and tidymodels versions are exactly CI's); a local R 4.6 with make install works too. Docker is needed for every track from Phase 6. Run scripts/doctor.sh at any time.

Windows 10/11 macOS Linux (Ubuntu/Debian)
Shell WSL2 with Ubuntu: wsl --install in an admin PowerShell, reboot Terminal (zsh) any
Git, make inside WSL: sudo apt install git make xcode-select --install sudo apt install git make
Python (all tracks) inside WSL: sudo apt install python3 python3-venv brew install python@3.12 sudo apt install python3 python3-venv
R track Docker Desktop with the WSL 2 engine (or R + RStudio) Docker Desktop or Colima (or R from CRAN + RStudio) Docker Engine (or r-base + pandoc)
Docker (Phase 6+) Docker Desktop with the WSL 2 engine Docker Desktop (or Colima) Docker Engine + compose plugin

Windows, important: clone and work inside the WSL file system (~/code/ds-case-study), not under /mnt/c/.... Hardware: 8 GB RAM is plenty; the export is small on purpose (a full run takes seconds).

Quick start (5 minutes)

git clone https://github.com/<you>/ds-case-study.git && cd ds-case-study   # your fork (Phase 0)
scripts/doctor.sh                        # check your tools
cd tracks/python                         # or tracks/sql
make install                             # dependencies (creates .venv)
make lint test                           # all tests are "pending" (skipped) until Phase 3
python3 ../../scripts/generate_data.py --check   # the export in data/raw is exactly the published one

R track (Docker): build the pinned toolchain once, then run any target inside it:

docker build -f tracks/r/Dockerfile --target dev -t churncast-r-dev:local .
docker run --rm -v "$PWD":/repo -w /repo/tracks/r churncast-r-dev:local make lint test

Every track offers the same commands:

Command What it does
make install install dependencies
make lint linter + formatter check (ruff / ruff + sqlfluff / lintr)
make test unit tests on the tiny hand-made fixture (seconds)
make coverage unit + data and model tests on the real export, with an 85 % line-coverage gate
make run write out/features.csv, metrics.json, model.json, experiments.jsonl (DATA=../../data/raw/2026-05)
make report execute the notebook / SQL report / R Markdown report top to bottom
make gate the evaluation gate of out/metrics.json (exit 1 if the model may not ship)

Compare a track with the golden file (or with another track): python3 scripts/parity.py contract/golden/metrics-2026-05.json tracks/python/out/metrics.json.

How the phases work

Phase Topic You end at tag
0 Setup and the brief phase-0-end
1 Problem framing: the decision, the label and success criteria phase-1-end
2 Design: data contract, timeline, modelling plan and ADRs phase-2-end
3 Implementation in five milestones: clean, check, features, models, evaluation phase-3-end
4 Testing data and models: golden files, invariants, leakage, parity phase-4-end
5 CI with GitHub Actions and the evaluation gate phase-5-end
6 Model card, prediction API, AI explainer, Docker and delivery phase-6-end = v1.0.0
7 Monitoring, drift, retraining and a post-mortem phase-7-end = v1.1.0
8 Retrospective and portfolio phase-8-end

Each phase page has the same 8 blocks: Why it matters · Objectives · Step-by-step instructions · AI-assist prompts · Deliverables · Self-check quiz · What you learned · Catch-up git commands.

Branches and tags

  • main: the starter. Every track has its structure, the full (pending) test suite and a TODO with a hint in every function or SQL view; the notebook and the R Markdown report are skeletons with the questions and no answers.
  • solution: the reference, one or more commits per phase, an annotated tag at the end of every phase (phase-0-end … phase-8-end) and the releases v1.0.0 and v1.1.0. Every tag passes the checks of all three tracks.

Stuck or behind? Start the next phase from the reference end of the previous one:

git fetch upstream --tags
git switch -c my-phase-4 phase-3-end

Compare with the reference: git diff phase-3-end -- tracks/python. Try first, then compare.

Pending tests

The complete test suite is already on main but switched off in one list per track (tests/conftest.py for Python and SQL, tests/testthat/helper-churncast.R for R). In Phase 3 you remove one entry at a time, watch the tests fail, and make them pass. Skipped is never reported as passed.

Working with AI

Use any assistant you like. Each phase has copy-paste prompts and a "Verify the output by" checklist; log what you asked and how you checked it in AI_LOG.md. Three rules matter most in modelling work: never paste row-level customer data or secrets into a prompt (aggregates and your own code are fine); never accept a metric, a formula or a claim about a model without reproducing it (the golden files, the metric cases and the other tracks are there to check it); and be most suspicious when a suggestion makes the model look better: that is what leakage looks like.

Coming from the old six-step guide?

The earlier version (Brief · Data · EDA · Modelling · Evaluation · Presentation) had the right story but no data, no tests and no way to know if your numbers were right. This version ships the data, a contract with every rule written down, golden files, three tracks that check each other, CI with an evaluation gate, a prediction API and a second release after the model meets the real world. The mapping of the old steps is in docs/README.md.

More

Questions or problems: open an issue using the templates, or contact us via https://www.aicancode.org/contact.

About

ChurnCast: an end-to-end data science case study on subscription churn (pandas + scikit-learn, DuckDB SQL or tidymodels): point-in-time features, leakage tests, calibrated models, three-track parity, CI with an evaluation gate, a Dockerised API with a guard-railed AI explainer, monitoring and retraining. main = starter, solution = reference.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages