ChurnCast: an end-to-end data science case study on subscription churn, with AI as your pair data scientist
ChurnCast is a guided project from AiCanCode, the capstone of the Data Science track. Over 9 phases (about 40 hours) you take one messy subscription dataset the whole way to a churn model that a retention team can use: requirements, a data contract and a modelling plan, a test-first implementation, data and model tests, CI, a prediction API in Docker, monitoring and retraining after the model goes live, and a retrospective. AI assistants help at every step, and you verify every number they produce.
The situation. TiffinBox is a meal-subscription startup in Pune, Bengaluru, Hyderabad, Nagpur and Indore. About 9 % of paying customers do not renew each month. The retention team can call 10 % of active customers a month with a discount offer, and today they choose whom to call with a rule of thumb (five risk flags). The Head of Growth asks: at the end of each month, who is most likely to cancel next month, and why? Your job is a model that beats the rule on the people the team can actually call, with probabilities honest enough for finance to budget with, and explanations a retention agent can act on.
The data. A 17-month export (January 2025 to May 2026) in four deliberately imperfect CSVs, with the problems
real exports have, and one trap: customers.cancellation_reason is filled in after a customer cancels.
| Table | Raw rows | Known problems |
|---|---|---|
customers.csv |
5,409 | exact duplicates, two date formats, impossible dates, Bangalore / bengaluru / BENGALURU, missing channel and age band, a leaky column |
invoices.csv |
28,505 | exact duplicates, amounts like ₹3,299 / 3,299.00, refunds and failed payments, orphans |
usage.csv |
27,036 | exact duplicates, negative counts, orphans |
tickets.csv |
6,165 | two timestamp formats, resolutions before creation, orphans |
What you deliver: a pipeline that turns any export into point-in-time features, compares two baselines with
two logistic regressions using grouped, out-of-time validation, and writes metrics.json and a language-neutral
model.json; a model card generated from the numbers; a FastAPI prediction API with explanations; and a report
(notebook or R Markdown) that runs top to bottom from a clean checkout. It is judged not on the highest AUC, but
on whether the evaluation is honest: no leakage, a fair comparison with the rule of thumb, calibrated
probabilities, a fairness check, and a clear statement of what the model cannot tell you.
Pick one track. The data, the contract, the golden files and the API are shared; only the pipeline code differs, and the three tracks must produce the same features to 1e-9 and the same metrics and model to 1e-6.
| Track | Stack | Folder |
|---|---|---|
| Python | Python 3.12 · pandas · scikit-learn · Jupyter · pytest · ruff | tracks/python |
| SQL | DuckDB SQL for cleaning, features, labels and every metric · sqlfluff · a thin Python runner (the logistic fit is scikit-learn) · pytest | tracks/sql |
| R | R 4.6 · tidymodels (parsnip, rsample, yardstick) · dplyr · R Markdown · testthat · covr · lintr, pinned in a Docker image | tracks/r |
The prediction API (serving/, from Phase 6) serves any track's model.json. Its optional AI explainer sees
only numbered facts (the score and the top contributions, never a customer id or a protected attribute) and is
checked by guardrails; offline, it falls back to a deterministic template. Everything runs on your laptop, with
no API key and no paid service.
Course repository: https://github.com/AICanCode-org/ds-case-study (main = starter, solution = reference).
.
├── README.md ← you are here
├── AI_LOG.md ← your log of AI prompts, outputs and how you verified them
├── docs/ ← the course: phase-0 … phase-8 + GitHub Actions explained
├── spec/ ← requirements, modelling plan, ADRs (you write them in Phases 1-2)
├── contract/ ← the shared contract: feature-spec.md (every rule), data-contract.yaml, the metrics
│ and model schemas, the tiny fixture, metric cases and the golden files
├── data/raw/2026-05/ ← the export (CSV + export.json), generated by scripts/generate_data.py
├── tracks/{python,sql,r}/ ← pipeline code + tests (same Makefile targets in each)
├── serving/ ← the prediction API + AI explainer (Phase 6)
├── models/ ← released model artefacts (Phase 6)
├── scripts/ ← doctor.sh, check-all.sh, verify-tags.sh, build-cms.py, parity.py, eval_gate.py,
│ model_card.py, validate_outputs.py, generate_data.py (+ smoke-test.sh)
├── cms/ ← machine-readable export of the course for the AiCanCode site
└── .github/ ← CI workflows, issue/PR templates
You need Git, make and Python 3.12+ (the scripts, the SQL runner and the API use it) plus the toolchain of
one track. The R track's toolchain is a Docker image (so your R, lintr and tidymodels versions are exactly
CI's); a local R 4.6 with make install works too. Docker is needed for every track from Phase 6. Run scripts/doctor.sh at any time.
| Windows 10/11 | macOS | Linux (Ubuntu/Debian) | |
|---|---|---|---|
| Shell | WSL2 with Ubuntu: wsl --install in an admin PowerShell, reboot |
Terminal (zsh) | any |
| Git, make | inside WSL: sudo apt install git make |
xcode-select --install |
sudo apt install git make |
| Python (all tracks) | inside WSL: sudo apt install python3 python3-venv |
brew install python@3.12 |
sudo apt install python3 python3-venv |
| R track | Docker Desktop with the WSL 2 engine (or R + RStudio) | Docker Desktop or Colima (or R from CRAN + RStudio) | Docker Engine (or r-base + pandoc) |
| Docker (Phase 6+) | Docker Desktop with the WSL 2 engine | Docker Desktop (or Colima) | Docker Engine + compose plugin |
Windows, important: clone and work inside the WSL file system (~/code/ds-case-study), not under
/mnt/c/.... Hardware: 8 GB RAM is plenty; the export is small on purpose (a full run takes seconds).
git clone https://github.com/<you>/ds-case-study.git && cd ds-case-study # your fork (Phase 0)
scripts/doctor.sh # check your tools
cd tracks/python # or tracks/sql
make install # dependencies (creates .venv)
make lint test # all tests are "pending" (skipped) until Phase 3
python3 ../../scripts/generate_data.py --check # the export in data/raw is exactly the published oneR track (Docker): build the pinned toolchain once, then run any target inside it:
docker build -f tracks/r/Dockerfile --target dev -t churncast-r-dev:local .
docker run --rm -v "$PWD":/repo -w /repo/tracks/r churncast-r-dev:local make lint testEvery track offers the same commands:
| Command | What it does |
|---|---|
make install |
install dependencies |
make lint |
linter + formatter check (ruff / ruff + sqlfluff / lintr) |
make test |
unit tests on the tiny hand-made fixture (seconds) |
make coverage |
unit + data and model tests on the real export, with an 85 % line-coverage gate |
make run |
write out/features.csv, metrics.json, model.json, experiments.jsonl (DATA=../../data/raw/2026-05) |
make report |
execute the notebook / SQL report / R Markdown report top to bottom |
make gate |
the evaluation gate of out/metrics.json (exit 1 if the model may not ship) |
Compare a track with the golden file (or with another track):
python3 scripts/parity.py contract/golden/metrics-2026-05.json tracks/python/out/metrics.json.
| Phase | Topic | You end at tag |
|---|---|---|
| 0 | Setup and the brief | phase-0-end |
| 1 | Problem framing: the decision, the label and success criteria | phase-1-end |
| 2 | Design: data contract, timeline, modelling plan and ADRs | phase-2-end |
| 3 | Implementation in five milestones: clean, check, features, models, evaluation | phase-3-end |
| 4 | Testing data and models: golden files, invariants, leakage, parity | phase-4-end |
| 5 | CI with GitHub Actions and the evaluation gate | phase-5-end |
| 6 | Model card, prediction API, AI explainer, Docker and delivery | phase-6-end = v1.0.0 |
| 7 | Monitoring, drift, retraining and a post-mortem | phase-7-end = v1.1.0 |
| 8 | Retrospective and portfolio | phase-8-end |
Each phase page has the same 8 blocks: Why it matters · Objectives · Step-by-step instructions · AI-assist prompts · Deliverables · Self-check quiz · What you learned · Catch-up git commands.
main: the starter. Every track has its structure, the full (pending) test suite and aTODOwith a hint in every function or SQL view; the notebook and the R Markdown report are skeletons with the questions and no answers.solution: the reference, one or more commits per phase, an annotated tag at the end of every phase (phase-0-end…phase-8-end) and the releasesv1.0.0andv1.1.0. Every tag passes the checks of all three tracks.
Stuck or behind? Start the next phase from the reference end of the previous one:
git fetch upstream --tags
git switch -c my-phase-4 phase-3-endCompare with the reference: git diff phase-3-end -- tracks/python. Try first, then compare.
The complete test suite is already on main but switched off in one list per track (tests/conftest.py for
Python and SQL, tests/testthat/helper-churncast.R for R). In Phase 3 you remove one entry at a time, watch the
tests fail, and make them pass. Skipped is never reported as passed.
Use any assistant you like. Each phase has copy-paste prompts and a "Verify the output by" checklist; log what
you asked and how you checked it in AI_LOG.md. Three rules matter most in modelling work: never
paste row-level customer data or secrets into a prompt (aggregates and your own code are fine); never accept a
metric, a formula or a claim about a model without reproducing it (the golden files, the metric cases and the
other tracks are there to check it); and be most suspicious when a suggestion makes the model look better:
that is what leakage looks like.
The earlier version (Brief · Data · EDA · Modelling · Evaluation · Presentation) had the right story but no data,
no tests and no way to know if your numbers were right. This version ships the data, a contract with every rule
written down, golden files, three tracks that check each other, CI with an evaluation gate, a prediction API and a
second release after the model meets the real world. The mapping of the old steps is in
docs/README.md.
- docs/README.md: index of all docs
- docs/github-actions-explained.md: every CI concept used here
- contract/feature-spec.md: every cleaning rule, feature, metric and gate check
- CONTRIBUTING.md · CHANGELOG.md · LICENSE (MIT)
Questions or problems: open an issue using the templates, or contact us via https://www.aicancode.org/contact.