Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,26 @@ Refactors, CI, and formatting land in the git history, not here.

## [Unreleased]

### Added

- **Release history: `GET /v3/releases/changes?version=N` and the MCP tool
`get_release_changes(version)`** — what changed in an IDC release versus the previous one:
series added / revised / removed (with size in TB and patients affected), collections that are
new / updated / removed, and analysis results that gained series. Any past release can be
described, computed from the served release's `index` + `prior_versions_index`; `version`
defaults to the served release.
- MCP prompt **`whats_new`** (optional `version`) — a ready-made "summarize this release" request.
- `GET /v3/tables` / `list_tables`: each table now carries **`notable_columns`**, a few columns
worth knowing before writing SQL (e.g. `series_init_idc_version` on `index`).

### Changed

- `prior_versions_index` is now documented: a table description and descriptions for
`min_idc_version`, `max_idc_version`, `crdc_series_uuid`, and `series_size_MB` (upstream ships
them empty).
- The `get_idc_version` tool description no longer implies the server can only speak to one
release; it points at `get_release_changes` and the version-history columns.

## [3.0.0b3] — 2026-08-10

Beta iteration: one shape for every cohort filter, and no request that silently answers with the
Expand Down
40 changes: 37 additions & 3 deletions docs/user-guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,7 +54,7 @@ understanding, because picking the right one makes everything else easy:

| Surface | Answers | REST | MCP tools |
|---|---|---|---|
| **Discovery** | "What exists? What can I filter on?" | `GET /v3/version`, `/v3/stats`, `/v3/collections`, `/v3/collections/{id}`, `/v3/analysis_results`, `/v3/attributes`, `/v3/attributes/{attr}/values` | `get_idc_version`, `get_stats`, `list_collections`, `get_collection`, `list_analysis_results`, `list_attributes`, `get_attribute_values` |
| **Discovery** | "What exists? What can I filter on? What changed in a release?" | `GET /v3/version`, `/v3/releases/changes`, `/v3/stats`, `/v3/collections`, `/v3/collections/{id}`, `/v3/analysis_results`, `/v3/attributes`, `/v3/attributes/{attr}/values` | `get_idc_version`, `get_release_changes`, `get_stats`, `list_collections`, `get_collection`, `list_analysis_results`, `list_attributes`, `get_attribute_values` |
| **Cohort** | "How big is *my* selection, and what's in it?" | `POST /v3/cohort/counts`, `POST /v3/cohort/manifest` | `build_cohort` |
| **Retrieval** | "Give me the download links" | `POST /v3/cohort/manifest.txt` | `get_cohort_urls` |
| **SQL** | "Run my custom query" + schema | `GET /v3/tables`, `/v3/tables/{table}`, `POST /v3/sql` | `list_tables`, `get_table_schema`, `run_sql` |
Expand Down Expand Up @@ -144,7 +144,7 @@ per-collection, keyed by `collection_id`).
| `index` | one row per **series** — the main table |
| `collections_index` | one row per collection (curated metadata) |
| `analysis_results_index` | one row per analysis result |
| `version_metadata_index` / `prior_versions_index` | IDC release versions / removed series |
| `version_metadata_index` / `prior_versions_index` | IDC release dates / superseded or removed series versions (`min_idc_version`–`max_idc_version`) — see [Release history](#release-history) |
| `seg_index`, `ann_index`, `ann_group_index`, `rtstruct_index` | segmentations / annotations / RT structures: **what was segmented** (`SegmentedPropertyType_CodeMeanings` — `BodyPartExamined` reflects the source acquisition, not this) and the **reference** to the image series they derive from (`segmented_SeriesInstanceUID` / `referenced_SeriesInstanceUID`) |
| `ct_index`, `mr_index`, `pt_index` | per-modality acquisition parameters (slice thickness, kVp, TE/TR, injected dose…) |
| `sm_index`, `sm_instance_index` | slide-microscopy (pathology) series / instance metadata |
Expand Down Expand Up @@ -215,6 +215,34 @@ WHERE i.collection_id = 'nlst' AND i.Modality = 'CT'
> `IDC_API_INCLUDE_INDICES=all` includes it; the clinical tools return a clear "not included"
> error otherwise).

### Release history

The server serves **one** IDC data release (see `GET /v3/version` / `get_idc_version`), but that
release carries the history of every earlier one, so "what's new in v24" or "what changed since
v20" are answerable without external release notes:

- On `index`, `series_init_idc_version` is the release a series first appeared in, and
`series_revised_idc_version` the release its current content dates from.
- `prior_versions_index` holds every **superseded or removed version** of a series — one row per
old `crdc_series_uuid`, valid from `min_idc_version` through `max_idc_version`. If its
`SeriesInstanceUID` is still in `index` the series was revised; otherwise it was removed.
- `version_metadata_index` dates each release.

`GET /v3/releases/changes?version=N` / `get_release_changes(version=N)` does the diff of release
N against N-1 for you: series **added / revised / removed** (with TB and patients affected),
collections that are **new / updated / removed**, and the analysis results that gained series.
Omit `version` for the served release. For per-series detail, query the columns above with SQL —
e.g. the series of a collection that changed since v20:

```sql
SELECT SeriesInstanceUID, series_init_idc_version, series_revised_idc_version
FROM index
WHERE collection_id = 'nlst' AND series_revised_idc_version > 20
```

The diff tells you *what* changed, not *why*; for the narrative, see the
[IDC release notes](https://learn.canceridc.dev/data/data-release-notes).

---

## 2. Using the REST API
Expand All @@ -230,13 +258,14 @@ uv run idc-api # http://127.0.0.1:8000 — Swagger UI at /v3/docs
| Method & path | Purpose |
|---|---|
| `GET /v3/version` | IDC data release served (e.g. `v24`) + pinned index version, **and** this server's own software version (`api_version`, plus `build` if the deploy stamped one) |
| `GET /v3/releases/changes?version=` | What changed in an IDC release vs. the previous one: series added / revised / removed, new / updated / removed collections, analysis results (default: the served release) |
| `GET /v3/stats` | Headline totals (collections, patients, studies, series, size_TB) |
| `GET /v3/collections` | List collections (datasets) |
| `GET /v3/collections/{id}` | Collection detail: counts, modalities, license breakdown |
| `GET /v3/analysis_results` | Derived datasets (segmentations/annotations) |
| `GET /v3/attributes` | Filterable attributes (name, type, term/range, categorical) |
| `GET /v3/attributes/{attr}/values?limit=` | Distinct values + counts for an attribute, plus a `note` caveat when one applies (e.g. `BodyPartExamined` ≠ segmented anatomy) |
| `GET /v3/tables` | Tables available to SQL |
| `GET /v3/tables` | Tables available to SQL, each with a few `notable_columns` |
| `GET /v3/tables/{table}` | Column schema for a table |
| `GET /v3/clinical/tables?collection_id=` | Per-collection clinical tables (optionally one collection) |
| `GET /v3/clinical/tables/{table}` | Clinical table columns + human-readable labels |
Expand Down Expand Up @@ -285,6 +314,8 @@ curl -s 'localhost:8000/v3/attributes/Modality/values?limit=10'

```bash
curl -s localhost:8000/v3/version # data release + this server's build
curl -s localhost:8000/v3/releases/changes # what's new in the served release
curl -s 'localhost:8000/v3/releases/changes?version=23' # …or in any earlier one
curl -s localhost:8000/v3/stats # headline totals
curl -s localhost:8000/v3/collections # list datasets
curl -s localhost:8000/v3/collections/nlst # one collection's detail
Expand Down Expand Up @@ -368,13 +399,16 @@ uv run idc-mcp --http --host 0.0.0.0 --port 8080 # hosted/shared

- **Discovery:** `get_idc_version`, `get_stats`, `list_collections`, `get_collection`,
`list_analysis_results`, `list_attributes`, `get_attribute_values`
- **Release history:** `get_release_changes`
- **Schema (for SQL):** `list_tables`, `get_table_schema`
- **Clinical data:** `list_clinical_tables`, `get_clinical_table_schema`, `get_clinical_table`
- **Cohort / query:** `build_cohort`, `run_sql`
- **Retrieval & side tools:** `get_cohort_urls`, `get_viewer_url`,
`get_citations`, `get_licenses`
- **Resources:** `idc://guide` (data model + recommended workflow), `idc://tables`,
`idc://schema/{table}`
- **Prompts:** `whats_new` (optional `version`) — a ready-made "summarize what's new in this
release" request for clients that show MCP prompts

Tool descriptions are prescriptive about *when* to call each one, and the server ships an
`idc://guide` resource with the same conceptual model as this document — so a capable agent can
Expand Down
2 changes: 2 additions & 0 deletions src/idc_api/core/context.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@
LicenseService,
ManifestService,
QueryService,
ReleaseService,
ViewerService,
)

Expand All @@ -22,6 +23,7 @@ def __init__(self, settings: Settings | None = None):
self.settings = settings or get_settings()
self.backend = DuckDBBackend(self.settings)
self.discovery = DiscoveryService(self.backend)
self.releases = ReleaseService(self.backend)
self.cohort = CohortService(self.backend, self.settings)
self.manifest = ManifestService(self.backend, self.settings)
self.query = QueryService(self.backend, self.settings)
Expand Down
62 changes: 62 additions & 0 deletions src/idc_api/core/models.py
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,63 @@ class Stats(BaseModel):
size_TB: float


class CollectionChange(BaseModel):
collection_id: str
status: str = Field(

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: status only ever holds 'new' | 'removed' | 'updated'. Literal["new", "removed", "updated"] would validate it and put the enum in /v3/openapi.json and the MCP output schema.

...,
description="'new' (no series in the previous release), 'removed' (no series in this "
"release), or 'updated' (present in both, with series added/revised/removed).",
)
series_added: int
series_revised: int
series_removed: int
size_TB_added: float
size_TB_revised: float
size_TB_removed: float
patients_affected: int = Field(
..., description="Distinct patients with at least one added, revised, or removed series."
)


class AnalysisResultChange(BaseModel):
analysis_result_id: str
is_new: bool = Field(..., description="True if this analysis result first appeared here.")
series_added: int


class ReleaseChanges(BaseModel):
"""What changed in one IDC release relative to the one before it."""

idc_version: str = Field(..., description="The release described, e.g. 'v24'.")
release_date: str | None = None
previous_version: str | None = Field(None, description="The release compared against.")
previous_release_date: str | None = None
current_version: str = Field(..., description="The release this server serves.")
series_added: int
series_revised: int
series_removed: int
size_TB_added: float
size_TB_revised: float
size_TB_removed: float
patients_affected: int
new_collections: list[str] = Field(
default_factory=list, description="Collections that first appeared in this release."
)
removed_collections: list[str] = Field(
default_factory=list, description="Collections with no series left in this release."
)
collections: list[CollectionChange] = Field(
default_factory=list, description="Per-collection change, largest additions first."
)
analysis_results: list[AnalysisResultChange] = Field(
default_factory=list,
description="Analysis results that gained series in this release. Counted from series "
"still present in the current release, so for an older release it omits series that were "
"later removed.",
)
note: str = ""


class CollectionSummary(BaseModel):
collection_id: str
collection_name: str | None = None
Expand Down Expand Up @@ -117,6 +174,11 @@ class TableInfo(BaseModel):
name: str
description: str = ""
column_count: int
notable_columns: list[str] = Field(
default_factory=list,
description="A few columns worth knowing about before writing SQL (not the full list — "
"get_table_schema has every column with its description).",
)


class TableList(BaseModel):
Expand Down
50 changes: 48 additions & 2 deletions src/idc_api/core/schema.py
Original file line number Diff line number Diff line change
Expand Up @@ -151,21 +151,67 @@ def _column_type(c: dict) -> str:
"(join to index on dicom_patient_id = index.PatientID). Use this table's column_label and "
"value mappings to interpret their often-cryptic coded columns."
),
# Upstream ships this table with no description at all.
"prior_versions_index": (
"Series versions IDC served in an earlier release but no longer serves: one row per "
"superseded or removed version of a series (identified by crdc_series_uuid), valid from "
"min_idc_version through max_idc_version. A SeriesInstanceUID that is also in `index` was "
"revised (its current version starts at index.series_revised_idc_version); one that is "
"not was removed. Together with `index` this reconstructs the content of any past release "
"— get_release_changes does that diff for you."
),
}

# Fill-ins for upstream column descriptions that are empty. Applied only where upstream has no
# text, so an upstream description takes over automatically once it lands.
COLUMN_DESCRIPTION_FILLINS: dict[str, dict[str, str]] = {
"prior_versions_index": {
"crdc_series_uuid": "Identifier of this specific version of the series (changes on "
"every revision); never equal to a crdc_series_uuid in `index`.",
"min_idc_version": "First IDC release (integer) that served this version of the series.",
"max_idc_version": "Last IDC release (integer) that served this version of the series; "

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The data contradicts these descriptions. prior_versions_index is not one row per version: 2,253 crdc_series_uuids have two contiguous rows that differ only by bucket. So "the next release revised or removed it" is wrong for those rows (see the comment on releases.py:38). The same claim is in the table description (line 155), _GUIDE, and docs/user-guide.md. An LLM that follows these docs in run_sql (e.g. max_idc_version = 19 AND SeriesInstanceUID IN index meaning "revised in v20") will count those 2,253 bucket moves as revisions.

Also consider adding crdc_series_uuid to NOTABLE_COLUMNS for this table, since it's the column that tells the two cases apart.

"the next release revised or removed it.",
"series_size_MB": "Size of this version of the series, in MB.",
},
}

# A handful of columns per table surfaced by list_tables, so a caller learns they exist without
# having to call get_table_schema first (a caller that guesses the obvious columns right never
# does). Not a full listing; names absent from a table's schema are dropped.
NOTABLE_COLUMNS: dict[str, list[str]] = {
"index": [
"collection_id",
"analysis_result_id",
"PatientID",
"SeriesInstanceUID",
"Modality",
"BodyPartExamined",
"SeriesDescription",
"license_short_name",
"series_size_MB",
"series_init_idc_version",
"series_revised_idc_version",
],
"version_metadata_index": ["idc_version", "version_timestamp"],
"prior_versions_index": ["SeriesInstanceUID", "min_idc_version", "max_idc_version"],
"seg_index": ["segmented_SeriesInstanceUID", "SegmentedPropertyType_CodeMeanings"],
}


@lru_cache(maxsize=None)
def table_schema(table: str) -> dict:
"""Return ``{name, description, columns:[{name,type,description}]}`` for a table,
sourced from the idc-index schema JSON shipped in INDEX_METADATA (table descriptions may be
repointed inward via ``TABLE_DESCRIPTION_OVERRIDES``)."""
repointed inward via ``TABLE_DESCRIPTION_OVERRIDES``; empty column descriptions are filled
from ``COLUMN_DESCRIPTION_FILLINS``)."""
meta = idc_index_data.INDEX_METADATA[metadata_key(table)]
schema = meta.get("schema", {}) or {}
fillins = COLUMN_DESCRIPTION_FILLINS.get(table, {})
columns = [
{
"name": c["name"],
"type": _column_type(c),
"description": c.get("description", "") or "",
"description": c.get("description", "") or fillins.get(c["name"], ""),
}
for c in schema.get("columns", [])
]
Expand Down
2 changes: 2 additions & 0 deletions src/idc_api/core/services/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@
from .licenses import LicenseService
from .manifest import ManifestService
from .query import QueryService
from .releases import ReleaseService
from .viewer import ViewerService

__all__ = [
Expand All @@ -21,5 +22,6 @@
"LicenseService",
"ManifestService",
"QueryService",
"ReleaseService",
"ViewerService",
]
2 changes: 2 additions & 0 deletions src/idc_api/core/services/query.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,11 +17,13 @@ def list_tables(self) -> TableList:
tables = []
for name in self.backend.list_tables():
sch = schema.table_schema(name)
names = {c["name"] for c in sch["columns"]}
tables.append(
TableInfo(
name=name,
description=sch["description"],
column_count=len(sch["columns"]),
notable_columns=[c for c in schema.NOTABLE_COLUMNS.get(name, []) if c in names],
)
)
return TableList(tables=tables)
Expand Down
Loading
Loading