Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,12 @@

## Unreleased — documentation and integrity audit

`GET /api/v1/publications/export/manifest` now exports bounded persisted publication metadata with origin labels, revisions, content hashes, retrieval timestamps, matching totals, and a SHA-256 fingerprint covering the manifest and exported items. Each page and count use one PostgreSQL snapshot; separate page requests remain independent. This metadata inventory includes explicitly labeled synthetic records and does not grant citation eligibility. Tests cover fingerprint verification, metadata changes, database requirements, pagination bounds, and persisted revision/freshness behavior.

`GET /api/v1/evidence/export/citation/schema` now exposes the citation-export contract for clients, including schema version, scoring method, required human-review metadata, known exclusion reasons, manifest fields, and the research disclaimer. This gives frontends and integrations a read-only capability-discovery path before requesting an export.

`create_app(database_url=None)` now explicitly disables database-backed repositories, while `create_app()` without a database argument still reads `DATABASE_URL` from the environment. Fixture-only tests can therefore state their intended boundary directly instead of deleting CI environment variables, reducing accidental PostgreSQL coupling in tests that are not exercising persistence.

Citation export now accepts `scoring_as_of` as an ISO 8601 query parameter, matching the evidence-list scoring contract. The normalized UTC time flows into the citation manifest's observed scoring-time set, allowing clients to recreate the same temporal scoring basis and compare SHA-256 export fingerprints without depending on the day the request is executed. Invalid scoring timestamps return the existing structured `INVALID_SCORING_AS_OF` error.

Persisted publication responses now expose an explicit `origin` contract with `unknown`, `manual`, `provider`, and `synthetic` states plus a tri-state `synthetic` interpretation. Synthetic seeds and `SYN-*` or `SEED-*` identifiers remain synthetic even when their payload has provider-shaped metadata. Legacy rows without a documented origin are normalized conservatively as `unknown` rather than being presented as real provider retrievals. Revision history carries the normalized origin fields in each payload.
Expand Down
70 changes: 69 additions & 1 deletion docs/API.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,13 +7,14 @@ pip install -e ".[dev,api,db]"
uvicorn 'openlongevity.api:create_app' --factory --reload
```

The service separates persisted publications from synthetic evidence demonstrations. This reference describes baseline `9fddcbb`; the running application's OpenAPI schema remains the interface to inspect for a different commit. PostgreSQL must be configured and migrated for publication operations. The operator ingestion key is an implemented control; complete user-role management and a production operating environment remain separate requirements.
The service separates persisted publications from synthetic evidence demonstrations. This reference describes baseline `9fddcbb`; the running application's OpenAPI schema remains the interface to inspect for a different commit. PostgreSQL must be configured and migrated for publication operations. Application construction now distinguishes omitted database configuration from explicit fixture-only operation: `create_app()` reads `DATABASE_URL` from the environment, while `create_app(database_url=None)` disables database-backed repositories even if the environment contains a database URL. The operator ingestion key is an implemented control; complete user-role management and a production operating environment remain separate requirements.

## Read-only exploration

```bash
curl http://localhost:8000/api/v1/health
curl 'http://localhost:8000/api/v1/evidence?topic=senescence'
curl 'http://localhost:8000/api/v1/evidence/export/citation/schema'
curl 'http://localhost:8000/api/v1/evidence/export/citation?topic=senescence&scoring_as_of=2021-01-01T00:00:00Z'
curl 'http://localhost:8000/api/v1/search?query=senescence&page=1&page_size=10'
curl 'http://localhost:8000/api/v1/research-gaps?topic=senescence'
Expand Down Expand Up @@ -63,6 +64,71 @@ Publication responses include two related interpretation fields: `origin` and `s

Publication history returns the normalized origin fields inside each revision payload. The revision content hash continues to identify the stored revision content at the time it was written; the response may also annotate legacy payloads with conservative origin fields for client clarity. Clients should therefore compare revision numbers and hashes for local history, and use `origin` and `synthetic` for interpretation. A change from `unknown` to `manual`, `provider`, or `synthetic` is a content-contract change that can create a new revision when saved through the repository.

### Publication metadata export manifest

`GET /api/v1/publications/export/manifest` exports a bounded page of persisted
publication metadata under schema `publication-export-v1`. The route accepts
`query`, `page`, and `page_size` with the same bounds and title-only filtering as
the publication list. For example:

```bash
curl --get 'http://localhost:8000/api/v1/publications/export/manifest' \
--data-urlencode 'query=senescence' \
--data-urlencode 'page=1' --data-urlencode 'page_size=20'
```

The response mode is `publication-metadata-export`. Each item includes its local
identifier, title, provider, source identifier, origin, tri-state synthetic label,
revision, stored content hash, and first and last retrieval timestamps. Unknown
origin retains `synthetic: null`; synthetic records remain explicitly marked and
are included in this metadata inventory. Inclusion does not establish human
review, citation eligibility, or scientific validity. The evidence citation
export has a separate review and exclusion policy.

The manifest records the normalized query, pagination parameters, total matching
rows, exported count, ordered identifiers, revision map, and content-hash map.
Top-level `total` counts only the exported items. An out-of-range page returns an
empty list while retaining the matching total. Rows are ordered by identifier.
The count and rows for one request use a PostgreSQL repeatable-read snapshot;
separate page requests do not share a snapshot. Ingestion during a multi-page
export can therefore change membership. This endpoint does not yet implement a
frozen whole-corpus export or provide the complete publication payload needed to
restore a database.

The SHA-256 `export_fingerprint` covers both the manifest without its fingerprint
field and every exported item. The exact verification procedure is:

```python
import hashlib
import json

# export is the decoded JSON response from this endpoint.
manifest = dict(export["manifest"])
expected = manifest.pop("export_fingerprint")
canonical = json.dumps(
{"manifest": manifest, "items": export["items"]},
sort_keys=True, separators=(",", ":"), ensure_ascii=True,
).encode("utf-8")
assert hashlib.sha256(canonical).hexdigest() == expected
```

Repeated requests with identical metadata and parameters produce the same
fingerprint. An unchanged publication retrieved again may keep its revision and
content hash while changing `last_retrieved_at`; that freshness change changes the
export fingerprint. A revision hash identifies repository content using the
repository's own serialization contract. It is not the hash of this smaller
metadata projection. The export fingerprint detects accidental changes when
compared against a trusted saved value; it is not a signature or proof of source
authenticity. Archive the response with the repository commit and retrieval
context for later comparisons.

Without database configuration the route returns `503 DATABASE_NOT_CONFIGURED`.
A storage failure returns `503 DATABASE_UNAVAILABLE`; invalid pagination or an
oversized query returns `422 INVALID_REQUEST`. Unit coverage verifies independent
fingerprint recomputation and metadata changes. PostgreSQL integration coverage
checks pagination, origin preservation, revision history, and refreshed retrieval
metadata against actual persisted rows.

## Ingestion request and transaction semantics

The ingestion body requires a nonempty query of at most two hundred characters and a limit between one and twenty-five, defaulting to five. Additional body fields are rejected by the request model. Whitespace-only queries are rejected when constructing the provider search query. The request should be sent by an authorized operator, and the server must have both the ingestion key and a configured publication repository before useful work can proceed.
Expand All @@ -77,6 +143,8 @@ The ingestion response uses the publication-page shape, but its total is the num

Evidence, evidence detail, and research-gap routes operate on synthetic fixtures at this baseline. Their behavior is useful for interface and heuristic tests. It is not a live extraction of persisted publications. The graph response is also illustrative. A frontend should label these demonstrations wherever the results appear, including copied summaries and exports, because the origin distinction can otherwise be lost when a response is separated from its route.

`GET /api/v1/evidence/export/citation/schema` exposes the export contract without running an export. It reports the citation-export schema version, accepted source-mode labels, the navigation-score method version, required human-review status and metadata fields, known exclusion reasons, manifest fields, and the research disclaimer. Clients should use this route for capability discovery and validation messages instead of hard-coding policy text into a frontend. The route is read-only and does not indicate that any specific record is citation eligible.

`GET /api/v1/evidence/export/citation` is the first executable export boundary for the fixture corpus. It returns `mode: citation-eligible`, an `items` list, an `excluded` list, totals for both lists, a schema version, and the research disclaimer. Under the current fixture-only evidence mode, `SYN-*` records are excluded with reason `synthetic_fixture`, so the citation-eligible item list is empty for the bundled cellular-senescence demonstration. Non-synthetic records must also carry `review_status: verified` plus reviewer identity, review timestamp, and review notes before they can enter the citation-eligible item list. This is intentional: the route proves that the platform can reject demonstration data and unverified evidence rather than allowing attractive records to leak into citation workflows.

The citation export route should not be described as a complete publication export system. It does not yet produce bibliographic formats, human-review certificates, or provider-backed evidence bundles. It establishes a narrow behavior that was previously documented only as a policy: synthetic fixtures are not observations and are excluded by default from citation-eligible evidence export. The response now includes a deterministic manifest with included and excluded identifiers, exclusion reasons, scoring metadata, and a SHA-256 `export_fingerprint` so repeated exports can be compared. The route accepts the same optional `scoring_as_of` ISO 8601 query parameter used by the evidence list, normalizes it to UTC, and carries the observed value set into `manifest.scoring_as_of`. Supplying the timestamp is the recommended way to produce a reproducible export fingerprint across different days. The `excluded` list remains a compact audit summary with identifier, title, and exclusion reason; the manifest is the comparison surface for temporal scoring metadata. Future work can extend the same contract to persisted publication records once review status and source authenticity are implemented for that path.
Expand Down
30 changes: 27 additions & 3 deletions src/openlongevity/api.py
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,8 @@
from .exports import (
EVIDENCE_SCORE_METHOD_VERSION,
build_citation_export,
build_publication_export_manifest,
citation_export_schema,
evidence_record_payload,
)
from .gaps import ResearchGapDetector
Expand Down Expand Up @@ -113,13 +115,19 @@ class RevisionResponse(BaseModel):
retrieved_at: str


_DATABASE_URL_UNSET = object()


def create_app(
database_url: str | None = None, provider: PubMedProvider | None = None,
ingestion_key: str | None = None, review_key: str | None = None,
database_url: str | None | object = _DATABASE_URL_UNSET,
provider: PubMedProvider | None = None,
ingestion_key: str | None = None,
review_key: str | None = None,
) -> FastAPI:
from .repository import EvidenceReviewRepository, PublicationRepository

database_url = database_url or getenv("DATABASE_URL")
if database_url is _DATABASE_URL_UNSET:
database_url = getenv("DATABASE_URL")
database = Database(database_url) if database_url else None
repository = PublicationRepository(database) if database else None
review_repository = EvidenceReviewRepository(database) if database else None
Expand Down Expand Up @@ -236,6 +244,18 @@ async def publications(
return PublicationPage(items=[PublicationResponse.model_validate(item) for item in items],
total=total, page=page, page_size=page_size)

@app.get("/api/v1/publications/export/manifest")
async def publication_export_manifest(
query: str = Query(default="", max_length=200),
page: int = Query(default=1, ge=1, le=10000),
page_size: int = Query(default=20, ge=1, le=100),
) -> dict[str, Any]:
query = query.strip()
items, total = await require_repository().export_rows(query, page, page_size)
return build_publication_export_manifest(
items, query=query, page=page, page_size=page_size, total_matching=total,
)

@app.get("/api/v1/publications/{identifier}", response_model=PublicationResponse)
async def publication(identifier: str) -> PublicationResponse:
if len(identifier) > 160:
Expand Down Expand Up @@ -287,6 +307,10 @@ async def evidence(
for r in records], "mode": "fixture-only",
"summary": summary, "disclaimer": DISCLAIMER}

@app.get("/api/v1/evidence/export/citation/schema")
async def citation_export_contract() -> dict[str, Any]:
return citation_export_schema()

@app.get("/api/v1/evidence/export/citation")
async def citation_export(
topic: str = Query(default="", max_length=120),
Expand Down
79 changes: 79 additions & 0 deletions src/openlongevity/exports.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,14 +11,93 @@
from .models import EvidenceRecord, ReviewStatus

CITATION_EXPORT_SCHEMA_VERSION = "citation-export-v1"
PUBLICATION_EXPORT_SCHEMA_VERSION = "publication-export-v1"
EVIDENCE_SCORE_METHOD_VERSION = "navigation-score-v1"


CITATION_EXPORT_EXCLUSION_REASONS = {
"synthetic_fixture": (
"Synthetic fixtures and demonstration records are excluded from "
"citation-eligible exports."
),
"not_human_verified": "Records require verified human review metadata before citation export.",
}
CITATION_EXPORT_REQUIRED_REVIEW_FIELDS = (
"review_status",
"reviewed_by",
"reviewed_at",
"review_notes",
)


def citation_export_schema() -> dict[str, Any]:
"""Describe the citation export contract for API clients and reviewers."""
return {
"schema_version": CITATION_EXPORT_SCHEMA_VERSION,
"mode": "citation-eligible",
"source_modes": ["fixture-only", "unspecified"],
"score_method": EVIDENCE_SCORE_METHOD_VERSION,
"required_review_status": ReviewStatus.VERIFIED.value,
"required_review_fields": list(CITATION_EXPORT_REQUIRED_REVIEW_FIELDS),
"exclusion_reasons": dict(CITATION_EXPORT_EXCLUSION_REASONS),
"manifest_fields": [
"schema_version",
"source_mode",
"input_records",
"included_records",
"excluded_records",
"included_identifiers",
"excluded_identifiers",
"exclusion_reasons",
"score_methods",
"scoring_as_of",
"export_fingerprint",
],
"disclaimer": DISCLAIMER,
}


def _stable_fingerprint(payload: Mapping[str, Any]) -> str:
encoded = json.dumps(payload, sort_keys=True, separators=(",", ":")).encode("utf-8")
return hashlib.sha256(encoded).hexdigest()


def build_publication_export_manifest(
records: list[Mapping[str, Any]],
*,
query: str,
page: int,
page_size: int,
total_matching: int,
) -> dict[str, Any]:
"""Fingerprint a page of persisted metadata, including freshness and origin."""
fields = (
"identifier", "title", "provider", "source_identifier", "origin", "synthetic",
"revision", "content_hash", "first_retrieved_at", "last_retrieved_at",
)
items = [{field: record[field] for field in fields} for record in records]
manifest = {
"schema_version": PUBLICATION_EXPORT_SCHEMA_VERSION,
"source_mode": "persisted-publications",
"query": query,
"page": page,
"page_size": page_size,
"total_matching": total_matching,
"exported_records": len(items),
"exported_identifiers": [item["identifier"] for item in items],
"revision_map": {item["identifier"]: item["revision"] for item in items},
"content_hashes": {item["identifier"]: item["content_hash"] for item in items},
}
manifest["export_fingerprint"] = _stable_fingerprint({"manifest": manifest, "items": items})
return {
"schema_version": PUBLICATION_EXPORT_SCHEMA_VERSION,
"mode": "publication-metadata-export",
"manifest": manifest,
"items": items,
"total": len(items),
"disclaimer": DISCLAIMER,
}


def evidence_record_payload(
record: EvidenceRecord,
Expand Down
29 changes: 29 additions & 0 deletions src/openlongevity/repository.py
Original file line number Diff line number Diff line change
Expand Up @@ -88,6 +88,35 @@ async def list(self, query: str, page: int, page_size: int) -> tuple[list[dict[s
).offset((page - 1) * page_size).limit(page_size))
return [self.serialize(row) for row in rows], int(total or 0)

async def export_rows(
self, query: str, page: int, page_size: int,
) -> tuple[list[dict[str, Any]], int]:
"""Read a bounded page and its count from one PostgreSQL snapshot."""
async with self.database.sessions() as session:
await session.connection(execution_options={"isolation_level": "REPEATABLE READ"})
condition = PublicationRow.title.icontains(query, autoescape=True)
total = await session.scalar(
select(func.count()).select_from(PublicationRow).where(condition)
)
rows = await session.scalars(
select(PublicationRow).where(condition).order_by(PublicationRow.identifier)
.offset((page - 1) * page_size).limit(page_size)
)
return [
{
"identifier": row.identifier,
"title": row.title,
"provider": row.provider,
"source_identifier": row.source_identifier,
**publication_origin_fields(row.payload),
"revision": row.revision,
"content_hash": row.content_hash,
"first_retrieved_at": row.first_retrieved_at,
"last_retrieved_at": row.last_retrieved_at,
}
for row in rows
], int(total or 0)

async def history(self, identifier: str) -> list[dict[str, Any]]:
async with self.database.sessions() as session:
rows = await session.scalars(select(PublicationRevisionRow).where(
Expand Down
Loading
Loading