diff --git a/CHANGELOG.md b/CHANGELOG.md index 1f48c06..95cc3ef 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,12 @@ ## Unreleased — documentation and integrity audit +`GET /api/v1/publications/export/manifest` now exports bounded persisted publication metadata with origin labels, revisions, content hashes, retrieval timestamps, matching totals, and a SHA-256 fingerprint covering the manifest and exported items. Each page and count use one PostgreSQL snapshot; separate page requests remain independent. This metadata inventory includes explicitly labeled synthetic records and does not grant citation eligibility. Tests cover fingerprint verification, metadata changes, database requirements, pagination bounds, and persisted revision/freshness behavior. + +`GET /api/v1/evidence/export/citation/schema` now exposes the citation-export contract for clients, including schema version, scoring method, required human-review metadata, known exclusion reasons, manifest fields, and the research disclaimer. This gives frontends and integrations a read-only capability-discovery path before requesting an export. + +`create_app(database_url=None)` now explicitly disables database-backed repositories, while `create_app()` without a database argument still reads `DATABASE_URL` from the environment. Fixture-only tests can therefore state their intended boundary directly instead of deleting CI environment variables, reducing accidental PostgreSQL coupling in tests that are not exercising persistence. + Citation export now accepts `scoring_as_of` as an ISO 8601 query parameter, matching the evidence-list scoring contract. The normalized UTC time flows into the citation manifest's observed scoring-time set, allowing clients to recreate the same temporal scoring basis and compare SHA-256 export fingerprints without depending on the day the request is executed. Invalid scoring timestamps return the existing structured `INVALID_SCORING_AS_OF` error. Persisted publication responses now expose an explicit `origin` contract with `unknown`, `manual`, `provider`, and `synthetic` states plus a tri-state `synthetic` interpretation. Synthetic seeds and `SYN-*` or `SEED-*` identifiers remain synthetic even when their payload has provider-shaped metadata. Legacy rows without a documented origin are normalized conservatively as `unknown` rather than being presented as real provider retrievals. Revision history carries the normalized origin fields in each payload. diff --git a/docs/API.md b/docs/API.md index b61d93a..42c8112 100644 --- a/docs/API.md +++ b/docs/API.md @@ -7,13 +7,14 @@ pip install -e ".[dev,api,db]" uvicorn 'openlongevity.api:create_app' --factory --reload ``` -The service separates persisted publications from synthetic evidence demonstrations. This reference describes baseline `9fddcbb`; the running application's OpenAPI schema remains the interface to inspect for a different commit. PostgreSQL must be configured and migrated for publication operations. The operator ingestion key is an implemented control; complete user-role management and a production operating environment remain separate requirements. +The service separates persisted publications from synthetic evidence demonstrations. This reference describes baseline `9fddcbb`; the running application's OpenAPI schema remains the interface to inspect for a different commit. PostgreSQL must be configured and migrated for publication operations. Application construction now distinguishes omitted database configuration from explicit fixture-only operation: `create_app()` reads `DATABASE_URL` from the environment, while `create_app(database_url=None)` disables database-backed repositories even if the environment contains a database URL. The operator ingestion key is an implemented control; complete user-role management and a production operating environment remain separate requirements. ## Read-only exploration ```bash curl http://localhost:8000/api/v1/health curl 'http://localhost:8000/api/v1/evidence?topic=senescence' +curl 'http://localhost:8000/api/v1/evidence/export/citation/schema' curl 'http://localhost:8000/api/v1/evidence/export/citation?topic=senescence&scoring_as_of=2021-01-01T00:00:00Z' curl 'http://localhost:8000/api/v1/search?query=senescence&page=1&page_size=10' curl 'http://localhost:8000/api/v1/research-gaps?topic=senescence' @@ -63,6 +64,71 @@ Publication responses include two related interpretation fields: `origin` and `s Publication history returns the normalized origin fields inside each revision payload. The revision content hash continues to identify the stored revision content at the time it was written; the response may also annotate legacy payloads with conservative origin fields for client clarity. Clients should therefore compare revision numbers and hashes for local history, and use `origin` and `synthetic` for interpretation. A change from `unknown` to `manual`, `provider`, or `synthetic` is a content-contract change that can create a new revision when saved through the repository. +### Publication metadata export manifest + +`GET /api/v1/publications/export/manifest` exports a bounded page of persisted +publication metadata under schema `publication-export-v1`. The route accepts +`query`, `page`, and `page_size` with the same bounds and title-only filtering as +the publication list. For example: + +```bash +curl --get 'http://localhost:8000/api/v1/publications/export/manifest' \ + --data-urlencode 'query=senescence' \ + --data-urlencode 'page=1' --data-urlencode 'page_size=20' +``` + +The response mode is `publication-metadata-export`. Each item includes its local +identifier, title, provider, source identifier, origin, tri-state synthetic label, +revision, stored content hash, and first and last retrieval timestamps. Unknown +origin retains `synthetic: null`; synthetic records remain explicitly marked and +are included in this metadata inventory. Inclusion does not establish human +review, citation eligibility, or scientific validity. The evidence citation +export has a separate review and exclusion policy. + +The manifest records the normalized query, pagination parameters, total matching +rows, exported count, ordered identifiers, revision map, and content-hash map. +Top-level `total` counts only the exported items. An out-of-range page returns an +empty list while retaining the matching total. Rows are ordered by identifier. +The count and rows for one request use a PostgreSQL repeatable-read snapshot; +separate page requests do not share a snapshot. Ingestion during a multi-page +export can therefore change membership. This endpoint does not yet implement a +frozen whole-corpus export or provide the complete publication payload needed to +restore a database. + +The SHA-256 `export_fingerprint` covers both the manifest without its fingerprint +field and every exported item. The exact verification procedure is: + +```python +import hashlib +import json + +# export is the decoded JSON response from this endpoint. +manifest = dict(export["manifest"]) +expected = manifest.pop("export_fingerprint") +canonical = json.dumps( + {"manifest": manifest, "items": export["items"]}, + sort_keys=True, separators=(",", ":"), ensure_ascii=True, +).encode("utf-8") +assert hashlib.sha256(canonical).hexdigest() == expected +``` + +Repeated requests with identical metadata and parameters produce the same +fingerprint. An unchanged publication retrieved again may keep its revision and +content hash while changing `last_retrieved_at`; that freshness change changes the +export fingerprint. A revision hash identifies repository content using the +repository's own serialization contract. It is not the hash of this smaller +metadata projection. The export fingerprint detects accidental changes when +compared against a trusted saved value; it is not a signature or proof of source +authenticity. Archive the response with the repository commit and retrieval +context for later comparisons. + +Without database configuration the route returns `503 DATABASE_NOT_CONFIGURED`. +A storage failure returns `503 DATABASE_UNAVAILABLE`; invalid pagination or an +oversized query returns `422 INVALID_REQUEST`. Unit coverage verifies independent +fingerprint recomputation and metadata changes. PostgreSQL integration coverage +checks pagination, origin preservation, revision history, and refreshed retrieval +metadata against actual persisted rows. + ## Ingestion request and transaction semantics The ingestion body requires a nonempty query of at most two hundred characters and a limit between one and twenty-five, defaulting to five. Additional body fields are rejected by the request model. Whitespace-only queries are rejected when constructing the provider search query. The request should be sent by an authorized operator, and the server must have both the ingestion key and a configured publication repository before useful work can proceed. @@ -77,6 +143,8 @@ The ingestion response uses the publication-page shape, but its total is the num Evidence, evidence detail, and research-gap routes operate on synthetic fixtures at this baseline. Their behavior is useful for interface and heuristic tests. It is not a live extraction of persisted publications. The graph response is also illustrative. A frontend should label these demonstrations wherever the results appear, including copied summaries and exports, because the origin distinction can otherwise be lost when a response is separated from its route. +`GET /api/v1/evidence/export/citation/schema` exposes the export contract without running an export. It reports the citation-export schema version, accepted source-mode labels, the navigation-score method version, required human-review status and metadata fields, known exclusion reasons, manifest fields, and the research disclaimer. Clients should use this route for capability discovery and validation messages instead of hard-coding policy text into a frontend. The route is read-only and does not indicate that any specific record is citation eligible. + `GET /api/v1/evidence/export/citation` is the first executable export boundary for the fixture corpus. It returns `mode: citation-eligible`, an `items` list, an `excluded` list, totals for both lists, a schema version, and the research disclaimer. Under the current fixture-only evidence mode, `SYN-*` records are excluded with reason `synthetic_fixture`, so the citation-eligible item list is empty for the bundled cellular-senescence demonstration. Non-synthetic records must also carry `review_status: verified` plus reviewer identity, review timestamp, and review notes before they can enter the citation-eligible item list. This is intentional: the route proves that the platform can reject demonstration data and unverified evidence rather than allowing attractive records to leak into citation workflows. The citation export route should not be described as a complete publication export system. It does not yet produce bibliographic formats, human-review certificates, or provider-backed evidence bundles. It establishes a narrow behavior that was previously documented only as a policy: synthetic fixtures are not observations and are excluded by default from citation-eligible evidence export. The response now includes a deterministic manifest with included and excluded identifiers, exclusion reasons, scoring metadata, and a SHA-256 `export_fingerprint` so repeated exports can be compared. The route accepts the same optional `scoring_as_of` ISO 8601 query parameter used by the evidence list, normalizes it to UTC, and carries the observed value set into `manifest.scoring_as_of`. Supplying the timestamp is the recommended way to produce a reproducible export fingerprint across different days. The `excluded` list remains a compact audit summary with identifier, title, and exclusion reason; the manifest is the comparison surface for temporal scoring metadata. Future work can extend the same contract to persisted publication records once review status and source authenticity are implemented for that path. diff --git a/src/openlongevity/api.py b/src/openlongevity/api.py index 5ce73c2..56b5c6c 100644 --- a/src/openlongevity/api.py +++ b/src/openlongevity/api.py @@ -22,6 +22,8 @@ from .exports import ( EVIDENCE_SCORE_METHOD_VERSION, build_citation_export, + build_publication_export_manifest, + citation_export_schema, evidence_record_payload, ) from .gaps import ResearchGapDetector @@ -113,13 +115,19 @@ class RevisionResponse(BaseModel): retrieved_at: str +_DATABASE_URL_UNSET = object() + + def create_app( - database_url: str | None = None, provider: PubMedProvider | None = None, - ingestion_key: str | None = None, review_key: str | None = None, + database_url: str | None | object = _DATABASE_URL_UNSET, + provider: PubMedProvider | None = None, + ingestion_key: str | None = None, + review_key: str | None = None, ) -> FastAPI: from .repository import EvidenceReviewRepository, PublicationRepository - database_url = database_url or getenv("DATABASE_URL") + if database_url is _DATABASE_URL_UNSET: + database_url = getenv("DATABASE_URL") database = Database(database_url) if database_url else None repository = PublicationRepository(database) if database else None review_repository = EvidenceReviewRepository(database) if database else None @@ -236,6 +244,18 @@ async def publications( return PublicationPage(items=[PublicationResponse.model_validate(item) for item in items], total=total, page=page, page_size=page_size) + @app.get("/api/v1/publications/export/manifest") + async def publication_export_manifest( + query: str = Query(default="", max_length=200), + page: int = Query(default=1, ge=1, le=10000), + page_size: int = Query(default=20, ge=1, le=100), + ) -> dict[str, Any]: + query = query.strip() + items, total = await require_repository().export_rows(query, page, page_size) + return build_publication_export_manifest( + items, query=query, page=page, page_size=page_size, total_matching=total, + ) + @app.get("/api/v1/publications/{identifier}", response_model=PublicationResponse) async def publication(identifier: str) -> PublicationResponse: if len(identifier) > 160: @@ -287,6 +307,10 @@ async def evidence( for r in records], "mode": "fixture-only", "summary": summary, "disclaimer": DISCLAIMER} + @app.get("/api/v1/evidence/export/citation/schema") + async def citation_export_contract() -> dict[str, Any]: + return citation_export_schema() + @app.get("/api/v1/evidence/export/citation") async def citation_export( topic: str = Query(default="", max_length=120), diff --git a/src/openlongevity/exports.py b/src/openlongevity/exports.py index cf7060b..1c1c3dc 100644 --- a/src/openlongevity/exports.py +++ b/src/openlongevity/exports.py @@ -11,14 +11,93 @@ from .models import EvidenceRecord, ReviewStatus CITATION_EXPORT_SCHEMA_VERSION = "citation-export-v1" +PUBLICATION_EXPORT_SCHEMA_VERSION = "publication-export-v1" EVIDENCE_SCORE_METHOD_VERSION = "navigation-score-v1" +CITATION_EXPORT_EXCLUSION_REASONS = { + "synthetic_fixture": ( + "Synthetic fixtures and demonstration records are excluded from " + "citation-eligible exports." + ), + "not_human_verified": "Records require verified human review metadata before citation export.", +} +CITATION_EXPORT_REQUIRED_REVIEW_FIELDS = ( + "review_status", + "reviewed_by", + "reviewed_at", + "review_notes", +) + + +def citation_export_schema() -> dict[str, Any]: + """Describe the citation export contract for API clients and reviewers.""" + return { + "schema_version": CITATION_EXPORT_SCHEMA_VERSION, + "mode": "citation-eligible", + "source_modes": ["fixture-only", "unspecified"], + "score_method": EVIDENCE_SCORE_METHOD_VERSION, + "required_review_status": ReviewStatus.VERIFIED.value, + "required_review_fields": list(CITATION_EXPORT_REQUIRED_REVIEW_FIELDS), + "exclusion_reasons": dict(CITATION_EXPORT_EXCLUSION_REASONS), + "manifest_fields": [ + "schema_version", + "source_mode", + "input_records", + "included_records", + "excluded_records", + "included_identifiers", + "excluded_identifiers", + "exclusion_reasons", + "score_methods", + "scoring_as_of", + "export_fingerprint", + ], + "disclaimer": DISCLAIMER, + } + + def _stable_fingerprint(payload: Mapping[str, Any]) -> str: encoded = json.dumps(payload, sort_keys=True, separators=(",", ":")).encode("utf-8") return hashlib.sha256(encoded).hexdigest() +def build_publication_export_manifest( + records: list[Mapping[str, Any]], + *, + query: str, + page: int, + page_size: int, + total_matching: int, +) -> dict[str, Any]: + """Fingerprint a page of persisted metadata, including freshness and origin.""" + fields = ( + "identifier", "title", "provider", "source_identifier", "origin", "synthetic", + "revision", "content_hash", "first_retrieved_at", "last_retrieved_at", + ) + items = [{field: record[field] for field in fields} for record in records] + manifest = { + "schema_version": PUBLICATION_EXPORT_SCHEMA_VERSION, + "source_mode": "persisted-publications", + "query": query, + "page": page, + "page_size": page_size, + "total_matching": total_matching, + "exported_records": len(items), + "exported_identifiers": [item["identifier"] for item in items], + "revision_map": {item["identifier"]: item["revision"] for item in items}, + "content_hashes": {item["identifier"]: item["content_hash"] for item in items}, + } + manifest["export_fingerprint"] = _stable_fingerprint({"manifest": manifest, "items": items}) + return { + "schema_version": PUBLICATION_EXPORT_SCHEMA_VERSION, + "mode": "publication-metadata-export", + "manifest": manifest, + "items": items, + "total": len(items), + "disclaimer": DISCLAIMER, + } + def evidence_record_payload( record: EvidenceRecord, diff --git a/src/openlongevity/repository.py b/src/openlongevity/repository.py index 65426c1..5940a34 100644 --- a/src/openlongevity/repository.py +++ b/src/openlongevity/repository.py @@ -88,6 +88,35 @@ async def list(self, query: str, page: int, page_size: int) -> tuple[list[dict[s ).offset((page - 1) * page_size).limit(page_size)) return [self.serialize(row) for row in rows], int(total or 0) + async def export_rows( + self, query: str, page: int, page_size: int, + ) -> tuple[list[dict[str, Any]], int]: + """Read a bounded page and its count from one PostgreSQL snapshot.""" + async with self.database.sessions() as session: + await session.connection(execution_options={"isolation_level": "REPEATABLE READ"}) + condition = PublicationRow.title.icontains(query, autoescape=True) + total = await session.scalar( + select(func.count()).select_from(PublicationRow).where(condition) + ) + rows = await session.scalars( + select(PublicationRow).where(condition).order_by(PublicationRow.identifier) + .offset((page - 1) * page_size).limit(page_size) + ) + return [ + { + "identifier": row.identifier, + "title": row.title, + "provider": row.provider, + "source_identifier": row.source_identifier, + **publication_origin_fields(row.payload), + "revision": row.revision, + "content_hash": row.content_hash, + "first_retrieved_at": row.first_retrieved_at, + "last_retrieved_at": row.last_retrieved_at, + } + for row in rows + ], int(total or 0) + async def history(self, identifier: str) -> list[dict[str, Any]]: async with self.database.sessions() as session: rows = await session.scalars(select(PublicationRevisionRow).where( diff --git a/tests/test_api.py b/tests/test_api.py index fc91330..c063b4e 100644 --- a/tests/test_api.py +++ b/tests/test_api.py @@ -12,11 +12,8 @@ TEST_DATABASE_URL = os.getenv("TEST_DATABASE_URL") -def test_evidence_without_database_retains_unreviewed_fixture( - monkeypatch: pytest.MonkeyPatch, -) -> None: - monkeypatch.delenv("DATABASE_URL", raising=False) - with TestClient(create_app()) as client: +def test_evidence_without_database_retains_unreviewed_fixture() -> None: + with TestClient(create_app(database_url=None)) as client: detail = client.get("/api/v1/evidence/SYN-001").json()["item"] listed = client.get("/api/v1/evidence").json()["items"][0] for volatile_field in ("navigation_score", "score_method", "scoring_as_of"): @@ -62,6 +59,21 @@ def test_health_contract() -> None: assert health.json()["version"] == "0.3.0" +def test_explicit_database_url_none_ignores_environment( + monkeypatch: pytest.MonkeyPatch, +) -> None: + monkeypatch.setenv( + "DATABASE_URL", + "postgresql+asyncpg://openlongevity:openlongevity@localhost:5432/openlongevity_test", + ) + + with TestClient(create_app(database_url=None)) as client: + health = client.get("/api/v1/health") + + assert health.status_code == 200 + assert health.json()["database"] == "not_configured" + + @pytest.mark.postgres @pytest.mark.skipif( not TEST_DATABASE_URL, @@ -127,6 +139,23 @@ def test_evidence_detail_includes_navigation_score_metadata() -> None: assert payload["item"]["scoring_as_of"] == payload["summary"]["scoring_as_of"] +def test_citation_export_schema_describes_contract() -> None: + client = TestClient(create_app(database_url=None)) + + response = client.get("/api/v1/evidence/export/citation/schema") + + assert response.status_code == 200 + payload = response.json() + assert payload["schema_version"] == "citation-export-v1" + assert payload["mode"] == "citation-eligible" + assert payload["score_method"] == "navigation-score-v1" + assert payload["required_review_status"] == "verified" + assert "reviewed_at" in payload["required_review_fields"] + assert "synthetic_fixture" in payload["exclusion_reasons"] + assert "export_fingerprint" in payload["manifest_fields"] + assert "medical advice" in payload["disclaimer"] + + def test_citation_export_excludes_synthetic_fixtures() -> None: client = TestClient(create_app()) response = client.get("/api/v1/evidence/export/citation", params={"topic": "senescence"}) @@ -147,12 +176,8 @@ def test_citation_export_excludes_synthetic_fixtures() -> None: assert payload["excluded"][0]["reason"] == "synthetic_fixture" -def test_citation_export_accepts_reproducible_scoring_time( - monkeypatch: pytest.MonkeyPatch, -) -> None: - monkeypatch.delenv("DATABASE_URL", raising=False) - monkeypatch.delenv("TEST_DATABASE_URL", raising=False) - client = TestClient(create_app()) +def test_citation_export_accepts_reproducible_scoring_time() -> None: + client = TestClient(create_app(database_url=None)) response = client.get( "/api/v1/evidence/export/citation", @@ -206,15 +231,8 @@ def test_review_endpoint_is_disabled_without_review_key() -> None: assert response.status_code == 503 assert response.json()["error"]["code"] == "REVIEW_DISABLED" -def test_review_endpoint_requires_database_for_persistence( - monkeypatch: pytest.MonkeyPatch, -) -> None: - # Ascundem variabilele de mediu doar pentru acest test, - # astfel incat baza de date sa para neconfigurata - monkeypatch.delenv("DATABASE_URL", raising=False) - monkeypatch.delenv("TEST_DATABASE_URL", raising=False) - - client = TestClient(create_app(review_key="review-secret")) +def test_review_endpoint_requires_database_for_persistence() -> None: + client = TestClient(create_app(database_url=None, review_key="review-secret")) response = client.post( "/api/v1/evidence/SYN-001/review", headers={"X-Review-Key": "review-secret"}, @@ -243,3 +261,20 @@ def test_review_endpoint_rejects_machine_status_as_human_review() -> None: ) assert response.status_code == 422 assert response.json()["error"]["code"] == "INVALID_REVIEW" + + +def test_publication_manifest_requires_explicit_database() -> None: + with TestClient(create_app(database_url=None)) as client: + response = client.get("/api/v1/publications/export/manifest") + assert response.status_code == 503 + assert response.json()["error"]["code"] == "DATABASE_NOT_CONFIGURED" + + +@pytest.mark.parametrize("params", [ + {"page": 0}, {"page": 10001}, {"page_size": 0}, {"page_size": 101}, {"query": "x" * 201}, +]) +def test_publication_manifest_validates_bounds_before_storage(params: dict) -> None: + with TestClient(create_app(database_url=None)) as client: + response = client.get("/api/v1/publications/export/manifest", params=params) + assert response.status_code == 422 + assert response.json()["error"]["code"] == "INVALID_REQUEST" diff --git a/tests/test_exports.py b/tests/test_exports.py index 065ecd4..cab42b4 100644 --- a/tests/test_exports.py +++ b/tests/test_exports.py @@ -1,6 +1,10 @@ from dataclasses import asdict, replace -from openlongevity.exports import build_citation_export, evidence_record_payload +from openlongevity.exports import ( + build_citation_export, + build_publication_export_manifest, + evidence_record_payload, +) from openlongevity.models import EvidenceRecord, ReviewStatus, StudyType @@ -99,3 +103,52 @@ def test_citation_export_fingerprint_changes_with_export_boundary() -> None: "not_human_verified": 1, "synthetic_fixture": 1, } + + +def test_publication_manifest_fingerprints_all_exported_metadata() -> None: + import hashlib + import json + + publication = { + "identifier": "PMID:123", "title": "Metadata test", "provider": "pubmed", + "source_identifier": "123", "origin": "unknown", "synthetic": None, + "revision": 1, "content_hash": "a" * 64, + "first_retrieved_at": "2026-09-22T00:00:00+00:00", + "last_retrieved_at": "2026-09-22T00:00:00+00:00", + } + options = {"query": "Metadata", "page": 1, "page_size": 20, "total_matching": 1} + result = build_publication_export_manifest([publication], **options) + assert result == build_publication_export_manifest([publication], **options) + assert result["mode"] == "publication-metadata-export" + assert result["items"][0]["synthetic"] is None + manifest = dict(result["manifest"]) + fingerprint = manifest.pop("export_fingerprint") + encoded = json.dumps( + {"manifest": manifest, "items": result["items"]}, sort_keys=True, separators=(",", ":"), + ).encode("utf-8") + assert hashlib.sha256(encoded).hexdigest() == fingerprint + assert manifest["revision_map"] == {"PMID:123": 1} + assert manifest["content_hashes"] == {"PMID:123": "a" * 64} + # Freshness or origin can change independently of a stored content hash. + for field, value in ( + ("revision", 2), ("content_hash", "b" * 64), ("title", "Changed title"), + ("last_retrieved_at", "2026-09-23T00:00:00+00:00"), + ("origin", "manual"), ("synthetic", True), + ): + changed = build_publication_export_manifest([{**publication, field: value}], **options) + assert changed["manifest"]["export_fingerprint"] != fingerprint + changed_query = build_publication_export_manifest( + [publication], **{**options, "query": "test"}, + ) + assert changed_query["manifest"]["export_fingerprint"] != fingerprint + + +def test_publication_manifest_empty_page_preserves_matching_total() -> None: + result = build_publication_export_manifest( + [], query="study", page=3, page_size=20, total_matching=25, + ) + assert result["total"] == 0 + assert result["items"] == [] + assert result["manifest"]["total_matching"] == 25 + assert result["manifest"]["revision_map"] == {} + assert result["manifest"]["exported_identifiers"] == [] diff --git a/tests/test_publication_origin.py b/tests/test_publication_origin.py index f9881b5..504432c 100644 --- a/tests/test_publication_origin.py +++ b/tests/test_publication_origin.py @@ -142,3 +142,66 @@ async def cleanup(database: Database, identifiers: list[str]) -> None: PublicationRow.identifier.in_(identifiers) )) await database.close() + + +async def test_publication_manifest_pagination_revisions_and_freshness() -> None: + assert TEST_DATABASE_URL is not None + token = uuid4().hex + database = Database(TEST_DATABASE_URL) + repository = PublicationRepository(database) + identifiers = [f"TEST-EXPORT-{token}-{index}" for index in range(2)] + query = f"Export {token} 100%_" + initial = publication( + identifier=identifiers[0], source_identifier=identifiers[0], + title=query, origin=PublicationOrigin.UNKNOWN, + ) + app = create_app(database_url=TEST_DATABASE_URL) + try: + await repository.save(initial) + await repository.save(replace( + initial, identifier=identifiers[1], + provenance=replace(initial.provenance, source_identifier=identifiers[1]), + origin=PublicationOrigin.SYNTHETIC, + )) + async with app.router.lifespan_context(app): + async with AsyncClient( + transport=ASGITransport(app=app), base_url="http://test", + ) as client: + route = "/api/v1/publications/export/manifest" + params = {"query": f" {query} ", "page_size": 1} + response = await client.get(route, params=params) + assert response.status_code == 200 + first = response.json() + assert first["manifest"]["query"] == query + assert first["manifest"]["total_matching"] == 2 + assert first["manifest"]["exported_identifiers"] == identifiers[:1] + assert first["items"][0]["synthetic"] is None + assert first == (await client.get(route, params=params)).json() + second = (await client.get(route, params={**params, "page": 2})).json() + assert second["items"][0]["identifier"] == identifiers[1] + assert second["items"][0]["synthetic"] is True + empty = (await client.get(route, params={**params, "page": 3})).json() + assert empty["items"] == [] + assert empty["manifest"]["total_matching"] == 2 + + await repository.save(replace(initial, title=f"{query} revised")) + revised = (await client.get(route, params=params)).json() + assert revised["items"][0]["revision"] == 2 + history = await repository.history(identifiers[0]) + assert revised["items"][0]["content_hash"] == history[-1]["content_hash"] + assert (revised["manifest"]["export_fingerprint"] + != first["manifest"]["export_fingerprint"]) + + await repository.save(replace( + initial, title=f"{query} revised", provenance=replace( + initial.provenance, retrieved_at="2026-09-23T00:00:00+00:00", + ), + )) + refreshed = (await client.get(route, params=params)).json() + assert refreshed["items"][0]["revision"] == 2 + assert (refreshed["manifest"]["content_hashes"] + == revised["manifest"]["content_hashes"]) + assert (refreshed["manifest"]["export_fingerprint"] + != revised["manifest"]["export_fingerprint"]) + finally: + await cleanup(database, identifiers)