Audit date: 2026-10-02. Target: production presets, checkout
896a76555145449cdaeacba7d07918ba96137e60. Except for the final default-enabled validation section, the results below
were produced before promotion; the GA implementation follows in the same
product PR. The earlier sections are historical baseline evidence.
The existing exact-tree architecture is suitable for graduation. CLI and MCP already share request normalization, service queries, projection, and error mapping; their intentionally different text formats need no redesign. At the baseline, graduation still required default CLI registration, public MCP composition, stable routing guidance, smoke inventory updates, and release delivery.
The baseline stable guide said package targets scope to package subpaths. That
is correct for indexed navigation and wrong for raw diffs. The diff descriptor
and local appendix explain the repository-wide exception. Promotion must add
that exception to the stable guide and its exact public skill copy together.
The stable routing table also needs an explicit raw-source-comparison route;
pkg_upgrade_review remains the route for upgrade assessment.
The descriptor's first sentence already fits the discovery length contract: “Compare source across exact package versions or public repository refs.” Preserve its distinct job, separate endpoints, repository scope, bounded content, preview limits, and compatibility caveat when removing Experimental. Descriptor-only discovery must be evaluated independently of explicit intent.
Run from the repository root with existing production authentication:
bun run scripts/code-diff-audit.ts .agent-eval/code-diff-ga/liveThe script uses real source CLI subprocesses and a real local stdio MCP server,
production endpoint presets and isolated CLI/MCP config. The baseline runs
used experimental opt-in; the current script explicitly sets tools = false
for both surfaces and passes no MCP experimental override. It does not modify the user's config. It writes
public response artifacts and a per-cell result matrix to the specified local
directory; it does not write credentials or request headers. The verified
local authentication source was macOS Keychain. This is a local Bun audit,
not a CI live-data gate or a cross-platform runtime claim.
| Coverage | Cases and assertions |
|---|---|
| All four views, both surfaces, text and JSON | Express npm versions, GitHub tags and SHAs, reverse and identical ranges, a matching glob and an empty glob, explicit file truncation, React repository-wide and package-path-filtered results |
| All twelve registry syntaxes | npm, PyPI, crates, RubyGems, Go, Hex, NuGet, Maven, Packagist, Swift, Zig, vcpkg; additional registries use bounded inventory and CLI/MCP JSON equality |
| Repository providers and addressing | GitHub, Codeberg, GitLab, compact targets, explicit --repo-url, tag/SHA identities; HEAD is validated independently per call because it is mutable |
| Omitted-view defaults | CLI patch output and MCP name-status inventory, each checked through JSON and text |
| Content and bounds | Exact commit identity, inventory completeness, mode-dependent fields, path filters, file truncation warnings, non-applicable patch suppression, partial patch budgets, JSON parity |
| Failures | Unpublished version, missing ref, embedded version/ref, empty endpoint, out-of-range file bounds, traversal and unsupported brace globs, unsupported repository metadata, unavailable historical package data, unregistered package |
Initial matrix: 99 passed, zero failed, before the additional registry and
provider cases were incorporated. Expanded matrix: 121 passed, zero failed;
the final strengthened matrix passed 123 cells, zero failed, including
omitted-view defaults, known commit identities, exact name/name-status text
paths, stat paths, and preservation of applicable patch content. Final artifacts
are in .agent-eval/code-diff-ga/final-live/; earlier runs remain retained.
Upper numeric limits, unusual path bytes, binary and
mode-only changes, content filtering/failure, malformed responses, refresh,
and GraphQL over-fetch controls also have deterministic test coverage. They
are not all manufactured on live repositories.
Independent patch check: read all 271 lines of lib/utils.js at both exact
Express SHAs using GitHits read --json; verify the served SHAs; apply the
returned scoped patch with git apply --check and git apply in a disposable
directory. The applied file exactly matched the independently read target.
Evidence: .agent-eval/code-diff-ga/patch-oracle.json, the two read-*.json
files, and live/glob-patch-cli.json. This verifies one real text patch,
not every patch, binary file, or metadata transition.
- Maven
org.apache.commons:commons-lang3resolves package metadata tohttps://gitbox.apache.org/repos/asf/commons-lang, outside the supported repository providers. Diff returns non-retryable BACKEND_ERROR with the explicit unsupported-repository message. A separate verified Guava Maven comparison succeeds. No inferred repository substitution was added. - Packagist
symfony/http-foundation6.4.0..6.4.1returns VERSION_NOT_FOUND on the ending version, with published-version and proven-ref alternatives. A separate8.1.7..8.1.8comparison selected from those alternatives succeeds. This is not evidence that the historical version was never published upstream; it establishes the production service's limitation. - The original Zig matrix fixture
gh/ziglibs/known-foldersreports one version0.7.0; its self-comparison succeeds. Follow-up verification found thatgh/ziglang/zigwas an unsuitable package fixture: the compiler's 0.14.0 manifest explicitly says it is not intended to be consumed as a package. Its NOT_FOUND response is not evidence of broken indexing on access. A valid library,zig:gh/hejsil/zig-clap, succeeds for0.11.0..0.12.0with 12 changed files, exact distinct tag SHAs, and identical bounded CLI/MCP name-status JSON. Both upstream manifests declare the requested package versions. Evidence is retained in.agent-eval/code-diff-ga/zig-clap-check.jsonandzig-clap-mcp-check.json. This additional comparison closes changed-version Zig coverage; it is separate from the original 123-cell matrix. It does not verify indexing on access for a previously unseen valid package. Valid packages are expected to index on access; absence from existing package data must not by itself be accepted as a product limitation. - React
19.0.0..19.1.0reports 2,530 repository-wide changed files; the first three ranked files are outsidepackages/react/. Apackages/react/**filter reports 45 changed files. This directly exercises the scope caveat.
These limitations are explicit negative cases, not disguised successful comparisons. No backend modifications or new fallback machinery were made.
There was one neutral diff workload. Four were added under
eval/agentic/workloads/ and classified experimental in suites.json:
| Workload suffix | Behavioral question |
|---|---|
experimental-code-diff |
Exact Express changed paths and source content; separate source evidence from upgrade safety |
-repository |
Repository tags, full commit identities, direction, and an identical-ref control |
-monorepo |
Bounded repository-wide overview followed by a package-path filter, with explicit scope and completeness |
-recovery |
Unavailable package version, evidence-backed alternatives, and a separate successful range without silently replacing the request |
-bounded |
File-count, aggregate patch-byte, and displayed preview limits; full returned content versus omitted content |
Executed 14 cells with the existing bun run agent:e2e harness:
- Claude and Codex: the original workload with
--guidance-profile descriptorsand neutral intent (two discovery cells). - Both agents: all five workloads with descriptors and
--intent-profile githits(ten intent cells). - Both agents: monorepo workload with
--guidance-profile fulland neutral intent (two full-guidance cells). - All cells used
--surface mcp --server local --experimental-tools, a 300-second per-workload timeout, and production endpoints. Intent batches used--concurrency 2. Model selection was the harness default: Claude resolved toclaude-sonnet-5-5; Codex requestedgpt-6-lunaat high effort. Authentication was injected through existing secure local stores; no token values were written into commands or artifacts.
All 14 cells produced valid success reports and no isolation-violation files.
Actual calls were inspected in tool-calls.json, answers and confidence in
final.json, derived usage in metrics.json, and returned evidence in the
redacted traces. Both agents used code_diff in all ten intent cells and both
full cells. Claude used it in neutral discovery; Codex used npm/GitHub evidence
and made zero GitHits calls. That Codex cell is not tool acceptance evidence.
Observed answers retained repository scope, truncation, missing-endpoint distinctions, and compatibility caveats. The bounded Codex run reported that no patch fit its selected 1024-byte requests; Claude explicitly raised a separate focused budget and obtained a full patch. These are trace observations, not graded usefulness, a discovery rate, or a performance comparison. Claude's adapter does not derive logical-call or token/cost metrics; do not compare its raw started events with Codex's logical counts.
Raw evidence is local and ignored under
.agent-eval/code-diff-ga/{claude,codex}-{discovery,intent,full}. Workload prompts
and the direct matrix remain committed for reproducibility. CLI-only agent
and published/hosted evaluations belong to release preparation and delivery.
The baseline experimental eval override supported only local MCP.
| Exact command | Result |
|---|---|
bun test src/commands/code/diff.test.ts src/tools/code-diff-parity.test.ts packages/mcp/src/tools/code-diff.test.ts packages/mcp/src/shared/code-diff-request.test.ts packages/mcp/src/shared/code-diff-response.test.ts packages/mcp/src/shared/code-diff-text.test.ts packages/core-internal/src/services/code-navigation-service.test.ts |
211 passed, zero failed |
bun test |
5,325 passed, zero failed across 230 files |
bun test scripts/agent-eval.test.ts scripts/agent-eval-suite.test.ts |
194 passed, zero failed |
bun run typecheck |
Pass |
bun run build |
Pass |
bun run smoke:cli --mode unauthenticated |
Pass |
bun run smoke:mcp --mode registration |
Pass |
bun run smoke:cli:built |
Pass under Node |
bun run smoke:mcp:built |
Pass under Node |
bun run scripts/code-diff-audit.ts .agent-eval/code-diff-ga/final-live |
123 passed, zero failed |
bunx --no-install biome lint scripts/code-diff-audit.ts |
Pass |
The first full unit run exposed two audit-edit integration omissions: manifest count expectations and explicit workload filenames in README. Both were fixed; the complete suite then passed. The prior README also understated the existing stable inventory as 25 rather than the manifest's 31; this was corrected. The config policy doc's stale 15-tool count was removed; the current public descriptor inventory was 12 at the baseline; GA promotion changes it to 13. Its exact inventory belongs to catalog tests. Authenticated validation was deliberately the diff-only matrix. Unrelated live Research and package-tool suites were not run. No Windows runtime, published GA package, hosted GA server, or graded answer-quality claim is made.
No optimization was attempted, so no performance benchmark is required or claimed. No production refactor is needed: reuse the shared diff components and move registration/guidance into the stable surface.
Internal review found two audit-proof gaps: parity alone could accept a shared wrong commit, and explicit views left live defaults untested. Both were fixed and the complete revised delta was re-reviewed clean; the final live matrix then passed. Claude round 1 found only minor plan/doc corrections: specify the audit script's stable-config migration and lifecycle, record the final matrix, and classify the future public service-contract release impact in Phase 1. All were accepted and applied; the doc-only round counts as clean.
The GA implementation makes CLI and public MCP diff available by default,
renames all five workloads to stable code-diff* names, and adds the base
comparison to smoke. At GA validation, the manifest had 41 workloads: 36 stable,
one stateful, and four experimental; smoke selects seven. Guidance and its public
MCP skill copy preserve exact parity and explain repository-wide raw scope.
Implementation checks passed: 5,327 unit tests across 230 files; 353 focused
MCP/catalog/guide/smoke/parity tests across 12 files; typecheck; build;
plugin generation/check; public-package validation including outside-root
typed provider and real packed SDK diff invocation; source unauthenticated CLI
and registration MCP smoke; and both built smoke modes under Node. The 353
focused tests made 2,082 assertions. Local logs are retained under
.agent-eval/code-diff-ga/stable-*.log.
The default-enabled production matrix passed 125 cells, zero failures at
27019e28f0ea45e2d016e94ffd5948d6e89e35f0. Both CLI and local stdio MCP
used an isolated experimental.tools = false config, with no MCP override.
This includes the distinct-version valid Zig library comparison and its
identical-version control. Exact command:
bun run scripts/code-diff-audit.ts .agent-eval/code-diff-ga/resumed-liveAll 14 default-enabled agent cells produced successful final reports,
with zero isolation violations. Both agents used code_diff in all ten
descriptor-intent cells and both full-guidance monorepo cells. Claude also
used it in neutral discovery. Codex neutral discovery used public npm/GitHub
evidence and made zero GitHits calls; that cell is not tool acceptance
evidence. Actual calls, answers/confidence, derived metrics and isolation
artifacts were inspected. Answers retained repository scope, unavailable
endpoint distinctions, incomplete patch evidence and compatibility caveats.
Claude's bounded run raised a separate focused budget to obtain a complete
returned patch; Codex retained the original 1024-byte cap and correctly
reported that the source service omitted all requested patches. Codex's
monorepo intent run explicitly distinguished complete backend coverage from
client display truncation. These are trace observations, not graded answer
quality or a discovery/performance rate. Claude logical-call/token/cost
metrics remain unavailable; Codex metrics must not be compared with Claude
raw event counts.
The cells use --surface mcp --server local, with no experimental flag and
an isolated false config. For each of Claude and Codex, reproduce the three
profiles using existing secure local agent authentication:
bun run agent:e2e --agent <claude|codex> --surface mcp --server local --guidance-profile descriptors --intent-profile githits --concurrency 2 --out <intent-output> --workload eval/agentic/workloads/code-diff.md --workload eval/agentic/workloads/code-diff-repository.md --workload eval/agentic/workloads/code-diff-monorepo.md --workload eval/agentic/workloads/code-diff-recovery.md --workload eval/agentic/workloads/code-diff-bounded.md
bun run agent:e2e --agent <claude|codex> --surface mcp --server local --guidance-profile descriptors --intent-profile neutral --concurrency 2 --out <discovery-output> --workload eval/agentic/workloads/code-diff.md
bun run agent:e2e --agent <claude|codex> --surface mcp --server local --guidance-profile full --intent-profile neutral --concurrency 2 --out <full-output> --workload eval/agentic/workloads/code-diff-monorepo.mdArtifacts remain local in .agent-eval/code-diff-ga/resumed-evals/, including
acceptance-inspection.json; every metrics record reports
experimentalTools: false. Earlier Keychain-blocked runs and diagnostic
evidence remain preserved. After the user returned, local access succeeded;
no credential, production data, timeout, or retry mechanism was changed.
Implementation review completed cleanly after minor documentation/test-helper corrections, with a fresh-context final check. The affected Resolve parity suite passed 31 tests, and unauthenticated CLI smoke passed again. CI at the validated commit passed build/checks, Linux/Windows unit suites, MCP package validation, and compatibility checks for Bun and Node 20/22/24/26. Publication was skipped. Published and hosted GA delivery remain separate release/adoption steps.
On 2026-10-03, package-upgrade-poor-changelog.md was added to stable-full
to test source evidence as a supplement to upgrade assessment. It asks about
the fixed npm Lodash 4.17.20 to 4.17.21 upgrade without naming tools or a
source-comparison method. The manifest now has 42 workloads, including 37
stable workloads; canary and smoke membership are unchanged.
Production fixture probes verified that upgrade review returns
fallback: package_versions, hasReleaseNoteBodies: false, and zero entries
with bodies. Source comparison resolves commits
f2e7063ee409ff40a60b14370c58dceee1a2efd4 and
c6e281b878b315c7a10d90f9c2af4cdb112d9625, with 14 changed files. This is a
gap in the available package evidence, not proof that upstream has no
changelog. Recheck that condition before interpreting future runs.
Six baseline cells ran configured for production local MCP against unchanged runtime
instructions at e3d0e9a3911188540d6e8041dd50c20ad44aaa27, with the workload
and inventory edits uncommitted and experimental tools explicitly disabled.
All produced successful final reports; no isolation violations were reported.
Calls, final answers/confidence, derived metrics, and provider traces were
inspected. Descriptor profiles exposed full descriptions and allowed
quick_start; these are not first-80-only discovery tests.
| Agent | Guidance / intent | Observed routing | Duration |
|---|---|---|---|
| Claude | descriptors / neutral | quick_start, upgrade review, advisories, source inventory and two focused patches |
22.9 s |
| Claude | descriptors / GitHits | Same evidence tools, including source inventory and two focused patches | 26.0 s |
| Claude | full / neutral | Upgrade review, advisories, source inventory and two focused patches | 19.8 s |
| Codex | descriptors / neutral | No GitHits calls; used web search and opened GitHub tag comparison, citing upstream commits, changelog and advisories | 85.9 s |
| Codex | descriptors / GitHits | Upgrade review, changelog, advisories, seven source comparisons and two exact helper reads | 65.8 s |
| Codex | full / neutral | Upgrade review, changelog, advisories, source inventory and JavaScript patch | 67.1 s |
Five cells combined pkg_upgrade_review with code_diff on the requested
endpoints. Answers described template-variable validation and whitespace
trimming changes, retained outstanding target-version advisories, and kept
newer remediation versions separate. Codex full guidance explicitly reported
partial patch coverage. Claude neutral discovery overclaimed that there were
no official release notes; the verified fact was missing returned bodies.
Codex neutral discovery found upstream changelog evidence and is not GitHits
tool acceptance evidence. Provider advisory records and upstream records also
differed in their coverage/aliases; these runs do not independently grade their
accuracy or answer quality. No service call sites or runtime tests were
available to establish application compatibility.
These observations support retaining the workload for guidance comparisons; one fixture and one run per cell do not establish a discovery success rate. Claude token/cost and logical-call metrics remain unavailable. Codex logical calls were 14 in descriptor intent, seven with full guidance, and zero GitHits calls in neutral discovery. Do not compare those counts to Claude raw events.
Reproduce each agent with the same workload and separate output directories:
GITHITS_ENV=prod bun run agent:e2e --agent <claude|codex> --surface mcp --server local --guidance-profile descriptors --intent-profile neutral --workload eval/agentic/workloads/package-upgrade-poor-changelog.md --out <discovery-output>
GITHITS_ENV=prod bun run agent:e2e --agent <claude|codex> --surface mcp --server local --guidance-profile descriptors --intent-profile githits --workload eval/agentic/workloads/package-upgrade-poor-changelog.md --out <intent-output>
GITHITS_ENV=prod bun run agent:e2e --agent <claude|codex> --surface mcp --server local --guidance-profile full --intent-profile neutral --workload eval/agentic/workloads/package-upgrade-poor-changelog.md --out <full-output>Secure local agent authentication and an isolated false experimental config
were used without printing credentials. Artifacts remain in
.agent-eval/code-diff-upgrade/baseline/; fixture probes are alongside it.
bun test scripts/agent-eval-suite.test.ts passed 49 tests with 320 assertions.
No runtime, descriptor, quick-start, or public skill changes were made in
this workload increment.