Skip to content

Bound sink execution per instance and preserve recovery evidence - #2548

Merged
HypForge merged 22 commits into
masterfrom
integration/sink-instance-bound-2226
Oct 8, 2026
Merged

HypForge merged 22 commits into
masterfrom
integration/sink-instance-bound-2226

Conversation

@HypForge

@HypForge HypForge commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Slow or stalled exports currently admit overlapping work and can delay independent destinations and daemon bookkeeping. This repair gives each sink instance one active run and one coalesced follow-up, keeps sequential manual sync truthful, and cancels central transport across its complete 330-second chunk deadline.

Successful full completion records recovery time, including recovery after restart. Diagnostic history is capped at 100 recognized records per instance while protecting the newly published record and latest actionable failure. Cleanup preserves cache payloads, cursors, unknown files, symlinks and other destinations. No configuration or schema migration is required.

Fixes #2226
Change-Set: sink-instance-bound-2226

Final accepted head: 5765858 against master 9e53cf4. Round 3 independently accepted the test-only CI wait correction with 27 focused tests, typecheck, failing-before/passing-after filesystem-delay proof and a finite failure-bound proof. Product bytes match the round-2 accepted head; the following product checks remain attributable to that prior head:

  • Independent Astra round 2 accepted; F1-F3 fixed and independently verified; shipping risk low with executable safety proof.
  • Independent npm test: 7,829 passed, 4 skipped, 0 failed; typecheck passed.
  • Six disposable CLI/runtime journeys independently passed, including stalled transport, shared cancellation, manual contention, restart recovery and bounded history containment.
  • CPU/memory inspection and bounded history/retainer probes passed.

The historical source-details boot-snapshot timing failure remains recorded and reproduced on the unchanged base; this repair does not change that test boundary. Installed services and production Cloud behavior were not exercised. The deadline proof uses controlled time. Separate CLI processes do not share execution state, and an uncooperative third-party plugin can keep shutdown pending. PR #1034 destination-history work remains deferred and outside this repair.

The first GitHub CI run exposed a new test helper that counted immediate turns before asynchronous filesystem preflight settled. The test-only correction preserves every scheduling assertion and waits with a bounded elapsed deadline. Original failed CI remains visible; fresh final-head CI is required before readiness or landing.

@HypForge

HypForge commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Sink #2226 independent review, round 2

Verdict: accepted at a71830a. F1-F3 are fixed and independently verified. No new finding. Independent shipping risk: low. This is scoped candidate acceptance and risk evidence, not publication or landing authorization.

Reviewer team-sink-instance-reviewer@hypforge, native node 01M4CB7HW4B4PNAMKNQQ4471HF; child sink2226-review-r2-20261007, existing parent nn-maint-sink-tick-overlap-20261007. Target remains 9e53cf4c096de091b4ff700dc2ce2c384fcf0849; previous reviewed head is a7eadd69f342887b21fee687b7ed4c288d8ccd15. Both are verified ancestors of this candidate. Round 2 of 2 began 2026-10-07T23:51:10.958Z; checkpoint 2026-10-08T01:21:10.958Z. Completed before checkpoint; no extension granted.

Verification used clean detached .review-validation-r2 beneath the owned managed review worktree, Node v24.2.0, and the existing compatible ignored dependency link. review-r2-provenance.json records actual cwd, native address, SHA, clean status and ancestry/check exits. The generated parent overlay, original round 1 checkout, baseline checkout and all historical evidence remain intact. No product or shared dependency edits occurred.

The valid incremental scope is all eight changed files from a7eadd6: runtime, driver, diagnostic history, driver state types, their three regression suites and CONFIGURATION prose (196 additions, 15 deletions). I inspected their surrounding lifecycle/status/selection interactions and the final explanatory wording. Round 1 whole-candidate inspection and its unaffected proofs carry forward. LLP 0471 instance ownership, manual work, daemon work and diagnostic history, LLP 0453 warning rules, and the prior LLP 0472 plan govern these fixes. The new conservative selector resolves the existing warning-preservation constraint; the documentation describes that behavior without changing accepted design intent.

Numbered finding ledger

  1. F1 / UX-01, P2: FIX verified. Boot now streams existing same-instance diagnostic metadata once and retains an unresolved timestamp, considering valid nonfuture prior success. Only a strictly later full completion removes that boundary and logs ordinary recovery. The original failed restart probe now passes against this candidate in review-r2-journeys.test.mjs; its observations record an ordinary z-local recovery after restart. The committed regression, independently executed in the full suite, additionally asserts exactly one recovery and no duplicate after another restart followed by another completion. Partial work, another instance and equal timestamps cannot satisfy the own-instance recovery rule. Original failure and UX-01 disposition remain in round 1 evidence.
  2. F2, P2: FIX verified. Every accepted scheduled dispatch assigns the requesting driver's reusable function to the one pending slot, including busy dispatch. The original unchanged review-r1-shared-driver.mjs now reports two actual exports and discovery/completion owners [manual, daemon]. The executed committed regression covers 1,000 busy fires with zero Promises, alternating requesting drivers, explicit manual priority/fresh progress and peak concurrency one. Existing fresh discovery/hold/stop/registry replacement cases also pass. There is no additional callback queue or per-fire closure. Original failure and REVIEW-F2-DISPOSITION remain preserved.
  3. F3, P2: FIX verified. The finite selector protects both new publication and the newest currently nonfuture record, using a single cutoff during reconciliation. The original unchanged future-cohort probe now retains the sole actionable file and warning with exactly 100 records and zero successes. The executed committed regression covers initial cleanup refusal, retry, new future publication, restarted manager, equal success and strictly later recovery. Unknown files, symlinks and other destinations retain their protections. CONFIGURATION now accurately describes the two protected records and remaining newest history. Original failure and REVIEW-F3-DISPOSITION remain preserved.

No finding is silently dropped, deferred or rejected. Original steward supplied committed fixes; the reviewer made none.

Independent execution

<evidence> is this report's directory. All commands below ran from the new exact candidate checkout; external probes are review artifacts. The copied journey/cost probes change candidate import paths and observation output only, preserving their assertions. Replaying all six disposable journeys gives direct evidence for the cross-cutting runtime/driver/history interactions and shipping safety. Running the repository's full standard gate covers their existing callers; this does not replace incremental inspection with a new whole-change review.

Command Result Evidence
npm test Exit 0, 7,833 total / 7,829 pass / 4 skip / 0 fail, 63.0 s review-r2-npm-test.log
npm run typecheck Exit 0 review-r2-typecheck.log
git diff a7eadd69 a71830a9 --check Exit 0 review-r2-provenance.json
node --test <evidence>/review-r2-journeys.test.mjs Exit 0, all six pass review-r2-journeys.log, review-r2-journey-observations.json
node <evidence>/review-r1-shared-driver.mjs Exit 0, requesting context correct review-r2-shared-driver.log
node <evidence>/review-r1-future-history.mjs Exit 0, actionable warning preserved review-r2-future-history.log
node --expose-gc --max-old-space-size=32 <evidence>/review-r2-history-cost.mjs Exit 0 review-r2-history-cost.json, .err

The actual CLI pipe and owned PTY await both destinations, display three acknowledged rows per destination and preserve explicit partial truth/exit 0. Same-process contention maintains peak one under 5,000 fires and gives truthful refusal/stop receipts. A running disposable daemon continues local exports, sweeps, source probes and heartbeat during a real loopback central stall, writes no unacknowledged central cursor, and aborts the transport before a held source shutdown is released. Restart preserves warning until full recovery. After 131 observed real fixture failures, 100 diagnostic records remain while cache payload, cursor, other destination, unknown-file and telemetry sentinels remain intact. Shared auth cancellation preserves live consumers and prevents cancelled identity publication.

The full suite independently executes the new richer regressions plus actual loopback header/body/upload/auth/retry cancellation, controlled elapsed deadlines, stable retry IDs/bytes, mid-partition watermark failure, privacy/hold and cleanup containment contracts, reference and hygiene gates. Round 1's three hermetic smokes and central retainer proof remain valid evidence for unchanged paths, clearly identified as prior-head checks. Read GUARDIAN-ACCEPTANCE-FINAL.md and the owner correction checkpoint; guardian independently accepted all six final-SHA journeys. Their claims are supplemental to the execution above.

CPU and memory

No additional CPU or memory concern requiring repair found in the changed and affected paths. Boot adds one streaming metadata pass per live sink, with a maximum 4 KiB read per recognized file and one numeric boundary per instance. It adds startup I/O proportional to old directory size, without retaining inventories or introducing a per-tick/per-completion rescan. Large legacy histories can therefore delay boot; no universal startup latency bound is claimed.

The requesting-driver slot stores one reusable function, not a function or Promise per timer fire. Shared state, manual receipts, stop/drain and replacement ownership remain finite. History selection holds 100 scalar records, transiently 101; the added actionable lookup is bounded by that cap. My repeated-failure probe used 5,000 old records, 1,000 preserved unknown entries and 500 subsequent failures: exactly two directory passes, 5,500 metadata reads of at most 4,096 bytes, about 779 ms initial work and 189 ms subsequent work. Subsequent live heap samples were about 4.5-4.7 MB under a 32 MiB V8 limit. RSS was higher, around 90 MB; V8 heap limit is not a process-memory limit. This is synthetic local filesystem evidence.

Central Error/upload-buffer lifetime code is unchanged by this delta. The round 1 18-iteration WeakRef/small-heap proof remains applicable, with zero failed Errors/buffers alive during the next partition wait; native transport settlement is separately executed above. Initial boot/status/cleanup remain O(directory size), discovery scales with current partitions, filesystem refusal can exceed the on-disk cap, and disabled maintenance adds no retry timer. Separate CLI processes do not share the gate. An uncooperative generic plugin can retain its one operation and keep shutdown pending. No hidden new uptime-growing collection or busy loop was found.

Historical failures and limits

Preserve guardian-full-a7eadd69.log: that natural full run failed the source-details boot-snapshot timing assertion. Reviewer round 1 and this round's green runs do not replace it. The controlled 100 ms maintenance-import delay independently reproduced the unchanged assertion on exact base9e53 and candidatea7; see review-r1-source-delay-base.log, review-r1-source-delay-candidate.log and T3-SOURCE-REFRESH-DISPOSITION.md. This establishes the existing unsynchronized boundary, not natural incidence or an all-runs-green claim. No unrelated timing fix was included. The historical pre-review comparator failure in integration-T4-full.log likewise remains visible; its committed correction predates round 1.

No installed service, production Cloud receiver, owner enrollment or real user data was exercised. CLI row fixtures are synthetic. The 330-second native deadline is tested with controlled time, not a wall-clock soak. No GitHub CI, current remote target, branch-policy eligibility or landing was verified by this local review.

Independent shipping assessment

Low risk at this exact head against the named target. Affected users are daemon/manual sink users, especially installations with slow central transport, multiple destinations or old failure history. Potential unintended impact would be missed export scheduling, misleading recovery, cancelled shared credentials, premature cursor progress or overbroad diagnostic cleanup. The executed specific safety proofs above demonstrate finite scheduling with fresh requesting ownership, preserved live auth consumers, no cursor for the stalled unacknowledged upload, preserved cache/cursor/privacy contracts and cleanup confined to recognized diagnostic evidence while retaining actionable warnings.

Behavioral changes are limited to execution ownership, cancellation, completion diagnostics and bounded history; there is no configuration/schema migration, enrollment or public plugin contract change. Reverting product code restores prior scheduling/diagnostic behavior without migrating user state. Pruned old diagnostic records cannot be restored by rollback: this is an intentional accepted history bound, not reversible evidence deletion. The actionable-warning and sentinel/containment proofs are therefore essential to the low unintended-impact assessment; a green suite alone would not suffice. Export payload and retry progress are outside that deletion policy and remain independently protected.

Acceptance is now supported with every finding closed by verified fix. The responsible owner retains publication, current-head CI/policy/delegation checks and any landing decision. This report grants none of those authorities. Return evidence to the live existing parent and preserve the review checkouts/ref until released under FLOW. Default two-round budget is exhausted; any new repair requiring another round must return to owner triage.

@HypForge

HypForge commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

The first GitHub CI run failed on Node24 in sink-driver-scheduling.test.js:169: the wait helper exhausted 100 immediate turns before the second destination reached export. The other Node jobs were cancelled by matrix fail-fast. Typecheck and LLP checks passed; required CI has not passed.

A controlled30ms delay on the second first-sync filesystem read reproduces the same assertion at this exact head. The isolated natural run passes. A test-only correction03996bf5adf6b9222d7471295afee448caf3b610 uses a bounded monotonic five-second deadline and five-millisecond yield, retaining every scheduling assertion. Its controlled probe,12 scheduling tests, typecheck and committed full suite pass (7829passed,4skipped,0failed). Production files are unchanged.

That correction awaits serial integration and an explicitly authorized incremental independent review. This PR remains a draft at the original accepted head until new-head review and CI gates pass. The original failed run remains evidence; there is no merge-queue request.

@HypForge

HypForge commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Sink #2226 independent review, round 3

Accepted at 5765858b802be30d95fe90e1b0889033fa0167f1. No new findings. Independent shipping risk remains low; fresh GitHub CI and owner delivery gates remain required.

Reviewer: team-sink-instance-reviewer@hypforge, native seat 01M4CB7HW4B4PNAMKNQQ4471HF. Child: sink2226-review-r3-20261008; existing parent: nn-maint-sink-tick-overlap-20261007. Target unchanged: 9e53cf4c096de091b4ff700dc2ce2c384fcf0849; previous accepted head: a71830a91d252410395bced5e3a7b30409f89f28.

Owner transition 2601 expressly grants one extra incremental round beyond two, for this demonstrated CI synchronization defect; steward acknowledgment 2603 is retained in review-r3-parent-transitions.json. This is round 3 of the authorized total 3, not a reviewer extension. Started 2026-10-08T00:09:44.228Z; checkpoint 2026-10-08T01:09:44.228Z. Completed before checkpoint.

Scope and exactness

Used a new clean detached .review-validation-r3 inside the owned managed review worktree, with the existing compatible ignored dependency link, Node v24.2.0. Native identity/cwd, clean status and HEAD verified before checks and again at completion. All earlier checkouts and reports remain intact; no product edits or public mutations.

Verified prior reviewed head and target ancestry, integration tree equality with correction 03996bf5adf6b9222d7471295afee448caf3b610, and exactly one changed file: test/core/sink-driver-scheduling.test.js, four additions/two deletions. Its only changes add the timer import and replace the helper's 100 immediate turns with a monotonic five-second deadline and five-millisecond timer yield. All test bodies are byte-identical; every other tracked file, including product, docs, configuration and dependencies, is unchanged from accepted a718. Previous review artifact hashes and all author correction artifact hashes verified.

Inspected every helper use and its surrounding assertions: 22 synchronous predicates checking call counts or discovery admission, across shared gates, coalescing, sequential manual destinations, hold/cron rechecks, drain, discovery stop, registry replacement and requesting-driver ownership. They do not return Promises or block synchronously. The longer wait permits asynchronous filesystem preflight to settle; it does not relax expected counts, peak concurrency, progress ownership, stop behavior, fresh discovery or zero-Promise firestorm assertions. Remaining immediate turns still serve the original specific observation points. LLP 0471 instance-ownership/manual-work and prior LLP coverage remain applicable; no design change or new LLP is needed.

Findings carried forward

  1. F1 / UX-01, P2: FIX verified, carried forward. Runtime boot recovery and its regression are unchanged from independently accepted round 2. Its restart/full-success/no-duplicate proof and observations remain applicable.
  2. F2, P2: FIX verified, carried forward. Product requesting-driver ownership is unchanged. The corrected wait helper now supports the same unchanged shared-context regression, independently rerun successfully this round, including alternating drivers, manual priority and zero busy-fire Promises.
  3. F3, P2: FIX verified, carried forward. History selector, status logic, documentation and all warning/refusal/retry/restart regressions are unchanged. Round 2 original-probe and richer regression evidence remains applicable.

The CI-reported helper synchronization defect is corrected and independently reproduced before/after. No additional finding, deferral or unresolved blocker was found in this delta. No historical finding is dropped.

Executed checks

Commands below ran in clean final head unless explicitly marked prior head. <evidence> denotes this directory.

Command Result Evidence
node --import <evidence>/delivery-ci-preflight-delay.mjs --test --test-name-pattern='manual destinations stay sequential' test/core/sink-driver-scheduling.test.js at prior a718 Expected exit 1, same assertion at original line 169 review-r3-delay-before.log
Same command at final 5765858 Exit 0, one pass review-r3-delay-after.log
node --test test/core/sink-driver-scheduling.test.js test/core/llp-ref-hygiene.test.js test/core/tracked-files.test.js Exit 0, 27 pass: 12 scheduling and 15 hygiene checks review-r3-focused.log
npm run typecheck Exit 0 review-r3-typecheck.log
git diff a71830a9 HEAD --check Exit 0 Executed directly
Exact committed helper extracted and executed with a never-true predicate Exit 0, expected useful assertion after 5006.9 ms, 814 predicate calls review-r3-helper-bound.json

For the last proof, node --input-type=module read the candidate file, extracted async function until(predicate) through its closing brace, evaluated that exact function with node:vm and real performance/timer/assert bindings, then used assert.rejects on an always-false predicate. It asserted the error code/message, elapsed time 5-7 seconds and fewer than 1,200 calls. This tests the real helper's failure bound without editing product/test files. The source command is retained in the native transcript.

The injected delay affects only the second first-sync-hold.json read by 30 ms. It reproduces the exact old failing assertion and establishes the causal flaw in counting immediate turns. It does not prove the original CI runner experienced exactly that delay. The original failed CI log remains delivery-ci-failed.log (Node 24, 7,830 pass, two skip, one failure of 7,833); it is not relabeled green. The natural isolated author probe passed and is preserved separately.

No broad repeat was warranted for a byte-verified test-only delta. The author's committed full suite (7,829 pass/four skip), prior independent full suite and six running journeys are historical evidence, not newly executed final-head tests. Owner integration and author correction logs remain separately attributable. Original source-refresh timing failure and controlled base/candidate reproduction, and the earlier comparator gate failure, remain distinct historical records in review-round-1/2.

CPU, memory and shipping risk

No new CPU or memory concern. The test helper has one deadline, one predicate and at most one pending timer per invocation. It yields between polls and retains no growing collection or detached timeout. The failure probe used about 58 ms total CPU over five seconds. The monotonic deadline avoids wall-clock adjustment; as with any timer, a frozen event loop can delay the assertion, so this is not a process-wide hard timeout. All actual predicates are cheap and synchronous. Production CPU/memory behavior is byte-identical to the explicit round 2 assessment.

Low shipping risk at this exact final head and target. The delta changes only test synchronization and can be reverted by reverting one test helper; it changes no user behavior, persisted state, access, credentials, payload handling or export progress. Its focused executable safety proof is the same delayed-I/O scenario failing before/passing after while semantic assertions remain identical, plus verified finite failure behavior and all scheduling checks.

For the complete candidate, the affected sink users and unintended-impact analysis from round 2 remain applicable by identical product bytes. Its independently executed six disposable CLI/PTY/daemon/loopback journeys prove stalled transport leaves cache/cursors intact, unrelated work proceeds, live auth consumers survive cancellation, cleanup preserves unrelated sentinels and actionable warnings, and recovery is truthful. These are reused historical proofs at a718, not new executions at 5765858. Intentional old diagnostic pruning remains irreversible; the warning/payload/cursor containment proofs remain essential to low unintended-impact risk. Product rollback still requires no schema/config migration.

Read GITHUB.md shipping policy and retained previous assessment. Local acceptance does not establish fresh GitHub matrix success, installed-service health, production receiver behavior or landing eligibility. The assignment reports PR2548 draft with old a718 CI failure; no fresh external read or public mutation was performed here. The responsible owner must push/verify the exact candidate under its authority and satisfy current CI, delegation, policy and queue requirements before any landing. This report grants no publication or landing authority.

Return to the existing live steward-owned parent, verified blocked on this review child. Preserve all review worktrees and evidence. The one extra round is now consumed; further repair/review requires owner triage.

@HypForge
HypForge marked this pull request as ready for review October 8, 2026 00:16
@HypForge HypForge added the neutral:approved neutral reviewed this and holds it for a maintainer merge (own or adopted PR; LLP 0025/0030) label Oct 8, 2026
@HypForge
HypForge added this pull request to the merge queue Oct 8, 2026
Merged via the queue into master with commit 3c2238c Oct 8, 2026
8 checks passed
@HypForge
HypForge deleted the integration/sink-instance-bound-2226 branch October 8, 2026 00:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

neutral:approved neutral reviewed this and holds it for a maintainer merge (own or adopted PR; LLP 0025/0030)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bound stalled sink work per instance without blocking other exports or recording sweeps

1 participant