Skip to content

feat: record why a job went back to Scheduled and show Retrying jobs - #10

Merged
pdevito3 merged 4 commits into
mainfrom
feat/retrying-jobs
Oct 7, 2026
Merged

pdevito3 merged 4 commits into
mainfrom
feat/retrying-jobs

Conversation

@pdevito3

@pdevito3 pdevito3 commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Summary

A Scheduled job now records why it went back to Scheduled. Retrying means Scheduled with a retry cause. An operator sees the jobs that have a problem, and a deploy does not show as a wave of problems.

 job outcome / sweep / operator action      retry_cause
 Failure, retry policy reschedules          = HandlerFailed   (single and batched reports)
 lease expires, sweep reschedules           = LeaseExpired
 requeue                                    = null
 claim / relinquish / terminal / cancel     unchanged (sticky)
 new job                                    = null
Retrying := state = Scheduled AND retry_cause IS NOT NULL

Why not Attempt > 0: a clean-stop relinquish gives back claimed jobs with Attempt > 0. Every deploy would look like a problem.

Surfaces:

BackWaveMonitor   JobQuery.Retrying, JobSnapshot.RetryCause
Dashboard         Failures > Retrying tab (third tab, Dead-Lettered stays the default)
                  Jobs > State filter > "Retrying" (under Scheduled)
                  Job detail > "Retry cause" row
Pro MCP           search_jobs { retrying: true }
Storage contract  conformance tests for each set / clear rule

Schema: one additive migration for each SQL adapter. There is no backfill. A job that is Scheduled before the upgrade has no cause, so it does not show as Retrying.

Adapter Script Version Index
Postgres 0002_retry_cause.sql 1 -> 2 partial ix_backwave_jobs_retrying
SQL Server 0003_retry_cause.sql 2 -> 3 filtered ix_backwave_jobs_retrying
SQLite 0003_retry_cause.sql 2 -> 3 partial ix_backwave_jobs_retrying
Oracle 0002_retry_cause.sql 1 -> 2 ix_bw_jobs_retrying (retry_cause)

Two fixes that the work found:

  • Dashboard SSE on shutdown. An open live dashboard tab kept the host in its shutdown wait until the 30 s timeout. The worker groups then could not relinquish, and the leases expired later as LeaseExpired. That is a false alarm on deploy. The stream now also ends on ApplicationStopping.
  • Oracle migration lock wait. On a cold boot, the new ALTER TABLE failed with ORA-00054 while another node ran the v1 script. The migrator now sets ddl_lock_timeout = 30 on its own unpooled session.

There is also one flake fix: CapturingLogger was not safe when two threads wrote to it, so LeaseReclaimLogTests failed 3 times in 25 runs.

Evidence

Tests (local, private compose ports):

Suite Result
Core (BackWave.Tests) green, pinned simulation fingerprint battery unchanged
Postgres 261/261
SQL Server 268/268
Oracle 294/294 (migration + upgrade filter 3/3 runs)
SQLite 301/301
Upgrade harness 3/3
SchemaGate 10/10
Dashboard 112/112
Pro.Mcp 140/140
Pro.Dashboard / Pro / Hosting / EF / Testing 29 / 12 / 86 / 16 / 29
SourceGenerators / Benchmarks 43 / 89
  • Before: LiveView_EndsTheSseStream_WhenTheApplicationStartsToStop reads the stream until the 15 s timeout.
    After: the stream ends right after StopApplication().
  • Before: LeaseReclaimLogTests failed 3 of 25 runs ("Collection was modified").
    After: 0 of 40 runs.

E2E (Sample.Api on SQLite, real workers, dashboard clicks):

Job Path Stored result Shows as Retrying
bd6f407b flaky handler failed, retry scheduled state 0, attempt 1, cause 1 yes, "Handler failed"
e1a720c6 process kill -9 while leased, lease expired state 0, attempt 2, cause 2 yes, "Lease expired"
facdfd89 greet SIGTERM with live tab open, relinquished state 0, attempt 1, no cause no
40bc05c0 greet fresh enqueue state 0, attempt 0, no cause no
ee6c80a9 flaky dead-lettered after 3 attempts state 5, cause 1 kept no (Dead-Lettered tab)
  • On SIGTERM with /backwave/executing open, the app stopped in under 1 s. app.log: "Worker group 'strict-priority' relinquished 2 lease(s) on shutdown."
  • MCP search_jobs {retrying: true} returned only bd6f407b and e1a720c6.

Screenshots (kept local, not uploaded):

  • failures-retrying: Retrying tab, count 2, rows for the 2 problem jobs with Next Attempt and Retry Cause.
  • failures-retrying-light: the same page in the light theme.
  • jobs-scheduled: the Scheduled filter shows all 4 Scheduled jobs.
  • jobs-retrying: the Retrying filter shows only the 2 problem jobs.
  • detail-e1a720c6 / detail-bd6f407b: "Retry cause" row reads "Lease expired" / "Handler failed".
  • detail-facdfd89: the relinquished job has no Retry cause row.

Merge Danger

Door: one-way

Each SQL adapter gets a schema version bump. The column is additive and nullable, and the N-1 fleet and in-place upgrade tests pass. A rollback of the code still leaves the column and the new version stamp. An older build then refuses the newer schema version.

Blast Radius: schema

All four SQL adapters migrate on the next start. Writes to retry_cause are on the outcome, expiry and requeue paths. The Oracle round-trip budget is unchanged except for the job page window bytes. The dashboard Failures page gets a third tab. The Jobs table gets two different columns only when the Retrying filter is set.

The server waits for open requests before it stops the hosted services.
An open dashboard tab held its live stream until the host shutdown
timeout ended. The worker groups then had no time to give their leases
back on a clean stop, and the leases expired later.

The stream now also ends on ApplicationStopping.
An abandoned handler can unwind on a thread-pool thread and log while
the test reads the records. The read then failed with "Collection was
modified". The scope stack was also shared across threads.

The capture now locks its records and returns a copy. Scopes use
LoggerExternalScopeProvider, which keeps one scope stack per async flow.
LeaseReclaimLogTests failed 3 times in 25 runs before and 0 times in 40
runs after.
A Scheduled job did not tell an operator if an attempt went wrong. A
job with Attempt > 0 is not a good signal, because a clean-stop
relinquish also gives back a claimed job. A deploy then looks like a
wave of problems.

Each job now records a retry cause:

- HandlerFailed: the handler failed and the retry policy rescheduled
  the job (single and batched outcome reports).
- LeaseExpired: the lease expired and the sweep rescheduled the job.

A requeue clears the cause. A claim, a relinquish, a terminal outcome,
a cancel and a parent latch keep it. A new job has no cause.

Retrying means Scheduled with a retry cause. It is available as:

- JobQuery.Retrying and JobSnapshot.RetryCause on the Monitor API.
- A Retrying tab on the Failures page, and a Retrying option in the
  Jobs state filter. Both show the next attempt and the cause.
- A Retry cause row on the job detail page.
- A "retrying" filter on the MCP search_jobs tool.

Each SQL adapter gets one additive migration: a nullable retry_cause
column and an index for the Retrying query. Postgres and Oracle go to
version 2, SQL Server and SQLite to version 3. There is no backfill.
Jobs that are Scheduled before the upgrade have no cause, so they do
not show as Retrying.

The Oracle migrator now sets ddl_lock_timeout on its own unpooled
session. Without it, the new ALTER TABLE failed with ORA-00054 when
other nodes ran the first script at the same time on a cold boot.
# Conflicts:
#	src/BackWave.Conformance/ConformanceSuite.cs
#	src/BackWave.Pro.Mcp/Tools/JobTools.cs
@pdevito3
pdevito3 merged commit 1c5334b into main Oct 7, 2026
3 of 4 checks passed
@pdevito3
pdevito3 deleted the feat/retrying-jobs branch October 7, 2026 21:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant