Version
workerd 2026-09-01 (npm @cloudflare/workerd-darwin-arm64, workerd --version prints workerd 2026-09-01), macOS arm64.
- The code path is the same on
main and in the latest release, v1.20260924.1 (src/workerd/io/io-context.c++, IoContext::startDeleteQueueSignalTask and IoContext::getCurrentTraceSpan).
Summary
Take a Durable Object that has no request in flight. If a promise created in its IoContext settles while a different request in the same isolate is running, workerd schedules the continuation on the DO's delete queue. The delete-queue task then calls context->run(...). For an actor, runSingle waits on the input gate with getCurrentTraceSpan(). Outside a JS lock that falls back to getMetrics() → getCurrentIncomingRequest(), which KJ_REQUIREs a current request. The task's catch (...) calls context->abort(e), so the idle actor is aborted. Every later call through a stub that reached that actor fails with internal error and the same exception. A stub obtained afterwards reaches a new instance.
This happens in practice with Effect's HTTP client, which is isolate-wide. effect/unstable/http/HttpClient registers every response in a module-level FinalizationRegistry whose callback calls controller.abort(). When V8 GC runs during another DO's request, the callback aborts a controller created by the idle DO, the abort settles the DO's promise, and the DO is killed. In our app, one forced gc() in a Worker fetch 2 s after the DO's last call broke it in 2 of 2 runs, against 0 of 2 control runs. Replacing the registry callback with a no-op made the same forced GC harmless in 2 of 2 runs.
Minimal repro
config.capnp:
using Workerd = import "/workerd/workerd.capnp";
const config :Workerd.Config = (
services = [ (name = "main", worker = .worker) ],
sockets = [ (name = "http", address = "127.0.0.1:18787", http = (), service = "main") ],
v8Flags = ["--expose-gc"],
);
const worker :Workerd.Worker = (
modules = [ (name = "main.js", esModule = embed "main.js") ],
compatibilityDate = "2026-09-01",
durableObjectNamespaces = [
(className = "Callee", uniqueKey = "callee"),
(className = "Caller", uniqueKey = "caller"),
],
durableObjectStorage = (inMemory = void),
bindings = [
(name = "CALLEE", durableObjectNamespace = "Callee"),
(name = "CALLER", durableObjectNamespace = "Caller"),
],
);
main.js:
import { DurableObject } from "cloudflare:workers";
// Isolate-wide, like effect/unstable/http/HttpClient's response registry.
const registry = new FinalizationRegistry((controller) => controller.abort());
const pending = [];
export class Callee extends DurableObject {
// Leaves a promise owned by this IoContext that settles when the controller is aborted.
async park(mode) {
let settle;
new Promise((resolve) => (settle = resolve)).then(() => {});
if (mode === "gc") {
const controller = new AbortController();
controller.signal.addEventListener("abort", () => settle());
registry.register({}, controller);
} else {
pending.push(settle);
}
return "parked";
}
async ping() {
return "pong";
}
}
export class Caller extends DurableObject {
// One Callee stub per Caller activation.
held = this.env.CALLEE.getByName("callee");
async heldPing() {
return await this.held.ping();
}
async gc() {
globalThis.gc();
return "gc";
}
async resolve() {
for (const settle of pending.splice(0)) settle();
return "resolved";
}
}
export default {
async fetch(request, env) {
const url = new URL(request.url);
const callee = env.CALLEE.getByName("callee");
const caller = env.CALLER.getByName("caller");
try {
switch (url.pathname) {
case "/park": return new Response(await callee.park(url.searchParams.get("mode")));
case "/held": return new Response(await caller.heldPing());
case "/fresh": return new Response(await callee.ping());
case "/gc": return new Response(await caller.gc());
case "/resolve": return new Response(await caller.resolve());
}
} catch (e) {
return new Response(`ERR ${e.message}`, { status: 500 });
}
return new Response("not found", { status: 404 });
},
};
repro.sh gc|resolve|control:
W=${WORKERD:-workerd}
$W serve --experimental config.capnp > workerd.log 2>&1 &
PID=$!; sleep 3
curl -s localhost:18787/held >/dev/null # Caller creates and uses its Callee stub
case $1 in control) m=gc;; *) m=$1;; esac
curl -s "localhost:18787/park?mode=$m" >/dev/null # Callee leaves a pending promise, then goes idle
sleep 2 # Callee idle (< 10 s, not evicted)
case $1 in gc) curl -s localhost:18787/gc >/dev/null;; resolve) curl -s localhost:18787/resolve >/dev/null;; esac
sleep 1
echo "held stub: $(curl -s localhost:18787/held)"
echo "held stub: $(curl -s localhost:18787/held)"
echo "fresh stub: $(curl -s localhost:18787/fresh)"
kill $PID
Results, 3 runs each:
| mode |
what settles Callee's promise while Callee is idle |
held stub |
fresh stub |
gc |
gc() inside Caller's request runs the FinalizationRegistry → abort() |
ERR internal error (3/3), forever |
pong |
resolve |
Caller calls the parked resolver directly |
ERR internal error (3/3), forever |
pong |
control |
nothing (same park, no GC) |
pong |
pong |
workerd.log:
workerd/jsg/util.c++:391: error: e = workerd/io/_virtual_includes/io/workerd/io/io-context.h:1301: failed: expected !incomingRequests.empty(); the IoContext has no current IncomingRequest; lastDeliveredLocation = src/workerd/api/worker-rpc.c++:2411:3 in run
stack: 1023f2497 102de8213 102df3d9f 102dff0e3 1022192c7 1021d3707 10221a22f 1033cd5d8 1033cfa13 1022cc688 10358e27c 10358d00c 1031cbc28 1031cc224; sentryErrorContext = jsgInternalError; wdErrId = …
workerd/io/worker.c++:2584: info: uncaught exception; source = Uncaught (in promise); stack = Error: internal error; reference = …
at async Caller.heldPing (main.js:…)
Symbolized with atos -p <workerd pid>:
workerd::IoContext::getMetrics() + 131
workerd::IoContext::getCurrentTraceSpan() + 191
workerd::IoContext::runSingle<workerd::IoContext::startDeleteQueueSignalTask(workerd::IoContext*)::$_0>(…, kj::Maybe<workerd::InputGate::Lock>) + 771
workerd::IoContext::startDeleteQueueSignalTask(workerd::IoContext*) (.resume) + 151
workerd::server::Server::ActorNamespace::ActorContainer::monitorOnBroken(workerd::Worker::Actor&) (.resume) + 35
workerd::server::Server::ActorNamespace::ActorContainer::getActor() + 899
workerd::server::Server::ActorNamespace::ActorContainer::startRequest(workerd::IoChannelFactory::SubrequestMetadata) (.resume) + 67
workerd::(anonymous namespace)::PromisedWorkerInterface::customEvent(…) (.resume) + 51
capnp::QueuedClient::call(…)::'lambda1'(…)
capnp::LocalRequest::sendImpl(bool)::'lambda'()
capnp::Request<workerd::rpc::JsRpcTarget::CallParams, workerd::rpc::JsRpcTarget::CallResults>::send()::'lambda'(…)
We see the same stack, byte for byte in the frame offsets, in the real app (Alchemy local dev, one isolate with a Worker and three DO classes).
Root cause
- A promise created while the Callee's IoContext is current carries its promise context tag (
IoContext::getPromiseContextTag, an IoCrossContextExecutor over the delete queue). Settling it from another context goes through DeleteQueue::scheduleAction → tryScheduleAction, which fulfills crossThreadFulfiller.
IoContext::startDeleteQueueSignalTask wakes up and calls co_await context->run(...).
- For an actor with no input lock,
runSingle calls a.getInputGate().wait(getCurrentTraceSpan()) (io-context.h, runSingle).
getCurrentTraceSpan() finds no currentLock (the task runs outside the JS lock), so it takes the fallback return getMetrics().getSpan();. getMetrics() is *getCurrentIncomingRequest().metrics, and getCurrentIncomingRequest() does KJ_REQUIRE(!incomingRequests.empty(), "the IoContext has no current IncomingRequest", ...).
- The exception reaches
catch (...) { context->abort(kj::getCaughtExceptionAsKj()); } in startDeleteQueueSignalTask. The actor is aborted, ActorContainer::monitorOnBroken records brokenReason, and every stub channel bound to that container rethrows it.
The actor had done nothing wrong. It was only idle when a queued continuation arrived. addTask and addWaitUntil already handle an empty incomingRequests, and getCurrentUserTraceSpan() returns SpanParent(nullptr) in the same situation. getCurrentTraceSpan() has no such guard, and the delete-queue task is the one caller that reaches it with no request by design.
Suggested fix
The delete-queue task reaches two request-bound calls in runSingle for an actor with no request, so guarding only one of them moves the failure to the other:
a.getInputGate().wait(getCurrentTraceSpan()), where getCurrentTraceSpan() falls back to getMetrics().
takeAsyncLockWhenActorCacheReady(now(), a, getMetrics()) (io-context.h, runSingle), which calls getMetrics() directly.
Suggested fix: have startDeleteQueueSignalTask take the input lock and the async lock without request metrics (no trace span, takeAsyncLockWithoutRequest-style), the same way addTask handles the no-request case, so draining the delete queue never needs an IncomingRequest.
A narrower alternative is to make both call sites request-optional: return SpanParent(nullptr) from getCurrentTraceSpan() when incomingRequests.empty(), as getCurrentUserTraceSpan() already does, and pass no metrics to takeAsyncLockWhenActorCacheReady in that case:
if (incomingRequests.empty()) return SpanParent(nullptr);
return getMetrics().getSpan();
Also worth considering: the delete-queue task should probably not abort() the whole actor when draining fails. A failed continuation is a bug in that continuation, not a reason to break every stub.
Impact
- With the
FinalizationRegistry pattern, which Effect's HttpClient uses for every response, any GC that happens during another DO's or the Worker's request can kill whichever idle DO created those responses. The result looks random and depends on load.
- Callers that keep a DO stub for longer than one call see a permanent outage until the caller is recreated.
Our workarounds: we removed the trigger by patching Effect's OTLP exporter to read every response body, successful or not, so no response stays in HttpClient's FinalizationRegistry (with it, the forced-GC repro in our app breaks 0 of 3 runs). As a safety net, callers get a new stub for every call and retry once on RpcCallError. Neither addresses the underlying abort: any promise an idle DO created that settles from another request still kills it.
Version
workerd 2026-09-01(npm@cloudflare/workerd-darwin-arm64,workerd --versionprintsworkerd 2026-09-01), macOS arm64.mainand in the latest release,v1.20260924.1(src/workerd/io/io-context.c++,IoContext::startDeleteQueueSignalTaskandIoContext::getCurrentTraceSpan).Summary
Take a Durable Object that has no request in flight. If a promise created in its IoContext settles while a different request in the same isolate is running, workerd schedules the continuation on the DO's delete queue. The delete-queue task then calls
context->run(...). For an actor,runSinglewaits on the input gate withgetCurrentTraceSpan(). Outside a JS lock that falls back togetMetrics()→getCurrentIncomingRequest(), whichKJ_REQUIREs a current request. The task'scatch (...)callscontext->abort(e), so the idle actor is aborted. Every later call through a stub that reached that actor fails withinternal errorand the same exception. A stub obtained afterwards reaches a new instance.This happens in practice with Effect's HTTP client, which is isolate-wide.
effect/unstable/http/HttpClientregisters every response in a module-levelFinalizationRegistrywhose callback callscontroller.abort(). When V8 GC runs during another DO's request, the callback aborts a controller created by the idle DO, the abort settles the DO's promise, and the DO is killed. In our app, one forcedgc()in a Workerfetch2 s after the DO's last call broke it in 2 of 2 runs, against 0 of 2 control runs. Replacing the registry callback with a no-op made the same forced GC harmless in 2 of 2 runs.Minimal repro
config.capnp:main.js:repro.sh gc|resolve|control:Results, 3 runs each:
gcgc()inside Caller's request runs the FinalizationRegistry →abort()ERR internal error(3/3), foreverpongresolveERR internal error(3/3), foreverpongcontrolpark, no GC)pongpongworkerd.log:Symbolized with
atos -p <workerd pid>:We see the same stack, byte for byte in the frame offsets, in the real app (Alchemy local dev, one isolate with a Worker and three DO classes).
Root cause
IoContext::getPromiseContextTag, anIoCrossContextExecutorover the delete queue). Settling it from another context goes throughDeleteQueue::scheduleAction→tryScheduleAction, which fulfillscrossThreadFulfiller.IoContext::startDeleteQueueSignalTaskwakes up and callsco_await context->run(...).runSinglecallsa.getInputGate().wait(getCurrentTraceSpan())(io-context.h,runSingle).getCurrentTraceSpan()finds nocurrentLock(the task runs outside the JS lock), so it takes the fallbackreturn getMetrics().getSpan();.getMetrics()is*getCurrentIncomingRequest().metrics, andgetCurrentIncomingRequest()doesKJ_REQUIRE(!incomingRequests.empty(), "the IoContext has no current IncomingRequest", ...).catch (...) { context->abort(kj::getCaughtExceptionAsKj()); }instartDeleteQueueSignalTask. The actor is aborted,ActorContainer::monitorOnBrokenrecordsbrokenReason, and every stub channel bound to that container rethrows it.The actor had done nothing wrong. It was only idle when a queued continuation arrived.
addTaskandaddWaitUntilalready handle an emptyincomingRequests, andgetCurrentUserTraceSpan()returnsSpanParent(nullptr)in the same situation.getCurrentTraceSpan()has no such guard, and the delete-queue task is the one caller that reaches it with no request by design.Suggested fix
The delete-queue task reaches two request-bound calls in
runSinglefor an actor with no request, so guarding only one of them moves the failure to the other:a.getInputGate().wait(getCurrentTraceSpan()), wheregetCurrentTraceSpan()falls back togetMetrics().takeAsyncLockWhenActorCacheReady(now(), a, getMetrics())(io-context.h,runSingle), which callsgetMetrics()directly.Suggested fix: have
startDeleteQueueSignalTasktake the input lock and the async lock without request metrics (no trace span,takeAsyncLockWithoutRequest-style), the same wayaddTaskhandles the no-request case, so draining the delete queue never needs anIncomingRequest.A narrower alternative is to make both call sites request-optional: return
SpanParent(nullptr)fromgetCurrentTraceSpan()whenincomingRequests.empty(), asgetCurrentUserTraceSpan()already does, and pass no metrics totakeAsyncLockWhenActorCacheReadyin that case:Also worth considering: the delete-queue task should probably not
abort()the whole actor when draining fails. A failed continuation is a bug in that continuation, not a reason to break every stub.Impact
FinalizationRegistrypattern, which Effect'sHttpClientuses for every response, any GC that happens during another DO's or the Worker's request can kill whichever idle DO created those responses. The result looks random and depends on load.Our workarounds: we removed the trigger by patching Effect's OTLP exporter to read every response body, successful or not, so no response stays in HttpClient's
FinalizationRegistry(with it, the forced-GC repro in our app breaks 0 of 3 runs). As a safety net, callers get a new stub for every call and retry once onRpcCallError. Neither addresses the underlying abort: any promise an idle DO created that settles from another request still kills it.