Skip to content

Idle Durable Object is aborted with "the IoContext has no current IncomingRequest" when one of its promises settles from another request #7517

Description

@FNDEVVE

Version

  • workerd 2026-09-01 (npm @cloudflare/workerd-darwin-arm64, workerd --version prints workerd 2026-09-01), macOS arm64.
  • The code path is the same on main and in the latest release, v1.20260924.1 (src/workerd/io/io-context.c++, IoContext::startDeleteQueueSignalTask and IoContext::getCurrentTraceSpan).

Summary

Take a Durable Object that has no request in flight. If a promise created in its IoContext settles while a different request in the same isolate is running, workerd schedules the continuation on the DO's delete queue. The delete-queue task then calls context->run(...). For an actor, runSingle waits on the input gate with getCurrentTraceSpan(). Outside a JS lock that falls back to getMetrics() → getCurrentIncomingRequest(), which KJ_REQUIREs a current request. The task's catch (...) calls context->abort(e), so the idle actor is aborted. Every later call through a stub that reached that actor fails with internal error and the same exception. A stub obtained afterwards reaches a new instance.

This happens in practice with Effect's HTTP client, which is isolate-wide. effect/unstable/http/HttpClient registers every response in a module-level FinalizationRegistry whose callback calls controller.abort(). When V8 GC runs during another DO's request, the callback aborts a controller created by the idle DO, the abort settles the DO's promise, and the DO is killed. In our app, one forced gc() in a Worker fetch 2 s after the DO's last call broke it in 2 of 2 runs, against 0 of 2 control runs. Replacing the registry callback with a no-op made the same forced GC harmless in 2 of 2 runs.

Minimal repro

config.capnp:

using Workerd = import "/workerd/workerd.capnp";

const config :Workerd.Config = (
  services = [ (name = "main", worker = .worker) ],
  sockets = [ (name = "http", address = "127.0.0.1:18787", http = (), service = "main") ],
  v8Flags = ["--expose-gc"],
);

const worker :Workerd.Worker = (
  modules = [ (name = "main.js", esModule = embed "main.js") ],
  compatibilityDate = "2026-09-01",
  durableObjectNamespaces = [
    (className = "Callee", uniqueKey = "callee"),
    (className = "Caller", uniqueKey = "caller"),
  ],
  durableObjectStorage = (inMemory = void),
  bindings = [
    (name = "CALLEE", durableObjectNamespace = "Callee"),
    (name = "CALLER", durableObjectNamespace = "Caller"),
  ],
);

main.js:

import { DurableObject } from "cloudflare:workers";

// Isolate-wide, like effect/unstable/http/HttpClient's response registry.
const registry = new FinalizationRegistry((controller) => controller.abort());
const pending = [];

export class Callee extends DurableObject {
  // Leaves a promise owned by this IoContext that settles when the controller is aborted.
  async park(mode) {
    let settle;
    new Promise((resolve) => (settle = resolve)).then(() => {});
    if (mode === "gc") {
      const controller = new AbortController();
      controller.signal.addEventListener("abort", () => settle());
      registry.register({}, controller);
    } else {
      pending.push(settle);
    }
    return "parked";
  }
  async ping() {
    return "pong";
  }
}

export class Caller extends DurableObject {
  // One Callee stub per Caller activation.
  held = this.env.CALLEE.getByName("callee");
  async heldPing() {
    return await this.held.ping();
  }
  async gc() {
    globalThis.gc();
    return "gc";
  }
  async resolve() {
    for (const settle of pending.splice(0)) settle();
    return "resolved";
  }
}

export default {
  async fetch(request, env) {
    const url = new URL(request.url);
    const callee = env.CALLEE.getByName("callee");
    const caller = env.CALLER.getByName("caller");
    try {
      switch (url.pathname) {
        case "/park": return new Response(await callee.park(url.searchParams.get("mode")));
        case "/held": return new Response(await caller.heldPing());
        case "/fresh": return new Response(await callee.ping());
        case "/gc": return new Response(await caller.gc());
        case "/resolve": return new Response(await caller.resolve());
      }
    } catch (e) {
      return new Response(`ERR ${e.message}`, { status: 500 });
    }
    return new Response("not found", { status: 404 });
  },
};

repro.sh gc|resolve|control:

W=${WORKERD:-workerd}
$W serve --experimental config.capnp > workerd.log 2>&1 &
PID=$!; sleep 3
curl -s localhost:18787/held >/dev/null            # Caller creates and uses its Callee stub
case $1 in control) m=gc;; *) m=$1;; esac
curl -s "localhost:18787/park?mode=$m" >/dev/null  # Callee leaves a pending promise, then goes idle
sleep 2                                            # Callee idle (< 10 s, not evicted)
case $1 in gc) curl -s localhost:18787/gc >/dev/null;; resolve) curl -s localhost:18787/resolve >/dev/null;; esac
sleep 1
echo "held stub:  $(curl -s localhost:18787/held)"
echo "held stub:  $(curl -s localhost:18787/held)"
echo "fresh stub: $(curl -s localhost:18787/fresh)"
kill $PID

Results, 3 runs each:

mode what settles Callee's promise while Callee is idle held stub fresh stub
gc gc() inside Caller's request runs the FinalizationRegistry → abort() ERR internal error (3/3), forever pong
resolve Caller calls the parked resolver directly ERR internal error (3/3), forever pong
control nothing (same park, no GC) pong pong

workerd.log:

workerd/jsg/util.c++:391: error: e = workerd/io/_virtual_includes/io/workerd/io/io-context.h:1301: failed: expected !incomingRequests.empty(); the IoContext has no current IncomingRequest; lastDeliveredLocation = src/workerd/api/worker-rpc.c++:2411:3 in run
stack: 1023f2497 102de8213 102df3d9f 102dff0e3 1022192c7 1021d3707 10221a22f 1033cd5d8 1033cfa13 1022cc688 10358e27c 10358d00c 1031cbc28 1031cc224; sentryErrorContext = jsgInternalError; wdErrId = …
workerd/io/worker.c++:2584: info: uncaught exception; source = Uncaught (in promise); stack = Error: internal error; reference = …
    at async Caller.heldPing (main.js:…)

Symbolized with atos -p <workerd pid>:

workerd::IoContext::getMetrics() + 131
workerd::IoContext::getCurrentTraceSpan() + 191
workerd::IoContext::runSingle<workerd::IoContext::startDeleteQueueSignalTask(workerd::IoContext*)::$_0>(…, kj::Maybe<workerd::InputGate::Lock>) + 771
workerd::IoContext::startDeleteQueueSignalTask(workerd::IoContext*) (.resume) + 151
workerd::server::Server::ActorNamespace::ActorContainer::monitorOnBroken(workerd::Worker::Actor&) (.resume) + 35
workerd::server::Server::ActorNamespace::ActorContainer::getActor() + 899
workerd::server::Server::ActorNamespace::ActorContainer::startRequest(workerd::IoChannelFactory::SubrequestMetadata) (.resume) + 67
workerd::(anonymous namespace)::PromisedWorkerInterface::customEvent(…) (.resume) + 51
capnp::QueuedClient::call(…)::'lambda1'(…)
capnp::LocalRequest::sendImpl(bool)::'lambda'()
capnp::Request<workerd::rpc::JsRpcTarget::CallParams, workerd::rpc::JsRpcTarget::CallResults>::send()::'lambda'(…)

We see the same stack, byte for byte in the frame offsets, in the real app (Alchemy local dev, one isolate with a Worker and three DO classes).

Root cause

  1. A promise created while the Callee's IoContext is current carries its promise context tag (IoContext::getPromiseContextTag, an IoCrossContextExecutor over the delete queue). Settling it from another context goes through DeleteQueue::scheduleAction → tryScheduleAction, which fulfills crossThreadFulfiller.
  2. IoContext::startDeleteQueueSignalTask wakes up and calls co_await context->run(...).
  3. For an actor with no input lock, runSingle calls a.getInputGate().wait(getCurrentTraceSpan()) (io-context.h, runSingle).
  4. getCurrentTraceSpan() finds no currentLock (the task runs outside the JS lock), so it takes the fallback return getMetrics().getSpan();. getMetrics() is *getCurrentIncomingRequest().metrics, and getCurrentIncomingRequest() does KJ_REQUIRE(!incomingRequests.empty(), "the IoContext has no current IncomingRequest", ...).
  5. The exception reaches catch (...) { context->abort(kj::getCaughtExceptionAsKj()); } in startDeleteQueueSignalTask. The actor is aborted, ActorContainer::monitorOnBroken records brokenReason, and every stub channel bound to that container rethrows it.

The actor had done nothing wrong. It was only idle when a queued continuation arrived. addTask and addWaitUntil already handle an empty incomingRequests, and getCurrentUserTraceSpan() returns SpanParent(nullptr) in the same situation. getCurrentTraceSpan() has no such guard, and the delete-queue task is the one caller that reaches it with no request by design.

Suggested fix

The delete-queue task reaches two request-bound calls in runSingle for an actor with no request, so guarding only one of them moves the failure to the other:

  • a.getInputGate().wait(getCurrentTraceSpan()), where getCurrentTraceSpan() falls back to getMetrics().
  • takeAsyncLockWhenActorCacheReady(now(), a, getMetrics()) (io-context.h, runSingle), which calls getMetrics() directly.

Suggested fix: have startDeleteQueueSignalTask take the input lock and the async lock without request metrics (no trace span, takeAsyncLockWithoutRequest-style), the same way addTask handles the no-request case, so draining the delete queue never needs an IncomingRequest.

A narrower alternative is to make both call sites request-optional: return SpanParent(nullptr) from getCurrentTraceSpan() when incomingRequests.empty(), as getCurrentUserTraceSpan() already does, and pass no metrics to takeAsyncLockWhenActorCacheReady in that case:

if (incomingRequests.empty()) return SpanParent(nullptr);
return getMetrics().getSpan();

Also worth considering: the delete-queue task should probably not abort() the whole actor when draining fails. A failed continuation is a bug in that continuation, not a reason to break every stub.

Impact

  • With the FinalizationRegistry pattern, which Effect's HttpClient uses for every response, any GC that happens during another DO's or the Worker's request can kill whichever idle DO created those responses. The result looks random and depends on load.
  • Callers that keep a DO stub for longer than one call see a permanent outage until the caller is recreated.

Our workarounds: we removed the trigger by patching Effect's OTLP exporter to read every response body, successful or not, so no response stays in HttpClient's FinalizationRegistry (with it, the forced-GC repro in our app breaks 0 of 3 runs). As a safety net, callers get a new stub for every call and retry once on RpcCallError. Neither addresses the underlying abort: any promise an idle DO created that settles from another request still kills it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions