Skip to content

Async persistence restored before attach() never rejoins an in-flight run #1639

Description

@AlanChauchet

TanStack AI version

@tanstack/ai-client 0.36.1 and 0.37.0 — reproduced against clean npm installations of both. Originally observed through @tanstack/ai-react 0.29.4.

Framework/Library version

Originally observed in React Native with an AsyncStorage-backed ChatClientPersistence adapter. The isolated reproduction below uses Node.js 24.14.1, with no React, native runtime, server, or model API required.

Describe the bug and the steps to reproduce it

An in-flight chat does not reconnect after a cold start when the asynchronous persistence read completes before the view calls client.attach().

The transcript restores successfully, but connection.joinRun() is never called. The partial assistant message therefore remains frozen even though the server still owns the generation and its replay log. Attaching before the same async read resolves works.

  1. Persist a combined { messages, resume: { resumeState: { threadId, runId } } } record for an in-flight run, with no pending interrupts.
  2. Construct a fresh ChatClient with an asynchronous getItem and a connection supporting joinRun.
  3. Let getItem resolve before mounting/attaching the view.
  4. Call client.attach().
  5. Observe that the messages are restored, but no rejoin occurs.

Expected: construction/hydration performs no network I/O before attachment; after attachment, the client joins the restored run exactly once, regardless of whether storage or attachment happened first.

Actual: only the attachment-first ordering calls joinRun. Storage-first restores the transcript but permanently skips the rejoin for that mount.

Your Minimal, Reproducible Example - (Sandbox Highly Recommended)

Self-contained executable reproduction below. It does not require application code, API keys, or a running backend.

In an empty directory:

npm init -y
npm install @tanstack/ai-client@0.37.0
# Save the following as repro.mjs
node repro.mjs

The same failure occurs with @tanstack/ai-client@0.36.1.

import assert from 'node:assert/strict';
import { setTimeout as delay } from 'node:timers/promises';
import { ChatClient } from '@tanstack/ai-client';

async function check(restoreBeforeAttach) {
  const joined = [];
  const snapshot = {
    messages: [{ id: 'user-1', role: 'user', parts: [{ type: 'text', content: 'Hello' }] }],
    resume: { resumeState: { threadId: 'thread-1', runId: 'run-1' } },
  };
  const client = new ChatClient({
    threadId: 'thread-1',
    persistence: {
      getItem: async () => structuredClone(snapshot),
      setItem: async () => {},
      removeItem: async () => {},
    },
    connection: {
      async *connect() { throw new Error('A restore must not start a new generation'); },
      async *joinRun(runId) {
        joined.push(runId);
        yield { type: 'RUN_STARTED', threadId: 'thread-1', runId, timestamp: Date.now() };
        yield { type: 'TEXT_MESSAGE_START', messageId: 'reply-1', role: 'assistant', timestamp: Date.now() };
        yield { type: 'TEXT_MESSAGE_CONTENT', messageId: 'reply-1', delta: 'Resumed reply', timestamp: Date.now() };
        yield { type: 'TEXT_MESSAGE_END', messageId: 'reply-1', timestamp: Date.now() };
        yield { type: 'RUN_FINISHED', threadId: 'thread-1', runId, timestamp: Date.now() };
      },
    },
  });
  try {
    if (restoreBeforeAttach) {
      await delay(0); // Async storage resolves before the view's mount effect.
      assert.equal(client.getMessages().length, 1); // Transcript restored successfully.
    }
    client.attach();
    await delay(50);
    assert.deepEqual(joined, ['run-1'], `restoreBeforeAttach=${restoreBeforeAttach}`);
    console.log(`PASS restoreBeforeAttach=${restoreBeforeAttach}`);
  } finally {
    client.detach();
    client.dispose();
  }
}

await check(false); // Control: attach first, then storage resolves. Passes.
await check(true);  // Bug: storage resolves first. joinRun is never called.

Observed output on both unpatched versions:

PASS restoreBeforeAttach=false
AssertionError [ERR_ASSERTION]: restoreBeforeAttach=true
+ actual - expected
+ []
- [ 'run-1' ]

Root cause

In packages/ai-client/src/chat-client.ts:

  • Constructor-time rejoinRunId is populated from synchronously restored state (or initialResumeSnapshot). An async persistence read leaves it unset.
  • applyPersistedResume() receives the async snapshot and calls maybeRejoinInFlight(runId).
  • If the view has not attached yet, maybeRejoinInFlight() returns because tailing is false. This guard is correct: an unmounted/discarded client must not open a connection.
  • The async run ID is not retained for a later attachment. attach() checks the constructor-time rejoinRunId, so it has no pending rejoin to perform.

Suggested fix and validation

Retain a bare in-flight run ID when async hydration completes, so attach() can consume it later. The core change tested locally is:

- private readonly rejoinRunId: string | null | undefined
+ private rejoinRunId: string | null | undefined

  private applyPersistedResume(snapshot: ChatResumeSnapshot): void {
    this.applyResumeSnapshot(snapshot)
    const hasInterrupts =
      Array.isArray(snapshot.pendingInterrupts) &&
      snapshot.pendingInterrupts.length > 0
    const runId = snapshot.resumeState?.runId
+   this.rejoinRunId = hasInterrupts ? null : runId
    if (!hasInterrupts && runId) {
      this.maybeRejoinInFlight(runId)
    }
  }

I verified that adding the equivalent assignment to the published 0.37.0 runtime makes both cases in the standalone reproduction pass. A complete fix should also keep this remembered ID current and clear it when the run becomes terminal, rather than leaving a stale constructor-era ID for later reattachments. Interrupt-only snapshots should retain their existing behavior instead of being automatically tailed.

The application-level regression test also recreates the client from a frozen JSON storage snapshot and routes the real client through a durable NDJSON backend. It checks both storage/attachment orderings, replay without a second generation POST, and no duplicate assistant text. Model output is simulated.

Related reports checked

Screenshots or Videos (Optional)

The executable assertions above reproduce the failure without a UI.

Do you intend to try to help solve this bug with your own PR?

A tested local workaround and reproduction are included above; no PR is being opened with this report.

Terms & Code of Conduct

  • I agree to follow this project's Code of Conduct.
  • I understand that a bug without a reliable, debuggable reproduction may not be fixed and may be closed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions