Skip to content

openai-base: Chat Completions reports premature SSE EOF as a successful stop #1603

Description

@L-1ngg

TanStack AI version

Repository commit 0f737ac7a60334c53d5178bc9d47d4ce540a9a2a, which matched main on 2026-10-02.

  • @tanstack/ai: 0.63.0
  • @tanstack/ai-openai: 0.25.1
  • @tanstack/openai-base: 0.12.1

These are checkout package versions. The reproduction uses builds from this commit, not separately installed npm releases.

Framework/Library version

Node.js 24.14.1, pnpm 11.9.0, OpenAI SDK 6.41.0. No UI framework.

Describe the bug and the steps to reproduce it

A Chat Completions response can reach EOF before protocol completion, yet chat() reports an ordinary successful stop. It emits RUN_FINISHED and calls middleware onFinish with the partial text.

The reproduction returns an HTTP 200 SSE body containing one text delta, partial, followed by EOF. The body has no non-null finish_reason, usage chunk, or [DONE] marker. It does not throw a read error or receive a cancellation request.

Steps

  1. Use a fresh repository checkout at the commit above.
  2. Install dependencies and build the relevant packages from the repository root:
pnpm install --frozen-lockfile
NX_DAEMON=false pnpm exec nx run-many --targets=build --projects=@tanstack/ai,@tanstack/openai-base,@tanstack/ai-openai --parallel=2 --skip-nx-cache
  1. Save the standalone reproduction below as repro.mjs in the repository root.
  2. Run node repro.mjs.

The successful control passes. The premature-EOF assertion fails with premature EOF called onFinish (1 !== 0).

Actual behavior

Response body Terminal event onFinish / onError calls Result
Text, finish_reason: "stop", usage-only chunk, [DONE] RUN_FINISHED 1 / 0 stop; all 12 usage tokens retained
Text, then EOF without any completion marker RUN_FINISHED 1 / 0 stop; partial text reported as complete

The second case emits:

RUN_STARTED
TEXT_MESSAGE_START
TEXT_MESSAGE_CONTENT
TEXT_MESSAGE_END
RUN_FINISHED

Its middleware result is:

{
  "onFinish": [{ "finishReason": "stop", "content": "partial" }],
  "onError": []
}

Consumers that use onFinish to accept a completed answer receive no indication that the response lacks a completion signal.

Expected behavior

For the incomplete response above, retain the partial text and report an explicit incomplete-stream error. Use RUN_ERROR and onError, without an ordinary successful RUN_FINISHED or onFinish.

Keep normal lifecycle cleanup and collect usage that arrives after a valid finish_reason. The existing incomplete-stream error code is a possible fit.

Cause and compatibility constraints

The EOF drain deliberately closes unfinished lifecycles. The finish-reason mapping also defaults an absent reason to stop, then emits RUN_FINISHED.

This fallback has compatibility history that a fix must address:

The OpenAI SDK consumes [DONE] without yielding it. When both streams lack finish_reason, iterator EOF cannot distinguish bare EOF from [DONE]. A fix needs an explicit completion policy at the appropriate layer. A missing-reason guard changes existing compatibility behavior and needs separate assessment.

Related reports checked on 2026-10-02

  • #1447, fixed by #1494, concerns Responses streams. That fix explicitly leaves the Chat Completions fallback in place.
  • #1497 concerns structured Responses streams.
  • #1601 and #1602 concern malformed tool arguments. This reproduction contains no tool calls or argument parsing.

The search found related reports above, but no separate report for this text-only Chat Completions EOF case.

Evidence boundary

The reproduction exercises the real OpenAI SDK decoder and public chat() API with a controlled SSE body. It proves the success classification after premature protocol EOF. It does not measure live provider failures or prove that historical provider exceptions still occur. No real model requests ran because no credentials were available.

Your Minimal, Reproducible Example - (Sandbox Highly Recommended)

Complete reproduction files and commands.

The complete local reproduction is also below. It needs no API key or network request. Run it from the repository root after the build step above.

import assert from 'node:assert/strict'
import { pathToFileURL } from 'node:url'

// Run from the repository root after building the three affected packages.
const root = pathToFileURL(`${process.cwd()}/`)
const { chat } = await import(new URL('packages/ai/dist/esm/index.js', root))
const { createOpenaiChatCompletions } = await import(
  new URL('packages/ai-openai/dist/esm/index.js', root)
)

const chunk = (delta, finish_reason = null) => ({
  id: 'chatcmpl-repro',
  object: 'chat.completion.chunk',
  created: 1,
  model: 'gpt-4o',
  choices: [{ index: 0, delta, finish_reason }],
})

async function run(complete) {
  const chunks = [chunk({ role: 'assistant', content: 'partial' })]
  if (complete) {
    chunks.push(chunk({}, 'stop'), {
      ...chunk({}),
      choices: [],
      usage: { prompt_tokens: 10, completion_tokens: 2, total_tokens: 12 },
    })
  }
  const body =
    chunks.map((value) => `data: ${JSON.stringify(value)}\n\n`).join('') +
    (complete ? 'data: [DONE]\n\n' : '')
  const result = { events: [], onFinish: [], onError: [] }
  const adapter = createOpenaiChatCompletions('gpt-4o', 'local-placeholder', {
    baseURL: 'https://provider.example/v1',
    maxRetries: 0,
    // Exercise the real SDK decoder without network access or a real key.
    fetch: async () =>
      new Response(body, {
        headers: { 'content-type': 'text/event-stream' },
      }),
  })
  for await (const event of chat({
    adapter,
    messages: [{ role: 'user', content: 'Hello' }],
    debug: false,
    middleware: [
      {
        name: 'observe-completion',
        onFinish(_context, info) {
          result.onFinish.push({
            finishReason: info.finishReason,
            content: info.content,
            usage: info.usage,
          })
        },
        onError(_context, info) {
          result.onError.push(info.error.message)
        },
      },
    ],
  })) {
    result.events.push(event.type)
  }
  return result
}

const control = await run(true)
assert.equal(control.onFinish.length, 1)
assert.equal(control.onFinish[0].finishReason, 'stop')
assert.equal(control.onFinish[0].usage.totalTokens, 12)
assert.equal(control.onError.length, 0)
assert.equal(control.events.includes('RUN_ERROR'), false)

const prematureEof = await run(false)
console.log(JSON.stringify({ control, prematureEof }, null, 2))

// Expected contract: partial output must not become an ordinary success.
// This assertion fails on the reported baseline.
assert.equal(prematureEof.onFinish.length, 0, 'premature EOF called onFinish')
assert.equal(prematureEof.events.includes('RUN_FINISHED'), false)
assert.equal(prematureEof.events.includes('TEXT_MESSAGE_CONTENT'), true)
assert.equal(prematureEof.events.includes('RUN_ERROR'), true)
assert.equal(prematureEof.onError.length, 1)

Screenshots or Videos (Optional)

Not applicable. The script prints both cases and the failed assertion.

Do you intend to try to help solve this bug with your own PR?

Yes, I am also opening a PR that solves the problem along side this issue

Terms & Code of Conduct

  • I agree to follow this project's Code of Conduct
  • I understand that if my bug cannot be reliable reproduced in a debuggable environment, it will probably not be fixed and this issue may even be closed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

has-prAn open PR references this issuewaiting-on: maintainerThe ball is in the maintainers’ court

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions