Skip to content

Represent reasoning as transcript entries with Anthropic support - #264

Open
qoli wants to merge 5 commits into
huggingface:mainfrom
qoli:codex/display-reasoning
Open

qoli wants to merge 5 commits into
huggingface:mainfrom
qoli:codex/display-reasoning

Conversation

@qoli

@qoli qoli commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Reasoning belongs in the transcript, separate from person-facing answer content. This revision replaces the original Response.reasoning / Snapshot.reasoning proposal with the Foundation Models 27 shape: Transcript.Entry.reasoning(Transcript.Reasoning), containing a stable ID, segments, opaque Data? signature, and metadata. Existing cumulative transcriptEntries carry it; response and snapshot initializer selectors remain unchanged.

Use case

I maintain AIReasoningCore and SwiftChat. AIReasoningCore implements a custom PiAILanguageModel: it maps pi-ai-swift's normalized reasoning events and terminal reasoning blocks into AnyLanguageModel sessions. SwiftChat displays provider-supplied reasoning in a separate collapsible area while answer text is still empty, accumulates it across tool rounds, and persists/restores it with the conversation. Reasoning transcript entries meet this use case and also retain the opaque state needed for provider replay. No separate response property or alternate session abstraction is needed.

The initial PR only exposed a field for custom models. This version also implements the built-in Anthropic path: thinking and redacted-thinking entries in streaming and nonstreaming responses, stable streamed IDs, opaque signature preservation, and replay after Codable transcript restoration. Redacted payloads have no display segments. A README example shows configuring Anthropic thinking and reading reasoning and answer separately.

Scope

  • Preserve existing nonstreaming Anthropic tool behavior; streamed tool rounds carry cumulative reasoning entries. This does not add a new nonstreaming orchestration loop.
  • Cross-provider conversations remain portable: adapters omit reasoning they cannot replay from provider requests, while retaining it in the transcript for display and persistence. Anthropic skips foreign or unmarked reasoning and still validates native Anthropic signatures. CoreML keeps its existing prompt-only behavior. Native Foundation Models 27 reasoning bridging remains outside this change.
  • Structured snapshots still require a representable partial value; scalar outputs may defer reasoning until one exists.
  • Cancellation changes and their tests are removed from this PR and split into Make cancelled stream transcript retention explicit #266, based independently on main. The session diff here is documentation only.
  • Includes current upstream main (0626c0f) without rewriting the existing PR branch history.

Review follow-up and validation

Addresses the review in #264 (comment): unsupported replay no longer blocks model switching, and the initializer-compatibility test uses a watchOS 27 availability guard. The new provider-request fixtures compare against each adapter's existing no-reasoning request projection; original Codable history remains intact.

  • Default offline suite: 595 tests in 60 suites passed.
  • MLX,Llama,CoreML traits: 750 tests in 69 suites passed.
  • iOS Simulator build, watchOS device build, and watchOS Simulator test-target build passed.
  • Changed-file strict formatting and whitespace checks passed.

Live-provider credentials were removed and CI=1 was used. No live model requests were made. The watchOS result is compilation of the actual tests, not just the library.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Built-in providers never populate reasoning, and the modified public initializer selectors break source compatibility.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 2 High severity · 1 Medium severity

Open (3)
What changed in this PR

Adds display-only cumulative reasoning to session responses and streaming snapshots.

Changes:

  • Propagates reasoning through response collection and structured wrappers.
  • Prevents canceled partial streams from committing empty responses.
  • Adds reasoning and cancellation regression tests.
File Description
LanguageModelSession.swift Adds reasoning APIs, propagation, and cancellation handling.
LocalGenerationUsage.swift Preserves reasoning in structured snapshots.
ReasoningTests.swift Tests reasoning behavior and cancellation.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread Sources/AnyLanguageModel/LanguageModelSession.swift Outdated
Comment thread Sources/AnyLanguageModel/LanguageModelSession.swift Outdated
Comment thread Sources/AnyLanguageModel/LanguageModelSession.swift Outdated
@mattt

mattt commented Sep 24, 2026

Copy link
Copy Markdown
Collaborator

Hi @qoli. Thanks for this, and for the tests, especially the one that caught a cancelled stream committing a partial response.

Before we go further, can you tell us how you're using reasoning? None of the built-in providers fill it in yet (Anthropic already collects thinking deltas, for example, but they don't reach the new field), so as written it's only reachable from a custom LanguageModel.

We'd also like whatever we add here to line up with Foundation Models 27, which represents reasoning as a transcript entry, Transcript.Entry.reasoning(_:), rather than as a property on the response. AnyLanguageModel 2.0 is going to follow that API (#210), so a reasoning property that stays out of the transcript pulls in a different direction. If your use case works with reasoning as transcript entries, that's the shape we'd rather build toward.

Could you also split the cancellation change into its own PR? It's a good catch, but it changes behavior for every cancelled stream, not only ones with reasoning: skipping the commit also drops lastSnapshot.transcriptEntries, so if a tool already ran, its call and output never reach the transcript. We should decide what to keep in that case on its own.

@qoli qoli changed the title Expose cumulative display reasoning on session responses Represent reasoning as transcript entries with Anthropic support Sep 24, 2026
@mattt
mattt requested a balanced review from Copilot September 24, 2026 23:21

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

CoreML still silently drops restored reasoning despite the documented explicit-rejection guarantee.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
Resolved since last review (3)

Comment thread README.md Outdated
Comment on lines +476 to +478
The Anthropic adapter can replay its own reasoning entries after Codable restoration.
Other adapters currently reject reasoning replay explicitly rather than flattening
it into answer text or silently dropping it. For structured scalar outputs that
@mattt

mattt commented Sep 25, 2026

Copy link
Copy Markdown
Collaborator

Hi @qoli. Thank you for turning this around so quickly, and for being so flexible about the shape. Moving reasoning into the transcript, matching Foundation Models' Transcript.Reasoning, and wiring up Anthropic is exactly what we were hoping for, and SwiftChat makes the use case clear.

Two things before we merge:

  • Right now a transcript with a reasoning entry makes every provider except Anthropic throw ReasoningReplayError.unsupportedProvider, and Anthropic throws for reasoning from another provider. So once a conversation has reasoning in it, it can't move to a different model, which is a big part of what AnyLanguageModel is for. Since reasoning is for display, could providers that can't replay it skip those entries instead?
  • The watchOS job fails to build ReasoningTests.swift: String's Generable conformance needs watchOS 27 there.

@qoli

qoli commented Sep 26, 2026

Copy link
Copy Markdown
Contributor Author

Thanks @mattt — addressed both points in 748539e.

  • Adapters now omit reasoning they cannot replay from their provider requests, while keeping the original entries in the transcript for display and persistence. Anthropic also skips foreign or unmarked reasoning; validation of its own replay signatures remains in place. CoreML retains its existing prompt-only behavior, and the documentation now makes that explicit.
  • The initializer-compatibility test now guards the watchOS 27 availability requirement. I verified both the watchOS device build and the watchOS Simulator test-target build.

Added deterministic streaming/nonstreaming request fixtures for Anthropic, OpenAI Chat Completions/Responses, OpenResponses, Gemini, and Ollama. They verify that unsupported reasoning/signatures are omitted without changing each adapter's existing request projection, and that Codable history remains intact.

Validation: 595 default offline tests passed; 750 tests passed with the MLX, Llama, and CoreML traits. iOS Simulator build and strict formatting checks also passed. No live model calls were used for this verification.

This change and the separate #266 refactor are now adopted together in my maintained fork (4d088af) and consumed by AIReasoningCore and SwiftChat Alpha 14. The combined fork passed 763 offline tests, with 47 Core and 29 SwiftChat coordinator tests passing against the published dependency pins. The two upstream PRs remain independently reviewable.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Restored nonstreaming tool histories can produce an empty assistant text block that Anthropic rejects.

Review effort: Balanced
Findings: 1 High severity · 1 Medium severity

Open (2)

Comment on lines +1007 to +1013
func appendAssistant(_ content: [AnthropicContent]) {
if let last = messages.last, last.role == .assistant {
messages[messages.count - 1] = .init(role: .assistant, content: last.content + content)
} else {
messages.append(.init(role: .assistant, content: content))
}
}
@mattt

mattt commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

Hi @qoli. Sorry, I stepped on this with #271. It fixed the empty text block that Copilot found here, in the same .response case of toAnthropicMessages(), and that's now the only conflict with main. To keep both changes, filter the blocks and then call appendAssistant:

case .response(let response):
    // Anthropic rejects text blocks without non-whitespace text,
    // such as the empty response of a turn that only called tools.
    let content = convertSegmentsToAnthropicContent(response.segments).filter { block in
        guard case .text(let text) = block else { return true }
        return !text.text.allSatisfy(\.isWhitespace)
    }
    guard !content.isEmpty else { continue }
    appendAssistant(content)

With that, the full test suite passes for me locally. Copilot's other open note, about Core ML dropping reasoning, is the behavior I asked for, so you can ignore it. Once you've merged main, I'll ask Copilot for another review, and we'll aim to get this into 0.15.0.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants