Skip to content

fix(core): recover when the router stops a repetitive generation - #38

Merged
TaewoooPark merged 2 commits into
mainfrom
fix/router-repetition-retry
Oct 4, 2026
Merged

TaewoooPark merged 2 commits into
mainfrom
fix/router-repetition-retry

Conversation

@TaewoooPark

Copy link
Copy Markdown
Owner

Problem

In a long interactive session (the SP3 micromagnetics benchmark), the conversation ended twice with nothing on screen. Both times Motif-3 had been reasoning for more than 10k tokens when Infron stopped the generation for repetition. The router does this inside a 200 response: it writes an error chunk into the SSE stream, then closes with an ordinary stop.

{"choices":[],"error":{"code":"repetition_detected","message":"Repetition was detected in the model's output and generation was stopped. Retry, or send a `repetition_penalty` above 1.","type":"generation_error"}}

readStream read usage and cost from chunks without choices and dropped the rest, so the error vanished. What was left looked like a finished reply with no text and no tool call. In a chat, a reply without an action ends the task, so the task ended with no summary, and the reasoning was hidden.

Changes

  • fix(core): retry a generation the router stops for repetition
    • An error chunk in the stream becomes a retryable TransportError of kind generation, with the server's code.
    • The loop samples the same request again at once, up to twice. Only the retry carries repetition_penalty: 1.05, as the router asks. The cut attempt is dropped.
    • If both retries are cut off too, the turn is handed back through the no-action path instead of ending the session.
    • Every other request is unchanged, in body and in recorded hash.
  • fix(core): keep a chat task going after an empty reply
    • A turn in which the server sent no text is handed back like any turn without an action, instead of ending the conversation silently.
    • A chat UI test whose fake model replied with nothing now replies ok.

Evidence

  • Frequency: 2 of 305 recorded responses were cut off, and both were over 10k completion tokens.
  • The two cut-off requests were replayed against Infron, twice per setting. A tool call came back in:
setting tool call reached
unchanged 3 of 4
repetition_penalty 1.05 4 of 4
repetition_penalty 1.1 3 of 4
  • Seeds are not deterministic on this endpoint, so even a plain retry draws a new sample.

Tests

pnpm typecheck, pnpm test (1644 passed), pnpm lint:tools and pnpm build all pass. New tests cover:

  • the error chunk in the stream
  • the body and hash staying unchanged without a penalty
  • the retry with the penalty
  • the fallback when the retries are cut off too
  • an empty chat reply

Treat an error chunk in the stream (Infron's repetition_detected) as a retryable generation error: resample at once, up to twice, with repetition_penalty 1.05 on the retry only, then hand the turn back instead of ending the session.
A turn in which the server sent no text is handed back like any turn without an action, instead of ending the conversation silently.
@TaewoooPark

Copy link
Copy Markdown
Owner Author

Maintainer review completed for 99daaed8e7bcc8022db1806c7a1efd92cb4d168f.

Live Motif model testing is currently unavailable because the Infron endpoint is unavailable to us. I did not repeat the historical live experiments described in this PR. The changed SSE parsing, retry request, retry limits, cancellation, empty-reply recovery, and record/replay behavior can be verified deterministically without a live model, so I have decided to merge based on those checks. This does not establish current endpoint health or that a penalty of 1.05 improves current real-model output quality.

An independent eight-case suite exercised the actual HttpTransport and runLoop with fragmented SSE byte streams: all eight passed at this head; seven failed on the unchanged baseline, while the unchanged HTTP/JSON error-compatibility case passed. Checks covered discarded partial tool calls, retry-only penalty and reset on the next turn, bounded exhaustion, abort before another request, whitespace-only responses, EOF error chunks, and recording/replay hashes. The existing 93 loop/transport/stream tests and 24 replay tests also passed, along with frozen install, typecheck, build, and tool-schema lint.

The regeneration budget is two retries per model turn; exhaustion returns through the existing bounded no-action path. All five head CI jobs passed: https://github.com/TaewoooPark/Motifcode/actions/runs/36267843721 . The current-main combination of #41, #40, and this PR also passed all 1,653 tests plus two cross-path checks combining paste normalization, generation recovery, and normal/cancelled tool batches. Python CI's five optional GPU-only skips remain outside this validation.

No blocking defect or mandatory follow-up implementation commit found. Merging as a squash commit after the preceding main checks complete. This is a source merge, not an npm release; no unrelated issue is being closed for this PR.

@TaewoooPark
TaewoooPark merged commit ad88ff9 into main Oct 4, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant