A statement says its mood - #745
Merged
Merged
Conversation
…e's subject Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: WaylandYang <wayland0916@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: WaylandYang <wayland0916@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: WaylandYang <wayland0916@gmail.com>
…ubject rule is withdrawn Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: WaylandYang <wayland0916@gmail.com>
WaylandYang
force-pushed
the
feat/a-statement-says-its-mood
branch
from
September 17, 2026 15:03
a9f74dd to
b385dda
Compare
This was referenced Sep 17, 2026
Merged
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pitfall 2 of
docs/design/prior-work.md: a triple the sentence did not assert (OLLIE 2012; speculation is a property of the tuple, Dong et al. 2023).What lands
A statement says its mood. A statement the passage requires, plans, expects, forecasts or makes conditional carries a qualifier keyed
moodwhose value is the passage's own words for that ("要", "should", "will", "is expected to", "if the merger closes"); the phrase stays as written. A verb that reports (announced, said, revealed) is not a mood: what was announced is stated. Alignment (0044 cut 2) will read the qualifier and never materialize a statement with a mood as a typed fact; nothing in the ledger changes, the qualifier is astatement_qualifiersrow like any other.On the NVDA batch the model writes 44 moods on 901 statements:
willon the plans of the PORTS-Pike release,Outlookon every figure of the guidance table ("NVIDIA —GAAP operating expenses→ $9.2 {mood: Outlook}"),are expected to beon the outlook prose,subject to definitive agreementson the acquisition terms,If an emerging growth companyon the cover-page election. No reporting verb was written as a mood.Tried and withdrawn in the same change: a rule giving a subjectless sentence the addressee or the speaker the passage names as its subject. Under it the model returned an empty reply for a fifth of a filing's chunks, tables included, and the bullets of a release already take the company as their subject without it. The trade-in notice's directive paragraphs, whose addressee is outside the paragraph unit, stay an open item.
Table rows quoted without their leading pipe now still find their period in the chunk: the check looks at the line the located quote sits on, not at the quote's first character (the model often drops the
|at the start of a row).Measured
Same four NVDA documents, the same hour, the extraction endpoint with reasoning on (see below), the same fast judge:
The extra misworded statements are the securities boilerplate ("Section 21E —created→ safe harbor", "partners —are→ forward-looking statements"), not the moods; within the 3-point band this branch set for itself.
The endpoint, not the contract, was the variable all afternoon. Three earlier runs of this change measured 10% misworded and a fifth of the chunks empty. A control with the #744 binary in the same hour reproduced that (460 statements, 7 empty table chunks, 33% misworded on the exhibit) while the proxy's
thinking: disabledwas honoured, and reproduced #744's own numbers (864 statements, 3%) once reasoning was switched on explicitly. Every earlier measurement of #731 to #744 had reasoning on by accident. Extraction with this model needs reasoning on; the cost is 17k to 40k completion tokens a call rather than 700, and one to two minutes a call. The bench notes record it; a run's completion-token count is now the first thing to check before comparing two runs.Not in this change
Alignment reading
mood; the subject of a directive sentence whose addressee is outside the chunk; a bench that measures run-to-run stability (ATOM 2025) as its own number.🤖 Generated with Claude Code