Skip to content

fix(score): sum scores across indexed columns - #18

Merged
eeeebbbbrrrr merged 6 commits into
fix-score-snapshot-memofrom
fix-multi-field-scores
Oct 6, 2026
Merged

eeeebbbbrrrr merged 6 commits into
fix-score-snapshot-memofrom
fix-multi-field-scores

Conversation

@claude

@claude claude Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Before: when a query searched several tin-indexed columns, tin.score, tin.full_score, and tin.max_score scored only the first searched column. WHERE title ==> 'ruby' OR body ==> 'ruby' gave body-only matches a zero score, and swapping the predicates moved the zero to the title-only matches. AND had the same order dependence.

After: each indexed column's searches score on their own, against that index's corpus, and a row's score is the sum over the columns whose searches it matches, as in tin. The issue's fixture scores 0.87138504 and 0.6931472 in either order. max_score is the highest summed score among matching documents. A row that only a non-search qual admits scores NULL, as in tin. max_score's policy, and combined score/full_score calls, follow tin too.

How this fixes #15

score_support used to bind the first ==> operand that had a tin index. It now groups the searches by index, and score_bound takes every group in one call. It sums the scores of the groups that match the row exactly, so the result does not depend on group order. In max_score mode it makes one corpus pass over the summed scores. Partial-index eligibility is checked per column.

Stacked on #17.

Fixes #15

Scoring bound only the first searched column that had a tin index, so a
search over several indexed columns scored one of them and the result
depended on the order of the predicates. Group the searches by the tin
index each binds to, score every group, and sum the groups that match the
row, as tin does. max_score takes the highest summed score among
documents that match every group the quals require.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012mKGtrg7DEhVPgqYrJtHGg
@claude
claude Bot requested review from eeeebbbbrrrr and piki October 6, 2026 18:07
tin cannot scan a disjunction in which a search without a tin index is
an alternative to an indexed one, so it refuses to score that query and
names the unindexed search. Do the same, require a group for max_score
whenever the quals imply a match of it, and keep an empty sum at +0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012mKGtrg7DEhVPgqYrJtHGg
Claude and others added 2 commits October 6, 2026 18:19
Pin the scores tin returns for the order-swapped searches, the summed
scores and max_score of rows matching several columns, and the shapes
with a partial index that the quals do not imply.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012mKGtrg7DEhVPgqYrJtHGg
Scoring and highlighting treated any operator named ==> with two
arguments as a search, so a ==> from another schema or for other types
bound to them. A ==>(text, int) put an int query into the text[] of
searches and crashed the backend. Compare the operator OID with Lead's
pg_catalog.==>(text, text) instead, looked up on each use.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012mKGtrg7DEhVPgqYrJtHGg
@eeeebbbbrrrr
eeeebbbbrrrr added this pull request to stack #20 October 6, 2026 18:25
Claude and others added 2 commits October 6, 2026 18:40
A row that only a non-search qual admits, such as the id = 8 of
body ==> 'gems' OR id = 8, scored 0. tin returns NULL for it, so return
NULL when none of the row's indexed columns matches its searches.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012mKGtrg7DEhVPgqYrJtHGg
tin scores a relation with one policy. max_score takes the policy of the
relation's tin.score calls, sharing their dense_ratio, term_add, and
term_replace, and otherwise reports the full_score maximum, including
when it is the only scoring call. Under the tin.score policy tin computes
no maximum for several indexes or a disjunction with other quals, and
returns NULL. It refuses tin.score beside tin.full_score on one relation,
and tin.score calls whose scan arguments differ. Do the same.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012mKGtrg7DEhVPgqYrJtHGg
@eeeebbbbrrrr
eeeebbbbrrrr marked this pull request as ready for review October 6, 2026 20:17
@eeeebbbbrrrr
eeeebbbbrrrr merged commit 50f8c04 into main Oct 6, 2026
4 checks passed
@eeeebbbbrrrr
eeeebbbbrrrr deleted the fix-multi-field-scores branch October 6, 2026 20:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Multi-field relevance scores depend on predicate order

1 participant