Skip to content

semantically_similar_to is scoped to one chunk, so similarity across chunks is never emitted, and plain string metrics miss it in a reconciliation pass #3948

Description

@jnrod03-rgb

The subagent prompt in references/extraction-spec.md (unchanged through
v0.9.72) says:

Semantic similarity: if two concepts in this chunk solve the same problem ...

Chunks are dispatched in parallel and no subagent sees another's output, so a
similarity whose two ends land in different chunks cannot be emitted by
construction. Agents are not missing these edges, they are unable to produce
them, and each agent's output still reads as having followed the instruction.

Related: #65 (same diagnosis; closed after v0.3.17 added same-directory
grouping in Step B1, while "a full global linking pass remains on the table"),
#7 (local embeddings, open) and its PR #1126 (closed without merging), #537.
This report adds a concrete case from the current pipeline and what worked when
I built the controller-side pass.

Concrete case

A project's most consequential design rule is stated in two different data files
(two different directories, so same-directory grouping does not help). Both were
extracted correctly as rationale nodes, but in different chunks. Each has one
edge, to its own file; they are not connected to each other and ended up in
different communities. The graph holds the same fact twice and does not know it.

The more the corpus is split, the more of the cross-cutting edges fall between
chunks, so more parallelism means a sparser semantic layer.

Proposal

  1. Reconciliation pass after Step B3 collection, before Part C. With all
    chunks in hand, compare rationale/concept node text across different
    source files, and emit semantically_similar_to edges marked INFERRED with a
    score derived from the measured similarity. Report the candidate pairs rather
    than writing silently, since this is inference over content the controller
    did not author.
  2. Amend the spec wording so agents know the cross-chunk case is handled
    elsewhere.

What I learned implementing (1)

I checked the pass against the known pair above, chosen before picking a
metric:

  • Word-shingle Jaccard found 5 pairs and missed the known pair. Shingles
    measure near-verbatim copying; two texts stating one rule in different words
    share almost no trigrams.
  • TF-IDF cosine found 32 pairs and still missed it (0.371 against a 0.42
    threshold). Both texts are long and surround the shared core with their own
    material, which dilutes the cosine.
  • What identified it was the number of rare terms in common (nine
    low-frequency terms). Accepting a pair on either a strong cosine or a high
    rare-term overlap with a moderate cosine caught it and added 12 more true
    pairs, with no false positives on manual review.

Two things seem worth carrying into any implementation, embeddings included:
score paraphrase by shared low-frequency vocabulary, not string overlap, with a
second acceptance path for long texts whose similarity is concentrated; and
validate against a known pair fixed in advance, because 5 pairs and 32 pairs
both looked like success.

Environment: observed on graphifyy 0.9.55, graphify install --platform claude,
Claude Code on Windows 11; now on 0.9.65. Spec wording and absence of a linking
pass checked in v0.9.72.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions