Skip to content
25 changes: 17 additions & 8 deletions bench/ab.py
Original file line number Diff line number Diff line change
Expand Up @@ -104,15 +104,24 @@ def run_ab( # noqa: PLR0913 - the suite's own selectors, plus the revision and
entry = write(name, corpus_root, scale)
for config in configs:
sides: dict[str, list[Measurement]] = {'A': [], 'B': []}
try:
for _ in range(rounds):
sides['A'].append(spawn(name, config, entry, repeat, cwd=tree))
unavailable: str | None = None
for _ in range(rounds):
if unavailable is None:
try:
sides['A'].append(spawn(name, config, entry, repeat, cwd=tree))
except RuntimeError as err:
# A corpus for a feature the revision lacks can
# fail there; that side has no number.
unavailable = str(err)
# B still runs: a broken current worker is a failure,
# not missing evidence.
try:
sides['B'].append(spawn(name, config, entry, repeat, cwd=_ROOT))
except RuntimeError:
# A corpus for a feature the revision lacks can fail
# there; that side has no number, which is the result.
failed = 'A' if len(sides['A']) == len(sides['B']) else 'B'
print(f'{name}/{config:<20} fails in {failed}; not comparable')
except RuntimeError as err:
print(f'{name}/{config} FAIL in B (this tree):\n{err}')
return 1
if unavailable is not None:
print(f'{name}/{config} fails in A; not comparable (B succeeded):\n{unavailable}')
continue
print(_row(f'{name}/{config}', sides['A'], sides['B']))
finally:
Expand Down
15 changes: 2 additions & 13 deletions docs/12-performance.md
Original file line number Diff line number Diff line change
Expand Up @@ -302,26 +302,15 @@ a change to a hot path, or to any code a claim is made about:
and this tree over one corpus. A time delta counts only if it is larger than
the reported noise. Record the allocation delta whether or not the time
moved, because it is nearly deterministic. A field added to every shape
shows there and nowhere else.
shows there and nowhere else. A workload the base revision cannot run is
reported as not comparable; one this tree cannot run fails the command.
3. Where the question is a leaf function's constant factor, add or run a
`bench micro` case at sizes that cover the function's range.
4. For a new feature, the base revision has no comparable number. Run
`python -m bench linearity --bench NAME` for time and memory instead.

Include the workload, both deltas, and the noise in the commit message.

A service workload's measured region is a cold snapshot: the parse, plus the
first query of each kind, then reuse of the built indices. Read its delta in
two parts. The parse side is where the base revision built and retained every
file's tree and the comparison does not. The first-use side is where a query
composes on demand the tree the base retained at parse time, one composition
per file per snapshot, about a third of a parse of the file; repeated
queries of the same kind are cached and show no delta. The first-use cost is
avoidable only by retaining the trees, the per-snapshot memory cost
(docs/15 § 2). When a change moves composition from parse time to first use,
name which side the delta is on before attributing it to the change's hot
path (docs/21 § 2, docs/21 § 4).

`bench compare` against the committed baseline is only a coarse check for
large regressions: the baseline and the comparison run at different times, and
the machine drifts in between. Follow § 5.1 before optimizing and § 5.2 when
Expand Down
5 changes: 5 additions & 0 deletions docs/13-public-api.md
Original file line number Diff line number Diff line change
Expand Up @@ -136,6 +136,11 @@ Default lint no longer needs whole trees merely to detect `schema:` and
retention. A consumer needing only that compact fact should read it rather
than retaining the whole source tree.

`Raml.written_sections` holds, per file, each section key an entity wrote
there (`types:`, a method's `headers:`, a type's `facets:`), with its owner's
ID and where its value ends, also independent of source retention. A section
a template contributed is not recorded (docs/21 § 4).

A configuration file's `parser:` section is `ParserConfig`. `limits(options)`
applies its `max_include_size`, `max_depth` and `regex_engine`; its
`workspace_root` and `remote` are left to the host, which weighs them against
Expand Down
10 changes: 7 additions & 3 deletions docs/15-implementation-plan.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,8 +32,11 @@ The shared-index/outline bundle is the next isolated recovery against merged
but mixed retained allocation rises 7.5%, chiefly the cached outline. An explicit
acceptance of that new retention trade-off is recorded; the scope, measurements and
ownership evidence are in the [shared-index report](reports/2026-10-10/service-shared-indices.md).
Remaining candidates are cursor-local source lookup and accurate authored section
ranges, each measured independently.
Cursor-local source lookup replaces hover's whole-file key index
([cursor hover report](reports/2026-10-11/service-cursor-hover.md)), and
outline sections are placed at the keys the parser records
([section report](reports/2026-10-11/service-section-ranges.md)); the
recorded sections retain up to 1.65% more on `endpoints`.
The record-backed representation replacement remains parked; its recurring
rebuild and first-use costs are recorded in the
[service cost review](reports/2026-10-10/service-authoring-cost-review.md).
Expand Down Expand Up @@ -88,7 +91,8 @@ the media-type fixes and the type-walk work landed together:
[integration report](reports/2026-10-10/service-baseline-integration.md).
Model-backed declaration inlays remove that first-use work from hints alone;
their accepted request-order peak trade-off is recorded in the inlay recovery
report. Cursor-local source lookup remains a separately measured candidate.
report. Hover reads source keys along the cursor's path and keeps no index of
them.

## 3. Potential future work

Expand Down
10 changes: 6 additions & 4 deletions docs/18-linting.md
Original file line number Diff line number Diff line change
Expand Up @@ -226,10 +226,12 @@ fastraml lint [--config FILE] [--severity S] [--ruleset NAME]
```

`--list-rules` and `--explain` require no input document. Other invocations
parse each file with `unwrap=True`, `validate=False`, and `retain_source=True`.
Validation remains the `validate` command's responsibility. A parse failure
produces no findings for that file, reports the parser error, and causes a
nonzero result; remaining files are still attempted.
parse each file with `unwrap=True`, `validate=False` and `retain_text=True`,
which source suppression reads (§ 4); `retain_source=True` only when
`Linter.requires_source` (§ 6). Validation remains the `validate` command's
responsibility. A parse failure produces no findings for that file, reports
the parser error, and causes a nonzero result; remaining files are still
attempted.

Formats:

Expand Down
52 changes: 31 additions & 21 deletions docs/21-language-service.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ and several views, and only `cli/` imports it (`docs/02` § 2;
| `service/text.py` | converting a column between fastRAML and a protocol |
| `service/queries.py` | the queries, in fastRAML positions |
| `service/outline.py` | the outline, over the authorship view (`docs/16` § 10) |
| `service/hover.py`, `service/hoverdocs.py` | author-facing hover, source-key indices and explanatory prose (§ 4.2) |
| `service/hover.py`, `service/hoverdocs.py` | author-facing hover, source-key lookup and explanatory prose (§ 4.2) |
| `service/datahover.py` | typed `DataNode` key and value spans, shared across value-bearing sites (§ 4.2) |
| `service/index.py` | lazy declaration, ID and reverse-hierarchy lookups shared by snapshot queries (§ 4) |
| `service/lenses.py` | code-lens sites and on-demand effective RAML type rendering (§ 4.3) |
Expand Down Expand Up @@ -197,10 +197,18 @@ The first outline request for a URI caches its complete result on the snapshot,
including an empty outline. Later requests borrow the same list and symbols;
callers must treat them as read-only. Root/dependency edits create a new snapshot
and outline cache. A held older snapshot continues to answer from its older model.
Caching does not add authored section positions or populate source grammar.

The model keeps no position for a section's key (`types:`, a method's
`headers:`), so a section spans its entries and selects the first.
Caching does not populate source grammar.

A section is placed at the key its owner wrote (`types:`, a method's
`headers:`), spanning the key and its value, and is listed even when empty.
The decoders record these keys in `Raml.written_sections` (docs/13 § 2): the
root's, a type's `facets:`, a resource's `uriParameters:`, a method's or a
`describedBy:`'s `headers:`, `queryParameters:` and `body:`, a response's
`headers:` and `body:`, and a scheme's `describedBy:`. A key is recorded only
inside its owner's span in its file, and not under a method or response a
template wrote, which no outline lists. A table written `schemas:` is named
so. A section the parser recorded no key for spans its entries and selects
the first.

What each entry lists is the authorship view's (`docs/16` § 10): a type's own
members, not those it inherits, so an inherited property is outlined under the
Expand All @@ -216,8 +224,10 @@ inside a resource or method (`docs/16` § 10).
A file outlines what it wrote, selected by `location` over the model its
snapshot parsed. An Extension or Overlay lists the types and other
declarations it added to the master's tables, and, under a master resource's
path, the methods and resources it added there: a section spanning them, since
the resource's key it wrote is not in the model. The master lists its own.
path, the methods and resources it added there. The merge keeps the master's
key where both wrote one, so each document's root section keys and resource
paths are recorded from its own tree before the merge: the Extension's
`types:` and restated `/a:` are placed at its keys. The master lists its own.

A trait or resource type is listed by name alone. Its body is decoded only
where it is applied (`docs/08` § 5), and the model keeps it undecoded, so there
Expand Down Expand Up @@ -326,14 +336,18 @@ extent is not a fallback for an unknown child. Inline JSON has only the
encoded scalar's root span; hover does not invent spans for decoded children.

Source-only primitive tokens, including those
in an unapplied template, are read through the type-expression parser and
checked against the source text. No source query binds a reference. Explanatory
in an unapplied template, are read through the type-expression parser. A
token's offset is its column only where the scalar's one-line span is its
text, or its text in quotes; where an escape or a tag shifts the columns, no
token is found. No source query binds a reference. Explanatory
prose is a documentation catalogue, not a table that accepts fields or overrides
parser diagnostics.

Hover indices and formatted subjects are lazy per snapshot. Source keys are indexed once per queried
file from retained nodes, or a composition of its retained text when source
trees were not retained. These nodes and indices die with the snapshot.
Hover indices and formatted subjects are lazy per snapshot. Source keys are not
indexed: each hover reads the path to its cursor (`syntax.keys_at`), one entry
per mapping level found by binary search, in the file's retained nodes, or a
composition of its retained text when source trees were not retained. Those
nodes die with the snapshot.
The typed-data token index is owned by the snapshot, lazy and built once; formatted
data targets are cached by hover. The same typed-data targets supply go-to-definition
for nested field keys and scalar values. A definition request may populate that
Expand Down Expand Up @@ -501,12 +515,6 @@ value, so an integer larger than a double reaches a JavaScript client as
written. It is the preview's source in `contrib/fastraml-vscode`
(`docs/17` § 4).

**Latency.** On `large`, an edit costs 475 ms and allocates 48.8 MB before
its parser diagnostics, against 366 ms for a plain `unwrap+validate` parse
(`python -m bench run --bench large --config service`). About a quarter of
the parse composes the unchanged libraries: the most a compose cache (G8)
could save.

## 6. Verification

- `test_service_text.py`: conversion in each encoding, both ways.
Expand All @@ -518,12 +526,14 @@ could save.
lifetime, retained original trees and distinct JSON-include normalization.
- `test_service_queries.py`: each query on one document with a library, a
DataType include, a trait and a resource type; every query on a parse
stopped at each stage; an Extension's outline; and, over the TCK, that every
outline entry holds its selection and lies in its parent.
stopped at each stage; an Extension's outline and its own section keys;
empty, `schemas:` and template-supplied sections; and, over the TCK, that
every outline entry holds its selection and lies in its parent.
- `test_service_hover.py`: contextual field meanings, full Markdown prose,
authored and inherited summaries, aliases, optional versus nullable values,
source-only template help, token boundaries, opaque data, educational examples,
and user-defined facet descriptions at declarations and supplied keys.
and user-defined facet descriptions at declarations and supplied keys; over
the TCK, the cursor path reaches every key the grammar walk yields.
- `test_service_data_hover.py`: shared nested-field and scalar-value help for
custom facets, annotations, examples, defaults and enums; array items,
inheritance, patterns, discriminated and ambiguous unions, recursive types,
Expand Down
3 changes: 3 additions & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,9 @@ architecture. The implementation and these documents must agree.
for explaining past decisions. It is not normative.
- `research/` contains open investigations. Nothing there defines parser
behavior or may be required by the implementation.
- `reports/YYYY-MM-DD/` contains dated benchmark measurements, audits and other
point-in-time records. Results stay in those reports; the numbered documents
describe current contracts and behavior, not individual runs.

Current parser work is complete. Deferred work is listed in the non-normative
[status and roadmap](15-implementation-plan.md). Conformance status belongs in
Expand Down
54 changes: 54 additions & 0 deletions docs/reports/2026-10-11/service-cursor-hover.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,54 @@
# Cursor-local hover source lookup

Date: 2026-10-11. Branch: `fix/parked-transfers`, from master `069b41f`;
measured against its parent `131b08e`. This is the third isolated experiment
in the [recovery plan](../../research/2026-10-10/service-recovery-plan.md).

## 1. Goal, workload and cache boundaries

The goal is to remove the whole-file source-key index from the first hover in
a file, without slowing repeated hovers or changing which token is explained.

Before: the first hover in a URI sorted every key `syntax.keys` yields and
parsed every type value for built-in tokens, checked against a split of the
file's text. Both lists lived as long as the snapshot: 15,203 keys and about
2.6 MB on the hover corpus (one 13,204-line file). Later hovers bisected them.

After: each hover reads `syntax.keys_at`, the keys enclosing the cursor, one
per mapping level, found by binary search over that mapping's entries, which
are in document order. Sequences are scanned, because an alias item keeps its
anchor's earlier position. Only the type value the cursor is in is parsed for
built-in tokens. Nothing is cached; the composed tree is still the snapshot's
or the shared source owner's (docs/21 § 4), so the cost of composition is
unchanged. Include-context discovery for headerless files is unchanged.

A token's offset is taken as its column where the scalar's one-line span is
its text, or its text in quotes, rather than by comparing a line of the source.
An escape or a tag lengthens the span, so those tokens stay unexplained, as
before. A type expression folded over several lines now gets no token help;
before, a token on its first line could.

## 2. Correctness and reach

Over every TCK `.raml` file, `keys_at` at each key's start returns that key
with the same site and table as the whole-file walk (more than 10,000 keys).
A quoted union explains `nil` at its exact span; an escaped `n\x69l` explains
nothing. The existing hover suite, including unapplied-template built-ins,
opaque data, expanded aliases and inherited include contexts, passes
unchanged. The inlay reach test now fails if `Hover._keys_at` is called.

## 3. Results

`python -m bench ab 131b08e --bench hover --bench service-session --config unwrap --rounds 5`:

| workload | time | noise | peak | retained |
|---|---|---|---|---|
| `hover` | 332.9 → 337.3 ms, noise | 16.2 % | unchanged | 26.29 → 23.71 MB (−9.8 %) |
| `service-session` | 1260.5 → 1142.8 ms (−9.3 %) | 7.4 % | unchanged | 27.50 → 24.87 MB (−9.6 %) |

`hover` runs one probe per declaration, so its time is dominated by repeated
hovers; the removed index was paid once. `service-session` hovers sparsely
after each of three edits, so it pays the index once per edit and gains.

`python -m bench linearity --bench hover`: time 0.992, peak 0.992,
retained 0.998.
Loading
Loading