Skip to content

chore(bench): add parse-phase benchmark - #126

Merged
luojiyin1987 merged 2 commits into
masterfrom
chore/bench-parse
Sep 26, 2026
Merged

luojiyin1987 merged 2 commits into
masterfrom
chore/bench-parse

Conversation

@luojiyin1987

@luojiyin1987 luojiyin1987 commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Add scripts/bench-parse.mjs.
  • Add the bench:parse script.
  • Add a Node 22 CI smoke step.
  • No parser runtime change.

Phases

The benchmark measures four phases on the same input:

phase call
A parseMd
B bare fromMarkdown + parser extensions
B2 B + recordingExtension
C parseMdWithSourceMap

Derived values: A-B (remark wrapper), B2-B (recording), C-B2 (post-parse source-map setup / indexing), C-B (total source-map increment).
indexNode is likely most of C-B2, but this harness does not time it alone.
C-B2 also covers RecordingState / WeakMap creation, the recordingExtension(state) call, parse-time snapshots, and the sourceMap object with its closures.

Each measured sample runs in its own child process.
The parent only aggregates.
A single process builds a large multiline AST with heavy GC churn, so an in-process median flips between runs.

Findings (256 KiB, local)

  • multiline-hmd: base parse is almost the whole cost.
    recording and post-parse setup stay near zero.
  • mixed-markdown: base parse is about 85 percent.
    recording is about 190 ms.
    post-parse setup is about 26 ms.

The benchmark does not identify one slow micromark tokenizer yet.

Smoke design

  • Check that every phase runs and returns a finite wall time.
  • Check AST parity across A, B, B2, and C.
  • Check 4x input growth with a ratio budget.
  • Do not check absolute wall time.

Note: multiline-hmd grows super-linearly at large sizes.
The min ratio from 64 KiB to 256 KiB is about 15x on a 4x input.
The smoke uses 16 KiB to 64 KiB and a 10x budget instead.
This keeps the check stable and still catches a gross regression.

Validation

  • pnpm run lint
  • pnpm run build
  • pnpm run test
  • node scripts/bench-parse.mjs --smoke
  • node scripts/bench-parse.mjs --shapes single-line,entity-dense,escape-dense,many-paragraphs,large-code-block --sizes 65536 --runs 1

Closes #125

Split parse cost into four phases on the same input:
A parseMd, B bare fromMarkdown, B2 B + recordingExtension,
C parseMdWithSourceMap.

Report derived A-B, B2-B, C-B2, and C-B.

Run each measured sample in its own child process. The parent only
aggregates. A single process builds a large multiline AST with heavy
GC churn, so an in-process median flips between runs.

Add bench:parse and a Node 22 smoke step.
The smoke checks finite output, AST parity, and 4x input growth.
It does not check absolute wall time.
Rename C-B2 to post-parse source-map setup / indexing. C-B2 also
covers RecordingState / WeakMap creation, the recordingExtension(state)
call, parse-time snapshots, and the sourceMap object with its closures.
The harness does not time indexNode alone.

Relax and explain the smoke growth comment. A 4x input can grow more
than 5x and stay linear. The 10x budget only catches severe regressions.

Delete a temp bundle directory only when this script created it.
A caller-supplied BENCH_PARSE_BUNDLE is no longer removed.
@luojiyin1987
luojiyin1987 merged commit ffb6e75 into master Sep 26, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

chore(bench): add parse-phase benchmark

1 participant