benchmark suite and results site for JS syntax highlighters
overview #
syntax-highlighter-bench is a benchmark suite and results site for JS syntax highlighters. It compares fuz_code, Twinkleplop, Prism, and Shiki, with Shiki's JS engine and its Oniguruma wasm engine counted as two libraries.
A run produces one results file holding many metrics, and the site is a view over that file where a reader picks the libraries, languages, and metrics to compare: the results page.
AI disclosure: this is an LLM-generated repo guided by a person.
What exists #
The harness that times the libraries, warm and from a cold start, the deterministic measurements, and the pages that render a results file are built.
- an adapter for each library, and the check that guards against a plain-text fallback, in
bench/libraries/andbench/check_output.ts— see adapters - the inputs and their manifest, in
bench/corpus/— see corpus - the language ids and the feature sets, in
bench/langs.tsandbench/sets.ts— see languages and sets - the timed harness, which measures each library alone in a process of its own and writes a
results file, in
bench/run.tsandbench/cell.ts— see method - cold start, part of the same run, where each sample is a fresh process that imports a
library and stops its own clock, in
bench/coldstart.ts— see method - the deterministic measurements, which are bundle size over each feature set, install
footprint, retained heap, token counts, output sizes, and coverage, in
bench/deterministic.ts, with their numbers committed asresults/deterministic.json— see deterministic metrics - the schema of the results file, in results_schema.ts — see results file
- the landing page and the results page, which render one validated results file, in
src/routes/andsrc/lib/— see results file for which file, and design for what is still to come
What is measured #
- Throughput, on one named machine: how long a warm library takes to tokenize an input, and to render it as HTML, with the peak memory of its process beside it.
- Cold start, on the same machine: how long a fresh Node process takes to import a library alone, with one language, and with a documentation site's languages, and to produce a first highlight. Each is measured for the bundled chunk and for the library imported unbundled, beside a bare Node baseline to subtract, with the process's peak memory.
- Deterministic numbers, the same anywhere: bundle size over each feature set, install footprint, retained heap, token counts, output sizes, and coverage.
Layout #
bench/— the harness. It runs on Node directly and is never imported by the site.results/— where the harness writes runs. The deterministic results are committed there, and a published timed run goes beside them asresults/latest.json.src/lib/results_schema.ts— the schema the harness and the site share, withsrc/lib/bench_constants.ts, the id lists it builds its enums from. The rest ofsrc/lib/turns a results file into the tables of the site.src/routes/— this site, built as static pages.
Neutrality #
The maintainer of this benchmark also maintains fuz_code, one of the measured libraries, so fairness has to come from the structure rather than from trust:
- the maintainer's library is flagged in the results file and on the site
- every input is labeled by its source, and real code that no measured library authored is the primary set; a snippet or a stress input written for this benchmark is labeled as that, never as neutral, and no stress input is tuned to a library
- each library is driven through one adapter with notes on the entry points it uses, so its maintainers can review or replace it
- the method and the commands to reproduce a run are published with the numbers
adapters #
Each library is measured through one adapter, and the adapter is the extension point. The
contract is in bench/adapter.ts, the adapters are under bench/libraries/, and bench/libraries.ts lists them.
interface HighlighterAdapter<TTokens = unknown> {
id: string;
label: string;
package: string; // its manifest is read for the version
by_maintainer: boolean;
note: string | null;
extra_packages?: ReadonlyArray<string>; // separately versioned packages whose versions are recorded too
langs: Partial<Record<Lang, string>>; // harness id to library id; absent means unsupported
via: Partial<Record<Lang, string>>; // a coverage footnote for a claimed language
output: 'classes' | 'inline_styles';
engine: 'scanner' | 'grammar_vm' | 'regex' | 'textmate';
theme_css: ReadonlyArray<string> | null; // the default theme's stylesheets; null when styles are inline
theme_schemes: ReadonlyArray<'light' | 'dark'>; // the color schemes that theme covers
wasm?: ReadonlyArray<{import: string; binary: string}>; // the wasm binaries it loads, counted apart from the JS
setup(langs: ReadonlyArray<Lang>): Promise<void>; // once per process, before any timing
tokenize(src: string, lang: Lang): TTokens; // timed: the library's own result
count_tokens(tokens: TTokens): number; // never timed
html(src: string, lang: Lang): string; // timed: the string consumers ship
bundle_entry(langs: ReadonlyArray<Lang>): string; // module source importing exactly these languages
bundle_html(exports: Record<string, any>, src: string, lang: Lang): string; // highlights through a built bundle's exports
bundle_html_source(lang: Lang): string; // that same call as an expression over `exports` and `src`
}Rules #
- Same stage.
tokenizeandhtmlare the narrowest public entries doing the same work as the other libraries': tokens out, and the HTML string a consumer ships. - Explicit languages. A library is measured only on the languages its adapter lists. A call for any other language, or for one that wasn't set up, throws. A library that returned escaped plain text for a language it didn't know would benchmark as very fast.
- Token counts are recorded, not equalised. Libraries split the same source differently, so each count is the library's own, and counting happens outside the timed call.
- Synchronous timed paths. A library with an async API does its loading in
setupand is timed through its synchronous core. - Browser-safe. An adapter uses no Node-only API, so it can run in a page unchanged. The harness reads versions, stylesheets, and wasm binaries from disk on its behalf.
bundle_entry and bundle_html serve the bundle sizes of deterministic metrics: the first writes the module a bundle is built from, and the
second highlights through that bundle's exports, which is how each bundle is shown to hold its
languages. bundle_html_source is the second as text: a cold-start process, which method describes, must load nothing but the module it measures, so it
can't load the adapter to make the call. The check of every bundle evaluates the text and
requires the same HTML.
What each adapter calls #
| library | tokenize | html | one token is |
|---|---|---|---|
| fuz_code | SyntaxStyler#lex | SyntaxStyler#stylize | a typed leaf or container |
| Twinkleplop | the language package's tokenize() | the language package's language() | a typed token after reclassification |
| Prism | Prism.tokenize | Prism.highlight | a Token at any depth |
| Shiki, both engines | codeToTokensBase | codeToHtml | a themed token, whitespace included |
Each adapter has a notes.md beside it with the entry points, the import shape,
how its tokens are counted, and its caveats, written so the library's maintainers can review
or replace it.
One default is changed. Shiki stops tokenizing a line after half a second and returns the rest of it as one token, so what a call returns depends on how fast the machine is at that moment: a first call on a busy machine can come back with fewer tokens than the next. Both Shiki adapters raise that limit out of reach rather than turning it off, so every call does the same work, and still reads the clock as Shiki's default does.
fuz_code's tokenize result is valid only until its next call: for performance, it
lexes into one buffer that every call reuses, where most other libraries return a result the
caller keeps. A consumer that keeps one copies it, and the timed call doesn't. Its notes.md says how the harness reads the result in time.
What the HTML looks like differs, and the results say so:
- fuz_code and Prism return bare
<span>elements with classes, and the consumer supplies the wrapper - Twinkleplop returns a
<pre>with a span for each line and classes on each token - Shiki returns a
<pre>with a span for each line and an inline color on each token, so it needs no stylesheet and its tokenizing includes resolving colors
Coverage #
Every adapter claims the timed languages. xml is claimed by all but Twinkleplop,
which has no XML language. Where support is indirect, the coverage matrix carries a footnote:
- fuz_code runs its TypeScript lexer for JS
- Prism's Svelte grammar is the third-party
prism-svelte, and its XML is its markup grammar under another name - Shiki's XML grammar loads the Java grammar it embeds
setup loads exactly the languages it is given, plus what a language embeds
structurally, like the script and style languages of HTML and Svelte. A timed process sets up
only the language it measures, so every library has the same languages loaded.
That decides how Markdown is measured: without highlighting inside fenced code blocks.
fuz_code, Prism, and Shiki highlight a fenced block only when its language is loaded, and
Twinkleplop never does, so loading one language puts them on the same footing. Two exceptions
come from what loads with Markdown itself: Prism's Markdown depends on its markup grammar, so
Prism still highlights a fence tagged html, xml, svg,
or markup, and fuz_code and Prism highlight a fence tagged as Markdown.
The output check #
Before a library is measured on a language, check_output runs it on the
language's harness snippet, a few lines written dense in syntax. The library must produce
tokens, and HTML with several kinds of styled span. A plain-text fallback has none, or only
structural ones: a wrapper for each line, and in Shiki one span in the default color. A
library that fails is recorded as excluded with the reason, which the results file keeps apart from unsupported.
Each input is then checked against a lower floor. It is lower because prose-heavy Markdown legitimately has few kinds of span, which is why the snippet decides and the floor only guards.
The tests run both for every adapter on every input of the corpus in every language it claims.
Adding a library #
- Add the library as a devDependency.
- Write
bench/libraries/<id>/adapter.ts, exporting an adapter, andnotes.mdbeside it. - Add the adapter to
ADAPTERSinbench/libraries.ts.
The tests then cover it, and a results file carries the roster of libraries, so the site needs no change.
corpus #
The corpus is the set of files the libraries are run on. It lives in bench/corpus/, and its manifest records what each input is and where it came
from. Every input is listed at the end of this page, with its upstream commit where it has
one.
Sources #
Every input is labeled by its source, in its path and in every result:
| source | what it holds |
|---|---|
neutral | real files from projects no measured library authored, copied unmodified at a pinned commit: the primary set |
shiki | the sample files Shiki's own benchmark uses, at a pinned commit |
home | a measured library's own samples, labeled with the library: fuz_code's test fixtures, and the corpus Twinkleplop benchmarks itself on |
harness | one short snippet for each language, written for this benchmark by its maintainer, who also maintains fuz_code |
stress | inputs that look for each library's worst cases: neutral inputs minified, and pathological files generated here by the maintainer, who also maintains fuz_code |
The neutral set has a small, a medium, and a large file for each timed language. The other sources are kept so a reader can see whether a result holds on inputs a library chose for itself, or on hostile ones, and the results page filters by source. No summary pools two sources. A home input is always shown under the library whose samples it is, and libraries' home inputs are never summarized together: they contributed very different amounts.
Stress inputs #
The stress inputs are generated by a script in the repository, bench/corpus_stress.ts, and none is tuned to a library. A library that fails on
one, or is far slower there than anywhere else, is a finding: it is shown, as an exclusion
with its reason or as the slow number, never smoothed over by changing the input. They come in
two kinds:
- derived — a neutral input as real code ships: the JS and the CSS minified by esbuild, and the JSON serialized with no whitespace, each one long line. It keeps its source's pins, is marked synthetic, and names how it was made, with the esbuild version, since another version may minify to other bytes
- written — a pathological file for each timed language: nesting far deeper than code is written, a line of TypeScript tens of KB long, an unterminated string and block comment near the top of a file, template literals nested in their substitutions, Markdown delimiters that never close, HTML tags that never close and attributes without quotes, and a selector list on one long line. Each says in its first comment what it stresses
Shiki is timed with its per-line time limit out of reach, as on every input, so on a stress input it highlights every line in full. With its default limit a line that slow comes back with its rest as plain text, so a consumer with default options gets that fallback there rather than the stress number.
Size tiers #
An input's tier comes from its size in bytes, never from its name:
| tier | size |
|---|---|
micro | under 600 bytes |
small | 600 bytes to under 3,200 |
medium | 3,200 bytes to under 32,000 |
large | 32,000 bytes and over |
The upper three center on roughly 1KB, 10KB, and 100KB. The micro tier is the docs-site case: a snippet of a few lines, where the fixed cost of a call outweighs the scanning. The large tier has no upper bound, and the neutral large inputs range from about 90KB to several hundred KB, so compare throughput in MB/s across languages, not time per call.
Synthetic inputs #
An input is marked synthetic when it is not one real file as published but was built by joining or repeating files to reach its size, or derived from one, as the minified stress inputs are. Real files are preferred. Most of Twinkleplop's own corpus is built this way and is marked.
One neutral input is synthetic: the large Svelte file. Svelte components of around 100KB are rare, so a script builds one from several real components of other projects. It splits each component into its scripts, markup, and style, and writes one of each block holding the bodies in order, unchanged. The result is shaped like one large component and is lexically well formed, but it is not a working component: unrelated scripts share one scope, so their names collide and Svelte's parser rejects it. No source is rewritten to avoid that. The manifest lists every component it is built from.
The manifest #
bench/corpus/manifest.json holds one BenchInput for
each input: its path, source, language, tier, size in bytes and lines, SHA-256, whether it is
synthetic, and its provenance. A results file carries the same entries for the inputs it
measured, described in results file.
Provenance is a list of BenchProvenance pins, each a repository, a commit, a path, a hash, and a license:
- none, for a harness snippet or a stress input written here
- one, for a verbatim copy, whose hash must equal the input's; the copy is still marked synthetic when its upstream built it by joining files
- one, for a derived stress input, whose
derivationnames the transformation that made it from that file - several, for a synthetic input, naming what it was built from
The manifest also holds a hash over the whole corpus. A run records it, so two runs over different inputs are never compared as one.
Checking and regenerating #
npm run corpus:check # the manifest matches the files, and every copy its pin
npm run corpus:manifest # regenerate the manifest and the table of inputs
npm run corpus:fetch # download neutral inputs by their pins, verifying each
npm run corpus:harvest -- --repos <dir> --check # the synthetic Svelte input matches its sources
npm run corpus:stress # regenerate the stress inputs and the manifest
npm run corpus:stress -- --check # the stress inputs are what generates Sizes, line counts, hashes, and tiers are derived from the files. The synthetic flag, the provenance, and the derivation are written by hand, or by the stress generator for its own inputs, and carried forward. The tests rebuild the manifest from disk and compare it to the committed one, so a file added, removed, or changed without regenerating fails. The tests also regenerate the stress inputs and compare them with the files.
The corpus is data: no formatter, linter, or type checker runs on it, since a reformatted file is no longer the upstream file.
Inputs #
Every input of the results this site renders, by source. The corpus hash of those results is e649d70d435efac06e26c753d9424bbf30122fcae8e8037396b3cc340ea1c101. The repository holds the same list as a table generated from the manifest.
| input | language | tier | bytes | lines | license | upstream, at its pinned commit |
|---|---|---|---|---|---|---|
neutral/bash/headerscheck | bash | medium | 10,792 | 280 | PostgreSQL | postgres/postgres@cc053b6e12 src/tools/pginclude/headerscheck |
neutral/bash/nvm.sh | bash | large | 179,635 | 5,416 | MIT | nvm-sh/nvm@a8bb497400 nvm.sh |
neutral/bash/postgres-wrapper.sh | bash | small | 707 | 25 | PostgreSQL | postgres/postgres@cc053b6e12 contrib/start-scripts/macos/postgres-wrapper.sh |
neutral/css/bulma.css | css | large | 763,923 | 21,564 | MIT | jgthms/bulma@c5cd04baf0 css/bulma.css |
neutral/css/default.css | css | medium | 12,501 | 535 | LicenseRef-W3C-Software-and-Document-2023 | w3c/csswg-drafts@0da42c8afd shared/style/default.css |
neutral/css/text.css | css | small | 875 | 48 | MIT | sveltejs/svelte.dev@a90ebe2304 packages/site-kit/src/lib/styles/text.css |
neutral/html/index.html | html | small | 1,526 | 51 | MIT | prettier/prettier@1dcd0b05d0 website/index.html |
neutral/html/Overview.src.html | html | large | 107,712 | 2,694 | LicenseRef-W3C-Software-and-Document-2023 | w3c/csswg-drafts@0da42c8afd selectors-3/Overview.src.html |
neutral/html/results.template.html | html | medium | 18,139 | 741 | MIT | sveltejs/svelte@7bc0a70fe6 benchmarking/compare/results.template.html |
neutral/js/class.js | js | medium | 9,760 | 371 | MIT | prettier/prettier@1dcd0b05d0 src/language-js/print/class.js |
neutral/js/client.js | js | large | 91,407 | 3,160 | MIT | sveltejs/kit@55ac0d5264 packages/kit/src/runtime/client/client.js |
neutral/js/parse.js | js | small | 963 | 41 | MIT | prettier/prettier@1dcd0b05d0 src/main/parse.js |
neutral/json/ast.json | json | large | 175,037 | 5,502 | Apache-2.0 | microsoft/typescript-go@89d5d5b284 _scripts/ast.json |
neutral/json/kit.tsconfig.json | json | small | 1,048 | 29 | MIT | sveltejs/kit@55ac0d5264 packages/kit/tsconfig.json |
neutral/json/typesMap.json | json | medium | 17,285 | 497 | Apache-2.0 | microsoft/TypeScript@637d5746b7 src/server/typesMap.json |
neutral/md/02-state.md | md | medium | 10,036 | 352 | MIT | sveltejs/svelte@7bc0a70fe6 documentation/docs/02-runes/02-$state.md |
neutral/md/Explainer.md | md | large | 169,115 | 3,504 | Apache-2.0 | WebAssembly/component-model@a25fc0b372 design/mvp/Explainer.md |
neutral/md/postgres.README.md | md | small | 989 | 21 | PostgreSQL | postgres/postgres@cc053b6e12 README.md |
neutral/rust/dent.rs | rust | medium | 11,211 | 352 | Unlicense OR MIT | BurntSushi/walkdir@4f26be4d45 src/dent.rs |
neutral/rust/raw_query.rs | rust | small | 852 | 37 | MIT | tokio-rs/axum@c59208c86f axum/src/extract/raw_query.rs |
neutral/rust/table.rs | rust | large | 100,960 | 3,197 | MIT OR Apache-2.0 | rust-lang/hashbrown@c62a63a61b src/table.rs |
neutral/svelte/Breadcrumb.svelte | svelte | small | 1,141 | 44 | MIT | techniq/svelte-ux@11c79b8c6e packages/svelte-ux/src/lib/components/Breadcrumb.svelte |
neutral/svelte/harvest.svelte synthetic | svelte | large | 106,723 | 3,186 | MIT | techniq/layerchart@ec05fb1b84 packages/layerchart/src/lib/components/Legend.sveltetechniq/layerchart@ec05fb1b84 packages/layerchart/src/lib/components/layers/Canvas.sveltetechniq/svelte-ux@11c79b8c6e packages/svelte-ux/src/lib/components/Button.sveltetechniq/svelte-ux@11c79b8c6e packages/svelte-ux/src/lib/components/TextField.sveltethemesberg/flowbite-svelte@85f20a048e src/lib/datepicker/Datepicker.sveltethemesberg/flowbite-svelte@85f20a048e src/lib/clipboard-manager/ClipboardManager.svelte |
neutral/svelte/Layer.svelte | svelte | medium | 9,455 | 343 | MIT | dimfeld/svelte-maplibre@05dd0b2a76 src/lib/Layer.svelte |
neutral/ts/core.ts | ts | large | 92,419 | 2,595 | Apache-2.0 | microsoft/TypeScript@637d5746b7 src/compiler/core.ts |
neutral/ts/options.ts | ts | medium | 10,080 | 251 | MIT | sveltejs/language-tools@bf2993e192 packages/svelte-check/src/options.ts |
neutral/ts/transform.ts | ts | small | 1,160 | 27 | Apache-2.0 | microsoft/TypeScript@637d5746b7 src/services/transform.ts |
| input | language | tier | bytes | lines | license | upstream, at its pinned commit |
|---|---|---|---|---|---|---|
shiki/bash/shellscript.sample | bash | small | 661 | 28 | MIT | shikijs/textmate-grammars-themes@45e292ef67 samples/shellscript.sample |
shiki/css/css.sample | css | micro | 578 | 46 | MIT | shikijs/textmate-grammars-themes@45e292ef67 samples/css.sample |
shiki/html/html.sample | html | small | 1,596 | 52 | MIT | shikijs/textmate-grammars-themes@45e292ef67 samples/html.sample |
shiki/js/javascript.sample | js | medium | 3,946 | 151 | MIT | shikijs/textmate-grammars-themes@45e292ef67 samples/javascript.sample |
shiki/json/json.sample | json | small | 880 | 38 | MIT | shikijs/textmate-grammars-themes@45e292ef67 samples/json.sample |
shiki/md/markdown.sample | md | medium | 3,725 | 170 | MIT | shikijs/textmate-grammars-themes@45e292ef67 samples/markdown.sample |
shiki/rust/rust.sample | rust | small | 1,011 | 39 | MIT | shikijs/textmate-grammars-themes@45e292ef67 samples/rust.sample |
shiki/svelte/svelte.sample | svelte | small | 714 | 28 | MIT | shikijs/textmate-grammars-themes@45e292ef67 samples/svelte.sample |
shiki/ts/typescript.sample | ts | small | 2,205 | 77 | MIT | shikijs/textmate-grammars-themes@45e292ef67 samples/typescript.sample |
| input | language | tier | bytes | lines | license | upstream, at its pinned commit |
|---|---|---|---|---|---|---|
home/fuz_code/bash/sample_complex.sh | bash | small | 2,749 | 190 | MIT | fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.sh |
home/fuz_code/css/sample_complex.css | css | micro | 502 | 45 | MIT | fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.css |
home/fuz_code/html/sample_complex.html | html | small | 733 | 42 | MIT | fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.html |
home/fuz_code/json/sample_complex.json | json | micro | 327 | 14 | MIT | fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.json |
home/fuz_code/md/sample_complex.md | md | small | 2,799 | 219 | MIT | fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.md |
home/fuz_code/rust/sample_complex.rs | rust | small | 3,012 | 118 | MIT | fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.rs |
home/fuz_code/svelte/sample_complex.svelte | svelte | small | 2,517 | 156 | MIT | fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.svelte |
home/fuz_code/ts/sample_complex.ts | ts | small | 2,883 | 152 | MIT | fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.ts |
| input | language | tier | bytes | lines | license | upstream, at its pinned commit |
|---|---|---|---|---|---|---|
home/twinkleplop/bash/bash.large.sh synthetic | bash | large | 98,856 | 5,045 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/bash.large.sh |
home/twinkleplop/bash/bash.medium.sh synthetic | bash | medium | 10,941 | 557 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/bash.medium.sh |
home/twinkleplop/bash/bash.small.sh synthetic | bash | small | 1,054 | 46 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/bash.small.sh |
home/twinkleplop/bash/deploy.sh | bash | medium | 5,567 | 192 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/deploy.sh |
home/twinkleplop/css/css.large.css synthetic | css | large | 97,622 | 4,814 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/css.large.css |
home/twinkleplop/css/css.medium.css synthetic | css | medium | 10,562 | 670 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/css.medium.css |
home/twinkleplop/css/css.small.css synthetic | css | small | 1,246 | 76 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/css.small.css |
home/twinkleplop/css/site.css synthetic | css | large | 40,864 | 1,901 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/site.css |
home/twinkleplop/html/app.html | html | small | 656 | 18 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/app.html |
home/twinkleplop/html/html.large.html synthetic | html | large | 99,865 | 4,426 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/html.large.html |
home/twinkleplop/html/html.medium.html synthetic | html | medium | 10,262 | 455 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/html.medium.html |
home/twinkleplop/html/html.small.html synthetic | html | small | 1,327 | 52 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/html.small.html |
home/twinkleplop/js/bench-suite.js synthetic | js | medium | 26,389 | 1,062 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/bench-suite.js |
home/twinkleplop/js/javascript.large.js synthetic | js | large | 102,836 | 4,843 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/javascript.large.js |
home/twinkleplop/js/javascript.medium.js synthetic | js | medium | 10,597 | 297 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/javascript.medium.js |
home/twinkleplop/js/javascript.small.js | js | small | 813 | 21 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/javascript.small.js |
home/twinkleplop/json/compiled-grammar.json | json | large | 124,022 | 14,193 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/compiled-grammar.json |
home/twinkleplop/json/json.large.json synthetic | json | large | 99,831 | 4,192 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/json.large.json |
home/twinkleplop/json/json.medium.json synthetic | json | medium | 10,131 | 420 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/json.medium.json |
home/twinkleplop/json/json.small.json synthetic | json | small | 974 | 41 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/json.small.json |
home/twinkleplop/md/architecture.md | md | medium | 25,141 | 341 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/architecture.md |
home/twinkleplop/md/grammar-docs.md synthetic | md | medium | 20,842 | 549 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/grammar-docs.md |
home/twinkleplop/md/markdown.large.md synthetic | md | large | 101,226 | 2,572 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/markdown.large.md |
home/twinkleplop/md/markdown.medium.md synthetic | md | medium | 9,733 | 809 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/markdown.medium.md |
home/twinkleplop/md/markdown.small.md synthetic | md | small | 1,186 | 72 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/markdown.small.md |
home/twinkleplop/rust/pipeline.rs | rust | medium | 8,806 | 326 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/pipeline.rs |
home/twinkleplop/rust/rust.large.rs synthetic | rust | large | 101,709 | 4,830 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/rust.large.rs |
home/twinkleplop/rust/rust.medium.rs synthetic | rust | medium | 12,710 | 602 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/rust.medium.rs |
home/twinkleplop/rust/rust.small.rs synthetic | rust | small | 1,260 | 64 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/rust.small.rs |
home/twinkleplop/svelte/site-codepanel.svelte | svelte | medium | 4,712 | 181 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/site-codepanel.svelte |
home/twinkleplop/svelte/site-inspectorpanel.svelte | svelte | medium | 3,947 | 190 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/site-inspectorpanel.svelte |
home/twinkleplop/svelte/svelte.large.svelte synthetic | svelte | large | 99,669 | 4,641 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/svelte.large.svelte |
home/twinkleplop/svelte/svelte.medium.svelte synthetic | svelte | medium | 10,090 | 436 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/svelte.medium.svelte |
home/twinkleplop/svelte/svelte.small.svelte | svelte | small | 714 | 31 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/svelte.small.svelte |
home/twinkleplop/ts/core-frames.ts synthetic | ts | large | 57,288 | 1,546 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/core-frames.ts |
home/twinkleplop/ts/core-runtime.ts synthetic | ts | large | 70,294 | 1,903 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/core-runtime.ts |
home/twinkleplop/ts/core-types.ts | ts | large | 58,047 | 1,405 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/core-types.ts |
home/twinkleplop/ts/typescript.large.ts synthetic | ts | large | 122,206 | 3,321 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/typescript.large.ts |
home/twinkleplop/ts/typescript.medium.ts synthetic | ts | medium | 7,628 | 393 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/typescript.medium.ts |
home/twinkleplop/ts/typescript.small.ts | ts | small | 799 | 27 | MIT | pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/typescript.small.ts |
| input | language | tier | bytes | lines | license | upstream, at its pinned commit |
|---|---|---|---|---|---|---|
harness/bash/snippet.sh | bash | micro | 435 | 21 | written for this benchmark | |
harness/css/snippet.css | css | micro | 412 | 24 | written for this benchmark | |
harness/html/snippet.html | html | micro | 496 | 21 | written for this benchmark | |
harness/js/snippet.js | js | micro | 465 | 14 | written for this benchmark | |
harness/json/snippet.json | json | micro | 318 | 18 | written for this benchmark | |
harness/md/snippet.md | md | micro | 383 | 21 | written for this benchmark | |
harness/rust/snippet.rs | rust | micro | 572 | 24 | written for this benchmark | |
harness/svelte/snippet.svelte | svelte | micro | 540 | 29 | written for this benchmark | |
harness/ts/snippet.ts | ts | micro | 509 | 19 | written for this benchmark |
| input | language | tier | bytes | lines | license | upstream, at its pinned commit |
|---|---|---|---|---|---|---|
stress/bash/nested.sh | bash | small | 2,233 | 58 | written for this benchmark | |
stress/css/default.min.css synthetic | css | medium | 8,359 | 1 | LicenseRef-W3C-Software-and-Document-2023 | minified by esbuild 0.28.1, from w3c/csswg-drafts@0da42c8afd shared/style/default.css |
stress/css/selectors.css | css | medium | 29,316 | 135 | written for this benchmark | |
stress/html/unclosed.html | html | medium | 20,358 | 80 | written for this benchmark | |
stress/js/client.min.js synthetic | js | medium | 28,413 | 11 | MIT | minified by esbuild 0.28.1, from sveltejs/kit@55ac0d5264 packages/kit/src/runtime/client/client.js |
stress/js/nested_templates.js | js | medium | 14,828 | 86 | written for this benchmark | |
stress/js/unterminated.js | js | medium | 17,013 | 409 | written for this benchmark | |
stress/json/ast.min.json synthetic | json | large | 68,385 | 1 | Apache-2.0 | minified: parsed and serialized again with no whitespace, from microsoft/typescript-go@89d5d5b284 _scripts/ast.json |
stress/json/nested.json | json | medium | 18,559 | 457 | written for this benchmark | |
stress/md/nested.md | md | medium | 7,494 | 114 | written for this benchmark | |
stress/rust/nested.rs | rust | medium | 5,441 | 139 | written for this benchmark | |
stress/svelte/nested_blocks.svelte | svelte | medium | 4,732 | 108 | written for this benchmark | |
stress/ts/long_line.ts | ts | medium | 28,857 | 14 | written for this benchmark | |
stress/ts/nested_blocks.ts | ts | medium | 11,864 | 197 | written for this benchmark |
Licenses #
Every copied file stays under its upstream license, which its provenance records. The corpus readme and the notice beside the neutral inputs give the attribution.
languages and sets #
The harness owns the language ids. Each library names its languages differently, so an adapter
maps these ids onto its library's own. bench/langs.ts is the source of truth, and
a results file carries the list, so the tables on this page are read from the results this
site renders.
The languages measured are fuz_code's built-in set, so the list favors it. Prism and Shiki each support hundreds of languages, and none of that shows here.
Language ids #
| id | language | timed |
|---|---|---|
ts | TypeScript | yes |
js | JS | yes |
css | CSS | yes |
html | HTML | yes |
xml | XML | no |
json | JSON | yes |
svelte | Svelte | yes |
md | Markdown | yes |
bash | Bash | yes |
rust | Rust | yes |
The timed languages are the ones every library in the starting roster supports, so a timed comparison is never missing a library by design.
A language that is an id but is not timed is one the roster doesn't share: XML is the case today, since Twinkleplop has no XML language. It appears in the coverage matrix and in feature sets. A cold-start scenario that loads one language loads a timed one, so it never names an untimed id.
The measured libraries support many more languages than these. An id is added when a feature set or a second library needs it.
Unsupported and excluded #
A library can be missing from a comparison for two different reasons:
- unsupported — the library's adapter doesn't claim the language. Adapters declare their languages explicitly and never fall back to plain text, which would otherwise benchmark as very fast.
- excluded — the adapter claims the language, but its output failed the check that runs before measuring, described in adapters.
Neither is recorded as a zero or counted in an average, and the results file keeps the two apart.
Feature sets #
A feature set is a named list of language ids that bundle size is measured over, and that a
cold-start scenario imports. The sets are declared in bench/sets.ts before any
numbers exist, so a set can't be chosen to flatter a result.
| id | label | languages | note |
|---|---|---|---|
core | core | none | |
docs | docs | js, ts, css, html, md, bash, json | |
web_rust | web + rust | ts, js, css, html, xml, json, svelte, md, bash, rust | fuz_code's full set (home turf) |
all | all | ts, js, css, html, xml, json, svelte, md, bash, rust | |
timed | every timed language | ts, js, css, html, json, svelte, md, bash, rust | |
single:<lang> | the language's name | one, for each language id | |
without:<lang> | timed less the language | every timed language but one, for each timed language |
What each is for:
core— the library with no languages registered: its runtime floor.single:<lang>— subtractingcoregives what that language adds, which the results page shows for every language.docs— what a documentation site imports, matching Twinkleplop's docs bundle. It is also the set a cold-start process loads in thesetscenario.web_rust— everything fuz_code supports, so it is home turf for that library, and its note says so. It is named by its content because the site is neutral.all— every language id. While it holds the same languages as another set, the two are one bundle, and the results page shows them as one row.timed— every timed language, which every library claims, andwithout:<lang>— the same less one of them. Subtracting the second from the first gives what one more language adds, without the runtime the first language pays for. deterministic metrics says how to read it.
A set is measured for a library only when the library supports every language in it. Otherwise the result names the languages it lacks. deterministic metrics describes how a set is built and measured.
method #
The timed harness measures throughput: how long each library takes to tokenize an input, and
to render it as HTML. It then measures cold start: what a library costs a process before it
has highlighted anything. bench/run.ts runs both and writes one results file.
The method builds on the comparison harness of Twinkleplop (MIT), and the anchor workload is ported from it. It differs in giving each library a process of its own, where Twinkleplop measures a cell's libraries together in one. The cold-start scenarios generalize Twinkleplop's cold-start harness across libraries.
Cells #
A cell is one input of the corpus in one mode, tokenize or html, measured across the libraries that claim its language. Its id is <file>:<mode>. A cell carries its input's source and size tier, so a
reader can compare on the neutral inputs alone, or see whether a result holds on a library's
home turf or on the stress inputs.
One process, one library #
A process measures one library on one cell and loads no other. Every number is therefore one library alone in a fresh process, and the roster can't move it: adding or removing a library changes no other library's result.
Two effects made that necessary, and both showed up as a library disagreeing with itself:
- Libraries that share code change how V8 optimizes that code for each other. Shiki's two engines share its tokenizer, and beside the wasm engine Shiki's JS engine measured far slower than it does alone.
- V8 sizes its young generation from recent allocation. In a shared process a library's time on a large input depended on which libraries ran beside it.
A cell is measured in passes. A pass is a process for every library, one after another, and the library a pass starts with changes from pass to pass and from cell to cell. The same code can settle at a slightly different speed from one process to the next, so a library's number pools the rounds of all its passes.
Node runs with its default flags, which is what a consumer gets. In a timed run one process runs at a time, and the parent does nothing while it measures: it blocks on the child and reads its output after it exits.
Inside a process #
Each step answers a way a simple timing loop goes wrong:
- One language loaded. The library sets up exactly the cell's language, once, before anything is timed. See adapters for what that means for Markdown.
- The output is checked first. The library must pass the output check on the language's harness snippet and then on the cell's input. A library that fails is recorded as excluded with the reason, and is not timed.
- Iterations calibrated for the library. Speeds differ by orders of magnitude, so a batch makes the number of calls that fills the target duration. A batch makes at least a few calls while they fit a cap, as many as fit when fewer do, and one when two no longer fit.
- A rewarm and a minor collection before every timed batch. The library runs untimed until it is back to speed, and a minor collection then clears what the rewarm allocated, so the collection it would cause doesn't land inside the batch. No full collection is forced between batches. One forced before every batch puts some libraries through a race each round: it can discard optimized code whose hidden classes only short-lived objects hold, so the library re-optimizes inside the clock, and it re-rolls how memory a library allocates outside the JS heap is reused, so a process settles at one of two speeds. A minor collection does neither.
- Results are kept and counted. A batch keeps its last result, and after the clock stops that result is counted. A count that differs from the checked one fails the cell: a tokenizer that silently stops early on a slow line must not be averaged in as a fast one.
- The HTML is read. Some libraries return a string the engine hasn't joined
up yet, which costs little to build, and every consumer then pays to flatten it on first
use. In
htmlmode the timed call reads one character of the string, for every library, so that cost is inside the clock. It depends on the size of the output, not on the library, so it weighs most on the fastest one, and it is free for a library whose string is already flat. Intokenizemode every library returns finished arrays, so there is nothing to force.
The forced minor collection needs the process started with --expose-gc, which bench/run.ts passes to each process.
A full run measures each library in 3 processes per cell, each timing 4 rounds of 20ms batches, after a 100ms warmup, with a 300ms rewarm before each batch. A library whose one call takes at least as long as the rewarm gets none: the batch before keeps it warm, and a rewarm would only repeat that call.
The rewarm is longer than the warmup on purpose. The engine still runs full collections of its own during a run, and one can leave code cold. Libraries recover from that at different speeds, Shiki more slowly than the others, and a library still recovering when its batch is timed reads slow, so the rewarm is long enough that every library in the roster has recovered.
What each number means #
Each library's numbers in a cell are a BenchCellValue:
| number | meaning |
|---|---|
ns_per_op | nanoseconds per call: the median across the rounds of every pass |
ops_per_sec, mb_per_sec | the same time as calls per second, and as megabytes of input per second, which is the one to compare across inputs of different sizes |
p10_ns, p90_ns | the 10th and 90th percentile across those rounds |
spread | the slowest round over the fastest, across passes as well as within them. Near 1 means they agreed, and a high spread says the cell was disturbed or the processes disagreed |
pass_medians_ns | each pass's own median over its rounds, in pass order: one number per process, so how far the processes disagreed can be read apart from the rounds |
iterations | calls per batch. Each process calibrates for itself, and this is the median across them |
rss_peak_bytes | the peak resident size of the library's process, the largest across its passes |
attempts | how many times the library was measured on the cell: 1 unless it was retried, as described below |
The token count and the size of the HTML come from the output check, under the same setup and
outside the clock. Libraries split the same source differently, so a token count is the
library's own and is recorded rather than equalised. Each is stated once for an input: the
token count on its tokenize cell and the output size on its html cell. deterministic metrics describes them, and the bundle sizes a run carries.
rss_peak_bytes is a whole process: Node itself, the harness, the library as its
adapter imports it, with one language set up, and the input. It is not what the library adds
to a page, and most of it is the same for every library. Compare it between libraries on the
same cell. On a small input it says little about footprint: a process warms up and is timed
for a fixed duration, so its peak follows how much the library allocated in that time more
than what it holds. What a library holds is the retained heap, described in deterministic metrics.
Cold start #
A timed cell measures a library that is already loaded and warm. Cold start measures the other end: a fresh Node process that loads the library and does one thing, once. Each sample is one process, and each BenchStartup entry is a scenario:
| scenario | what the process does |
|---|---|
bare | nothing: the Node baseline, with no library, which every other time includes |
core | imports the library with no language |
lang | imports the library with one language, for each timed language it claims |
set | imports the library with the languages of the docs feature set: what a
documentation site loads |
first_highlight | imports the library with one language and highlights that language's harness snippet once, to HTML: the time to a first highlight, for each timed language it claims |
What is imported is the module the adapter's bundle_entry writes
for a feature set (core, single:<lang>, docs):
the same module bundle size is measured over, so a time and a size describe one thing. The
adapter modules themselves are never loaded here. Some of them import every language, which
would charge a library for languages the scenario doesn't load. A first highlight makes the
call the adapter documents for that module (bundle_html_source), and the check of
every built bundle confirms it returns the same HTML as bundle_html.
Two forms of every scenario, apart from bare:
bundled— the built chunk whose size deterministic metrics reports, imported as one file. A library's wasm import stays outside its bundle, as it does for the size, so that one module still loads from the package.- unbundled — the same module's source, written to a file under
node_modules/.cache/in the checkout and imported from there. Node resolves each package from the checkout'snode_modulesand loads every file the library reaches, as a consumer without a bundler does.
The gap between the two is Node finding, reading, and compiling many files in place of one, and code the bundler removed, which a bundled page or server doesn't pay. Read the form that matches how the library would be shipped.
Start is the start of the process, by its own clock. The process reads performance.now() once, when its work is done, and Node counts that from the
moment the process began. So a time includes Node itself starting up, and leaves out what
happens outside the process: the parent creating it, the OS loading the Node binary, and the
exit. A clock in the parent would add those, with their own variance. The bare scenario is the same process with nothing imported, so a scenario's time less bare's is what the library added.
The process is a consumer's. It runs a generated module holding the import,
the one call of a first highlight, and a few lines that report, with no harness code and no
flags: no --expose-gc, which only the timed cells use. The source a first
highlight is given is a literal in that module, so no file is read inside the clock, and one
character of the HTML is read before the clock stops, as the timed cells do, so a string the
engine hasn't joined up is paid for.
The compile cache is off. V8 can keep compiled code on disk between
processes, and Node uses that only when asked. A second process would then skip compiling what
it loads, and would not be cold. Every cold-start process runs with NODE_DISABLE_COMPILE_CACHE=1 and without NODE_COMPILE_CACHE, each
sample reports that no cache directory is in use, and startup_compile_cache in BenchRun records it.
Interleaved rounds. A round runs every scenario once, one process at a time, with the parent blocked on each. The scenarios that differ only in the library run one after another, and the library that goes first rotates from round to round and from scenario to scenario. Because every round holds every scenario, a machine that slows down partway slows a sample of each of them, and not the scenarios that happened to come last.
The first rounds are discarded. They bring every file a scenario reads into the OS page cache, which a machine pays once and not at every start. A full run keeps 20 samples of each scenario, after 3 discarded rounds.
Each scenario is checked by its first process. The import has to work and the
module has to export something, and a first highlight's HTML has to pass the same check on the
harness snippet that a timed cell applies. A library that fails is recorded in startup_excluded with the reason, and has no number there: this is how a library
that can't be imported outside a bundler would show. A process that was ended, by a signal or
an abort, did not fail on its own: that fails the run, which can be resumed, and excludes
nothing. A scenario over languages a library doesn't claim is unsupported, and is never run.
Every later sample must agree with the first, and when the run carries the deterministic
results, a first highlight's HTML must be the size they have for that library on that snippet.
| number | meaning |
|---|---|
samples | how many processes were kept |
median_ms | the median across them, in milliseconds since the process started |
p10_ms, p90_ms | the 10th and 90th percentile |
spread | the slowest sample over the fastest. A process start has a long slow tail, so this is wider than a cell's, and the percentiles say more |
rss_peak_bytes | the median peak resident size of those processes, read when the clock stopped |
Memory. A cold-start process has loaded the library and done nothing else, so
its peak resident size is closer to a footprint than a timed cell's. It is still a whole
process, most of it Node: subtract the bare entry's. What is left is the code and
data the library loaded, and for a first highlight what one call allocated, at the sizes the
engine chose for a process of that age. The retained heap, described in deterministic metrics, is the steadier figure for what a library holds: read after
collecting garbage, less a bare process's, and the same from one process to the next within a
small tolerance.
Nothing is retried. A disturbance lands on one sample of every scenario, not
on one scenario, and the median over many processes sets a slow one aside. Sampling one
scenario again later would also take it out of the rounds that make the scenarios comparable.
There is no noise floor for cold start either: a calibration measures the timed cells only.
Without one, treat a difference between two scenarios that is inside either one's p10_ms to p90_ms band as no difference.
Part of the run. The cold-start processes follow the cells in the same run: under the same machine lock, between the same anchor readings, in the same journal, and written to the same results file. A smoke run takes one sample of every scenario.
Kinds of run #
A results file says what kind of run produced it, in BenchRun, with the passes, the rounds, the durations, and any filters:
full— the real measurement. Only a full run with no filters is publishable, and such a file must be whole: every library has a number or an exclusion in every cell whose language it claims, and in every cold-start scenario it supports, in both forms.smoke— every library on every cell, in one process each, for one round of the shortest batches, then one sample of every cold-start scenario, with several processes running at a time. It proves the harness and every adapter work, and its numbers mean nothing.calibrate— measures the machine's noise floor, described below. It measures the cells only, and no cold start.
npm run bench -- --machine <name> # a full run: results/<date>_<name>.json
npm run bench:calibrate -- --machine <name> # the noise floor: results/calibration_<name>.json
npm run bench:smoke # everything once, briefly and in parallel
node bench/run.ts --langs ts --sizes small # part of a run, to results/scratch/
node bench/run.ts --metrics startup # cold start alone, to results/scratch/
npm run bench -- --machine <name> --plan # what a run would do, measuring nothing
npm run bench -- --machine <name> --resume # continue a run that was stopped
node bench/run.ts --help | flag | what it does |
|---|---|
--smoke | a smoke run |
--calibrate | a calibration run |
--libraries, --langs, --sizes, --modes, --sources | measure only these, as comma-separated ids. An unknown id is an error, not an empty run. --libraries and --langs narrow the cold-start scenarios too;
the other three select inputs, which cold start has none of |
--metrics | measure only throughput (the timed cells) or only startup (cold start). Both by default. Naming both is no filter, and neither is throughput on a calibration, which measures nothing else |
--passes, --rounds, --target-ms | how many processes measure each library on a cell, how many batches each times, and how long a batch aims to run |
--retries | how many times a library whose rounds on a cell disagreed is measured again there. 0 turns it off |
--startup-samples | how many fresh processes sample each cold-start scenario, after the discarded rounds |
--resume | continue the run its journal records, which must be this same run |
--plan | print what the run would measure, where it would write, and an estimate of how long it would take, and measure nothing |
--machine | the name the machine is recorded under. Required for an unfiltered full or calibration run; otherwise the hostname |
--out | where to write the results file instead |
--help | print the flags |
A full run with no filters is written to results/, named by the UTC date it
started and its machine, and it won't overwrite a run already there. A smoke run or a filtered
one is written to results/scratch/, which git ignores, so it can't be committed
as if it were a result. A cell process that fails, or any process ended by a signal, fails the
run, and no results file is written; the journal keeps what was measured.
--plan makes every check a run makes before it measures, then prints the cells,
libraries, and processes, the cold-start scenarios with the processes they add and any a
library is unsupported for, the parameters, where the results and the journal would go, and an
estimate of the time, worked out from the parameters and not measured. It is refused where the
run would be.
Stopping and resuming #
A full run is one process for each library, pass, and cell, then one for each sample of each
cold-start scenario, one at a time, so it is long, and a calibration's cells take twice as
long. A timed run keeps a journal beside its results file: each process's result is appended
and synced to disk before the next process starts. The parent writes only between processes,
never while one is measuring. When a run stops, whether by a signal, a failed process, or a
crash, the journal stays, and --resume continues the run from it.
- What is recorded. A first line describing the run, then for each stretch of measuring (a segment) a line with its opening anchor reading and load average, one line for every process result, a cell's or a cold-start sample's, and a line with its closing anchor reading.
- What must match. The kind of run, its parameters and filters, the corpus
hash, every library's version and extra packages, the machine's name, CPU, threads, and
frequency settings (governor, driver, energy-performance preference, and boost), the OS
release, the Node version, the
NODE_OPTIONSthe processes inherit, and the commit with its dirty marker. A resume with any of them different is refused, naming what differs, so results from two setups are never mixed. With a dirty working tree the journal can't tell whether the uncommitted changes are the ones it started with, and a resume says so. - Nothing installed mid-run. Every cell process reports the versions of its
library's packages it found installed, and one that differs from what the run was planned
with stops the run, naming the package and both versions, before its result is journaled.
Cold-start processes report none, so the run reads the versions again before writing its
file, and on a change writes nothing and sets the journal aside as
<journal>.changed: its cold-start samples were never checked against the change, so the run starts over. - Never by accident. A journal already there stops a run that wasn't given
--resume: it is neither overwritten nor continued until you say which, and the refusal says what run the journal records. A journal holding a result the run would never ask for is refused too. - Never after an overlap. A run that lost the machine lock sets its journal
aside as
<journal>.overlapped, since another run may have overlapped it. It writes no results, and the journal can't be resumed. - What is skipped. The processes the journal holds are not run again. The rest run in the order planned, and the results are assembled exactly as an uninterrupted run assembles them.
- What the file says.
segments, in BenchRun, counts the stretches the run was measured in: 1 for an uninterrupted run.generated_at, and the date in a full run's file name, are when the first segment started.duration_msis the time spent measuring, without the time between segments, andload_at_startthe highest of the segments' load averages at their start. - A crash in the middle of a write. A last line with no newline is a write that didn't finish. It is dropped, and that process is measured again. A complete line that isn't a journal line is not a crash, and the journal is refused.
The journal is removed once the results file is written. A smoke run is short and keeps none.
Retries #
A library whose rounds on a cell are far apart was disturbed by something: another program, or
its own processes disagreeing. Once a cell's passes are done, a timed run measures each
library whose spread there is over 1.25 again, every pass of it, up to --retries more times (2 by default) and until an attempt is under the threshold.
The attempt kept is the one with the lowest spread, the earliest if they tie. It is never chosen by its time. Keeping the fastest attempt would pull every retried number toward fast, and would favor a noisier library, which is retried more often. Choosing on agreement alone takes the measurement that was least disturbed, whichever way it moved the number.
attempts records how many times a library was measured on a cell, so the numbers
that needed a retry can be told from the rest. A retried value's spread and percentiles are
those of the attempt that was kept, so they describe its calmest attempt, not everything that
was measured. The summary a run prints counts the retried cells of each library, and lists the
library-cells still over the threshold after the last attempt. What should be the same every
time (the token count, the output size, and whether the library passed the check) must agree
across attempts as it must across passes, or the run fails at that cell.
A smoke run doesn't retry, and cold-start scenarios are never retried.
The noise floor #
Before a machine's numbers are trusted, it is calibrated. A calibration measures every library as if it were two libraries: each gets twice the passes, split alternately into two arms, and each arm is exactly what a full run would measure. The two arms should be equal. How far apart they are is what the machine and the method produce when nothing differs, including what varies from one process to the next. BenchNoiseFloor holds three numbers:
per_cellis the 95th percentile of that difference, across cells and libraries. A difference between two libraries in one cell that is smaller than this is not a result.geomeantakes each library's two geometric means over all cells, and is their difference for the library where it was largest.per_processis the 95th percentile of the difference between two passes of one arm, over every pair of passes, arm, cell, and library: how far two single processes measuring the same thing come apart. It reads higher thanper_cell, which compares values that each pool several processes.
A difference is the larger time over the smaller, less one, so it reads the same whichever arm was slower.
A calibration is written to results/calibration_<machine>.json. A later
full run copies its noise floor if the calibration describes that run: unfiltered, with a
stable anchor, and on the same machine name, CPU, threads, and frequency settings, with the
same Node version, corpus, and parameters, retries included, and every library at the same
version with the same extra packages. The commit is not compared, so a commit that changes
nothing measured keeps the floor. Otherwise the run records no noise floor and says why, and
the site won't publish it.
A calibration retries too, each arm of a library on its own and by the same rule, so an arm
stays exactly what a full run would measure and the floor describes a full run with the same
retries. The rule looks only at an arm's own rounds, never at how far apart the two arms are,
so it can't steer the floor. A calibration file's values are the first arm's, and so are their attempts.
Every run with more than one pass also measures its own process noise, an A/A figure taken
from the run itself: each library on each cell was measured by several processes that should
agree, and every pair of a value's pass_medians_ns gives a difference. BenchProcessNoise records how many pairs there were and their
median, 95th percentile, and largest difference, as meta.process_noise. Pairs are
only ever taken within one library on one cell. Its 95th percentile is the run's counterpart
of the calibration's per_process, and the run's summary and the provenance line
on the results page print the two side by side. Both are the 95th percentile of one sample, so
either is higher about as often as not: a run's figure several times the calibration's is what
says its processes disagreed more than the calibration's did. The reader compares them, and
the site does not refuse a run over it.
The run this site shows was measured on laptop1, whose floor is 2.3% for one cell and 0.3% for a geometric mean. The results page marks a ratio inside it.
The anchor #
A run measures a fixed reference workload before its first process and after its last. The
workload involves no measured library, so a change between the two readings is the machine
changing state: heat, frequency scaling, another program. drift is the relative
change, in BenchAnchor, and a run is stable while it
stays under 3%.
The anchor only brackets the run. Something that disturbed the middle of it shows as a high spread, and the summary a run prints lists every library whose rounds on a cell
disagreed.
A run that was stopped and resumed takes a reading at the start and end of every segment, and its drift is that of the two readings furthest apart. A machine that changed state between segments, or within one, is then flagged and not hidden by readings that happen to agree at the two ends. A segment stopped by a signal or a failed process still takes its closing reading. One that was killed outright can't, and the next segment's opening reading stands in for it.
Before it spawns anything, a run reads the machine's load average over the last minute and
records it as load_at_start. On a quiet machine it is well under 1, and a timed
run warns when it is over 1, since something else is then running. It warns and goes on: the
anchor, the spread, and the retries are what judge the run.
One run at a time #
Two measurement runs at once make each other wrong, not just noisy. A run takes a machine-wide
lock, the file /tmp/fuz_benchmark.lock that names its process, and holds it for
as long as it measures, as does each segment of a resumed run. The path is fixed and ignores TMPDIR, and the lock is fuz_util's benchmark lock, so any other benchmark that
uses it takes the same one and never measures at the same time as this one. A second run is
refused and told who holds the lock. A lock whose process is gone is taken over.
Caveats #
- Features vary. The libraries have different features. Shiki highlights with the TextMate grammars and themes VS Code uses; Twinkleplop aims for a similar feature set, with its own themes, annotations, and Twoslash support; and both do more than fuz_code, which has the fewest features of the four: a single-pass lexer per language, with CSS classes for its theme. Their fidelity differs too: how finely each splits the same source, and how accurately it classifies each part. These numbers compare the work all of them share, not what one offers that another doesn't, and they time each library's output without scoring its quality.
- The languages are fuz_code's. The languages measured are fuz_code's built-in set, so the list favors it. Prism and Shiki each support hundreds of languages, and none of that shows here.
- Classes and inline styles. Shiki resolves colors while it tokenizes and
writes them inline, and the other libraries emit class names that a stylesheet colors. The
htmlmode measures the HTML each library returns by default, which is not the same work. - Token counts differ. A library that emits fewer tokens for the same source may be doing less work for each byte, not the same work faster.
- Synchronous paths. A library with an async API loads in its setup and is timed through its synchronous core.
- The HTML is read. In
htmlmode the timed call reads one character of the string, so a library that returns a string the engine hasn't joined up pays to flatten it inside the clock. A benchmark that discards the string doesn't charge for that, so these numbers can differ from a library's own. - Markdown without its fences. A timed process loads only the language it measures, so a fenced code block inside Markdown is not highlighted, with the exceptions adapters lists. A site that loads several languages pays more for Markdown than these cells show.
- Home turf. Some inputs are a measured library's own samples, which it was developed against. They are labeled with that library wherever they appear, summarized by owner, and left out of the default view and the headline figures, which use the neutral inputs.
- Stress inputs. Some inputs are built to find each library's worst case: minified code, deep nesting, very long lines, constructs that never close. They say little about typical code, so they are summarized on their own and left out of the default view and the headline figures. A slow call there can be far slower than anywhere else, and a run shows it as measured. Shiki is called with its per-line time limit out of reach, as everywhere in this benchmark, so a stress number is the cost of highlighting every line in full. With Shiki's default limit (500ms per line), a line that takes longer comes back with its rest as one plain token: a consumer with default options gets a fallback on that line, not this number.
- One bundler's sizes. A bundle size is what a page ships when Vite, with the Rolldown inside it, builds it and esbuild minifies it; the results file records the three versions. Another bundler can remove different code from the same entry, since what it can prove unused, whether it folds a constant argument into a function, and how it wraps CommonJS all vary. Built by Rollup, which Vite used before Rolldown, with the same options, the bundles differ unevenly: Rollup removes Twinkleplop's diagnostics code, which it proves is never called, and Rolldown keeps it; Rollup specializes a function fuz_code's Markdown language calls with a constant argument, and Rolldown doesn't; and Rolldown adds a small CommonJS interop helper to Prism's bundles.
- Synthetic inputs. A few inputs were built by joining or repeating files to reach a size, or derived from one by minifying it, and are marked. A repeated file is more regular than real code of the same size.
- Steady state. In a cell, every library is warm when it is timed. What a first call costs is in the cold-start scenarios.
- Cold start is Node's. A cold-start time is a Node process on one machine, with the files already in the page cache. A browser parses and compiles differently, and fetches over a network, where the bundle sizes matter instead.
- Unbundled depends on the install. The unbundled form loads the files a package manager laid out, so its time moves with the number of files a library ships and how deep they sit. It says what that library costs a Node process without a bundler, and nothing about a page.
- Collections. Every batch follows a rewarm (unless one call outlasts it) and a forced minor collection, and full collections happen only when the engine schedules them, so garbage from earlier batches carries into later ones and a full collection can still land inside a batch. A library pays for what it allocates (in the JS heap or outside it) as the engine collects it during the run, which is close to, but not the same as, how a page collects between highlights.
- Measured alone. A page that loads several highlighters, or a busy application, gives the engine a different history than a process holding one library.
Compared with Twinkleplop's benchmark #
Twinkleplop's own comparison harness is the one this method builds on, and the two answer different questions: it tracks Twinkleplop against other libraries across its own versions, and this one compares a fixed roster on shared inputs. Where they differ, as of Twinkleplop 0.3.1:
| Twinkleplop's | this benchmark | |
|---|---|---|
| libraries | Twinkleplop, Shiki (both engines), Prism, sugar-high, speed-highlight | fuz_code, Twinkleplop, Prism, Shiki (both engines) |
| languages | a broad set, most of what Twinkleplop supports | the timed languages every library here supports |
| inputs | generated inputs at three sizes per language, and Shiki's samples | pinned files no measured library wrote as the headline, with Shiki's samples, the libraries' own, snippets, and stress inputs kept apart |
| isolation | one process per cell, its libraries alternating within each round | one process per library and pass, in a rotating order |
| rounds | one pass of several rounds, a short rewarm, a full collection each round | several passes of a few rounds each, a longer rewarm, a minor collection only |
| statistics | median, minimum, p10 and p90, and a bootstrap interval | median, p10 and p90, spread, each pass's median, and the gap between processes |
| retries and noise | no retries; a calibrated noise floor for its A/B harness, not the comparison | a disturbed cell is measured again; a calibrated noise floor for the comparison itself |
| HTML timing | the returned string is discarded, so a library returning a rope never flattens it | one character of the string is read inside the clock, which flattens it |
| cold start | Twinkleplop alone, with scenarios about its own internals | every library, through the same bundle entries, in the results file |
| other metrics | none in the comparison | bundle size, install footprint, retained heap, token counts, output sizes |
| history | a history of published runs across versions | each published run stands alone |
| machine | a bare-metal server with boost off, the run confined to cores away from the system's | a laptop with boost, SMT, and ASLR left at their defaults |
Numbers from the two are not comparable with each other: the machines, the inputs, and the HTML timing differ.
Reproducing a run #
A results file records the commit, the Node version, the machine, the corpus hash, and every
library's version and extra packages. From a checkout of that commit, npm ci installs the same versions, and the command for the kind of run measures it again. The
deterministic numbers come out the same on any machine running the same Node version, the
retained heap within its tolerance, and a run stops if one of its processes disagrees with the
committed deterministic metrics. A full run with no filters is refused unless that
file describes it. Run the smoke run first: it checks every cell and every first highlight
against the file, before a long run can stop on a disagreement. The timed numbers depend on
the machine, so compare the ratios between libraries rather than the times.
A quiet machine matters more than a fast one: close other programs, and prefer a fixed CPU
frequency governor. Before trusting a run, check that its anchor is stable, that spread is near 1 for every library in every cell, and that nothing unexpected is
excluded, from a cell or from a cold-start scenario. bench/run.ts prints all
three when it finishes, with how many library-cells were retried and whether the run was
resumed.
deterministic metrics #
Some of what the benchmark reports is the same on any machine: how many bytes a library adds
to a page, what installing it adds to a project, how much memory it holds once loaded, how
many tokens it finds in an input, how large its HTML is, and which languages it covers. These
are measured by a command of their own, with nothing timed, and committed as results/deterministic.json.
npm run bench:deterministic # regenerate results/deterministic.json
npm run bench:deterministic -- --check # regenerate in memory, and fail if the committed file differs The file goes through the same schema as a timed run, described in results file, and regenerating it on an unchanged tree changes nothing.
Bundle size #
For each library and each feature set of languages and sets, the library's adapter
writes the module a consumer would write: the library with exactly the set's languages
registered, in the shape the library documents. bench/sizes.ts bundles that
module, minifies it, and reports its bytes in BenchBundleSizes:
raw— the minified JSgzip— those bytes gzipped at level 9, the highestbrotli— those bytes brotli-compressed at quality 11, the highest
How a bundle is built:
- Vite's build API in library mode bundles the entry into one ES module, as a production build for the browser, resolving from the versions installed here.
- Vite's library mode never removes the whitespace of an ES module, which suits a library and not a page. So the module is then minified once, whole, by esbuild, which keeps license comments: Prism's bundles carry the one in its core file.
- Compression is Node's zlib and brotli.
A set that holds a language the library doesn't claim is not measured for it. The entry names the languages it lacks instead, and is never a zero.
The shape each adapter writes, which its notes.md states in full:
| library | with languages | with none: core |
|---|---|---|
| fuz_code | a SyntaxStyler with the needed lexer_* modules added | the styler alone |
| Twinkleplop | the two factories of each language's package | the runtime and the grammar compiler of the core package, which every language package is built from |
| Prism | the core file and one component file for each grammar, in dependency order | the core file alone |
| Shiki, both engines | shiki/core, one engine, the github-light theme, and one module
for each language | the core, the engine, and the theme, since Shiki can't render without a theme |
A library's maintainers may know a leaner documented shape. The entry is one function of the
adapter, bundle_entry, and a change to it is a small pull request.
Wasm and theme CSS #
Two things stay out of the JS figures and have a section each, because folding them in would
hide a real difference between libraries. Both are reported as the JS is, as raw, gzip, and brotli, in BenchByteSizes:
wasm— the WebAssembly binary a library loads, and null for a library with none. Shiki's Oniguruma engine loads one. The documented import carries the binary inlined in a JS module as text, which a bundler would count as JS. The build leaves that one import out of the bundle, and the binary file itself is measured.css— each library's documented default stylesheet, minified by the same esbuild. Output that carries classes needs a stylesheet to show any color. Output that carries inline styles needs none, so for Shiki the section is null: its theme is in the JS figure, and its colors are in every page of HTML. A theme of several stylesheets is measured file by file, each compressed on its own, and summed.
Default themes don't cover the same ground. Each library's roster entry says which color
schemes its theme holds, in theme_schemes: the stylesheets of fuz_code and
Twinkleplop hold a light and a dark palette, and Prism's holds one. Read the stylesheet sizes
beside that.
Which stylesheets are a library's default theme is stated by its adapter, and a theme shipped as a package of its own has its version recorded with the library's.
Each bundle is checked #
A bundle that lost a language to tree shaking, or was written without one, would be reported as small. So every bundle is written to disk and imported in a process of its own, and each language of its set is highlighted through the bundle's exports, on that language's harness snippet.
- For a bundle of one language, the HTML must equal what the library produces through its adapter with that language set up. A language missing something it embeds, like HTML without its script and style languages, renders differently and fails.
- For a bundle of several languages, the HTML must have as many kinds of styled span as the snippet gate of adapters asks for. Equality doesn't hold there by design: a Markdown fence is highlighted when its language is in the bundle, and left plain in a process with one language loaded.
A bundle that doesn't import or exports nothing fails too. The check doesn't prove a bundle is the smallest that works.
What a language costs #
One set holds no languages, and for each language one set holds it alone, so the bytes a
language adds to a library are single:<lang> minus core, with
no extra measurement. The results page derives that view from the sets and never stores it, so
it can't drift from them. It is the default view of the bundle sizes there.
The sum of those differences over a set and the set's measured size differ when a library shares code between languages, or when one language brings another with it, as HTML brings the script and style languages. The results page shows the sum beside the measured size for that reason.
That difference is the price of the first language: it includes any runtime
the language needs that the floor leaves out, which a library pays once. So two more sets
measure the price of one more language: timed, every timed
language, and without:<lang>, the same less one. Their difference is what a
language adds to a bundle that already holds the others. It is read in minified bytes and
never compressed: compression works across a whole bundle, so the difference of two compressed
bundles doesn't isolate one language's share, and can come out negative. A language the others
already hold, like JS inside TypeScript, adds nothing there.
Token counts and output sizes #
For every input of the corpus and every library that claims its language, bench/outputs.ts records what one tokenize call and one html call return:
tokens— the library's own token countoutput— the HTML's size in UTF-8 bytes,html_bytes, and how many elements it has,spanscompressed— the same HTMLgzipped andbrotli-compressed at the levels bundle sizes use: what a server sends for a page that renders the input highlighted. The HTML is produced again for this, outside the check, and must come to the size the check saw.
They are taken as a timed cell takes them: a process sets up exactly one language in one library, the snippet gate runs, and each input is then checked. A library that fails is recorded as excluded with the reason, not counted.
Equal sizes don't mean equal output: two libraries, or Shiki's two engines, can return HTML of the same size that colors a token differently, and nothing here compares content.
Each number is stated once for an input, on the cell whose call returns it: the token count on
the input's tokenize cell and the output sizes, raw and compressed, on its html cell.
What the numbers don't say:
- Token counts are not comparable as work. Libraries split the same source differently. Shiki counts runs of whitespace and unstyled text, the others count only typed tokens, and a container and what it holds may each count. Fewer tokens is less output, not a faster library.
- Classes and inline styles are different HTML. A class name is short and
needs a stylesheet. An inline color is repeated on every token and needs none. Compare
html_byteswith the theme CSS column beside it. Compression narrows the gap, since a repeated color is what it removes best, which is why the compressed sizes are recorded too. - Wrappers differ. Twinkleplop and Shiki wrap the output in
<pre>and a span for each line, and fuz_code and Prism return bare token spans.spanscounts every element. - Prism's two calls can disagree on Markdown. Prism highlights the inside of
a fenced block in a hook that only
highlightruns, nevertokenize. For a fence whose language is loaded, the HTML holds elements the token count doesn't include.
Install footprint #
What npm install of a library adds to a project, in BenchInstall: bench/install.ts walks the runtime
dependency closure of the adapter's package and each of its separately versioned packages,
resolving each dependency from the package that needs it as Node resolves it, through this
repository's node_modules.
packages— the packages in the closure, each counted oncefiles— the files in those packages' directories, less any nestednode_modules, whose packages are counted through the closurebytes— the sum of those files' sizes as stored, not the disk blocks
What the closure follows:
dependencies, alwayspeerDependenciesnot marked optional, which npm installs as well: fuz_code's utility library is oneoptionalDependenciesthat are installed and not limited to some platforms byos,cpu, orlibc. A package with a binary for each platform installs one of them on each machine, and counting it would make the figure the machine's. No library in the roster has one today.- never
devDependencies
Each library is counted whole, on its own, as installing it alone would give: the two Shiki adapters install the same package and have the same footprint. npm installs what a package's tarball holds, so the footprint is everything the package ships, every language, theme, and type declaration included, and not what a page loads.
The versions are the ones this repository's lockfile installs, transitive ones included. A
fresh npm install of the same library may resolve a dependency to a newer version
within its range, and its footprint then differs. A package is told apart by name and version,
so one installed twice at the same version counts once. closure is a hash of the
sorted name@version list: a timed run carries the file only when each library's
closure is this tree's, and the site names a library whose dependencies moved since a run.
Retained heap #
What a page or a server holds once a library has loaded one language and highlighted once, in BenchHeap, for every timed language each library claims. bench/heap.ts runs a Node process that imports the library's bundle of that
language alone, the single: bundle whose size is reported, highlights the
language's harness snippet once through the adapter's bundle_html_source, lets go
of the HTML, and runs full collections until one changes the heap by under a kibibyte. It then
reads three numbers, each in KiB over a bare process of the same module without the import and
the call:
retained_kb— the bytes in use in V8's heap, less the two spaces that hold machine codecode_kb— the bytes in use in those two spaces. They hold the regular expressions the engine compiled to native code as well as the functions it optimized, so a library that lexes with regular expressions holds much of its memory here. Machine code is the processor's: the figures are for x64.external_kb—process.memoryUsage().external: memory outside the V8 heap that V8 is told about, like array buffers, and a WebAssembly engine's linear memory, counted at the size it has grown to, which it never gives back. The OS backs only the pages it has touched, so resident memory can be lower.
The process imports the bundle, not the adapter: some adapters import every language statically, and an adapter is TypeScript, which would load Node's type stripper into the reading. The generated module holds nothing of the harness, as a cold-start process doesn't.
It runs with three flags, for this reading only. The timed cells run with --expose-gc alone, for their minor collections, and cold start with none. --expose-gc gives it the collections. --max-semi-space-size=16 pins
the young generation, which V8 otherwise sizes from the machine's memory, and a small one
leaves some libraries holding a different amount. --single-threaded keeps the
engine from finishing compilation on another thread between the collections and the reading,
which otherwise moves Shiki's figures by kilobytes from one process to the next.
A theme in a stylesheet is held by the page, not the library, so a library that writes its styles inline holds its theme in these figures and one that writes classes doesn't, as with the bundle sizes. The figures are after one highlight of a small snippet; larger inputs leave more, most of all for regexp and wasm engines.
Coverage #
The coverage matrix is each adapter's language map: true for a language the
library claims, a footnote where the support is indirect, and no entry for a language it
doesn't claim, in BenchCoverageEntry. The footnotes are written
once, in the adapters, and adapters lists what they cover.
What the file records #
A deterministic file holds no timed value, so it records no machine, no time, no Node version, and no commit. The Node version does shape the compressed sizes and the heaps (below), but a file naming it would differ between machines that agree, and a commit can't name the tree it is itself committed in. What the numbers do depend on is in the file:
- each library's version, and the versions of its separately versioned packages
- each library's install closure, as a hash
- the corpus hash
- the versions of Vite, of the Rolldown inside it, and of esbuild, and the compression levels, in BenchBundler
Two dependencies are not recorded by version. The gzip and brotli sizes, of bundles and of HTML alike, come from the zlib and brotli inside Node, and a Node with another version of either can compress to a few bytes more or less. The file is generated on the Node version the checks run on. And the retained heaps follow V8's object layout: they are for the Node version the checks run on, on x64, without pointer compression, as official Node builds are. The install footprints follow the versions the lockfile installs, which a dependency update moves without moving a library's own version; the closure hash records them. A failed check says which of these it looks like.
A heap reading can also differ from the last by a little between two processes on the same
machine. Two readings agree to well under a kibibyte, and the check still compares them within
16.4 kB or 0.5% of the larger,
whichever is larger, and every other number exactly. Regenerating keeps a committed reading
that agrees, so an unchanged tree regenerates an unchanged file, and a change smaller than the
tolerance isn't recorded until it, with later ones, exceeds it. The site reads two numbers
within the tolerance as level, marked ≈.
Beside a timed run #
A timed run, described in method, takes its cells' token counts and output sizes from its own processes, by the same check under the same setup. When the deterministic file describes the run (the same library versions and extra packages, install closures, corpus, and feature sets), the run does two more things:
- it compares each library's first result on a cell with the file, and stops at a difference
- it copies the bundle sizes, the install footprints, the retained heaps, and the HTML's compressed sizes of its libraries from the file, and measures none of them (cold start rebuilds a missing cached bundle, checked against those sizes)
The compressed sizes are not compared on their own: the raw size each process reports is, and the file applies only at the same install closures, which pin every package that renders the HTML, transitive ones included. Compressing every process's HTML again would compress the same bytes in every process, for numbers the file already holds.
So a timed results file holds the deterministic numbers of its cells and the sizes of its libraries, equal to the deterministic file's. When the file doesn't describe the run, a smoke or filtered run says so and carries none of those sizes, and a full run with no filters is refused before it measures anything: regenerating the file is quick, and the run is long.
Kept current #
On every change the repository's check workflow runs gro check, bench:deterministic -- --check, and a smoke run of every library on every cell
and cold-start scenario. The deterministic check rebuilds every bundle and recounts every
input, and fails with the places that differ when the committed file is stale. Nothing in the
workflow is a measurement of speed.
results file #
A run of the benchmark is one JSON file. The harness writes it and the site reads it, and both go through the same schema, BenchResults in results_schema.ts, so a malformed run fails loudly instead of rendering blanks. The schema is strict: an unknown key is an error.
bench/run.ts writes one for a timed run, as method describes,
and bench/deterministic.ts writes the committed results/deterministic.json, which holds only the numbers that are the same on any
machine, as deterministic metrics describes.
Sections #
- BenchMeta — the machine, Node version, commit, and corpus hash;
what kind of run it was and how it measured, in BenchRun, with
the retries it allowed, how cold start was sampled, how many segments it was measured in,
and the load average it started at; the roster of libraries with their versions and the
color schemes their default themes cover; the excluded list for cells, and
startup_excludedfor cold-start scenarios, in BenchStartupExcluded; the anchor's drift; the machine's noise floor; the run's own process noise, in BenchProcessNoise, from every pair of a value's passes; what the bundle sizes were built with, and the compression levels every compressed size is taken at, in BenchBundler - BenchLang and BenchSet — the language ids and feature sets of the run, described in languages and sets
- BenchInput — the inputs the run measured, each with its source, size tier, hash, and the upstream commit it was copied from, as in the corpus manifest
- BenchCell — one input measured in one mode, with each library's
numbers. Its id is
<file>:<mode>. Atokenizecell states the token counts and anhtmlcell the output sizes, with the HTML's compressed sizes in BenchCellCompressed, so each is stated once for an input. Each value carriesattempts: how many times the library was measured there, more than 1 when its rounds disagreed and it was retried. A retried value's spread and percentiles are those of the attempt that was kept. Each value also carriespass_medians_ns, the kept attempt's median for each pass, so the processes' disagreement can be recomputed from the file - BenchStartup — one cold-start scenario: what a fresh process
loaded (
bare,core,lang,set, orfirst_highlight), for which library, with which language or feature set, and whether as one bundled chunk. Its numbers are the median time since the process started, the 10th and 90th percentile, the spread, how many samples were kept, and the processes' median peak resident size. The onebareentry, with no library, is the Node baseline every other entry includes - BenchBundleEntry — for each library over each feature set, its bundle sizes in BenchBundleSizes, or the languages of the set it lacks. The sizes are the JS's, with the wasm and the theme CSS each measured apart
- BenchInstall — for each library, what
npm installof it adds: the packages of its runtime dependency closure, their files, those files' bytes, and a hash of the closure'sname@versionlist - BenchHeap — for each library and each timed language it claims, what it holds once that language is loaded and highlighted once: the V8 heap less machine code, the machine code, and the external memory beside it, in KiB
- BenchCoverageEntry — which languages each library claims, with a footnote where the support is indirect
The file describes itself. The libraries, languages, sets, and inputs are lists inside it, and every other section refers to them by id, or by file for inputs, so the site needs no list of its own and adding a library changes no site code.
Two classes of number #
Machine-dependent numbers come from timing on one named machine and mean
nothing apart from it: throughput and memory in each cell's values, and the startup section.
Deterministic numbers come out the same anywhere: token counts and output
sizes in each cell's tokens, output, and compressed,
plus bundle, install, heap, and coverage.
The retained heaps are deterministic for a Node version and platform, within the tolerance deterministic metrics gives.
A file may hold only the deterministic ones. It then has no meta.run and no
anchor, and its generated_at, machine, node, and commit are null: they say where timed numbers came from, and are required as soon
as a file has any. A timed run carries the deterministic numbers too, so one published file
holds both.
Publishable or not #
meta.run tells a publishable run from the rest. Its kind is full, smoke, or calibrate, and its filters are null unless the run was restricted to some libraries, languages, sizes, modes, sources, or
metrics (the timed cells, or cold start). A smoke run's numbers mean nothing, and a
calibration exists for its noise floor.
bench_results_is_publishable is the test for a result: a full run with no filters and no gap in its cells or its cold start, a stable anchor, a noise floor carried from the machine's calibration, and a commit that records everything that ran.
No filters is what a run set out to cover. Whether the file then holds all of it is checked by bench_results_find_gaps, from the file alone:
- in every cell, each library of the roster that claims the cell's language has a value, or an entry in the excluded list for that cell
- every listed input of a timed language has a cell in each mode the file's cells use
A full run's cold start is checked the same way, by bench_results_find_startup_gaps. Each of these has an entry or an exclusion:
- the
barebaseline - for each library of the roster, unbundled and bundled:
core;langandfirst_highlightfor every timed language it claims; andsetfor every feature set the file's cold start names whose languages it all claims
A file with no filters and a gap doesn't parse, whether it is a full run or a calibration, so
a run that silently lost a library can't pass as complete. A calibration measures no cold
start, so only its cells are checked. A filtered or smoke run lists what it has and may have
gaps. Two things the file can't show are an input missing altogether and a whole mode missing:
its inputs and modes are the ones that were measured, and only the corpus hash names the
corpus they came from. A third is a cold-start set missing altogether: the sets a set scenario covers are read from the file too, so a file with every set entry removed still reads as complete.
A run that was stopped and resumed is a result like any other. meta.run.segments says how many stretches it was measured in, and its anchor covers every one of them, as method describes.
meta.commit ends in -dirty when the working tree had changes the
commit doesn't record, and such a run can't be reproduced from its commit.
What the site renders #
The site is built from one results file, parsed through the schema when the site is built, so a malformed file fails the build and no number reaches a page unchecked:
results/latest.json, the latest published timed run, when it is committed:npm run results:publishcopies a run there once it passes the test below. It carries its own bundle sizes, install footprints, retained heaps, token counts, and output sizes, so every number on a page is from one file.- otherwise
results/deterministic.json. The pages then show the deterministic numbers, and say that no timed results are published where a timed number would be. A deterministic file that holds a timed number fails the build.
A latest run must pass bench_results_is_publishable, or the build fails and says why. A smoke run, a calibration, a filtered run, a run with a gap, a run from a commit with uncommitted changes, a run whose anchor drifted, and a run with no noise floor are never shown as results. A drifted run is measured again, not published under a warning.
A published run is also held against the deterministic file. When the two describe the same library versions, the same install closures, the same corpus, and the same bundler and compression levels, their deterministic numbers must be equal: the feature sets, the coverage, the bundle sizes, the install footprints, and each cell's token counts and output sizes, raw and compressed. A difference fails the build and names the first place they differ. The retained heaps are compared within the check's tolerance instead, and never fail the build. When a library's version, its dependencies, the corpus, or what built and compressed the bundles (the versions of Vite, of the bundler inside it, and of esbuild, and the compression levels) has moved since the run, or the retained heaps differ by more, as they do when the Node version moves, the run is shown as the snapshot it is, and the line that says where the numbers came from also says what has moved since.
The fixture run the tests use holds invented numbers. A build renders it only when the BENCH_RESULTS_FIXTURE environment variable is 1, which is for
working on the pages before a run exists, and every page of such a build is marked as fixture
data. A build without the variable never reads the file.
Consistency checks #
Beyond the shape of each section, parsing checks that the sections agree:
- every library, language, set, and cell id is unique and refers to a listed entry, except the library a home input came from, which may be outside the run
- no two cells measure the same input in the same mode, and no cell covers an untimed language
- every cell names a listed input, describes it as that entry does, and is named after its file and mode
- an input written for the benchmark has no provenance, a verbatim copy has the hash its one pin records, and a derived stress input names its derivation, and one built from several files is marked synthetic
- a run with timed numbers records its anchor, when it started, and its machine, Node version, and commit, and how its numbers were measured
- token counts are on
tokenizecells only and output sizes onhtmlcells only, and a library with a timed value in a cell has that cell's deterministic number - a value's 10th percentile is at most its 90th, and an anchor is stable exactly when its drift is under the threshold
- a run's filters describe the file: every cell and cold-start entry is inside them, the libraries listed are exactly the ones filtered to, and no id repeats
- a smoke run carries no noise floor, allows no retries, and is one segment
- no value has more attempts than the run's retries allow
- a value has a pass median for each of the run's passes, and the process noise is there exactly when a cell has a value and the run measured more than one pass, counting every pair of passes of every value, with its median at most its 95th percentile and that at most its largest
- a run with no filters has no gap, as above
- each startup entry carries the fields its scenario needs and no others, for languages the library claims, and appears once, as a number or as an exclusion and never both
- a
langorfirst_highlightentry names a timed language. Asetentry names a feature set, which may hold an untimed one - every startup entry keeps the samples the run says, its median lies between its 10th and 90th percentile, and a calibration has none
- how cold start was sampled is recorded exactly when the run measured any, and is null for a calibration and for a run whose metrics filter leaves it out
- a library has no numbers for a language its coverage doesn't claim
- a library listed as excluded from a cell claims that cell's language and has no numbers there
- a bundle entry holds sizes only when the library claims every language of the set, and otherwise names exactly the languages it lacks
- bundle sizes cover every library over every set or are absent altogether, and what they were built with is recorded exactly when they are present
- the install footprints and the HTML's compressed sizes come with the bundle sizes: every
library of the roster has a footprint, and every library with an output on an
htmlcell has its compressed sizes, exactly when the bundle sizes are present. A footprint has at least a file for each package - the retained heaps come with the bundle sizes too, for every timed language each library claims and no other
- a compressed size is no larger than the raw one plus a compression format's own overhead, the HTML's included, and the theme CSS sizes are null exactly for a library that writes its styles inline
- a file with no timed numbers is complete: in every cell each library that claims the language has the number the cell states or is listed as excluded, and every listed input has a cell in each mode
Use parse_bench_results to validate a file and get every problem in one message.
design #
More views #
Each of these is a view over the same results file:
- a page for each library: its version, its adapter's notes, its sizes, and its cells
- a page for each language, with every library's rendering of the same sample side by side
- the change from one published run to the next, for every library
- the throughput cells run in the reader's own browser
More measurements #
- Markdown with the languages of its fenced blocks loaded, beside the cells that load Markdown alone
- other runtimes than Node
api #
Browse the full api docs.