syntax-highlighter-bench
logo for syntax-highlighter-bench

benchmark suite and results site for JS syntax highlighters

repo

overview #

syntax-highlighter-bench is a benchmark suite and results site for JS syntax highlighters. It compares fuz_code, Twinkleplop, Prism, and Shiki, with Shiki's JS engine and its Oniguruma wasm engine counted as two libraries.

A run produces one results file holding many metrics, and the site is a view over that file where a reader picks the libraries, languages, and metrics to compare: the results page.

AI disclosure: this is an LLM-generated repo guided by a person.

What exists
#

The harness that times the libraries, warm and from a cold start, the deterministic measurements, and the pages that render a results file are built.

  • an adapter for each library, and the check that guards against a plain-text fallback, in bench/libraries/ and bench/check_output.ts — see adapters
  • the inputs and their manifest, in bench/corpus/ — see corpus
  • the language ids and the feature sets, in bench/langs.ts and bench/sets.ts — see languages and sets
  • the timed harness, which measures each library alone in a process of its own and writes a results file, in bench/run.ts and bench/cell.ts — see method
  • cold start, part of the same run, where each sample is a fresh process that imports a library and stops its own clock, in bench/coldstart.ts — see method
  • the deterministic measurements, which are bundle size over each feature set, install footprint, retained heap, token counts, output sizes, and coverage, in bench/deterministic.ts, with their numbers committed as results/deterministic.json — see deterministic metrics
  • the schema of the results file, in results_schema.ts — see results file
  • the landing page and the results page, which render one validated results file, in src/routes/ and src/lib/ — see results file for which file, and design for what is still to come

What is measured
#

  • Throughput, on one named machine: how long a warm library takes to tokenize an input, and to render it as HTML, with the peak memory of its process beside it.
  • Cold start, on the same machine: how long a fresh Node process takes to import a library alone, with one language, and with a documentation site's languages, and to produce a first highlight. Each is measured for the bundled chunk and for the library imported unbundled, beside a bare Node baseline to subtract, with the process's peak memory.
  • Deterministic numbers, the same anywhere: bundle size over each feature set, install footprint, retained heap, token counts, output sizes, and coverage.

Layout
#

  • bench/ — the harness. It runs on Node directly and is never imported by the site.
  • results/ — where the harness writes runs. The deterministic results are committed there, and a published timed run goes beside them as results/latest.json.
  • src/lib/results_schema.ts — the schema the harness and the site share, with src/lib/bench_constants.ts, the id lists it builds its enums from. The rest of src/lib/ turns a results file into the tables of the site.
  • src/routes/ — this site, built as static pages.

Neutrality
#

The maintainer of this benchmark also maintains fuz_code, one of the measured libraries, so fairness has to come from the structure rather than from trust:

  • the maintainer's library is flagged in the results file and on the site
  • every input is labeled by its source, and real code that no measured library authored is the primary set; a snippet or a stress input written for this benchmark is labeled as that, never as neutral, and no stress input is tuned to a library
  • each library is driven through one adapter with notes on the entry points it uses, so its maintainers can review or replace it
  • the method and the commands to reproduce a run are published with the numbers

adapters #

Each library is measured through one adapter, and the adapter is the extension point. The contract is in bench/adapter.ts, the adapters are under bench/libraries/, and bench/libraries.ts lists them.

interface HighlighterAdapter<TTokens = unknown> { id: string; label: string; package: string; // its manifest is read for the version by_maintainer: boolean; note: string | null; extra_packages?: ReadonlyArray<string>; // separately versioned packages whose versions are recorded too langs: Partial<Record<Lang, string>>; // harness id to library id; absent means unsupported via: Partial<Record<Lang, string>>; // a coverage footnote for a claimed language output: 'classes' | 'inline_styles'; engine: 'scanner' | 'grammar_vm' | 'regex' | 'textmate'; theme_css: ReadonlyArray<string> | null; // the default theme's stylesheets; null when styles are inline theme_schemes: ReadonlyArray<'light' | 'dark'>; // the color schemes that theme covers wasm?: ReadonlyArray<{import: string; binary: string}>; // the wasm binaries it loads, counted apart from the JS setup(langs: ReadonlyArray<Lang>): Promise<void>; // once per process, before any timing tokenize(src: string, lang: Lang): TTokens; // timed: the library's own result count_tokens(tokens: TTokens): number; // never timed html(src: string, lang: Lang): string; // timed: the string consumers ship bundle_entry(langs: ReadonlyArray<Lang>): string; // module source importing exactly these languages bundle_html(exports: Record<string, any>, src: string, lang: Lang): string; // highlights through a built bundle's exports bundle_html_source(lang: Lang): string; // that same call as an expression over `exports` and `src` }

Rules
#

  • Same stage. tokenize and html are the narrowest public entries doing the same work as the other libraries': tokens out, and the HTML string a consumer ships.
  • Explicit languages. A library is measured only on the languages its adapter lists. A call for any other language, or for one that wasn't set up, throws. A library that returned escaped plain text for a language it didn't know would benchmark as very fast.
  • Token counts are recorded, not equalised. Libraries split the same source differently, so each count is the library's own, and counting happens outside the timed call.
  • Synchronous timed paths. A library with an async API does its loading in setup and is timed through its synchronous core.
  • Browser-safe. An adapter uses no Node-only API, so it can run in a page unchanged. The harness reads versions, stylesheets, and wasm binaries from disk on its behalf.

bundle_entry and bundle_html serve the bundle sizes of deterministic metrics: the first writes the module a bundle is built from, and the second highlights through that bundle's exports, which is how each bundle is shown to hold its languages. bundle_html_source is the second as text: a cold-start process, which method describes, must load nothing but the module it measures, so it can't load the adapter to make the call. The check of every bundle evaluates the text and requires the same HTML.

What each adapter calls
#

librarytokenizehtmlone token is
fuz_codeSyntaxStyler#lexSyntaxStyler#stylizea typed leaf or container
Twinkleplopthe language package's tokenize()the language package's language()a typed token after reclassification
PrismPrism.tokenizePrism.highlighta Token at any depth
Shiki, both enginescodeToTokensBasecodeToHtmla themed token, whitespace included

Each adapter has a notes.md beside it with the entry points, the import shape, how its tokens are counted, and its caveats, written so the library's maintainers can review or replace it.

One default is changed. Shiki stops tokenizing a line after half a second and returns the rest of it as one token, so what a call returns depends on how fast the machine is at that moment: a first call on a busy machine can come back with fewer tokens than the next. Both Shiki adapters raise that limit out of reach rather than turning it off, so every call does the same work, and still reads the clock as Shiki's default does.

fuz_code's tokenize result is valid only until its next call: for performance, it lexes into one buffer that every call reuses, where most other libraries return a result the caller keeps. A consumer that keeps one copies it, and the timed call doesn't. Its notes.md says how the harness reads the result in time.

What the HTML looks like differs, and the results say so:

  • fuz_code and Prism return bare <span> elements with classes, and the consumer supplies the wrapper
  • Twinkleplop returns a <pre> with a span for each line and classes on each token
  • Shiki returns a <pre> with a span for each line and an inline color on each token, so it needs no stylesheet and its tokenizing includes resolving colors

Coverage
#

Every adapter claims the timed languages. xml is claimed by all but Twinkleplop, which has no XML language. Where support is indirect, the coverage matrix carries a footnote:

  • fuz_code runs its TypeScript lexer for JS
  • Prism's Svelte grammar is the third-party prism-svelte, and its XML is its markup grammar under another name
  • Shiki's XML grammar loads the Java grammar it embeds

setup loads exactly the languages it is given, plus what a language embeds structurally, like the script and style languages of HTML and Svelte. A timed process sets up only the language it measures, so every library has the same languages loaded.

That decides how Markdown is measured: without highlighting inside fenced code blocks. fuz_code, Prism, and Shiki highlight a fenced block only when its language is loaded, and Twinkleplop never does, so loading one language puts them on the same footing. Two exceptions come from what loads with Markdown itself: Prism's Markdown depends on its markup grammar, so Prism still highlights a fence tagged html, xml, svg, or markup, and fuz_code and Prism highlight a fence tagged as Markdown.

The output check
#

Before a library is measured on a language, check_output runs it on the language's harness snippet, a few lines written dense in syntax. The library must produce tokens, and HTML with several kinds of styled span. A plain-text fallback has none, or only structural ones: a wrapper for each line, and in Shiki one span in the default color. A library that fails is recorded as excluded with the reason, which the results file keeps apart from unsupported.

Each input is then checked against a lower floor. It is lower because prose-heavy Markdown legitimately has few kinds of span, which is why the snippet decides and the floor only guards.

The tests run both for every adapter on every input of the corpus in every language it claims.

Adding a library
#

  1. Add the library as a devDependency.
  2. Write bench/libraries/<id>/adapter.ts, exporting an adapter, and notes.md beside it.
  3. Add the adapter to ADAPTERS in bench/libraries.ts.

The tests then cover it, and a results file carries the roster of libraries, so the site needs no change.

corpus #

The corpus is the set of files the libraries are run on. It lives in bench/corpus/, and its manifest records what each input is and where it came from. Every input is listed at the end of this page, with its upstream commit where it has one.

Sources
#

Every input is labeled by its source, in its path and in every result:

sourcewhat it holds
neutralreal files from projects no measured library authored, copied unmodified at a pinned commit: the primary set
shikithe sample files Shiki's own benchmark uses, at a pinned commit
homea measured library's own samples, labeled with the library: fuz_code's test fixtures, and the corpus Twinkleplop benchmarks itself on
harnessone short snippet for each language, written for this benchmark by its maintainer, who also maintains fuz_code
stressinputs that look for each library's worst cases: neutral inputs minified, and pathological files generated here by the maintainer, who also maintains fuz_code

The neutral set has a small, a medium, and a large file for each timed language. The other sources are kept so a reader can see whether a result holds on inputs a library chose for itself, or on hostile ones, and the results page filters by source. No summary pools two sources. A home input is always shown under the library whose samples it is, and libraries' home inputs are never summarized together: they contributed very different amounts.

Stress inputs
#

The stress inputs are generated by a script in the repository, bench/corpus_stress.ts, and none is tuned to a library. A library that fails on one, or is far slower there than anywhere else, is a finding: it is shown, as an exclusion with its reason or as the slow number, never smoothed over by changing the input. They come in two kinds:

  • derived — a neutral input as real code ships: the JS and the CSS minified by esbuild, and the JSON serialized with no whitespace, each one long line. It keeps its source's pins, is marked synthetic, and names how it was made, with the esbuild version, since another version may minify to other bytes
  • written — a pathological file for each timed language: nesting far deeper than code is written, a line of TypeScript tens of KB long, an unterminated string and block comment near the top of a file, template literals nested in their substitutions, Markdown delimiters that never close, HTML tags that never close and attributes without quotes, and a selector list on one long line. Each says in its first comment what it stresses

Shiki is timed with its per-line time limit out of reach, as on every input, so on a stress input it highlights every line in full. With its default limit a line that slow comes back with its rest as plain text, so a consumer with default options gets that fallback there rather than the stress number.

Size tiers
#

An input's tier comes from its size in bytes, never from its name:

tiersize
microunder 600 bytes
small600 bytes to under 3,200
medium3,200 bytes to under 32,000
large32,000 bytes and over

The upper three center on roughly 1KB, 10KB, and 100KB. The micro tier is the docs-site case: a snippet of a few lines, where the fixed cost of a call outweighs the scanning. The large tier has no upper bound, and the neutral large inputs range from about 90KB to several hundred KB, so compare throughput in MB/s across languages, not time per call.

Synthetic inputs
#

An input is marked synthetic when it is not one real file as published but was built by joining or repeating files to reach its size, or derived from one, as the minified stress inputs are. Real files are preferred. Most of Twinkleplop's own corpus is built this way and is marked.

One neutral input is synthetic: the large Svelte file. Svelte components of around 100KB are rare, so a script builds one from several real components of other projects. It splits each component into its scripts, markup, and style, and writes one of each block holding the bodies in order, unchanged. The result is shaped like one large component and is lexically well formed, but it is not a working component: unrelated scripts share one scope, so their names collide and Svelte's parser rejects it. No source is rewritten to avoid that. The manifest lists every component it is built from.

The manifest
#

bench/corpus/manifest.json holds one BenchInput for each input: its path, source, language, tier, size in bytes and lines, SHA-256, whether it is synthetic, and its provenance. A results file carries the same entries for the inputs it measured, described in results file.

Provenance is a list of BenchProvenance pins, each a repository, a commit, a path, a hash, and a license:

  • none, for a harness snippet or a stress input written here
  • one, for a verbatim copy, whose hash must equal the input's; the copy is still marked synthetic when its upstream built it by joining files
  • one, for a derived stress input, whose derivation names the transformation that made it from that file
  • several, for a synthetic input, naming what it was built from

The manifest also holds a hash over the whole corpus. A run records it, so two runs over different inputs are never compared as one.

Checking and regenerating
#

npm run corpus:check # the manifest matches the files, and every copy its pin npm run corpus:manifest # regenerate the manifest and the table of inputs npm run corpus:fetch # download neutral inputs by their pins, verifying each npm run corpus:harvest -- --repos <dir> --check # the synthetic Svelte input matches its sources npm run corpus:stress # regenerate the stress inputs and the manifest npm run corpus:stress -- --check # the stress inputs are what generates

Sizes, line counts, hashes, and tiers are derived from the files. The synthetic flag, the provenance, and the derivation are written by hand, or by the stress generator for its own inputs, and carried forward. The tests rebuild the manifest from disk and compare it to the committed one, so a file added, removed, or changed without regenerating fails. The tests also regenerate the stress inputs and compare them with the files.

The corpus is data: no formatter, linter, or type checker runs on it, since a reformatted file is no longer the upstream file.

Inputs
#

Every input of the results this site renders, by source. The corpus hash of those results is e649d70d435efac06e26c753d9424bbf30122fcae8e8037396b3cc340ea1c101. The repository holds the same list as a table generated from the manifest.

neutral — real code no measured library authored
inputlanguagetierbyteslineslicenseupstream, at its pinned commit
neutral/bash/headerscheck bashmedium10,792280PostgreSQL
postgres/postgres@cc053b6e12 src/tools/pginclude/headerscheck
neutral/bash/nvm.sh bashlarge179,6355,416MIT
neutral/bash/postgres-wrapper.sh bashsmall70725PostgreSQL
postgres/postgres@cc053b6e12 contrib/start-scripts/macos/postgres-wrapper.sh
neutral/css/bulma.css csslarge763,92321,564MIT
neutral/css/default.css cssmedium12,501535LicenseRef-W3C-Software-and-Document-2023
w3c/csswg-drafts@0da42c8afd shared/style/default.css
neutral/css/text.css csssmall87548MIT
sveltejs/svelte.dev@a90ebe2304 packages/site-kit/src/lib/styles/text.css
neutral/html/index.html htmlsmall1,52651MIT
prettier/prettier@1dcd0b05d0 website/index.html
neutral/html/Overview.src.html htmllarge107,7122,694LicenseRef-W3C-Software-and-Document-2023
w3c/csswg-drafts@0da42c8afd selectors-3/Overview.src.html
neutral/html/results.template.html htmlmedium18,139741MIT
sveltejs/svelte@7bc0a70fe6 benchmarking/compare/results.template.html
neutral/js/class.js jsmedium9,760371MIT
prettier/prettier@1dcd0b05d0 src/language-js/print/class.js
neutral/js/client.js jslarge91,4073,160MIT
sveltejs/kit@55ac0d5264 packages/kit/src/runtime/client/client.js
neutral/js/parse.js jssmall96341MIT
neutral/json/ast.json jsonlarge175,0375,502Apache-2.0
neutral/json/kit.tsconfig.json jsonsmall1,04829MIT
sveltejs/kit@55ac0d5264 packages/kit/tsconfig.json
neutral/json/typesMap.json jsonmedium17,285497Apache-2.0
microsoft/TypeScript@637d5746b7 src/server/typesMap.json
neutral/md/02-state.md mdmedium10,036352MIT
sveltejs/svelte@7bc0a70fe6 documentation/docs/02-runes/02-$state.md
neutral/md/Explainer.md mdlarge169,1153,504Apache-2.0
neutral/md/postgres.README.md mdsmall98921PostgreSQL
neutral/rust/dent.rs rustmedium11,211352Unlicense OR MIT
neutral/rust/raw_query.rs rustsmall85237MIT
tokio-rs/axum@c59208c86f axum/src/extract/raw_query.rs
neutral/rust/table.rs rustlarge100,9603,197MIT OR Apache-2.0
neutral/svelte/Breadcrumb.svelte sveltesmall1,14144MIT
techniq/svelte-ux@11c79b8c6e packages/svelte-ux/src/lib/components/Breadcrumb.svelte
neutral/svelte/harvest.svelte syntheticsveltelarge106,7233,186MIT
techniq/layerchart@ec05fb1b84 packages/layerchart/src/lib/components/Legend.svelte
techniq/layerchart@ec05fb1b84 packages/layerchart/src/lib/components/layers/Canvas.svelte
techniq/svelte-ux@11c79b8c6e packages/svelte-ux/src/lib/components/Button.svelte
techniq/svelte-ux@11c79b8c6e packages/svelte-ux/src/lib/components/TextField.svelte
themesberg/flowbite-svelte@85f20a048e src/lib/datepicker/Datepicker.svelte
themesberg/flowbite-svelte@85f20a048e src/lib/clipboard-manager/ClipboardManager.svelte
neutral/svelte/Layer.svelte sveltemedium9,455343MIT
neutral/ts/core.ts tslarge92,4192,595Apache-2.0
microsoft/TypeScript@637d5746b7 src/compiler/core.ts
neutral/ts/options.ts tsmedium10,080251MIT
sveltejs/language-tools@bf2993e192 packages/svelte-check/src/options.ts
neutral/ts/transform.ts tssmall1,16027Apache-2.0
microsoft/TypeScript@637d5746b7 src/services/transform.ts
shiki — Shiki's own benchmark samples
inputlanguagetierbyteslineslicenseupstream, at its pinned commit
shiki/bash/shellscript.sample bashsmall66128MIT
shiki/css/css.sample cssmicro57846MIT
shiki/html/html.sample htmlsmall1,59652MIT
shiki/js/javascript.sample jsmedium3,946151MIT
shiki/json/json.sample jsonsmall88038MIT
shiki/md/markdown.sample mdmedium3,725170MIT
shiki/rust/rust.sample rustsmall1,01139MIT
shiki/svelte/svelte.sample sveltesmall71428MIT
shiki/ts/typescript.sample tssmall2,20577MIT
home: fuz_code — home turf: fuz_code's own samples
inputlanguagetierbyteslineslicenseupstream, at its pinned commit
home/fuz_code/bash/sample_complex.sh bashsmall2,749190MIT
fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.sh
home/fuz_code/css/sample_complex.css cssmicro50245MIT
fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.css
home/fuz_code/html/sample_complex.html htmlsmall73342MIT
fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.html
home/fuz_code/json/sample_complex.json jsonmicro32714MIT
fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.json
home/fuz_code/md/sample_complex.md mdsmall2,799219MIT
fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.md
home/fuz_code/rust/sample_complex.rs rustsmall3,012118MIT
fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.rs
home/fuz_code/svelte/sample_complex.svelte sveltesmall2,517156MIT
fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.svelte
home/fuz_code/ts/sample_complex.ts tssmall2,883152MIT
fuzdev/fuz_code@2e2fafcf97 src/test/fixtures/samples/sample_complex.ts
home: twinkleplop — home turf: twinkleplop's own samples
inputlanguagetierbyteslineslicenseupstream, at its pinned commit
home/twinkleplop/bash/bash.large.sh syntheticbashlarge98,8565,045MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/bash.large.sh
home/twinkleplop/bash/bash.medium.sh syntheticbashmedium10,941557MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/bash.medium.sh
home/twinkleplop/bash/bash.small.sh syntheticbashsmall1,05446MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/bash.small.sh
home/twinkleplop/bash/deploy.sh bashmedium5,567192MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/deploy.sh
home/twinkleplop/css/css.large.css syntheticcsslarge97,6224,814MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/css.large.css
home/twinkleplop/css/css.medium.css syntheticcssmedium10,562670MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/css.medium.css
home/twinkleplop/css/css.small.css syntheticcsssmall1,24676MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/css.small.css
home/twinkleplop/css/site.css syntheticcsslarge40,8641,901MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/site.css
home/twinkleplop/html/app.html htmlsmall65618MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/app.html
home/twinkleplop/html/html.large.html synthetichtmllarge99,8654,426MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/html.large.html
home/twinkleplop/html/html.medium.html synthetichtmlmedium10,262455MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/html.medium.html
home/twinkleplop/html/html.small.html synthetichtmlsmall1,32752MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/html.small.html
home/twinkleplop/js/bench-suite.js syntheticjsmedium26,3891,062MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/bench-suite.js
home/twinkleplop/js/javascript.large.js syntheticjslarge102,8364,843MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/javascript.large.js
home/twinkleplop/js/javascript.medium.js syntheticjsmedium10,597297MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/javascript.medium.js
home/twinkleplop/js/javascript.small.js jssmall81321MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/javascript.small.js
home/twinkleplop/json/compiled-grammar.json jsonlarge124,02214,193MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/compiled-grammar.json
home/twinkleplop/json/json.large.json syntheticjsonlarge99,8314,192MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/json.large.json
home/twinkleplop/json/json.medium.json syntheticjsonmedium10,131420MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/json.medium.json
home/twinkleplop/json/json.small.json syntheticjsonsmall97441MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/json.small.json
home/twinkleplop/md/architecture.md mdmedium25,141341MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/architecture.md
home/twinkleplop/md/grammar-docs.md syntheticmdmedium20,842549MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/grammar-docs.md
home/twinkleplop/md/markdown.large.md syntheticmdlarge101,2262,572MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/markdown.large.md
home/twinkleplop/md/markdown.medium.md syntheticmdmedium9,733809MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/markdown.medium.md
home/twinkleplop/md/markdown.small.md syntheticmdsmall1,18672MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/markdown.small.md
home/twinkleplop/rust/pipeline.rs rustmedium8,806326MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/pipeline.rs
home/twinkleplop/rust/rust.large.rs syntheticrustlarge101,7094,830MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/rust.large.rs
home/twinkleplop/rust/rust.medium.rs syntheticrustmedium12,710602MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/rust.medium.rs
home/twinkleplop/rust/rust.small.rs syntheticrustsmall1,26064MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/rust.small.rs
home/twinkleplop/svelte/site-codepanel.svelte sveltemedium4,712181MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/site-codepanel.svelte
home/twinkleplop/svelte/site-inspectorpanel.svelte sveltemedium3,947190MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/site-inspectorpanel.svelte
home/twinkleplop/svelte/svelte.large.svelte syntheticsveltelarge99,6694,641MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/svelte.large.svelte
home/twinkleplop/svelte/svelte.medium.svelte syntheticsveltemedium10,090436MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/svelte.medium.svelte
home/twinkleplop/svelte/svelte.small.svelte sveltesmall71431MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/svelte.small.svelte
home/twinkleplop/ts/core-frames.ts synthetictslarge57,2881,546MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/core-frames.ts
home/twinkleplop/ts/core-runtime.ts synthetictslarge70,2941,903MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/core-runtime.ts
home/twinkleplop/ts/core-types.ts tslarge58,0471,405MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/real/core-types.ts
home/twinkleplop/ts/typescript.large.ts synthetictslarge122,2063,321MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/typescript.large.ts
home/twinkleplop/ts/typescript.medium.ts synthetictsmedium7,628393MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/typescript.medium.ts
home/twinkleplop/ts/typescript.small.ts tssmall79927MIT
pngwn/twinkleplop@f5418aaabf lib/bench/perf/corpus/sized/typescript.small.ts
harness — snippets written for this benchmark by its maintainer, who also maintains fuz_code
inputlanguagetierbyteslineslicenseupstream, at its pinned commit
harness/bash/snippet.sh bashmicro43521 written for this benchmark
harness/css/snippet.css cssmicro41224 written for this benchmark
harness/html/snippet.html htmlmicro49621 written for this benchmark
harness/js/snippet.js jsmicro46514 written for this benchmark
harness/json/snippet.json jsonmicro31818 written for this benchmark
harness/md/snippet.md mdmicro38321 written for this benchmark
harness/rust/snippet.rs rustmicro57224 written for this benchmark
harness/svelte/snippet.svelte sveltemicro54029 written for this benchmark
harness/ts/snippet.ts tsmicro50919 written for this benchmark
stress — pathological and minified inputs, generated for this benchmark by its maintainer, who also maintains fuz_code, to find worst cases; Shiki is timed with its per-line limit out of reach, so it highlights lines its default would cut off
inputlanguagetierbyteslineslicenseupstream, at its pinned commit
stress/bash/nested.sh bashsmall2,23358 written for this benchmark
stress/css/default.min.css syntheticcssmedium8,3591LicenseRef-W3C-Software-and-Document-2023
minified by esbuild 0.28.1, from
w3c/csswg-drafts@0da42c8afd shared/style/default.css
stress/css/selectors.css cssmedium29,316135 written for this benchmark
stress/html/unclosed.html htmlmedium20,35880 written for this benchmark
stress/js/client.min.js syntheticjsmedium28,41311MIT
minified by esbuild 0.28.1, from
sveltejs/kit@55ac0d5264 packages/kit/src/runtime/client/client.js
stress/js/nested_templates.js jsmedium14,82886 written for this benchmark
stress/js/unterminated.js jsmedium17,013409 written for this benchmark
stress/json/ast.min.json syntheticjsonlarge68,3851Apache-2.0
minified: parsed and serialized again with no whitespace, from
stress/json/nested.json jsonmedium18,559457 written for this benchmark
stress/md/nested.md mdmedium7,494114 written for this benchmark
stress/rust/nested.rs rustmedium5,441139 written for this benchmark
stress/svelte/nested_blocks.svelte sveltemedium4,732108 written for this benchmark
stress/ts/long_line.ts tsmedium28,85714 written for this benchmark
stress/ts/nested_blocks.ts tsmedium11,864197 written for this benchmark

Licenses
#

Every copied file stays under its upstream license, which its provenance records. The corpus readme and the notice beside the neutral inputs give the attribution.

languages and sets #

The harness owns the language ids. Each library names its languages differently, so an adapter maps these ids onto its library's own. bench/langs.ts is the source of truth, and a results file carries the list, so the tables on this page are read from the results this site renders.

The languages measured are fuz_code's built-in set, so the list favors it. Prism and Shiki each support hundreds of languages, and none of that shows here.

Language ids
#

The language ids of the results this site renders.
idlanguagetimed
tsTypeScriptyes
jsJSyes
cssCSSyes
htmlHTMLyes
xmlXMLno
jsonJSONyes
svelteSvelteyes
mdMarkdownyes
bashBashyes
rustRustyes

The timed languages are the ones every library in the starting roster supports, so a timed comparison is never missing a library by design.

A language that is an id but is not timed is one the roster doesn't share: XML is the case today, since Twinkleplop has no XML language. It appears in the coverage matrix and in feature sets. A cold-start scenario that loads one language loads a timed one, so it never names an untimed id.

The measured libraries support many more languages than these. An id is added when a feature set or a second library needs it.

Unsupported and excluded
#

A library can be missing from a comparison for two different reasons:

  • unsupported — the library's adapter doesn't claim the language. Adapters declare their languages explicitly and never fall back to plain text, which would otherwise benchmark as very fast.
  • excluded — the adapter claims the language, but its output failed the check that runs before measuring, described in adapters.

Neither is recorded as a zero or counted in an average, and the results file keeps the two apart.

Feature sets
#

A feature set is a named list of language ids that bundle size is measured over, and that a cold-start scenario imports. The sets are declared in bench/sets.ts before any numbers exist, so a set can't be chosen to flatter a result.

The feature sets of the results this site renders.
idlabellanguagesnote
corecorenone
docsdocsjs, ts, css, html, md, bash, json
web_rustweb + rustts, js, css, html, xml, json, svelte, md, bash, rustfuz_code's full set (home turf)
allallts, js, css, html, xml, json, svelte, md, bash, rust
timedevery timed languagets, js, css, html, json, svelte, md, bash, rust
single:<lang>the language's nameone, for each language id
without:<lang>timed less the languageevery timed language but one, for each timed language

What each is for:

  • core — the library with no languages registered: its runtime floor.
  • single:<lang> — subtracting core gives what that language adds, which the results page shows for every language.
  • docs — what a documentation site imports, matching Twinkleplop's docs bundle. It is also the set a cold-start process loads in the set scenario.
  • web_rust — everything fuz_code supports, so it is home turf for that library, and its note says so. It is named by its content because the site is neutral.
  • all — every language id. While it holds the same languages as another set, the two are one bundle, and the results page shows them as one row.
  • timed — every timed language, which every library claims, and without:<lang> — the same less one of them. Subtracting the second from the first gives what one more language adds, without the runtime the first language pays for. deterministic metrics says how to read it.

A set is measured for a library only when the library supports every language in it. Otherwise the result names the languages it lacks. deterministic metrics describes how a set is built and measured.

method #

The timed harness measures throughput: how long each library takes to tokenize an input, and to render it as HTML. It then measures cold start: what a library costs a process before it has highlighted anything. bench/run.ts runs both and writes one results file.

The method builds on the comparison harness of Twinkleplop (MIT), and the anchor workload is ported from it. It differs in giving each library a process of its own, where Twinkleplop measures a cell's libraries together in one. The cold-start scenarios generalize Twinkleplop's cold-start harness across libraries.

Cells
#

A cell is one input of the corpus in one mode, tokenize or html, measured across the libraries that claim its language. Its id is <file>:<mode>. A cell carries its input's source and size tier, so a reader can compare on the neutral inputs alone, or see whether a result holds on a library's home turf or on the stress inputs.

One process, one library
#

A process measures one library on one cell and loads no other. Every number is therefore one library alone in a fresh process, and the roster can't move it: adding or removing a library changes no other library's result.

Two effects made that necessary, and both showed up as a library disagreeing with itself:

  • Libraries that share code change how V8 optimizes that code for each other. Shiki's two engines share its tokenizer, and beside the wasm engine Shiki's JS engine measured far slower than it does alone.
  • V8 sizes its young generation from recent allocation. In a shared process a library's time on a large input depended on which libraries ran beside it.

A cell is measured in passes. A pass is a process for every library, one after another, and the library a pass starts with changes from pass to pass and from cell to cell. The same code can settle at a slightly different speed from one process to the next, so a library's number pools the rounds of all its passes.

Node runs with its default flags, which is what a consumer gets. In a timed run one process runs at a time, and the parent does nothing while it measures: it blocks on the child and reads its output after it exits.

Inside a process
#

Each step answers a way a simple timing loop goes wrong:

  • One language loaded. The library sets up exactly the cell's language, once, before anything is timed. See adapters for what that means for Markdown.
  • The output is checked first. The library must pass the output check on the language's harness snippet and then on the cell's input. A library that fails is recorded as excluded with the reason, and is not timed.
  • Iterations calibrated for the library. Speeds differ by orders of magnitude, so a batch makes the number of calls that fills the target duration. A batch makes at least a few calls while they fit a cap, as many as fit when fewer do, and one when two no longer fit.
  • A rewarm and a minor collection before every timed batch. The library runs untimed until it is back to speed, and a minor collection then clears what the rewarm allocated, so the collection it would cause doesn't land inside the batch. No full collection is forced between batches. One forced before every batch puts some libraries through a race each round: it can discard optimized code whose hidden classes only short-lived objects hold, so the library re-optimizes inside the clock, and it re-rolls how memory a library allocates outside the JS heap is reused, so a process settles at one of two speeds. A minor collection does neither.
  • Results are kept and counted. A batch keeps its last result, and after the clock stops that result is counted. A count that differs from the checked one fails the cell: a tokenizer that silently stops early on a slow line must not be averaged in as a fast one.
  • The HTML is read. Some libraries return a string the engine hasn't joined up yet, which costs little to build, and every consumer then pays to flatten it on first use. In html mode the timed call reads one character of the string, for every library, so that cost is inside the clock. It depends on the size of the output, not on the library, so it weighs most on the fastest one, and it is free for a library whose string is already flat. In tokenize mode every library returns finished arrays, so there is nothing to force.

The forced minor collection needs the process started with --expose-gc, which bench/run.ts passes to each process.

A full run measures each library in 3 processes per cell, each timing 4 rounds of 20ms batches, after a 100ms warmup, with a 300ms rewarm before each batch. A library whose one call takes at least as long as the rewarm gets none: the batch before keeps it warm, and a rewarm would only repeat that call.

The rewarm is longer than the warmup on purpose. The engine still runs full collections of its own during a run, and one can leave code cold. Libraries recover from that at different speeds, Shiki more slowly than the others, and a library still recovering when its batch is timed reads slow, so the rewarm is long enough that every library in the roster has recovered.

What each number means
#

Each library's numbers in a cell are a BenchCellValue:

numbermeaning
ns_per_opnanoseconds per call: the median across the rounds of every pass
ops_per_sec, mb_per_secthe same time as calls per second, and as megabytes of input per second, which is the one to compare across inputs of different sizes
p10_ns, p90_nsthe 10th and 90th percentile across those rounds
spreadthe slowest round over the fastest, across passes as well as within them. Near 1 means they agreed, and a high spread says the cell was disturbed or the processes disagreed
pass_medians_nseach pass's own median over its rounds, in pass order: one number per process, so how far the processes disagreed can be read apart from the rounds
iterationscalls per batch. Each process calibrates for itself, and this is the median across them
rss_peak_bytesthe peak resident size of the library's process, the largest across its passes
attemptshow many times the library was measured on the cell: 1 unless it was retried, as described below

The token count and the size of the HTML come from the output check, under the same setup and outside the clock. Libraries split the same source differently, so a token count is the library's own and is recorded rather than equalised. Each is stated once for an input: the token count on its tokenize cell and the output size on its html cell. deterministic metrics describes them, and the bundle sizes a run carries.

rss_peak_bytes is a whole process: Node itself, the harness, the library as its adapter imports it, with one language set up, and the input. It is not what the library adds to a page, and most of it is the same for every library. Compare it between libraries on the same cell. On a small input it says little about footprint: a process warms up and is timed for a fixed duration, so its peak follows how much the library allocated in that time more than what it holds. What a library holds is the retained heap, described in deterministic metrics.

Cold start
#

A timed cell measures a library that is already loaded and warm. Cold start measures the other end: a fresh Node process that loads the library and does one thing, once. Each sample is one process, and each BenchStartup entry is a scenario:

scenariowhat the process does
barenothing: the Node baseline, with no library, which every other time includes
coreimports the library with no language
langimports the library with one language, for each timed language it claims
setimports the library with the languages of the docs feature set: what a documentation site loads
first_highlightimports the library with one language and highlights that language's harness snippet once, to HTML: the time to a first highlight, for each timed language it claims

What is imported is the module the adapter's bundle_entry writes for a feature set (core, single:<lang>, docs): the same module bundle size is measured over, so a time and a size describe one thing. The adapter modules themselves are never loaded here. Some of them import every language, which would charge a library for languages the scenario doesn't load. A first highlight makes the call the adapter documents for that module (bundle_html_source), and the check of every built bundle confirms it returns the same HTML as bundle_html.

Two forms of every scenario, apart from bare:

  • bundled — the built chunk whose size deterministic metrics reports, imported as one file. A library's wasm import stays outside its bundle, as it does for the size, so that one module still loads from the package.
  • unbundled — the same module's source, written to a file under node_modules/.cache/ in the checkout and imported from there. Node resolves each package from the checkout's node_modules and loads every file the library reaches, as a consumer without a bundler does.

The gap between the two is Node finding, reading, and compiling many files in place of one, and code the bundler removed, which a bundled page or server doesn't pay. Read the form that matches how the library would be shipped.

Start is the start of the process, by its own clock. The process reads performance.now() once, when its work is done, and Node counts that from the moment the process began. So a time includes Node itself starting up, and leaves out what happens outside the process: the parent creating it, the OS loading the Node binary, and the exit. A clock in the parent would add those, with their own variance. The bare scenario is the same process with nothing imported, so a scenario's time less bare's is what the library added.

The process is a consumer's. It runs a generated module holding the import, the one call of a first highlight, and a few lines that report, with no harness code and no flags: no --expose-gc, which only the timed cells use. The source a first highlight is given is a literal in that module, so no file is read inside the clock, and one character of the HTML is read before the clock stops, as the timed cells do, so a string the engine hasn't joined up is paid for.

The compile cache is off. V8 can keep compiled code on disk between processes, and Node uses that only when asked. A second process would then skip compiling what it loads, and would not be cold. Every cold-start process runs with NODE_DISABLE_COMPILE_CACHE=1 and without NODE_COMPILE_CACHE, each sample reports that no cache directory is in use, and startup_compile_cache in BenchRun records it.

Interleaved rounds. A round runs every scenario once, one process at a time, with the parent blocked on each. The scenarios that differ only in the library run one after another, and the library that goes first rotates from round to round and from scenario to scenario. Because every round holds every scenario, a machine that slows down partway slows a sample of each of them, and not the scenarios that happened to come last.

The first rounds are discarded. They bring every file a scenario reads into the OS page cache, which a machine pays once and not at every start. A full run keeps 20 samples of each scenario, after 3 discarded rounds.

Each scenario is checked by its first process. The import has to work and the module has to export something, and a first highlight's HTML has to pass the same check on the harness snippet that a timed cell applies. A library that fails is recorded in startup_excluded with the reason, and has no number there: this is how a library that can't be imported outside a bundler would show. A process that was ended, by a signal or an abort, did not fail on its own: that fails the run, which can be resumed, and excludes nothing. A scenario over languages a library doesn't claim is unsupported, and is never run. Every later sample must agree with the first, and when the run carries the deterministic results, a first highlight's HTML must be the size they have for that library on that snippet.

numbermeaning
sampleshow many processes were kept
median_msthe median across them, in milliseconds since the process started
p10_ms, p90_msthe 10th and 90th percentile
spreadthe slowest sample over the fastest. A process start has a long slow tail, so this is wider than a cell's, and the percentiles say more
rss_peak_bytesthe median peak resident size of those processes, read when the clock stopped

Memory. A cold-start process has loaded the library and done nothing else, so its peak resident size is closer to a footprint than a timed cell's. It is still a whole process, most of it Node: subtract the bare entry's. What is left is the code and data the library loaded, and for a first highlight what one call allocated, at the sizes the engine chose for a process of that age. The retained heap, described in deterministic metrics, is the steadier figure for what a library holds: read after collecting garbage, less a bare process's, and the same from one process to the next within a small tolerance.

Nothing is retried. A disturbance lands on one sample of every scenario, not on one scenario, and the median over many processes sets a slow one aside. Sampling one scenario again later would also take it out of the rounds that make the scenarios comparable. There is no noise floor for cold start either: a calibration measures the timed cells only. Without one, treat a difference between two scenarios that is inside either one's p10_ms to p90_ms band as no difference.

Part of the run. The cold-start processes follow the cells in the same run: under the same machine lock, between the same anchor readings, in the same journal, and written to the same results file. A smoke run takes one sample of every scenario.

Kinds of run
#

A results file says what kind of run produced it, in BenchRun, with the passes, the rounds, the durations, and any filters:

  • full — the real measurement. Only a full run with no filters is publishable, and such a file must be whole: every library has a number or an exclusion in every cell whose language it claims, and in every cold-start scenario it supports, in both forms.
  • smoke — every library on every cell, in one process each, for one round of the shortest batches, then one sample of every cold-start scenario, with several processes running at a time. It proves the harness and every adapter work, and its numbers mean nothing.
  • calibrate — measures the machine's noise floor, described below. It measures the cells only, and no cold start.
npm run bench -- --machine <name> # a full run: results/<date>_<name>.json npm run bench:calibrate -- --machine <name> # the noise floor: results/calibration_<name>.json npm run bench:smoke # everything once, briefly and in parallel node bench/run.ts --langs ts --sizes small # part of a run, to results/scratch/ node bench/run.ts --metrics startup # cold start alone, to results/scratch/ npm run bench -- --machine <name> --plan # what a run would do, measuring nothing npm run bench -- --machine <name> --resume # continue a run that was stopped node bench/run.ts --help
flagwhat it does
--smokea smoke run
--calibratea calibration run
--libraries, --langs, --sizes, --modes, --sourcesmeasure only these, as comma-separated ids. An unknown id is an error, not an empty run. --libraries and --langs narrow the cold-start scenarios too; the other three select inputs, which cold start has none of
--metricsmeasure only throughput (the timed cells) or only startup (cold start). Both by default. Naming both is no filter, and neither is throughput on a calibration, which measures nothing else
--passes, --rounds, --target-mshow many processes measure each library on a cell, how many batches each times, and how long a batch aims to run
--retrieshow many times a library whose rounds on a cell disagreed is measured again there. 0 turns it off
--startup-sampleshow many fresh processes sample each cold-start scenario, after the discarded rounds
--resumecontinue the run its journal records, which must be this same run
--planprint what the run would measure, where it would write, and an estimate of how long it would take, and measure nothing
--machinethe name the machine is recorded under. Required for an unfiltered full or calibration run; otherwise the hostname
--outwhere to write the results file instead
--helpprint the flags

A full run with no filters is written to results/, named by the UTC date it started and its machine, and it won't overwrite a run already there. A smoke run or a filtered one is written to results/scratch/, which git ignores, so it can't be committed as if it were a result. A cell process that fails, or any process ended by a signal, fails the run, and no results file is written; the journal keeps what was measured.

--plan makes every check a run makes before it measures, then prints the cells, libraries, and processes, the cold-start scenarios with the processes they add and any a library is unsupported for, the parameters, where the results and the journal would go, and an estimate of the time, worked out from the parameters and not measured. It is refused where the run would be.

Stopping and resuming
#

A full run is one process for each library, pass, and cell, then one for each sample of each cold-start scenario, one at a time, so it is long, and a calibration's cells take twice as long. A timed run keeps a journal beside its results file: each process's result is appended and synced to disk before the next process starts. The parent writes only between processes, never while one is measuring. When a run stops, whether by a signal, a failed process, or a crash, the journal stays, and --resume continues the run from it.

  • What is recorded. A first line describing the run, then for each stretch of measuring (a segment) a line with its opening anchor reading and load average, one line for every process result, a cell's or a cold-start sample's, and a line with its closing anchor reading.
  • What must match. The kind of run, its parameters and filters, the corpus hash, every library's version and extra packages, the machine's name, CPU, threads, and frequency settings (governor, driver, energy-performance preference, and boost), the OS release, the Node version, the NODE_OPTIONS the processes inherit, and the commit with its dirty marker. A resume with any of them different is refused, naming what differs, so results from two setups are never mixed. With a dirty working tree the journal can't tell whether the uncommitted changes are the ones it started with, and a resume says so.
  • Nothing installed mid-run. Every cell process reports the versions of its library's packages it found installed, and one that differs from what the run was planned with stops the run, naming the package and both versions, before its result is journaled. Cold-start processes report none, so the run reads the versions again before writing its file, and on a change writes nothing and sets the journal aside as <journal>.changed: its cold-start samples were never checked against the change, so the run starts over.
  • Never by accident. A journal already there stops a run that wasn't given --resume: it is neither overwritten nor continued until you say which, and the refusal says what run the journal records. A journal holding a result the run would never ask for is refused too.
  • Never after an overlap. A run that lost the machine lock sets its journal aside as <journal>.overlapped, since another run may have overlapped it. It writes no results, and the journal can't be resumed.
  • What is skipped. The processes the journal holds are not run again. The rest run in the order planned, and the results are assembled exactly as an uninterrupted run assembles them.
  • What the file says. segments, in BenchRun, counts the stretches the run was measured in: 1 for an uninterrupted run. generated_at, and the date in a full run's file name, are when the first segment started. duration_ms is the time spent measuring, without the time between segments, and load_at_start the highest of the segments' load averages at their start.
  • A crash in the middle of a write. A last line with no newline is a write that didn't finish. It is dropped, and that process is measured again. A complete line that isn't a journal line is not a crash, and the journal is refused.

The journal is removed once the results file is written. A smoke run is short and keeps none.

Retries
#

A library whose rounds on a cell are far apart was disturbed by something: another program, or its own processes disagreeing. Once a cell's passes are done, a timed run measures each library whose spread there is over 1.25 again, every pass of it, up to --retries more times (2 by default) and until an attempt is under the threshold.

The attempt kept is the one with the lowest spread, the earliest if they tie. It is never chosen by its time. Keeping the fastest attempt would pull every retried number toward fast, and would favor a noisier library, which is retried more often. Choosing on agreement alone takes the measurement that was least disturbed, whichever way it moved the number.

attempts records how many times a library was measured on a cell, so the numbers that needed a retry can be told from the rest. A retried value's spread and percentiles are those of the attempt that was kept, so they describe its calmest attempt, not everything that was measured. The summary a run prints counts the retried cells of each library, and lists the library-cells still over the threshold after the last attempt. What should be the same every time (the token count, the output size, and whether the library passed the check) must agree across attempts as it must across passes, or the run fails at that cell.

A smoke run doesn't retry, and cold-start scenarios are never retried.

The noise floor
#

Before a machine's numbers are trusted, it is calibrated. A calibration measures every library as if it were two libraries: each gets twice the passes, split alternately into two arms, and each arm is exactly what a full run would measure. The two arms should be equal. How far apart they are is what the machine and the method produce when nothing differs, including what varies from one process to the next. BenchNoiseFloor holds three numbers:

  • per_cell is the 95th percentile of that difference, across cells and libraries. A difference between two libraries in one cell that is smaller than this is not a result.
  • geomean takes each library's two geometric means over all cells, and is their difference for the library where it was largest.
  • per_process is the 95th percentile of the difference between two passes of one arm, over every pair of passes, arm, cell, and library: how far two single processes measuring the same thing come apart. It reads higher than per_cell, which compares values that each pool several processes.

A difference is the larger time over the smaller, less one, so it reads the same whichever arm was slower.

A calibration is written to results/calibration_<machine>.json. A later full run copies its noise floor if the calibration describes that run: unfiltered, with a stable anchor, and on the same machine name, CPU, threads, and frequency settings, with the same Node version, corpus, and parameters, retries included, and every library at the same version with the same extra packages. The commit is not compared, so a commit that changes nothing measured keeps the floor. Otherwise the run records no noise floor and says why, and the site won't publish it.

A calibration retries too, each arm of a library on its own and by the same rule, so an arm stays exactly what a full run would measure and the floor describes a full run with the same retries. The rule looks only at an arm's own rounds, never at how far apart the two arms are, so it can't steer the floor. A calibration file's values are the first arm's, and so are their attempts.

Every run with more than one pass also measures its own process noise, an A/A figure taken from the run itself: each library on each cell was measured by several processes that should agree, and every pair of a value's pass_medians_ns gives a difference. BenchProcessNoise records how many pairs there were and their median, 95th percentile, and largest difference, as meta.process_noise. Pairs are only ever taken within one library on one cell. Its 95th percentile is the run's counterpart of the calibration's per_process, and the run's summary and the provenance line on the results page print the two side by side. Both are the 95th percentile of one sample, so either is higher about as often as not: a run's figure several times the calibration's is what says its processes disagreed more than the calibration's did. The reader compares them, and the site does not refuse a run over it.

The run this site shows was measured on laptop1, whose floor is 2.3% for one cell and 0.3% for a geometric mean. The results page marks a ratio inside it.

The anchor
#

A run measures a fixed reference workload before its first process and after its last. The workload involves no measured library, so a change between the two readings is the machine changing state: heat, frequency scaling, another program. drift is the relative change, in BenchAnchor, and a run is stable while it stays under 3%.

The anchor only brackets the run. Something that disturbed the middle of it shows as a high spread, and the summary a run prints lists every library whose rounds on a cell disagreed.

A run that was stopped and resumed takes a reading at the start and end of every segment, and its drift is that of the two readings furthest apart. A machine that changed state between segments, or within one, is then flagged and not hidden by readings that happen to agree at the two ends. A segment stopped by a signal or a failed process still takes its closing reading. One that was killed outright can't, and the next segment's opening reading stands in for it.

Before it spawns anything, a run reads the machine's load average over the last minute and records it as load_at_start. On a quiet machine it is well under 1, and a timed run warns when it is over 1, since something else is then running. It warns and goes on: the anchor, the spread, and the retries are what judge the run.

One run at a time
#

Two measurement runs at once make each other wrong, not just noisy. A run takes a machine-wide lock, the file /tmp/fuz_benchmark.lock that names its process, and holds it for as long as it measures, as does each segment of a resumed run. The path is fixed and ignores TMPDIR, and the lock is fuz_util's benchmark lock, so any other benchmark that uses it takes the same one and never measures at the same time as this one. A second run is refused and told who holds the lock. A lock whose process is gone is taken over.

Caveats
#

  • Features vary. The libraries have different features. Shiki highlights with the TextMate grammars and themes VS Code uses; Twinkleplop aims for a similar feature set, with its own themes, annotations, and Twoslash support; and both do more than fuz_code, which has the fewest features of the four: a single-pass lexer per language, with CSS classes for its theme. Their fidelity differs too: how finely each splits the same source, and how accurately it classifies each part. These numbers compare the work all of them share, not what one offers that another doesn't, and they time each library's output without scoring its quality.
  • The languages are fuz_code's. The languages measured are fuz_code's built-in set, so the list favors it. Prism and Shiki each support hundreds of languages, and none of that shows here.
  • Classes and inline styles. Shiki resolves colors while it tokenizes and writes them inline, and the other libraries emit class names that a stylesheet colors. The html mode measures the HTML each library returns by default, which is not the same work.
  • Token counts differ. A library that emits fewer tokens for the same source may be doing less work for each byte, not the same work faster.
  • Synchronous paths. A library with an async API loads in its setup and is timed through its synchronous core.
  • The HTML is read. In html mode the timed call reads one character of the string, so a library that returns a string the engine hasn't joined up pays to flatten it inside the clock. A benchmark that discards the string doesn't charge for that, so these numbers can differ from a library's own.
  • Markdown without its fences. A timed process loads only the language it measures, so a fenced code block inside Markdown is not highlighted, with the exceptions adapters lists. A site that loads several languages pays more for Markdown than these cells show.
  • Home turf. Some inputs are a measured library's own samples, which it was developed against. They are labeled with that library wherever they appear, summarized by owner, and left out of the default view and the headline figures, which use the neutral inputs.
  • Stress inputs. Some inputs are built to find each library's worst case: minified code, deep nesting, very long lines, constructs that never close. They say little about typical code, so they are summarized on their own and left out of the default view and the headline figures. A slow call there can be far slower than anywhere else, and a run shows it as measured. Shiki is called with its per-line time limit out of reach, as everywhere in this benchmark, so a stress number is the cost of highlighting every line in full. With Shiki's default limit (500ms per line), a line that takes longer comes back with its rest as one plain token: a consumer with default options gets a fallback on that line, not this number.
  • One bundler's sizes. A bundle size is what a page ships when Vite, with the Rolldown inside it, builds it and esbuild minifies it; the results file records the three versions. Another bundler can remove different code from the same entry, since what it can prove unused, whether it folds a constant argument into a function, and how it wraps CommonJS all vary. Built by Rollup, which Vite used before Rolldown, with the same options, the bundles differ unevenly: Rollup removes Twinkleplop's diagnostics code, which it proves is never called, and Rolldown keeps it; Rollup specializes a function fuz_code's Markdown language calls with a constant argument, and Rolldown doesn't; and Rolldown adds a small CommonJS interop helper to Prism's bundles.
  • Synthetic inputs. A few inputs were built by joining or repeating files to reach a size, or derived from one by minifying it, and are marked. A repeated file is more regular than real code of the same size.
  • Steady state. In a cell, every library is warm when it is timed. What a first call costs is in the cold-start scenarios.
  • Cold start is Node's. A cold-start time is a Node process on one machine, with the files already in the page cache. A browser parses and compiles differently, and fetches over a network, where the bundle sizes matter instead.
  • Unbundled depends on the install. The unbundled form loads the files a package manager laid out, so its time moves with the number of files a library ships and how deep they sit. It says what that library costs a Node process without a bundler, and nothing about a page.
  • Collections. Every batch follows a rewarm (unless one call outlasts it) and a forced minor collection, and full collections happen only when the engine schedules them, so garbage from earlier batches carries into later ones and a full collection can still land inside a batch. A library pays for what it allocates (in the JS heap or outside it) as the engine collects it during the run, which is close to, but not the same as, how a page collects between highlights.
  • Measured alone. A page that loads several highlighters, or a busy application, gives the engine a different history than a process holding one library.

Compared with Twinkleplop's benchmark
#

Twinkleplop's own comparison harness is the one this method builds on, and the two answer different questions: it tracks Twinkleplop against other libraries across its own versions, and this one compares a fixed roster on shared inputs. Where they differ, as of Twinkleplop 0.3.1:

Twinkleplop'sthis benchmark
librariesTwinkleplop, Shiki (both engines), Prism, sugar-high, speed-highlightfuz_code, Twinkleplop, Prism, Shiki (both engines)
languagesa broad set, most of what Twinkleplop supportsthe timed languages every library here supports
inputsgenerated inputs at three sizes per language, and Shiki's samplespinned files no measured library wrote as the headline, with Shiki's samples, the libraries' own, snippets, and stress inputs kept apart
isolationone process per cell, its libraries alternating within each roundone process per library and pass, in a rotating order
roundsone pass of several rounds, a short rewarm, a full collection each roundseveral passes of a few rounds each, a longer rewarm, a minor collection only
statisticsmedian, minimum, p10 and p90, and a bootstrap intervalmedian, p10 and p90, spread, each pass's median, and the gap between processes
retries and noiseno retries; a calibrated noise floor for its A/B harness, not the comparisona disturbed cell is measured again; a calibrated noise floor for the comparison itself
HTML timingthe returned string is discarded, so a library returning a rope never flattens itone character of the string is read inside the clock, which flattens it
cold startTwinkleplop alone, with scenarios about its own internalsevery library, through the same bundle entries, in the results file
other metricsnone in the comparisonbundle size, install footprint, retained heap, token counts, output sizes
historya history of published runs across versionseach published run stands alone
machinea bare-metal server with boost off, the run confined to cores away from the system'sa laptop with boost, SMT, and ASLR left at their defaults

Numbers from the two are not comparable with each other: the machines, the inputs, and the HTML timing differ.

Reproducing a run
#

A results file records the commit, the Node version, the machine, the corpus hash, and every library's version and extra packages. From a checkout of that commit, npm ci installs the same versions, and the command for the kind of run measures it again. The deterministic numbers come out the same on any machine running the same Node version, the retained heap within its tolerance, and a run stops if one of its processes disagrees with the committed deterministic metrics. A full run with no filters is refused unless that file describes it. Run the smoke run first: it checks every cell and every first highlight against the file, before a long run can stop on a disagreement. The timed numbers depend on the machine, so compare the ratios between libraries rather than the times.

A quiet machine matters more than a fast one: close other programs, and prefer a fixed CPU frequency governor. Before trusting a run, check that its anchor is stable, that spread is near 1 for every library in every cell, and that nothing unexpected is excluded, from a cell or from a cold-start scenario. bench/run.ts prints all three when it finishes, with how many library-cells were retried and whether the run was resumed.

deterministic metrics #

Some of what the benchmark reports is the same on any machine: how many bytes a library adds to a page, what installing it adds to a project, how much memory it holds once loaded, how many tokens it finds in an input, how large its HTML is, and which languages it covers. These are measured by a command of their own, with nothing timed, and committed as results/deterministic.json.

npm run bench:deterministic # regenerate results/deterministic.json npm run bench:deterministic -- --check # regenerate in memory, and fail if the committed file differs

The file goes through the same schema as a timed run, described in results file, and regenerating it on an unchanged tree changes nothing.

Bundle size
#

For each library and each feature set of languages and sets, the library's adapter writes the module a consumer would write: the library with exactly the set's languages registered, in the shape the library documents. bench/sizes.ts bundles that module, minifies it, and reports its bytes in BenchBundleSizes:

  • raw — the minified JS
  • gzip — those bytes gzipped at level 9, the highest
  • brotli — those bytes brotli-compressed at quality 11, the highest

How a bundle is built:

  • Vite's build API in library mode bundles the entry into one ES module, as a production build for the browser, resolving from the versions installed here.
  • Vite's library mode never removes the whitespace of an ES module, which suits a library and not a page. So the module is then minified once, whole, by esbuild, which keeps license comments: Prism's bundles carry the one in its core file.
  • Compression is Node's zlib and brotli.

A set that holds a language the library doesn't claim is not measured for it. The entry names the languages it lacks instead, and is never a zero.

The shape each adapter writes, which its notes.md states in full:

librarywith languageswith none: core
fuz_codea SyntaxStyler with the needed lexer_* modules addedthe styler alone
Twinkleplopthe two factories of each language's packagethe runtime and the grammar compiler of the core package, which every language package is built from
Prismthe core file and one component file for each grammar, in dependency orderthe core file alone
Shiki, both enginesshiki/core, one engine, the github-light theme, and one module for each languagethe core, the engine, and the theme, since Shiki can't render without a theme

A library's maintainers may know a leaner documented shape. The entry is one function of the adapter, bundle_entry, and a change to it is a small pull request.

Wasm and theme CSS
#

Two things stay out of the JS figures and have a section each, because folding them in would hide a real difference between libraries. Both are reported as the JS is, as raw, gzip, and brotli, in BenchByteSizes:

  • wasm — the WebAssembly binary a library loads, and null for a library with none. Shiki's Oniguruma engine loads one. The documented import carries the binary inlined in a JS module as text, which a bundler would count as JS. The build leaves that one import out of the bundle, and the binary file itself is measured.
  • css — each library's documented default stylesheet, minified by the same esbuild. Output that carries classes needs a stylesheet to show any color. Output that carries inline styles needs none, so for Shiki the section is null: its theme is in the JS figure, and its colors are in every page of HTML. A theme of several stylesheets is measured file by file, each compressed on its own, and summed.

Default themes don't cover the same ground. Each library's roster entry says which color schemes its theme holds, in theme_schemes: the stylesheets of fuz_code and Twinkleplop hold a light and a dark palette, and Prism's holds one. Read the stylesheet sizes beside that.

Which stylesheets are a library's default theme is stated by its adapter, and a theme shipped as a package of its own has its version recorded with the library's.

Each bundle is checked
#

A bundle that lost a language to tree shaking, or was written without one, would be reported as small. So every bundle is written to disk and imported in a process of its own, and each language of its set is highlighted through the bundle's exports, on that language's harness snippet.

  • For a bundle of one language, the HTML must equal what the library produces through its adapter with that language set up. A language missing something it embeds, like HTML without its script and style languages, renders differently and fails.
  • For a bundle of several languages, the HTML must have as many kinds of styled span as the snippet gate of adapters asks for. Equality doesn't hold there by design: a Markdown fence is highlighted when its language is in the bundle, and left plain in a process with one language loaded.

A bundle that doesn't import or exports nothing fails too. The check doesn't prove a bundle is the smallest that works.

What a language costs
#

One set holds no languages, and for each language one set holds it alone, so the bytes a language adds to a library are single:<lang> minus core, with no extra measurement. The results page derives that view from the sets and never stores it, so it can't drift from them. It is the default view of the bundle sizes there.

The sum of those differences over a set and the set's measured size differ when a library shares code between languages, or when one language brings another with it, as HTML brings the script and style languages. The results page shows the sum beside the measured size for that reason.

That difference is the price of the first language: it includes any runtime the language needs that the floor leaves out, which a library pays once. So two more sets measure the price of one more language: timed, every timed language, and without:<lang>, the same less one. Their difference is what a language adds to a bundle that already holds the others. It is read in minified bytes and never compressed: compression works across a whole bundle, so the difference of two compressed bundles doesn't isolate one language's share, and can come out negative. A language the others already hold, like JS inside TypeScript, adds nothing there.

Token counts and output sizes
#

For every input of the corpus and every library that claims its language, bench/outputs.ts records what one tokenize call and one html call return:

  • tokens — the library's own token count
  • output — the HTML's size in UTF-8 bytes, html_bytes, and how many elements it has, spans
  • compressed — the same HTML gzipped and brotli-compressed at the levels bundle sizes use: what a server sends for a page that renders the input highlighted. The HTML is produced again for this, outside the check, and must come to the size the check saw.

They are taken as a timed cell takes them: a process sets up exactly one language in one library, the snippet gate runs, and each input is then checked. A library that fails is recorded as excluded with the reason, not counted.

Equal sizes don't mean equal output: two libraries, or Shiki's two engines, can return HTML of the same size that colors a token differently, and nothing here compares content.

Each number is stated once for an input, on the cell whose call returns it: the token count on the input's tokenize cell and the output sizes, raw and compressed, on its html cell.

What the numbers don't say:

  • Token counts are not comparable as work. Libraries split the same source differently. Shiki counts runs of whitespace and unstyled text, the others count only typed tokens, and a container and what it holds may each count. Fewer tokens is less output, not a faster library.
  • Classes and inline styles are different HTML. A class name is short and needs a stylesheet. An inline color is repeated on every token and needs none. Compare html_bytes with the theme CSS column beside it. Compression narrows the gap, since a repeated color is what it removes best, which is why the compressed sizes are recorded too.
  • Wrappers differ. Twinkleplop and Shiki wrap the output in <pre> and a span for each line, and fuz_code and Prism return bare token spans. spans counts every element.
  • Prism's two calls can disagree on Markdown. Prism highlights the inside of a fenced block in a hook that only highlight runs, never tokenize. For a fence whose language is loaded, the HTML holds elements the token count doesn't include.

Install footprint
#

What npm install of a library adds to a project, in BenchInstall: bench/install.ts walks the runtime dependency closure of the adapter's package and each of its separately versioned packages, resolving each dependency from the package that needs it as Node resolves it, through this repository's node_modules.

  • packages — the packages in the closure, each counted once
  • files — the files in those packages' directories, less any nested node_modules, whose packages are counted through the closure
  • bytes — the sum of those files' sizes as stored, not the disk blocks

What the closure follows:

  • dependencies, always
  • peerDependencies not marked optional, which npm installs as well: fuz_code's utility library is one
  • optionalDependencies that are installed and not limited to some platforms by os, cpu, or libc. A package with a binary for each platform installs one of them on each machine, and counting it would make the figure the machine's. No library in the roster has one today.
  • never devDependencies

Each library is counted whole, on its own, as installing it alone would give: the two Shiki adapters install the same package and have the same footprint. npm installs what a package's tarball holds, so the footprint is everything the package ships, every language, theme, and type declaration included, and not what a page loads.

The versions are the ones this repository's lockfile installs, transitive ones included. A fresh npm install of the same library may resolve a dependency to a newer version within its range, and its footprint then differs. A package is told apart by name and version, so one installed twice at the same version counts once. closure is a hash of the sorted name@version list: a timed run carries the file only when each library's closure is this tree's, and the site names a library whose dependencies moved since a run.

Retained heap
#

What a page or a server holds once a library has loaded one language and highlighted once, in BenchHeap, for every timed language each library claims. bench/heap.ts runs a Node process that imports the library's bundle of that language alone, the single: bundle whose size is reported, highlights the language's harness snippet once through the adapter's bundle_html_source, lets go of the HTML, and runs full collections until one changes the heap by under a kibibyte. It then reads three numbers, each in KiB over a bare process of the same module without the import and the call:

  • retained_kb — the bytes in use in V8's heap, less the two spaces that hold machine code
  • code_kb — the bytes in use in those two spaces. They hold the regular expressions the engine compiled to native code as well as the functions it optimized, so a library that lexes with regular expressions holds much of its memory here. Machine code is the processor's: the figures are for x64.
  • external_kb — process.memoryUsage().external: memory outside the V8 heap that V8 is told about, like array buffers, and a WebAssembly engine's linear memory, counted at the size it has grown to, which it never gives back. The OS backs only the pages it has touched, so resident memory can be lower.

The process imports the bundle, not the adapter: some adapters import every language statically, and an adapter is TypeScript, which would load Node's type stripper into the reading. The generated module holds nothing of the harness, as a cold-start process doesn't.

It runs with three flags, for this reading only. The timed cells run with --expose-gc alone, for their minor collections, and cold start with none. --expose-gc gives it the collections. --max-semi-space-size=16 pins the young generation, which V8 otherwise sizes from the machine's memory, and a small one leaves some libraries holding a different amount. --single-threaded keeps the engine from finishing compilation on another thread between the collections and the reading, which otherwise moves Shiki's figures by kilobytes from one process to the next.

A theme in a stylesheet is held by the page, not the library, so a library that writes its styles inline holds its theme in these figures and one that writes classes doesn't, as with the bundle sizes. The figures are after one highlight of a small snippet; larger inputs leave more, most of all for regexp and wasm engines.

Coverage
#

The coverage matrix is each adapter's language map: true for a language the library claims, a footnote where the support is indirect, and no entry for a language it doesn't claim, in BenchCoverageEntry. The footnotes are written once, in the adapters, and adapters lists what they cover.

What the file records
#

A deterministic file holds no timed value, so it records no machine, no time, no Node version, and no commit. The Node version does shape the compressed sizes and the heaps (below), but a file naming it would differ between machines that agree, and a commit can't name the tree it is itself committed in. What the numbers do depend on is in the file:

  • each library's version, and the versions of its separately versioned packages
  • each library's install closure, as a hash
  • the corpus hash
  • the versions of Vite, of the Rolldown inside it, and of esbuild, and the compression levels, in BenchBundler

Two dependencies are not recorded by version. The gzip and brotli sizes, of bundles and of HTML alike, come from the zlib and brotli inside Node, and a Node with another version of either can compress to a few bytes more or less. The file is generated on the Node version the checks run on. And the retained heaps follow V8's object layout: they are for the Node version the checks run on, on x64, without pointer compression, as official Node builds are. The install footprints follow the versions the lockfile installs, which a dependency update moves without moving a library's own version; the closure hash records them. A failed check says which of these it looks like.

A heap reading can also differ from the last by a little between two processes on the same machine. Two readings agree to well under a kibibyte, and the check still compares them within 16.4 kB or 0.5% of the larger, whichever is larger, and every other number exactly. Regenerating keeps a committed reading that agrees, so an unchanged tree regenerates an unchanged file, and a change smaller than the tolerance isn't recorded until it, with later ones, exceeds it. The site reads two numbers within the tolerance as level, marked ≈.

Beside a timed run
#

A timed run, described in method, takes its cells' token counts and output sizes from its own processes, by the same check under the same setup. When the deterministic file describes the run (the same library versions and extra packages, install closures, corpus, and feature sets), the run does two more things:

  • it compares each library's first result on a cell with the file, and stops at a difference
  • it copies the bundle sizes, the install footprints, the retained heaps, and the HTML's compressed sizes of its libraries from the file, and measures none of them (cold start rebuilds a missing cached bundle, checked against those sizes)

The compressed sizes are not compared on their own: the raw size each process reports is, and the file applies only at the same install closures, which pin every package that renders the HTML, transitive ones included. Compressing every process's HTML again would compress the same bytes in every process, for numbers the file already holds.

So a timed results file holds the deterministic numbers of its cells and the sizes of its libraries, equal to the deterministic file's. When the file doesn't describe the run, a smoke or filtered run says so and carries none of those sizes, and a full run with no filters is refused before it measures anything: regenerating the file is quick, and the run is long.

Kept current
#

On every change the repository's check workflow runs gro check, bench:deterministic -- --check, and a smoke run of every library on every cell and cold-start scenario. The deterministic check rebuilds every bundle and recounts every input, and fails with the places that differ when the committed file is stale. Nothing in the workflow is a measurement of speed.

results file #

A run of the benchmark is one JSON file. The harness writes it and the site reads it, and both go through the same schema, BenchResults in results_schema.ts, so a malformed run fails loudly instead of rendering blanks. The schema is strict: an unknown key is an error.

bench/run.ts writes one for a timed run, as method describes, and bench/deterministic.ts writes the committed results/deterministic.json, which holds only the numbers that are the same on any machine, as deterministic metrics describes.

Sections
#

  • BenchMeta — the machine, Node version, commit, and corpus hash; what kind of run it was and how it measured, in BenchRun, with the retries it allowed, how cold start was sampled, how many segments it was measured in, and the load average it started at; the roster of libraries with their versions and the color schemes their default themes cover; the excluded list for cells, and startup_excluded for cold-start scenarios, in BenchStartupExcluded; the anchor's drift; the machine's noise floor; the run's own process noise, in BenchProcessNoise, from every pair of a value's passes; what the bundle sizes were built with, and the compression levels every compressed size is taken at, in BenchBundler
  • BenchLang and BenchSet — the language ids and feature sets of the run, described in languages and sets
  • BenchInput — the inputs the run measured, each with its source, size tier, hash, and the upstream commit it was copied from, as in the corpus manifest
  • BenchCell — one input measured in one mode, with each library's numbers. Its id is <file>:<mode>. A tokenize cell states the token counts and an html cell the output sizes, with the HTML's compressed sizes in BenchCellCompressed, so each is stated once for an input. Each value carries attempts: how many times the library was measured there, more than 1 when its rounds disagreed and it was retried. A retried value's spread and percentiles are those of the attempt that was kept. Each value also carries pass_medians_ns, the kept attempt's median for each pass, so the processes' disagreement can be recomputed from the file
  • BenchStartup — one cold-start scenario: what a fresh process loaded (bare, core, lang, set, or first_highlight), for which library, with which language or feature set, and whether as one bundled chunk. Its numbers are the median time since the process started, the 10th and 90th percentile, the spread, how many samples were kept, and the processes' median peak resident size. The one bare entry, with no library, is the Node baseline every other entry includes
  • BenchBundleEntry — for each library over each feature set, its bundle sizes in BenchBundleSizes, or the languages of the set it lacks. The sizes are the JS's, with the wasm and the theme CSS each measured apart
  • BenchInstall — for each library, what npm install of it adds: the packages of its runtime dependency closure, their files, those files' bytes, and a hash of the closure's name@version list
  • BenchHeap — for each library and each timed language it claims, what it holds once that language is loaded and highlighted once: the V8 heap less machine code, the machine code, and the external memory beside it, in KiB
  • BenchCoverageEntry — which languages each library claims, with a footnote where the support is indirect

The file describes itself. The libraries, languages, sets, and inputs are lists inside it, and every other section refers to them by id, or by file for inputs, so the site needs no list of its own and adding a library changes no site code.

Two classes of number
#

Machine-dependent numbers come from timing on one named machine and mean nothing apart from it: throughput and memory in each cell's values, and the startup section.

Deterministic numbers come out the same anywhere: token counts and output sizes in each cell's tokens, output, and compressed, plus bundle, install, heap, and coverage. The retained heaps are deterministic for a Node version and platform, within the tolerance deterministic metrics gives.

A file may hold only the deterministic ones. It then has no meta.run and no anchor, and its generated_at, machine, node, and commit are null: they say where timed numbers came from, and are required as soon as a file has any. A timed run carries the deterministic numbers too, so one published file holds both.

Publishable or not
#

meta.run tells a publishable run from the rest. Its kind is full, smoke, or calibrate, and its filters are null unless the run was restricted to some libraries, languages, sizes, modes, sources, or metrics (the timed cells, or cold start). A smoke run's numbers mean nothing, and a calibration exists for its noise floor.

bench_results_is_publishable is the test for a result: a full run with no filters and no gap in its cells or its cold start, a stable anchor, a noise floor carried from the machine's calibration, and a commit that records everything that ran.

No filters is what a run set out to cover. Whether the file then holds all of it is checked by bench_results_find_gaps, from the file alone:

  • in every cell, each library of the roster that claims the cell's language has a value, or an entry in the excluded list for that cell
  • every listed input of a timed language has a cell in each mode the file's cells use

A full run's cold start is checked the same way, by bench_results_find_startup_gaps. Each of these has an entry or an exclusion:

  • the bare baseline
  • for each library of the roster, unbundled and bundled: core; lang and first_highlight for every timed language it claims; and set for every feature set the file's cold start names whose languages it all claims

A file with no filters and a gap doesn't parse, whether it is a full run or a calibration, so a run that silently lost a library can't pass as complete. A calibration measures no cold start, so only its cells are checked. A filtered or smoke run lists what it has and may have gaps. Two things the file can't show are an input missing altogether and a whole mode missing: its inputs and modes are the ones that were measured, and only the corpus hash names the corpus they came from. A third is a cold-start set missing altogether: the sets a set scenario covers are read from the file too, so a file with every set entry removed still reads as complete.

A run that was stopped and resumed is a result like any other. meta.run.segments says how many stretches it was measured in, and its anchor covers every one of them, as method describes.

meta.commit ends in -dirty when the working tree had changes the commit doesn't record, and such a run can't be reproduced from its commit.

What the site renders
#

The site is built from one results file, parsed through the schema when the site is built, so a malformed file fails the build and no number reaches a page unchecked:

  • results/latest.json, the latest published timed run, when it is committed: npm run results:publish copies a run there once it passes the test below. It carries its own bundle sizes, install footprints, retained heaps, token counts, and output sizes, so every number on a page is from one file.
  • otherwise results/deterministic.json. The pages then show the deterministic numbers, and say that no timed results are published where a timed number would be. A deterministic file that holds a timed number fails the build.

A latest run must pass bench_results_is_publishable, or the build fails and says why. A smoke run, a calibration, a filtered run, a run with a gap, a run from a commit with uncommitted changes, a run whose anchor drifted, and a run with no noise floor are never shown as results. A drifted run is measured again, not published under a warning.

A published run is also held against the deterministic file. When the two describe the same library versions, the same install closures, the same corpus, and the same bundler and compression levels, their deterministic numbers must be equal: the feature sets, the coverage, the bundle sizes, the install footprints, and each cell's token counts and output sizes, raw and compressed. A difference fails the build and names the first place they differ. The retained heaps are compared within the check's tolerance instead, and never fail the build. When a library's version, its dependencies, the corpus, or what built and compressed the bundles (the versions of Vite, of the bundler inside it, and of esbuild, and the compression levels) has moved since the run, or the retained heaps differ by more, as they do when the Node version moves, the run is shown as the snapshot it is, and the line that says where the numbers came from also says what has moved since.

The fixture run the tests use holds invented numbers. A build renders it only when the BENCH_RESULTS_FIXTURE environment variable is 1, which is for working on the pages before a run exists, and every page of such a build is marked as fixture data. A build without the variable never reads the file.

Consistency checks
#

Beyond the shape of each section, parsing checks that the sections agree:

  • every library, language, set, and cell id is unique and refers to a listed entry, except the library a home input came from, which may be outside the run
  • no two cells measure the same input in the same mode, and no cell covers an untimed language
  • every cell names a listed input, describes it as that entry does, and is named after its file and mode
  • an input written for the benchmark has no provenance, a verbatim copy has the hash its one pin records, and a derived stress input names its derivation, and one built from several files is marked synthetic
  • a run with timed numbers records its anchor, when it started, and its machine, Node version, and commit, and how its numbers were measured
  • token counts are on tokenize cells only and output sizes on html cells only, and a library with a timed value in a cell has that cell's deterministic number
  • a value's 10th percentile is at most its 90th, and an anchor is stable exactly when its drift is under the threshold
  • a run's filters describe the file: every cell and cold-start entry is inside them, the libraries listed are exactly the ones filtered to, and no id repeats
  • a smoke run carries no noise floor, allows no retries, and is one segment
  • no value has more attempts than the run's retries allow
  • a value has a pass median for each of the run's passes, and the process noise is there exactly when a cell has a value and the run measured more than one pass, counting every pair of passes of every value, with its median at most its 95th percentile and that at most its largest
  • a run with no filters has no gap, as above
  • each startup entry carries the fields its scenario needs and no others, for languages the library claims, and appears once, as a number or as an exclusion and never both
  • a lang or first_highlight entry names a timed language. A set entry names a feature set, which may hold an untimed one
  • every startup entry keeps the samples the run says, its median lies between its 10th and 90th percentile, and a calibration has none
  • how cold start was sampled is recorded exactly when the run measured any, and is null for a calibration and for a run whose metrics filter leaves it out
  • a library has no numbers for a language its coverage doesn't claim
  • a library listed as excluded from a cell claims that cell's language and has no numbers there
  • a bundle entry holds sizes only when the library claims every language of the set, and otherwise names exactly the languages it lacks
  • bundle sizes cover every library over every set or are absent altogether, and what they were built with is recorded exactly when they are present
  • the install footprints and the HTML's compressed sizes come with the bundle sizes: every library of the roster has a footprint, and every library with an output on an html cell has its compressed sizes, exactly when the bundle sizes are present. A footprint has at least a file for each package
  • the retained heaps come with the bundle sizes too, for every timed language each library claims and no other
  • a compressed size is no larger than the raw one plus a compression format's own overhead, the HTML's included, and the theme CSS sizes are null exactly for a library that writes its styles inline
  • a file with no timed numbers is complete: in every cell each library that claims the language has the number the cell states or is listed as excluded, and every listed input has a cell in each mode

Use parse_bench_results to validate a file and get every problem in one message.

design #

More views
#

Each of these is a view over the same results file:

  • a page for each library: its version, its adapter's notes, its sizes, and its cells
  • a page for each language, with every library's rendering of the same sample side by side
  • the change from one published run to the next, for every library
  • the throughput cells run in the reader's own browser

More measurements
#

  • Markdown with the languages of its fenced blocks loaded, beside the cells that load Markdown alone
  • other runtimes than Node

api #

Browse the full api docs.

library #

syntax-highlighter-bench
benchmark suite and results site for JS syntax highlighters
homepage repo license