deterministic metrics #

Some of what the benchmark reports is the same on any machine: how many bytes a library adds to a page, what installing it adds to a project, how much memory it holds once loaded, how many tokens it finds in an input, how large its HTML is, and which languages it covers. These are measured by a command of their own, with nothing timed, and committed as results/deterministic.json.

npm run bench:deterministic # regenerate results/deterministic.json npm run bench:deterministic -- --check # regenerate in memory, and fail if the committed file differs

The file goes through the same schema as a timed run, described in results file, and regenerating it on an unchanged tree changes nothing.

Bundle size
#

For each library and each feature set of languages and sets, the library's adapter writes the module a consumer would write: the library with exactly the set's languages registered, in the shape the library documents. bench/sizes.ts bundles that module, minifies it, and reports its bytes in BenchBundleSizes:

  • raw — the minified JS
  • gzip — those bytes gzipped at level 9, the highest
  • brotli — those bytes brotli-compressed at quality 11, the highest

How a bundle is built:

  • Vite's build API in library mode bundles the entry into one ES module, as a production build for the browser, resolving from the versions installed here.
  • Vite's library mode never removes the whitespace of an ES module, which suits a library and not a page. So the module is then minified once, whole, by esbuild, which keeps license comments: Prism's bundles carry the one in its core file.
  • Compression is Node's zlib and brotli.

A set that holds a language the library doesn't claim is not measured for it. The entry names the languages it lacks instead, and is never a zero.

The shape each adapter writes, which its notes.md states in full:

librarywith languageswith none: core
fuz_codea SyntaxStyler with the needed lexer_* modules addedthe styler alone
Twinkleplopthe two factories of each language's packagethe runtime and the grammar compiler of the core package, which every language package is built from
Prismthe core file and one component file for each grammar, in dependency orderthe core file alone
Shiki, both enginesshiki/core, one engine, the github-light theme, and one module for each languagethe core, the engine, and the theme, since Shiki can't render without a theme

A library's maintainers may know a leaner documented shape. The entry is one function of the adapter, bundle_entry, and a change to it is a small pull request.

Wasm and theme CSS
#

Two things stay out of the JS figures and have a section each, because folding them in would hide a real difference between libraries. Both are reported as the JS is, as raw, gzip, and brotli, in BenchByteSizes:

  • wasm — the WebAssembly binary a library loads, and null for a library with none. Shiki's Oniguruma engine loads one. The documented import carries the binary inlined in a JS module as text, which a bundler would count as JS. The build leaves that one import out of the bundle, and the binary file itself is measured.
  • css — each library's documented default stylesheet, minified by the same esbuild. Output that carries classes needs a stylesheet to show any color. Output that carries inline styles needs none, so for Shiki the section is null: its theme is in the JS figure, and its colors are in every page of HTML. A theme of several stylesheets is measured file by file, each compressed on its own, and summed.

Default themes don't cover the same ground. Each library's roster entry says which color schemes its theme holds, in theme_schemes: the stylesheets of fuz_code and Twinkleplop hold a light and a dark palette, and Prism's holds one. Read the stylesheet sizes beside that.

Which stylesheets are a library's default theme is stated by its adapter, and a theme shipped as a package of its own has its version recorded with the library's.

Each bundle is checked
#

A bundle that lost a language to tree shaking, or was written without one, would be reported as small. So every bundle is written to disk and imported in a process of its own, and each language of its set is highlighted through the bundle's exports, on that language's harness snippet.

  • For a bundle of one language, the HTML must equal what the library produces through its adapter with that language set up. A language missing something it embeds, like HTML without its script and style languages, renders differently and fails.
  • For a bundle of several languages, the HTML must have as many kinds of styled span as the snippet gate of adapters asks for. Equality doesn't hold there by design: a Markdown fence is highlighted when its language is in the bundle, and left plain in a process with one language loaded.

A bundle that doesn't import or exports nothing fails too. The check doesn't prove a bundle is the smallest that works.

What a language costs
#

One set holds no languages, and for each language one set holds it alone, so the bytes a language adds to a library are single:<lang> minus core, with no extra measurement. The results page derives that view from the sets and never stores it, so it can't drift from them. It is the default view of the bundle sizes there.

The sum of those differences over a set and the set's measured size differ when a library shares code between languages, or when one language brings another with it, as HTML brings the script and style languages. The results page shows the sum beside the measured size for that reason.

That difference is the price of the first language: it includes any runtime the language needs that the floor leaves out, which a library pays once. So two more sets measure the price of one more language: timed, every timed language, and without:<lang>, the same less one. Their difference is what a language adds to a bundle that already holds the others. It is read in minified bytes and never compressed: compression works across a whole bundle, so the difference of two compressed bundles doesn't isolate one language's share, and can come out negative. A language the others already hold, like JS inside TypeScript, adds nothing there.

Token counts and output sizes
#

For every input of the corpus and every library that claims its language, bench/outputs.ts records what one tokenize call and one html call return:

  • tokens — the library's own token count
  • output — the HTML's size in UTF-8 bytes, html_bytes, and how many elements it has, spans
  • compressed — the same HTML gzipped and brotli-compressed at the levels bundle sizes use: what a server sends for a page that renders the input highlighted. The HTML is produced again for this, outside the check, and must come to the size the check saw.

They are taken as a timed cell takes them: a process sets up exactly one language in one library, the snippet gate runs, and each input is then checked. A library that fails is recorded as excluded with the reason, not counted.

Equal sizes don't mean equal output: two libraries, or Shiki's two engines, can return HTML of the same size that colors a token differently, and nothing here compares content.

Each number is stated once for an input, on the cell whose call returns it: the token count on the input's tokenize cell and the output sizes, raw and compressed, on its html cell.

What the numbers don't say:

  • Token counts are not comparable as work. Libraries split the same source differently. Shiki counts runs of whitespace and unstyled text, the others count only typed tokens, and a container and what it holds may each count. Fewer tokens is less output, not a faster library.
  • Classes and inline styles are different HTML. A class name is short and needs a stylesheet. An inline color is repeated on every token and needs none. Compare html_bytes with the theme CSS column beside it. Compression narrows the gap, since a repeated color is what it removes best, which is why the compressed sizes are recorded too.
  • Wrappers differ. Twinkleplop and Shiki wrap the output in <pre> and a span for each line, and fuz_code and Prism return bare token spans. spans counts every element.
  • Prism's two calls can disagree on Markdown. Prism highlights the inside of a fenced block in a hook that only highlight runs, never tokenize. For a fence whose language is loaded, the HTML holds elements the token count doesn't include.

Install footprint
#

What npm install of a library adds to a project, in BenchInstall: bench/install.ts walks the runtime dependency closure of the adapter's package and each of its separately versioned packages, resolving each dependency from the package that needs it as Node resolves it, through this repository's node_modules.

  • packages — the packages in the closure, each counted once
  • files — the files in those packages' directories, less any nested node_modules, whose packages are counted through the closure
  • bytes — the sum of those files' sizes as stored, not the disk blocks

What the closure follows:

  • dependencies, always
  • peerDependencies not marked optional, which npm installs as well: fuz_code's utility library is one
  • optionalDependencies that are installed and not limited to some platforms by os, cpu, or libc. A package with a binary for each platform installs one of them on each machine, and counting it would make the figure the machine's. No library in the roster has one today.
  • never devDependencies

Each library is counted whole, on its own, as installing it alone would give: the two Shiki adapters install the same package and have the same footprint. npm installs what a package's tarball holds, so the footprint is everything the package ships, every language, theme, and type declaration included, and not what a page loads.

The versions are the ones this repository's lockfile installs, transitive ones included. A fresh npm install of the same library may resolve a dependency to a newer version within its range, and its footprint then differs. A package is told apart by name and version, so one installed twice at the same version counts once. closure is a hash of the sorted name@version list: a timed run carries the file only when each library's closure is this tree's, and the site names a library whose dependencies moved since a run.

Retained heap
#

What a page or a server holds once a library has loaded one language and highlighted once, in BenchHeap, for every timed language each library claims. bench/heap.ts runs a Node process that imports the library's bundle of that language alone, the single: bundle whose size is reported, highlights the language's harness snippet once through the adapter's bundle_html_source, lets go of the HTML, and runs full collections until one changes the heap by under a kibibyte. It then reads three numbers, each in KiB over a bare process of the same module without the import and the call:

  • retained_kb — the bytes in use in V8's heap, less the two spaces that hold machine code
  • code_kb — the bytes in use in those two spaces. They hold the regular expressions the engine compiled to native code as well as the functions it optimized, so a library that lexes with regular expressions holds much of its memory here. Machine code is the processor's: the figures are for x64.
  • external_kb — process.memoryUsage().external: memory outside the V8 heap that V8 is told about, like array buffers, and a WebAssembly engine's linear memory, counted at the size it has grown to, which it never gives back. The OS backs only the pages it has touched, so resident memory can be lower.

The process imports the bundle, not the adapter: some adapters import every language statically, and an adapter is TypeScript, which would load Node's type stripper into the reading. The generated module holds nothing of the harness, as a cold-start process doesn't.

It runs with three flags, for this reading only. The timed cells run with --expose-gc alone, for their minor collections, and cold start with none. --expose-gc gives it the collections. --max-semi-space-size=16 pins the young generation, which V8 otherwise sizes from the machine's memory, and a small one leaves some libraries holding a different amount. --single-threaded keeps the engine from finishing compilation on another thread between the collections and the reading, which otherwise moves Shiki's figures by kilobytes from one process to the next.

A theme in a stylesheet is held by the page, not the library, so a library that writes its styles inline holds its theme in these figures and one that writes classes doesn't, as with the bundle sizes. The figures are after one highlight of a small snippet; larger inputs leave more, most of all for regexp and wasm engines.

Coverage
#

The coverage matrix is each adapter's language map: true for a language the library claims, a footnote where the support is indirect, and no entry for a language it doesn't claim, in BenchCoverageEntry. The footnotes are written once, in the adapters, and adapters lists what they cover.

What the file records
#

A deterministic file holds no timed value, so it records no machine, no time, no Node version, and no commit. The Node version does shape the compressed sizes and the heaps (below), but a file naming it would differ between machines that agree, and a commit can't name the tree it is itself committed in. What the numbers do depend on is in the file:

  • each library's version, and the versions of its separately versioned packages
  • each library's install closure, as a hash
  • the corpus hash
  • the versions of Vite, of the Rolldown inside it, and of esbuild, and the compression levels, in BenchBundler

Two dependencies are not recorded by version. The gzip and brotli sizes, of bundles and of HTML alike, come from the zlib and brotli inside Node, and a Node with another version of either can compress to a few bytes more or less. The file is generated on the Node version the checks run on. And the retained heaps follow V8's object layout: they are for the Node version the checks run on, on x64, without pointer compression, as official Node builds are. The install footprints follow the versions the lockfile installs, which a dependency update moves without moving a library's own version; the closure hash records them. A failed check says which of these it looks like.

A heap reading can also differ from the last by a little between two processes on the same machine. Two readings agree to well under a kibibyte, and the check still compares them within 16.4 kB or 0.5% of the larger, whichever is larger, and every other number exactly. Regenerating keeps a committed reading that agrees, so an unchanged tree regenerates an unchanged file, and a change smaller than the tolerance isn't recorded until it, with later ones, exceeds it. The site reads two numbers within the tolerance as level, marked ≈.

Beside a timed run
#

A timed run, described in method, takes its cells' token counts and output sizes from its own processes, by the same check under the same setup. When the deterministic file describes the run (the same library versions and extra packages, install closures, corpus, and feature sets), the run does two more things:

  • it compares each library's first result on a cell with the file, and stops at a difference
  • it copies the bundle sizes, the install footprints, the retained heaps, and the HTML's compressed sizes of its libraries from the file, and measures none of them (cold start rebuilds a missing cached bundle, checked against those sizes)

The compressed sizes are not compared on their own: the raw size each process reports is, and the file applies only at the same install closures, which pin every package that renders the HTML, transitive ones included. Compressing every process's HTML again would compress the same bytes in every process, for numbers the file already holds.

So a timed results file holds the deterministic numbers of its cells and the sizes of its libraries, equal to the deterministic file's. When the file doesn't describe the run, a smoke or filtered run says so and carries none of those sizes, and a full run with no filters is refused before it measures anything: regenerating the file is quick, and the run is long.

Kept current
#

On every change the repository's check workflow runs gro check, bench:deterministic -- --check, and a smoke run of every library on every cell and cold-start scenario. The deterministic check rebuilds every bundle and recounts every input, and fails with the places that differ when the committed file is stale. Nothing in the workflow is a measurement of speed.