results file #

A run of the benchmark is one JSON file. The harness writes it and the site reads it, and both go through the same schema, BenchResults in results_schema.ts, so a malformed run fails loudly instead of rendering blanks. The schema is strict: an unknown key is an error.

bench/run.ts writes one for a timed run, as method describes, and bench/deterministic.ts writes the committed results/deterministic.json, which holds only the numbers that are the same on any machine, as deterministic metrics describes.

Sections
#

  • BenchMeta — the machine, Node version, commit, and corpus hash; what kind of run it was and how it measured, in BenchRun, with the retries it allowed, how cold start was sampled, how many segments it was measured in, and the load average it started at; the roster of libraries with their versions and the color schemes their default themes cover; the excluded list for cells, and startup_excluded for cold-start scenarios, in BenchStartupExcluded; the anchor's drift; the machine's noise floor; the run's own process noise, in BenchProcessNoise, from every pair of a value's passes; what the bundle sizes were built with, and the compression levels every compressed size is taken at, in BenchBundler
  • BenchLang and BenchSet — the language ids and feature sets of the run, described in languages and sets
  • BenchInput — the inputs the run measured, each with its source, size tier, hash, and the upstream commit it was copied from, as in the corpus manifest
  • BenchCell — one input measured in one mode, with each library's numbers. Its id is <file>:<mode>. A tokenize cell states the token counts and an html cell the output sizes, with the HTML's compressed sizes in BenchCellCompressed, so each is stated once for an input. Each value carries attempts: how many times the library was measured there, more than 1 when its rounds disagreed and it was retried. A retried value's spread and percentiles are those of the attempt that was kept. Each value also carries pass_medians_ns, the kept attempt's median for each pass, so the processes' disagreement can be recomputed from the file
  • BenchStartup — one cold-start scenario: what a fresh process loaded (bare, core, lang, set, or first_highlight), for which library, with which language or feature set, and whether as one bundled chunk. Its numbers are the median time since the process started, the 10th and 90th percentile, the spread, how many samples were kept, and the processes' median peak resident size. The one bare entry, with no library, is the Node baseline every other entry includes
  • BenchBundleEntry — for each library over each feature set, its bundle sizes in BenchBundleSizes, or the languages of the set it lacks. The sizes are the JS's, with the wasm and the theme CSS each measured apart
  • BenchInstall — for each library, what npm install of it adds: the packages of its runtime dependency closure, their files, those files' bytes, and a hash of the closure's name@version list
  • BenchHeap — for each library and each timed language it claims, what it holds once that language is loaded and highlighted once: the V8 heap less machine code, the machine code, and the external memory beside it, in KiB
  • BenchCoverageEntry — which languages each library claims, with a footnote where the support is indirect

The file describes itself. The libraries, languages, sets, and inputs are lists inside it, and every other section refers to them by id, or by file for inputs, so the site needs no list of its own and adding a library changes no site code.

Two classes of number
#

Machine-dependent numbers come from timing on one named machine and mean nothing apart from it: throughput and memory in each cell's values, and the startup section.

Deterministic numbers come out the same anywhere: token counts and output sizes in each cell's tokens, output, and compressed, plus bundle, install, heap, and coverage. The retained heaps are deterministic for a Node version and platform, within the tolerance deterministic metrics gives.

A file may hold only the deterministic ones. It then has no meta.run and no anchor, and its generated_at, machine, node, and commit are null: they say where timed numbers came from, and are required as soon as a file has any. A timed run carries the deterministic numbers too, so one published file holds both.

Publishable or not
#

meta.run tells a publishable run from the rest. Its kind is full, smoke, or calibrate, and its filters are null unless the run was restricted to some libraries, languages, sizes, modes, sources, or metrics (the timed cells, or cold start). A smoke run's numbers mean nothing, and a calibration exists for its noise floor.

bench_results_is_publishable is the test for a result: a full run with no filters and no gap in its cells or its cold start, a stable anchor, a noise floor carried from the machine's calibration, and a commit that records everything that ran.

No filters is what a run set out to cover. Whether the file then holds all of it is checked by bench_results_find_gaps, from the file alone:

  • in every cell, each library of the roster that claims the cell's language has a value, or an entry in the excluded list for that cell
  • every listed input of a timed language has a cell in each mode the file's cells use

A full run's cold start is checked the same way, by bench_results_find_startup_gaps. Each of these has an entry or an exclusion:

  • the bare baseline
  • for each library of the roster, unbundled and bundled: core; lang and first_highlight for every timed language it claims; and set for every feature set the file's cold start names whose languages it all claims

A file with no filters and a gap doesn't parse, whether it is a full run or a calibration, so a run that silently lost a library can't pass as complete. A calibration measures no cold start, so only its cells are checked. A filtered or smoke run lists what it has and may have gaps. Two things the file can't show are an input missing altogether and a whole mode missing: its inputs and modes are the ones that were measured, and only the corpus hash names the corpus they came from. A third is a cold-start set missing altogether: the sets a set scenario covers are read from the file too, so a file with every set entry removed still reads as complete.

A run that was stopped and resumed is a result like any other. meta.run.segments says how many stretches it was measured in, and its anchor covers every one of them, as method describes.

meta.commit ends in -dirty when the working tree had changes the commit doesn't record, and such a run can't be reproduced from its commit.

What the site renders
#

The site is built from one results file, parsed through the schema when the site is built, so a malformed file fails the build and no number reaches a page unchecked:

  • results/latest.json, the latest published timed run, when it is committed: npm run results:publish copies a run there once it passes the test below. It carries its own bundle sizes, install footprints, retained heaps, token counts, and output sizes, so every number on a page is from one file.
  • otherwise results/deterministic.json. The pages then show the deterministic numbers, and say that no timed results are published where a timed number would be. A deterministic file that holds a timed number fails the build.

A latest run must pass bench_results_is_publishable, or the build fails and says why. A smoke run, a calibration, a filtered run, a run with a gap, a run from a commit with uncommitted changes, a run whose anchor drifted, and a run with no noise floor are never shown as results. A drifted run is measured again, not published under a warning.

A published run is also held against the deterministic file. When the two describe the same library versions, the same install closures, the same corpus, and the same bundler and compression levels, their deterministic numbers must be equal: the feature sets, the coverage, the bundle sizes, the install footprints, and each cell's token counts and output sizes, raw and compressed. A difference fails the build and names the first place they differ. The retained heaps are compared within the check's tolerance instead, and never fail the build. When a library's version, its dependencies, the corpus, or what built and compressed the bundles (the versions of Vite, of the bundler inside it, and of esbuild, and the compression levels) has moved since the run, or the retained heaps differ by more, as they do when the Node version moves, the run is shown as the snapshot it is, and the line that says where the numbers came from also says what has moved since.

The fixture run the tests use holds invented numbers. A build renders it only when the BENCH_RESULTS_FIXTURE environment variable is 1, which is for working on the pages before a run exists, and every page of such a build is marked as fixture data. A build without the variable never reads the file.

Consistency checks
#

Beyond the shape of each section, parsing checks that the sections agree:

  • every library, language, set, and cell id is unique and refers to a listed entry, except the library a home input came from, which may be outside the run
  • no two cells measure the same input in the same mode, and no cell covers an untimed language
  • every cell names a listed input, describes it as that entry does, and is named after its file and mode
  • an input written for the benchmark has no provenance, a verbatim copy has the hash its one pin records, and a derived stress input names its derivation, and one built from several files is marked synthetic
  • a run with timed numbers records its anchor, when it started, and its machine, Node version, and commit, and how its numbers were measured
  • token counts are on tokenize cells only and output sizes on html cells only, and a library with a timed value in a cell has that cell's deterministic number
  • a value's 10th percentile is at most its 90th, and an anchor is stable exactly when its drift is under the threshold
  • a run's filters describe the file: every cell and cold-start entry is inside them, the libraries listed are exactly the ones filtered to, and no id repeats
  • a smoke run carries no noise floor, allows no retries, and is one segment
  • no value has more attempts than the run's retries allow
  • a value has a pass median for each of the run's passes, and the process noise is there exactly when a cell has a value and the run measured more than one pass, counting every pair of passes of every value, with its median at most its 95th percentile and that at most its largest
  • a run with no filters has no gap, as above
  • each startup entry carries the fields its scenario needs and no others, for languages the library claims, and appears once, as a number or as an exclusion and never both
  • a lang or first_highlight entry names a timed language. A set entry names a feature set, which may hold an untimed one
  • every startup entry keeps the samples the run says, its median lies between its 10th and 90th percentile, and a calibration has none
  • how cold start was sampled is recorded exactly when the run measured any, and is null for a calibration and for a run whose metrics filter leaves it out
  • a library has no numbers for a language its coverage doesn't claim
  • a library listed as excluded from a cell claims that cell's language and has no numbers there
  • a bundle entry holds sizes only when the library claims every language of the set, and otherwise names exactly the languages it lacks
  • bundle sizes cover every library over every set or are absent altogether, and what they were built with is recorded exactly when they are present
  • the install footprints and the HTML's compressed sizes come with the bundle sizes: every library of the roster has a footprint, and every library with an output on an html cell has its compressed sizes, exactly when the bundle sizes are present. A footprint has at least a file for each package
  • the retained heaps come with the bundle sizes too, for every timed language each library claims and no other
  • a compressed size is no larger than the raw one plus a compression format's own overhead, the HTML's included, and the theme CSS sizes are null exactly for a library that writes its styles inline
  • a file with no timed numbers is complete: in every cell each library that claims the language has the number the cell states or is listed as excluded, and every listed input has a cell in each mode

Use parse_bench_results to validate a file and get every problem in one message.