method #
The timed harness measures throughput: how long each library takes to tokenize an input, and
to render it as HTML. It then measures cold start: what a library costs a process before it
has highlighted anything. bench/run.ts runs both and writes one results file.
The method builds on the comparison harness of Twinkleplop (MIT), and the anchor workload is ported from it. It differs in giving each library a process of its own, where Twinkleplop measures a cell's libraries together in one. The cold-start scenarios generalize Twinkleplop's cold-start harness across libraries.
Cells #
A cell is one input of the corpus in one mode, tokenize or html, measured across the libraries that claim its language. Its id is <file>:<mode>. A cell carries its input's source and size tier, so a
reader can compare on the neutral inputs alone, or see whether a result holds on a library's
home turf or on the stress inputs.
One process, one library #
A process measures one library on one cell and loads no other. Every number is therefore one library alone in a fresh process, and the roster can't move it: adding or removing a library changes no other library's result.
Two effects made that necessary, and both showed up as a library disagreeing with itself:
- Libraries that share code change how V8 optimizes that code for each other. Shiki's two engines share its tokenizer, and beside the wasm engine Shiki's JS engine measured far slower than it does alone.
- V8 sizes its young generation from recent allocation. In a shared process a library's time on a large input depended on which libraries ran beside it.
A cell is measured in passes. A pass is a process for every library, one after another, and the library a pass starts with changes from pass to pass and from cell to cell. The same code can settle at a slightly different speed from one process to the next, so a library's number pools the rounds of all its passes.
Node runs with its default flags, which is what a consumer gets. In a timed run one process runs at a time, and the parent does nothing while it measures: it blocks on the child and reads its output after it exits.
Inside a process #
Each step answers a way a simple timing loop goes wrong:
- One language loaded. The library sets up exactly the cell's language, once, before anything is timed. See adapters for what that means for Markdown.
- The output is checked first. The library must pass the output check on the language's harness snippet and then on the cell's input. A library that fails is recorded as excluded with the reason, and is not timed.
- Iterations calibrated for the library. Speeds differ by orders of magnitude, so a batch makes the number of calls that fills the target duration. A batch makes at least a few calls while they fit a cap, as many as fit when fewer do, and one when two no longer fit.
- A rewarm and a minor collection before every timed batch. The library runs untimed until it is back to speed, and a minor collection then clears what the rewarm allocated, so the collection it would cause doesn't land inside the batch. No full collection is forced between batches. One forced before every batch puts some libraries through a race each round: it can discard optimized code whose hidden classes only short-lived objects hold, so the library re-optimizes inside the clock, and it re-rolls how memory a library allocates outside the JS heap is reused, so a process settles at one of two speeds. A minor collection does neither.
- Results are kept and counted. A batch keeps its last result, and after the clock stops that result is counted. A count that differs from the checked one fails the cell: a tokenizer that silently stops early on a slow line must not be averaged in as a fast one.
- The HTML is read. Some libraries return a string the engine hasn't joined
up yet, which costs little to build, and every consumer then pays to flatten it on first
use. In
htmlmode the timed call reads one character of the string, for every library, so that cost is inside the clock. It depends on the size of the output, not on the library, so it weighs most on the fastest one, and it is free for a library whose string is already flat. Intokenizemode every library returns finished arrays, so there is nothing to force.
The forced minor collection needs the process started with --expose-gc, which bench/run.ts passes to each process.
A full run measures each library in 3 processes per cell, each timing 4 rounds of 20ms batches, after a 100ms warmup, with a 300ms rewarm before each batch. A library whose one call takes at least as long as the rewarm gets none: the batch before keeps it warm, and a rewarm would only repeat that call.
The rewarm is longer than the warmup on purpose. The engine still runs full collections of its own during a run, and one can leave code cold. Libraries recover from that at different speeds, Shiki more slowly than the others, and a library still recovering when its batch is timed reads slow, so the rewarm is long enough that every library in the roster has recovered.
What each number means #
Each library's numbers in a cell are a BenchCellValue:
| number | meaning |
|---|---|
ns_per_op | nanoseconds per call: the median across the rounds of every pass |
ops_per_sec, mb_per_sec | the same time as calls per second, and as megabytes of input per second, which is the one to compare across inputs of different sizes |
p10_ns, p90_ns | the 10th and 90th percentile across those rounds |
spread | the slowest round over the fastest, across passes as well as within them. Near 1 means they agreed, and a high spread says the cell was disturbed or the processes disagreed |
pass_medians_ns | each pass's own median over its rounds, in pass order: one number per process, so how far the processes disagreed can be read apart from the rounds |
iterations | calls per batch. Each process calibrates for itself, and this is the median across them |
rss_peak_bytes | the peak resident size of the library's process, the largest across its passes |
attempts | how many times the library was measured on the cell: 1 unless it was retried, as described below |
The token count and the size of the HTML come from the output check, under the same setup and
outside the clock. Libraries split the same source differently, so a token count is the
library's own and is recorded rather than equalised. Each is stated once for an input: the
token count on its tokenize cell and the output size on its html cell. deterministic metrics describes them, and the bundle sizes a run carries.
rss_peak_bytes is a whole process: Node itself, the harness, the library as its
adapter imports it, with one language set up, and the input. It is not what the library adds
to a page, and most of it is the same for every library. Compare it between libraries on the
same cell. On a small input it says little about footprint: a process warms up and is timed
for a fixed duration, so its peak follows how much the library allocated in that time more
than what it holds. What a library holds is the retained heap, described in deterministic metrics.
Cold start #
A timed cell measures a library that is already loaded and warm. Cold start measures the other end: a fresh Node process that loads the library and does one thing, once. Each sample is one process, and each BenchStartup entry is a scenario:
| scenario | what the process does |
|---|---|
bare | nothing: the Node baseline, with no library, which every other time includes |
core | imports the library with no language |
lang | imports the library with one language, for each timed language it claims |
set | imports the library with the languages of the docs feature set: what a
documentation site loads |
first_highlight | imports the library with one language and highlights that language's harness snippet once, to HTML: the time to a first highlight, for each timed language it claims |
What is imported is the module the adapter's bundle_entry writes
for a feature set (core, single:<lang>, docs):
the same module bundle size is measured over, so a time and a size describe one thing. The
adapter modules themselves are never loaded here. Some of them import every language, which
would charge a library for languages the scenario doesn't load. A first highlight makes the
call the adapter documents for that module (bundle_html_source), and the check of
every built bundle confirms it returns the same HTML as bundle_html.
Two forms of every scenario, apart from bare:
bundled— the built chunk whose size deterministic metrics reports, imported as one file. A library's wasm import stays outside its bundle, as it does for the size, so that one module still loads from the package.- unbundled — the same module's source, written to a file under
node_modules/.cache/in the checkout and imported from there. Node resolves each package from the checkout'snode_modulesand loads every file the library reaches, as a consumer without a bundler does.
The gap between the two is Node finding, reading, and compiling many files in place of one, and code the bundler removed, which a bundled page or server doesn't pay. Read the form that matches how the library would be shipped.
Start is the start of the process, by its own clock. The process reads performance.now() once, when its work is done, and Node counts that from the
moment the process began. So a time includes Node itself starting up, and leaves out what
happens outside the process: the parent creating it, the OS loading the Node binary, and the
exit. A clock in the parent would add those, with their own variance. The bare scenario is the same process with nothing imported, so a scenario's time less bare's is what the library added.
The process is a consumer's. It runs a generated module holding the import,
the one call of a first highlight, and a few lines that report, with no harness code and no
flags: no --expose-gc, which only the timed cells use. The source a first
highlight is given is a literal in that module, so no file is read inside the clock, and one
character of the HTML is read before the clock stops, as the timed cells do, so a string the
engine hasn't joined up is paid for.
The compile cache is off. V8 can keep compiled code on disk between
processes, and Node uses that only when asked. A second process would then skip compiling what
it loads, and would not be cold. Every cold-start process runs with NODE_DISABLE_COMPILE_CACHE=1 and without NODE_COMPILE_CACHE, each
sample reports that no cache directory is in use, and startup_compile_cache in BenchRun records it.
Interleaved rounds. A round runs every scenario once, one process at a time, with the parent blocked on each. The scenarios that differ only in the library run one after another, and the library that goes first rotates from round to round and from scenario to scenario. Because every round holds every scenario, a machine that slows down partway slows a sample of each of them, and not the scenarios that happened to come last.
The first rounds are discarded. They bring every file a scenario reads into the OS page cache, which a machine pays once and not at every start. A full run keeps 20 samples of each scenario, after 3 discarded rounds.
Each scenario is checked by its first process. The import has to work and the
module has to export something, and a first highlight's HTML has to pass the same check on the
harness snippet that a timed cell applies. A library that fails is recorded in startup_excluded with the reason, and has no number there: this is how a library
that can't be imported outside a bundler would show. A process that was ended, by a signal or
an abort, did not fail on its own: that fails the run, which can be resumed, and excludes
nothing. A scenario over languages a library doesn't claim is unsupported, and is never run.
Every later sample must agree with the first, and when the run carries the deterministic
results, a first highlight's HTML must be the size they have for that library on that snippet.
| number | meaning |
|---|---|
samples | how many processes were kept |
median_ms | the median across them, in milliseconds since the process started |
p10_ms, p90_ms | the 10th and 90th percentile |
spread | the slowest sample over the fastest. A process start has a long slow tail, so this is wider than a cell's, and the percentiles say more |
rss_peak_bytes | the median peak resident size of those processes, read when the clock stopped |
Memory. A cold-start process has loaded the library and done nothing else, so
its peak resident size is closer to a footprint than a timed cell's. It is still a whole
process, most of it Node: subtract the bare entry's. What is left is the code and
data the library loaded, and for a first highlight what one call allocated, at the sizes the
engine chose for a process of that age. The retained heap, described in deterministic metrics, is the steadier figure for what a library holds: read after
collecting garbage, less a bare process's, and the same from one process to the next within a
small tolerance.
Nothing is retried. A disturbance lands on one sample of every scenario, not
on one scenario, and the median over many processes sets a slow one aside. Sampling one
scenario again later would also take it out of the rounds that make the scenarios comparable.
There is no noise floor for cold start either: a calibration measures the timed cells only.
Without one, treat a difference between two scenarios that is inside either one's p10_ms to p90_ms band as no difference.
Part of the run. The cold-start processes follow the cells in the same run: under the same machine lock, between the same anchor readings, in the same journal, and written to the same results file. A smoke run takes one sample of every scenario.
Kinds of run #
A results file says what kind of run produced it, in BenchRun, with the passes, the rounds, the durations, and any filters:
full— the real measurement. Only a full run with no filters is publishable, and such a file must be whole: every library has a number or an exclusion in every cell whose language it claims, and in every cold-start scenario it supports, in both forms.smoke— every library on every cell, in one process each, for one round of the shortest batches, then one sample of every cold-start scenario, with several processes running at a time. It proves the harness and every adapter work, and its numbers mean nothing.calibrate— measures the machine's noise floor, described below. It measures the cells only, and no cold start.
npm run bench -- --machine <name> # a full run: results/<date>_<name>.json
npm run bench:calibrate -- --machine <name> # the noise floor: results/calibration_<name>.json
npm run bench:smoke # everything once, briefly and in parallel
node bench/run.ts --langs ts --sizes small # part of a run, to results/scratch/
node bench/run.ts --metrics startup # cold start alone, to results/scratch/
npm run bench -- --machine <name> --plan # what a run would do, measuring nothing
npm run bench -- --machine <name> --resume # continue a run that was stopped
node bench/run.ts --help | flag | what it does |
|---|---|
--smoke | a smoke run |
--calibrate | a calibration run |
--libraries, --langs, --sizes, --modes, --sources | measure only these, as comma-separated ids. An unknown id is an error, not an empty run. --libraries and --langs narrow the cold-start scenarios too;
the other three select inputs, which cold start has none of |
--metrics | measure only throughput (the timed cells) or only startup (cold start). Both by default. Naming both is no filter, and neither is throughput on a calibration, which measures nothing else |
--passes, --rounds, --target-ms | how many processes measure each library on a cell, how many batches each times, and how long a batch aims to run |
--retries | how many times a library whose rounds on a cell disagreed is measured again there. 0 turns it off |
--startup-samples | how many fresh processes sample each cold-start scenario, after the discarded rounds |
--resume | continue the run its journal records, which must be this same run |
--plan | print what the run would measure, where it would write, and an estimate of how long it would take, and measure nothing |
--machine | the name the machine is recorded under. Required for an unfiltered full or calibration run; otherwise the hostname |
--out | where to write the results file instead |
--help | print the flags |
A full run with no filters is written to results/, named by the UTC date it
started and its machine, and it won't overwrite a run already there. A smoke run or a filtered
one is written to results/scratch/, which git ignores, so it can't be committed
as if it were a result. A cell process that fails, or any process ended by a signal, fails the
run, and no results file is written; the journal keeps what was measured.
--plan makes every check a run makes before it measures, then prints the cells,
libraries, and processes, the cold-start scenarios with the processes they add and any a
library is unsupported for, the parameters, where the results and the journal would go, and an
estimate of the time, worked out from the parameters and not measured. It is refused where the
run would be.
Stopping and resuming #
A full run is one process for each library, pass, and cell, then one for each sample of each
cold-start scenario, one at a time, so it is long, and a calibration's cells take twice as
long. A timed run keeps a journal beside its results file: each process's result is appended
and synced to disk before the next process starts. The parent writes only between processes,
never while one is measuring. When a run stops, whether by a signal, a failed process, or a
crash, the journal stays, and --resume continues the run from it.
- What is recorded. A first line describing the run, then for each stretch of measuring (a segment) a line with its opening anchor reading and load average, one line for every process result, a cell's or a cold-start sample's, and a line with its closing anchor reading.
- What must match. The kind of run, its parameters and filters, the corpus
hash, every library's version and extra packages, the machine's name, CPU, threads, and
frequency settings (governor, driver, energy-performance preference, and boost), the OS
release, the Node version, the
NODE_OPTIONSthe processes inherit, and the commit with its dirty marker. A resume with any of them different is refused, naming what differs, so results from two setups are never mixed. With a dirty working tree the journal can't tell whether the uncommitted changes are the ones it started with, and a resume says so. - Nothing installed mid-run. Every cell process reports the versions of its
library's packages it found installed, and one that differs from what the run was planned
with stops the run, naming the package and both versions, before its result is journaled.
Cold-start processes report none, so the run reads the versions again before writing its
file, and on a change writes nothing and sets the journal aside as
<journal>.changed: its cold-start samples were never checked against the change, so the run starts over. - Never by accident. A journal already there stops a run that wasn't given
--resume: it is neither overwritten nor continued until you say which, and the refusal says what run the journal records. A journal holding a result the run would never ask for is refused too. - Never after an overlap. A run that lost the machine lock sets its journal
aside as
<journal>.overlapped, since another run may have overlapped it. It writes no results, and the journal can't be resumed. - What is skipped. The processes the journal holds are not run again. The rest run in the order planned, and the results are assembled exactly as an uninterrupted run assembles them.
- What the file says.
segments, in BenchRun, counts the stretches the run was measured in: 1 for an uninterrupted run.generated_at, and the date in a full run's file name, are when the first segment started.duration_msis the time spent measuring, without the time between segments, andload_at_startthe highest of the segments' load averages at their start. - A crash in the middle of a write. A last line with no newline is a write that didn't finish. It is dropped, and that process is measured again. A complete line that isn't a journal line is not a crash, and the journal is refused.
The journal is removed once the results file is written. A smoke run is short and keeps none.
Retries #
A library whose rounds on a cell are far apart was disturbed by something: another program, or
its own processes disagreeing. Once a cell's passes are done, a timed run measures each
library whose spread there is over 1.25 again, every pass of it, up to --retries more times (2 by default) and until an attempt is under the threshold.
The attempt kept is the one with the lowest spread, the earliest if they tie. It is never chosen by its time. Keeping the fastest attempt would pull every retried number toward fast, and would favor a noisier library, which is retried more often. Choosing on agreement alone takes the measurement that was least disturbed, whichever way it moved the number.
attempts records how many times a library was measured on a cell, so the numbers
that needed a retry can be told from the rest. A retried value's spread and percentiles are
those of the attempt that was kept, so they describe its calmest attempt, not everything that
was measured. The summary a run prints counts the retried cells of each library, and lists the
library-cells still over the threshold after the last attempt. What should be the same every
time (the token count, the output size, and whether the library passed the check) must agree
across attempts as it must across passes, or the run fails at that cell.
A smoke run doesn't retry, and cold-start scenarios are never retried.
The noise floor #
Before a machine's numbers are trusted, it is calibrated. A calibration measures every library as if it were two libraries: each gets twice the passes, split alternately into two arms, and each arm is exactly what a full run would measure. The two arms should be equal. How far apart they are is what the machine and the method produce when nothing differs, including what varies from one process to the next. BenchNoiseFloor holds three numbers:
per_cellis the 95th percentile of that difference, across cells and libraries. A difference between two libraries in one cell that is smaller than this is not a result.geomeantakes each library's two geometric means over all cells, and is their difference for the library where it was largest.per_processis the 95th percentile of the difference between two passes of one arm, over every pair of passes, arm, cell, and library: how far two single processes measuring the same thing come apart. It reads higher thanper_cell, which compares values that each pool several processes.
A difference is the larger time over the smaller, less one, so it reads the same whichever arm was slower.
A calibration is written to results/calibration_<machine>.json. A later
full run copies its noise floor if the calibration describes that run: unfiltered, with a
stable anchor, and on the same machine name, CPU, threads, and frequency settings, with the
same Node version, corpus, and parameters, retries included, and every library at the same
version with the same extra packages. The commit is not compared, so a commit that changes
nothing measured keeps the floor. Otherwise the run records no noise floor and says why, and
the site won't publish it.
A calibration retries too, each arm of a library on its own and by the same rule, so an arm
stays exactly what a full run would measure and the floor describes a full run with the same
retries. The rule looks only at an arm's own rounds, never at how far apart the two arms are,
so it can't steer the floor. A calibration file's values are the first arm's, and so are their attempts.
Every run with more than one pass also measures its own process noise, an A/A figure taken
from the run itself: each library on each cell was measured by several processes that should
agree, and every pair of a value's pass_medians_ns gives a difference. BenchProcessNoise records how many pairs there were and their
median, 95th percentile, and largest difference, as meta.process_noise. Pairs are
only ever taken within one library on one cell. Its 95th percentile is the run's counterpart
of the calibration's per_process, and the run's summary and the provenance line
on the results page print the two side by side. Both are the 95th percentile of one sample, so
either is higher about as often as not: a run's figure several times the calibration's is what
says its processes disagreed more than the calibration's did. The reader compares them, and
the site does not refuse a run over it.
The run this site shows was measured on laptop1, whose floor is 2.3% for one cell and 0.3% for a geometric mean. The results page marks a ratio inside it.
The anchor #
A run measures a fixed reference workload before its first process and after its last. The
workload involves no measured library, so a change between the two readings is the machine
changing state: heat, frequency scaling, another program. drift is the relative
change, in BenchAnchor, and a run is stable while it
stays under 3%.
The anchor only brackets the run. Something that disturbed the middle of it shows as a high spread, and the summary a run prints lists every library whose rounds on a cell
disagreed.
A run that was stopped and resumed takes a reading at the start and end of every segment, and its drift is that of the two readings furthest apart. A machine that changed state between segments, or within one, is then flagged and not hidden by readings that happen to agree at the two ends. A segment stopped by a signal or a failed process still takes its closing reading. One that was killed outright can't, and the next segment's opening reading stands in for it.
Before it spawns anything, a run reads the machine's load average over the last minute and
records it as load_at_start. On a quiet machine it is well under 1, and a timed
run warns when it is over 1, since something else is then running. It warns and goes on: the
anchor, the spread, and the retries are what judge the run.
One run at a time #
Two measurement runs at once make each other wrong, not just noisy. A run takes a machine-wide
lock, the file /tmp/fuz_benchmark.lock that names its process, and holds it for
as long as it measures, as does each segment of a resumed run. The path is fixed and ignores TMPDIR, and the lock is fuz_util's benchmark lock, so any other benchmark that
uses it takes the same one and never measures at the same time as this one. A second run is
refused and told who holds the lock. A lock whose process is gone is taken over.
Caveats #
- Features vary. The libraries have different features. Shiki highlights with the TextMate grammars and themes VS Code uses; Twinkleplop aims for a similar feature set, with its own themes, annotations, and Twoslash support; and both do more than fuz_code, which has the fewest features of the four: a single-pass lexer per language, with CSS classes for its theme. Their fidelity differs too: how finely each splits the same source, and how accurately it classifies each part. These numbers compare the work all of them share, not what one offers that another doesn't, and they time each library's output without scoring its quality.
- The languages are fuz_code's. The languages measured are fuz_code's built-in set, so the list favors it. Prism and Shiki each support hundreds of languages, and none of that shows here.
- Classes and inline styles. Shiki resolves colors while it tokenizes and
writes them inline, and the other libraries emit class names that a stylesheet colors. The
htmlmode measures the HTML each library returns by default, which is not the same work. - Token counts differ. A library that emits fewer tokens for the same source may be doing less work for each byte, not the same work faster.
- Synchronous paths. A library with an async API loads in its setup and is timed through its synchronous core.
- The HTML is read. In
htmlmode the timed call reads one character of the string, so a library that returns a string the engine hasn't joined up pays to flatten it inside the clock. A benchmark that discards the string doesn't charge for that, so these numbers can differ from a library's own. - Markdown without its fences. A timed process loads only the language it measures, so a fenced code block inside Markdown is not highlighted, with the exceptions adapters lists. A site that loads several languages pays more for Markdown than these cells show.
- Home turf. Some inputs are a measured library's own samples, which it was developed against. They are labeled with that library wherever they appear, summarized by owner, and left out of the default view and the headline figures, which use the neutral inputs.
- Stress inputs. Some inputs are built to find each library's worst case: minified code, deep nesting, very long lines, constructs that never close. They say little about typical code, so they are summarized on their own and left out of the default view and the headline figures. A slow call there can be far slower than anywhere else, and a run shows it as measured. Shiki is called with its per-line time limit out of reach, as everywhere in this benchmark, so a stress number is the cost of highlighting every line in full. With Shiki's default limit (500ms per line), a line that takes longer comes back with its rest as one plain token: a consumer with default options gets a fallback on that line, not this number.
- One bundler's sizes. A bundle size is what a page ships when Vite, with the Rolldown inside it, builds it and esbuild minifies it; the results file records the three versions. Another bundler can remove different code from the same entry, since what it can prove unused, whether it folds a constant argument into a function, and how it wraps CommonJS all vary. Built by Rollup, which Vite used before Rolldown, with the same options, the bundles differ unevenly: Rollup removes Twinkleplop's diagnostics code, which it proves is never called, and Rolldown keeps it; Rollup specializes a function fuz_code's Markdown language calls with a constant argument, and Rolldown doesn't; and Rolldown adds a small CommonJS interop helper to Prism's bundles.
- Synthetic inputs. A few inputs were built by joining or repeating files to reach a size, or derived from one by minifying it, and are marked. A repeated file is more regular than real code of the same size.
- Steady state. In a cell, every library is warm when it is timed. What a first call costs is in the cold-start scenarios.
- Cold start is Node's. A cold-start time is a Node process on one machine, with the files already in the page cache. A browser parses and compiles differently, and fetches over a network, where the bundle sizes matter instead.
- Unbundled depends on the install. The unbundled form loads the files a package manager laid out, so its time moves with the number of files a library ships and how deep they sit. It says what that library costs a Node process without a bundler, and nothing about a page.
- Collections. Every batch follows a rewarm (unless one call outlasts it) and a forced minor collection, and full collections happen only when the engine schedules them, so garbage from earlier batches carries into later ones and a full collection can still land inside a batch. A library pays for what it allocates (in the JS heap or outside it) as the engine collects it during the run, which is close to, but not the same as, how a page collects between highlights.
- Measured alone. A page that loads several highlighters, or a busy application, gives the engine a different history than a process holding one library.
Compared with Twinkleplop's benchmark #
Twinkleplop's own comparison harness is the one this method builds on, and the two answer different questions: it tracks Twinkleplop against other libraries across its own versions, and this one compares a fixed roster on shared inputs. Where they differ, as of Twinkleplop 0.3.1:
| Twinkleplop's | this benchmark | |
|---|---|---|
| libraries | Twinkleplop, Shiki (both engines), Prism, sugar-high, speed-highlight | fuz_code, Twinkleplop, Prism, Shiki (both engines) |
| languages | a broad set, most of what Twinkleplop supports | the timed languages every library here supports |
| inputs | generated inputs at three sizes per language, and Shiki's samples | pinned files no measured library wrote as the headline, with Shiki's samples, the libraries' own, snippets, and stress inputs kept apart |
| isolation | one process per cell, its libraries alternating within each round | one process per library and pass, in a rotating order |
| rounds | one pass of several rounds, a short rewarm, a full collection each round | several passes of a few rounds each, a longer rewarm, a minor collection only |
| statistics | median, minimum, p10 and p90, and a bootstrap interval | median, p10 and p90, spread, each pass's median, and the gap between processes |
| retries and noise | no retries; a calibrated noise floor for its A/B harness, not the comparison | a disturbed cell is measured again; a calibrated noise floor for the comparison itself |
| HTML timing | the returned string is discarded, so a library returning a rope never flattens it | one character of the string is read inside the clock, which flattens it |
| cold start | Twinkleplop alone, with scenarios about its own internals | every library, through the same bundle entries, in the results file |
| other metrics | none in the comparison | bundle size, install footprint, retained heap, token counts, output sizes |
| history | a history of published runs across versions | each published run stands alone |
| machine | a bare-metal server with boost off, the run confined to cores away from the system's | a laptop with boost, SMT, and ASLR left at their defaults |
Numbers from the two are not comparable with each other: the machines, the inputs, and the HTML timing differ.
Reproducing a run #
A results file records the commit, the Node version, the machine, the corpus hash, and every
library's version and extra packages. From a checkout of that commit, npm ci installs the same versions, and the command for the kind of run measures it again. The
deterministic numbers come out the same on any machine running the same Node version, the
retained heap within its tolerance, and a run stops if one of its processes disagrees with the
committed deterministic metrics. A full run with no filters is refused unless that
file describes it. Run the smoke run first: it checks every cell and every first highlight
against the file, before a long run can stop on a disagreement. The timed numbers depend on
the machine, so compare the ratios between libraries rather than the times.
A quiet machine matters more than a fast one: close other programs, and prefer a fixed CPU
frequency governor. Before trusting a run, check that its anchor is stable, that spread is near 1 for every library in every cell, and that nothing unexpected is
excluded, from a cell or from a cold-start scenario. bench/run.ts prints all
three when it finishes, with how many library-cells were retried and whether the run was
resumed.