overview #
syntax-highlighter-bench is a benchmark suite and results site for JS syntax highlighters. It compares fuz_code, Twinkleplop, Prism, and Shiki, with Shiki's JS engine and its Oniguruma wasm engine counted as two libraries.
A run produces one results file holding many metrics, and the site is a view over that file where a reader picks the libraries, languages, and metrics to compare: the results page.
AI disclosure: this is an LLM-generated repo guided by a person.
What exists #
The harness that times the libraries, warm and from a cold start, the deterministic measurements, and the pages that render a results file are built.
- an adapter for each library, and the check that guards against a plain-text fallback, in
bench/libraries/andbench/check_output.ts— see adapters - the inputs and their manifest, in
bench/corpus/— see corpus - the language ids and the feature sets, in
bench/langs.tsandbench/sets.ts— see languages and sets - the timed harness, which measures each library alone in a process of its own and writes a
results file, in
bench/run.tsandbench/cell.ts— see method - cold start, part of the same run, where each sample is a fresh process that imports a
library and stops its own clock, in
bench/coldstart.ts— see method - the deterministic measurements, which are bundle size over each feature set, install
footprint, retained heap, token counts, output sizes, and coverage, in
bench/deterministic.ts, with their numbers committed asresults/deterministic.json— see deterministic metrics - the schema of the results file, in results_schema.ts — see results file
- the landing page and the results page, which render one validated results file, in
src/routes/andsrc/lib/— see results file for which file, and design for what is still to come
What is measured #
- Throughput, on one named machine: how long a warm library takes to tokenize an input, and to render it as HTML, with the peak memory of its process beside it.
- Cold start, on the same machine: how long a fresh Node process takes to import a library alone, with one language, and with a documentation site's languages, and to produce a first highlight. Each is measured for the bundled chunk and for the library imported unbundled, beside a bare Node baseline to subtract, with the process's peak memory.
- Deterministic numbers, the same anywhere: bundle size over each feature set, install footprint, retained heap, token counts, output sizes, and coverage.
Layout #
bench/— the harness. It runs on Node directly and is never imported by the site.results/— where the harness writes runs. The deterministic results are committed there, and a published timed run goes beside them asresults/latest.json.src/lib/results_schema.ts— the schema the harness and the site share, withsrc/lib/bench_constants.ts, the id lists it builds its enums from. The rest ofsrc/lib/turns a results file into the tables of the site.src/routes/— this site, built as static pages.
Neutrality #
The maintainer of this benchmark also maintains fuz_code, one of the measured libraries, so fairness has to come from the structure rather than from trust:
- the maintainer's library is flagged in the results file and on the site
- every input is labeled by its source, and real code that no measured library authored is the primary set; a snippet or a stress input written for this benchmark is labeled as that, never as neutral, and no stress input is tuned to a library
- each library is driven through one adapter with notes on the entry points it uses, so its maintainers can review or replace it
- the method and the commands to reproduce a run are published with the numbers