Export Performance Envelope
Export Performance Envelope
Section titled “Export Performance Envelope”What a large export actually costs, measured — not estimated. The numbers below come from two benchmark tiers that are reported separately on purpose, because they answer different questions and conflating them would overstate what has been verified.
This page is the reference for capacity planning and the precondition for any future chapter-streaming work: you cannot argue that streaming is needed without a measured baseline to compare against.
On this page
Section titled “On this page”- The two tiers
- How the numbers are measured
- Measured envelope: engine tier
- Measured envelope: end-to-end tier
- What the numbers mean
- Reproducing a measurement
- Trend tracking in CI
- Troubleshooting
- Related topics
The two tiers
Section titled “The two tiers”| Engine tier | End-to-end tier | |
|---|---|---|
| Script | scripts/bench/run-bench.ts |
scripts/bench/run-e2e-bench.ts |
| Input | Pre-parsed ExportBlock[] |
Confluence storage XHTML |
| Exercises | composeChapters, DOCX serialize + zip, PDF serialize + real Typst WASM compile, image embedding |
everything in the engine tier plus the tree walk (fetchExportTree), per-page storageToBlocks, and the full MacroRendererRegistry resolver pass (live Jira renderer, draw.io renderer, scroll-* macros) |
| Does NOT exercise | storage parsing, macro resolution, tree traversal, asset fetch over the network | real HTTP (auth, pagination, rate limits, retry), real attachment downloads, a browser host (heap limits, module-worker transfer cost) |
The Chromium-hosted variant (browser heap, module-worker transfer cost) is out of scope for both tiers today. It is an open question, not a silently dropped requirement.
How the numbers are measured
Section titled “How the numbers are measured”Wall clock
Section titled “Wall clock”Each phase is timed in-process around the phase work only. Setup (fixture generation, wasm
and font loading, template loading) sits outside every reported ms. Reported values are
the median of 3 runs, each in a fresh process.
The end-to-end tier additionally reports cold and warm: the same work run twice inside one process.
Peak RSS
Section titled “Peak RSS”/usr/bin/time reports the peak resident set size of one whole process. There is no
way to ask it for “the peak RSS of phase 2”, and an in-process process.memoryUsage()
sample is not a substitute — it measures the heap at an instant, is contaminated by the
previous phase’s uncollected garbage, and misses the wasm linear memory the Typst compiler
owns.
So the runners do this instead:
wholeProcessPeakRssBytes— one child process runs every phase; its peak is the run-level number.peakRssBytesper phase — each phase also runs in its own child process, so its peak is still a whole-process number, phase-scoped only because the process was.
A per-phase number therefore includes the runtime baseline, the fixture, and (for the
PDF phase) the wasm and font bytes. The baseline phase — load everything, do no work —
is measured so that floor is visible rather than guessed at. No number on this page is
derived from in-process heap sampling.
/usr/bin/time comes in two incompatible flavours: GNU (-v, kbytes) on Linux/CI and BSD
(-l, bytes) on macOS. Every record carries which one produced it as rssMethod, and the
trend comparison refuses to compare across them.
Measured envelope: engine tier
Section titled “Measured envelope: engine tier”Fixture: 500 pages, 3,590 blocks, seed 0x9e3779b9 — per page ~3 headings, prose, one
list; every 10th page a 200-row table; every 25th page a code block and an embedded PNG.
Machine: Apple M5 Max (18 cores), macOS darwin 25.4.0 arm64, Bun 1.3.8,
typst.ts 0.7.0 / Typst 0.14.2, RSS via bsd-time-l. Median of 3.
| Phase | Wall clock | Peak RSS (whole process) | Output |
|---|---|---|---|
baseline (load only, no work) |
— | 183 MB | — |
compose |
9 ms | 152 MB | 5.0 MB (block JSON) |
docx (serialize + zip) |
222 ms | 336 MB | 523 KB |
pdf (serialize + Typst compile) |
7,952 ms | 3,035 MB | 16.2 MB |
| whole run (all phases, one process) | 8,142 ms | 3,114 MB | — |
Measured envelope: end-to-end tier
Section titled “Measured envelope: end-to-end tier”Fixture: the same 500 pages generated as storage XHTML, yielding 4,779 composed blocks.
Adds scroll-title/scroll-pagebreak macros (every 8th page), a resolvable Jira JQL macro
(every 12th), and a draw.io macro that settles on the placeholder floor (every 20th).
Machine: identical to above. Median of 3.
Read the Cold column as the export cost. Warm† is a cache-hit repeat of identical input, not throughput — see the warning under Wall clock.
| Phase | Cold | Warm† | Peak RSS (whole process) | Output |
|---|---|---|---|---|
baseline (load only, no work) |
— | — | 187 MB | — |
fetch (tree walk + storageToBlocks) |
33 ms | 26 ms | 197 MB | 5.1 MB |
resolve (macro registry pass) |
6 ms | 3 ms | 200 MB | 5.1 MB |
compose |
13 ms | 7 ms | 221 MB | 5.1 MB |
docx (serialize + zip) |
236 ms | 126 ms | 441 MB | 546 KB |
pdf (serialize + Typst compile) |
7,677 ms | 577 ms | 3,351 MB | 17.1 MB |
| whole run (all phases, one process) | 8,127 ms | — | 3,400 MB | — |
† Warm re-runs byte-identical input, so Typst’s incremental cache absorbs most of the work. It is a lower bound on one-time setup cost, never a steady-state export time.
What the numbers mean
Section titled “What the numbers mean”PDF compilation is the entire cost. It is ~94% of wall clock in both tiers. Optimising anything else — the parser, the resolver, compose — cannot move the total meaningfully at this scale. Work aimed at export speed should target the Typst compile or the document size handed to it.
The parsing/resolver gap is real but small — at this fixture. The three phases the
engine tier cannot see (fetch, resolve, compose) total ~52 ms of ~8.1 s: well under
1%. That is an empirical result for this corpus, not a licence to assume the engine tier
is always a good proxy. It notably does not hold for memory — the end-to-end tier peaks
~9% higher (3,400 MB vs 3,114 MB) because the parsed tree is resident alongside the
compile — and it says nothing about a real tenant, where network latency and attachment
downloads dominate everything measured here.
Peak RSS exceeds the planning placeholder. Spec 011’s provisional budget was
“500-page PDF compile < 2 GB RSS”. The measured figure is 3.0–3.4 GB. The placeholder
was a guess made before measurement; this page is the measurement. Any absolute budget
frozen into scripts/bench/budgets.json must start from these numbers, and a
memory-constrained host (a browser tab, a small CI runner, a container with a 2 GB limit)
should be assumed to fail a 500-page PDF export until proven otherwise.
The warm PDF number is a lower bound, not a second-export cost. Cold 7,677 ms vs warm 577 ms is a ~13× gap, most of which is Typst’s incremental compilation cache hitting on byte-identical input. Exporting a genuinely different document will not be 577 ms. Read the warm column as “how much of the cold number is one-time setup”, not as “steady-state throughput”.
Reproducing a measurement
Section titled “Reproducing a measurement”# Engine tier: 500 pages, median of 3 runs per phasebun run bench:engine -- --pages 500 --repeat 3
# End-to-end tierbun run bench:e2e -- --pages 500 --repeat 3Records land in scripts/bench/out/bench-engine.json and bench-e2e.json (both
gitignored). Each record carries the fixture digest, the Typst wasm digest, the font-set
digest, OS/arch/runner, Bun version, and rssMethod, so a number is always attributable to
an environment.
Useful flags:
| Flag | Default | Meaning |
|---|---|---|
--pages <n> |
500 |
Fixture size |
--seed <n> |
0x9e3779b9 |
PRNG seed; same seed → byte-identical fixture |
--repeat <n> |
1 |
Runs per phase; the median is reported |
--phase <name> |
— | Run one phase in this process (used internally by the parent) |
--out <path> |
scripts/bench/out/… |
Where to write the record |
A quick smoke run:
bun run bench:engine -- --pages 30 --repeat 1Trend tracking in CI
Section titled “Trend tracking in CI”.github/workflows/bench.yml runs both tiers nightly (and on manual dispatch). It is
continue-on-error: true and never runs per PR — CI wall clock is too noisy to gate
merges on, and a gate people re-run until it passes is not a gate.
The workflow restores the previous records via actions/cache, compares each phase against
the rolling median of comparable prior records, and emits a GitHub ::warning:: when
wall clock regresses >20% or peak RSS >15%.
“Comparable” is strict: same tier, page count, fixture digest, Typst wasm digest, font-set
digest, OS, arch, runner, and rssMethod. A deliberate compiler or font bump therefore
resets the baseline instead of firing a false alarm — and a warning names the phase, the
tier, and the machine, so a regression localizes rather than showing up as an unexplained
total.
The runner field identifies a runner class, never an instance. In CI it is built from
RUNNER_ENVIRONMENT, RUNNER_OS and RUNNER_ARCH (e.g. ci:github-hosted:Linux:X64);
locally it is the hostname. This matters more than it looks: GitHub-hosted runners set
RUNNER_NAME to a per-instance value (GitHub Actions 2, GitHub Actions 14, …) that
changes between runs, so keying on it would make every nightly non-comparable with every
other nightly — the trend would report “no comparable history yet” forever and never warn,
silently, behind continue-on-error: true. Set ATLCLI_BENCH_RUNNER to declare your own
class label on a self-hosted bench rig.
Absolute budgets stay unfrozen until roughly two weeks of trend data exist. Only then
should scripts/bench/budgets.json be written and the workflow flipped to failing.
Troubleshooting
Section titled “Troubleshooting”Peak RSS is null in the record.
Neither /usr/bin/time -v (GNU) nor /usr/bin/time -l (BSD) was available. The runners
report rssMethod: "unavailable" and record null rather than inventing a number. Timing
is unaffected. On Debian/Ubuntu, install time (the shell builtin is not enough).
A phase child fails with “Cannot find module ‘@atlcli/…’”.
Phase children need in-repo workspace resolution. Run through bun run bench:engine /
bench:e2e, which pass --conditions=development; invoking bun scripts/bench/run-bench.ts
directly without that flag resolves the packages’ unbuilt dist barrels.
Numbers swing between runs on a laptop.
Thermal throttling and background load both move wall clock by tens of percent. Raise
--repeat, and treat cross-machine comparisons as invalid — that is exactly what the
environment fingerprint on each record exists to prevent.
The nightly warns after a Typst or font bump.
It should not: a changed compilerWasmDigest or fontSetDigest makes prior records
non-comparable, so the trend restarts. If a warning does appear across such a bump, the
digest was not recorded — check the environment block of the record.
The nightly always says “no comparable history yet”.
Every run is being treated as a new environment. Compare the environment blocks of two
consecutive records and find the field that moved — most likely runner, if something
reintroduced an instance-specific value there (see above), or datasetDigest, if the
fixture generator stopped being deterministic.
Related topics
Section titled “Related topics”- PDF Export Engine — the compile path these numbers measure
- DOCX Export Engine — the serialize + zip path
- Export Asset Contract — how wasm and fonts are loaded
- Public Export API (v1) — the seams the benchmarks drive