Benchmarks¶
pixi run bench runs the criterion suite in benches/ and
writes a report to target/criterion/. pixi run bench-test runs every benchmark once without timing it, which
is what CI does: a shared runner's numbers say more about the runner than about the code, but a benchmark that no
longer compiles or panics on its inputs is a real break.
What is measured, and against what¶
Two lockfiles:
| Input | What it is | Packages in the document |
|---|---|---|
| reference | This repository's own pixi.lock — a real project's dependency graph, 289 locked entries across its environments |
72 for default / linux-64 |
| synthetic | Generated in the benchmark: 2000 conda packages in one environment, each depending on four of its predecessors, with a mix of license expressions | 2000 |
The synthetic one exists because the costs that matter at 2000 packages — the dependency graph, the license table, the writers' allocations — are invisible at 72.
| Group | What it times |
|---|---|
parse |
Lockfile text → parsed lockfile (rattler's YAML reader) |
model |
Parsed lockfile → the format-agnostic model: purls, the dependency graph, license normalization |
write |
Model → document bytes, once per format and spec version |
cyclonedx-stdout |
The same document written to a pipe rather than a file, because that is a different writer |
license/normalize |
One pass over eight expressions: a bare id, a compound, one needing a rewrite, one unparsable, and the empty case |
report |
Building and rendering a report, per kind and per output format |
enrich |
Reading conda licenses out of an extracted package cache — offline, from the recorded fixtures, so the number is the reading and parsing rather than the network |
Network enrichment is deliberately absent: a benchmark whose result depends on what PyPI felt like doing that morning measures the morning.
Numbers¶
Apple M4 (10 cores), macOS 26.2, rustc 1.98.1, release profile, criterion's median of its sample.
Reading and writing¶
| Benchmark | reference (72 packages) | synthetic (2000 packages) |
|---|---|---|
parse |
4.76 ms | 25.1 ms |
model |
255 µs | 6.84 ms |
write cyclonedx 1.6 |
454 µs | 12.3 ms |
write cyclonedx 1.7 |
452 µs | 12.3 ms |
write spdx 2.3 |
367 µs | 11.2 ms |
write spdx 3.0 |
525 µs | 15.1 ms |
Parsing dominates. On the reference lockfile it is 4.76 ms against 255 µs to build the model and under half a
millisecond to write the document: reading pixi.lock is roughly nine tenths of an offline run. That cost is
rattler's YAML reader, not this crate's code, and it scales with the whole file rather than with the environment
selected — which is also why --all-environments is much cheaper per document than the first one.
SPDX 3.0 is the most expensive writer (a JSON-LD graph, one node per element); SPDX 2.3 is the cheapest.
Reports and enrichment¶
| Benchmark | reference | synthetic |
|---|---|---|
report packages, table |
378 µs | 11.2 ms |
report packages, json |
44.0 µs | 1.06 ms |
report licenses, table |
174 µs | 7.28 ms |
report licenses, json |
46.9 µs | 1.08 ms |
report python, table |
2.83 µs | 5.37 µs |
report python, json |
1.03 µs | 3.19 µs |
report packages, tree |
112 µs | — |
license/normalize (8 expressions) |
1.53 µs | — |
enrich conda licenses from the package cache |
183 µs | — |
The table renderer costs an order of magnitude more than the JSON one at every size: that is comfy-table
measuring and wrapping every cell to fit a terminal, which the JSON form never does. The python report only
looks at packages that constrain the interpreter, so it barely grows.
A run that writes many documents¶
--all-environments --all-platforms on this repository writes 15 documents holding 853 package entries between
them, over 289 distinct packages. Cold cache, --fetch-licenses, measured end to end:
| wall | requests | |
|---|---|---|
| Before 0.10.0, ten at a time | 6.0 s | 325 |
Before 0.10.0, PIXI_SBOM_CONCURRENCY=24 |
4.17 s | 325 |
| Shared lookups, ten at a time | 5.08 s | 325 |
Shared lookups, PIXI_SBOM_CONCURRENCY=24 |
3.08 s | 325 |
The request count does not move, because the download cache already stopped the second document re-fetching what the first one downloaded. What changed is that the run no longer drains a small pool of requests per document before starting the next one: every document's packages are looked up together, so the pool stays full. That is worth 15% on its own, and it is what lets the concurrency setting pay off — the two together halve the run.
With a warm cache the whole batch takes about 0.05 s either way, so none of this is visible on a second run.
Memory¶
Documents are serialized straight from the writers' own structs to the output. Building a
serde_json::Value first — which is what happened before 0.10.0 — held the whole document twice: once as a tree
of boxed strings and maps, once as the bytes.
Peak resident set size writing one document from a generated lockfile, measured with /usr/bin/time -l:
| Lockfile | Format | Before | After |
|---|---|---|---|
| 2000 packages | CycloneDX 1.6 | 59.3 MB | 60.1 MB |
| 10000 packages | CycloneDX 1.6 | 227.6 MB | 151.5 MB |
| 10000 packages | SPDX 2.3 | 203.1 MB | 148.0 MB |
At 2000 packages the difference is inside the noise — the high-water mark there is the lockfile parser, not the writer. At 10000 it is a third of the peak. The output is byte-identical either way, which a unit test asserts by writing each format both ways and comparing.
Buffering, and a regression the benchmarks could not see¶
The writers serialize straight out in many small pieces. A file gets a BufWriter; std::io::stdout() is
line-buffered and flushes on every newline, so streaming into it unbuffered meant a syscall per line. Writing a
10000-package document took 0.39 s to a pipe against 0.17 s to a file until stdout was buffered by hand too.
Every benchmark scenario wrote to a file, so a 2.3× regression on --output - — a documented mode, and the one
people pipe into jq — survived a whole performance milestone unnoticed. There is now a cyclonedx-stdout
scenario for exactly that reason. When a change touches how bytes leave the process, the benchmark has to
exercise every way they leave it.
The license-text budget¶
--license-texts embeds the text of every licence file, and each file is capped at 1 MiB on its own. Across a
large environment that is still unbounded, so a document holds at most 64 MiB of licence text in total. Past
that the files are still listed by name, the run warns, and the document records it as
license-texts: N file(s) listed by name only in pixi:incomplete — the same bargain the per-file cap already
strikes, and visible in the document rather than silent.
The binary¶
Release profile: lto = true, codegen-units = 1, strip = true, panic = "abort", and mimalloc as the global
allocator. Measured on an Apple M4:
| size | 10000-package run | peak RSS | --version |
|
|---|---|---|---|---|
| 0.9.5 | 9.01 MB | 0.22 s | 151.5 MB | 2.6 ms |
panic = "abort" |
7.45 MB | 0.22 s | 151.5 MB | 2.6 ms |
| + mimalloc (shipped) | 7.61 MB | 0.16 s | 124.8 MB | 2.6 ms |
panic = "abort" drops the unwinding tables, which is 17% of the binary and costs nothing at run time: a panic
in a command-line tool ends the process either way, only now without unwinding first.
mimalloc is installed in the binary, not the library, so nothing that links pixi_sbom has an allocator
forced on it — which also means the criterion benchmarks above, which link the library, measure the system
allocator. The shipped binary is faster than they say.
What it costs, per platform. The Performance workflow measured mimalloc on its own, comparing the commit that added it against its parent on each runner:
| Platform | binary | wall time | peak memory |
|---|---|---|---|
| osx-arm64 | −15.6% | faster | −7% to −20% |
| osx-64 | −11.0% | faster | +1% to +11% |
| linux-64 | −9.6% | −25% to −27% | +34% to +44% |
| linux-aarch64 | −9.8% | −30% to −35% | +34% to +45% |
| win-64 | −29.6% | faster | +35% to +40% |
On macOS it is a straight win. On Linux and Windows it is a deliberate trade: a quarter to a third off the run, for a third to a half more peak memory — 95 MiB to 128 MiB on a 10000-package lockfile. That is the choice this project has made, not a regression that slipped through. If peak memory matters more than wall time where you run this, the escape hatch is to build without the allocator; there is no flag for it today, so say so in an issue and there will be.
It also means the memory win from writing documents straight out (above) is partly spent again on those
platforms: the two changes pull in opposite directions and roughly cancel for the document writers, while the
report path, which still builds a serde_json::Value, shows the allocator's cost in full.
opt-level = "s" was measured and not taken: 6.11 MB, 18% smaller again, but about 10% slower on the path
that dominates a run. Startup is 2.6 ms and none of these moved it.
What is not trimmable from here: rattler_lock depends on rattler_solve — a dependency solver a lockfile
reader never runs — to re-export two enums. It is already built with default-features = false.
Each release records the binary size per platform in the workflow's job summary, so a dependency that doubles the download is visible in the run that shipped it.
Watching it on every platform¶
pixi run bench and the numbers above are one machine. The Performance workflow
(.github/workflows/perf.yml) measures the binary on all five platforms a release is built for —
linux-64, linux-aarch64, osx-64, osx-arm64, win-64 — each on its own architecture, because an
aarch64 regression is invisible on x86-64 and the Windows allocator is not the macOS one.
It builds two refs on the same runner, minutes apart, and compares them. That matters: a hosted runner is a shared machine whose throughput varies by tens of percent between runs, so an absolute "0.16 s" from one run and "0.21 s" from the next say nothing, while the difference between two binaries measured back to back says a great deal.
| Metric | How steady | What the workflow does |
|---|---|---|
| Binary size | byte-exact | fails at +5% and more than 256 KiB |
| Peak memory | a few percent within one allocator | fails at +15% and more than 8 MiB |
| Wall time | tens of percent on a shared runner | reported in the job summary, never fails |
A gate needs both a share and an amount. A share on its own fails a build over mimalloc reserving an arena in a
2.5 MiB startup footprint — 15.2%, and 0.38 MiB, which is nobody's problem. An amount on its own misses a small
scenario doubling. pixi run perf-test covers what the comparison does with a given pair of numbers, and CI runs
it, because this gate decides whether a build fails.
The memory limit was 60% for one release, so that comparisons spanning the 0.9.5 → 0.10.0 allocator change — up
to +45.4%, deliberately — did not fail on a documented decision. Every comparison now starts from the latest release tag, which
perf.yml finds with git describe when no base is given — and every one of those has mimalloc, so both sides use
the same allocator and the real numbers are single digits again. The limit is back to 15%, where it catches far more.
Growth accepted on purpose. Because the base is the latest release, a release that legitimately grows the
binary past the gate would keep main red until it is tagged. scripts/perf-allowance.toml records that decision
instead of loosening the gate the way the 60% did: each entry names the release it is measured against
(base = "v1.6.0"), the raised limit (size_percent, memory_percent) and why. It applies only when the
comparison's base is that release, so it expires by itself once a newer tag (a release candidate included) becomes
the baseline, and the job summary prints it whenever it applied. Delete spent entries at release time. 1.7.0's six
new lockfile readers grew the binary +8.75% (+0.9 MiB on linux-64, no new crates), and the entry for v1.6.0
allows 10%.
Peak memory is the child process's own high-water mark: wait4 on Linux and macOS,
GetProcessMemoryInfo on the handle of the finished child on Windows. Not a poller, which would miss the peak.
The child is started with posix_spawn rather than a fork, because a forked child inherits the parent's page
tables and Linux counts those pages against it — which reported the measuring script's own footprint as the
binary's peak for every scenario smaller than it. The script records its own resident size beside the results so
that mistake is visible if it ever comes back.
It runs on demand (workflow_dispatch, with an optional base ref), weekly, and on pushes to main that touch
the source, comparing against the latest release tag. It deliberately does not run on pull requests: each
platform builds the binary twice with LTO. Run it by hand on the pull requests where performance is the point.
The same measurement runs locally, against any two builds:
pixi run perf --binary target/release/pixi-sbom --out head.json
# check out the other ref, rebuild, then
pixi run perf --binary target/release/pixi-sbom --out base.json
pixi run perf --compare base.json head.json
Measured that way, 0.9.5 against 0.10.0 on an Apple M4: binary −13.6%, wall time −29% to −45%, peak memory −35% to −46%.
Reading a regression¶
Criterion compares each run with the previous one in target/criterion/ and prints the change, so the useful
sequence is: run the suite on main, make the change, run it again, and read the percentages. A change above
about 5% on this machine is real; below that is noise unless it repeats.
The numbers above are the 0.9.5 baseline, recorded before the 0.10.0 performance work began, so what that milestone changes can be judged against something.