The benchmark runs examples/goblin-pizza.ts, which uses methods, derived stats, timers, worker leases, retries and fencing. It starts its own fresh local cluster and never touches an existing one.
LATEST MEASURED RUN ·
A busy day in the garden.
- Global customer calls / s
- 54,754
- Independent Raft groups
- 8 × 3 replicas
Reads2,303,662 calls
Writes988,978 calls
Show as a table
| Latency | Calls | Share | Cumulative |
|---|---|---|---|
| 0.0332 ms–0.0389 ms | 6 | <0.01% | 0.00% |
| 0.0389 ms–0.0456 ms | 123 | 0.01% | 0.01% |
| 0.0456 ms–0.0535 ms | 989 | 0.04% | 0.05% |
| 0.0535 ms–0.0628 ms | 3,471 | 0.15% | 0.20% |
| 0.0628 ms–0.0736 ms | 6,277 | 0.27% | 0.47% |
| 0.0736 ms–0.0863 ms | 9,286 | 0.40% | 0.87% |
| 0.0863 ms–0.101 ms | 13,559 | 0.59% | 1.46% |
| 0.101 ms–0.119 ms | 19,268 | 0.84% | 2.30% |
| 0.119 ms–0.139 ms | 26,215 | 1.14% | 3.44% |
| 0.139 ms–0.163 ms | 33,934 | 1.47% | 4.91% |
| 0.163 ms–0.191 ms | 41,885 | 1.82% | 6.73% |
| 0.191 ms–0.224 ms | 50,101 | 2.17% | 8.90% |
| 0.224 ms–0.263 ms | 58,569 | 2.54% | 11.45% |
| 0.263 ms–0.308 ms | 66,518 | 2.89% | 14.33% |
| 0.308 ms–0.362 ms | 74,820 | 3.25% | 17.58% |
| 0.362 ms–0.424 ms | 84,508 | 3.67% | 21.25% |
| 0.424 ms–0.497 ms | 95,783 | 4.16% | 25.41% |
| 0.497 ms–0.583 ms | 108,766 | 4.72% | 30.13% |
| 0.583 ms–0.684 ms | 120,714 | 5.24% | 35.37% |
| 0.684 ms–0.802 ms | 129,799 | 5.63% | 41.00% |
| 0.802 ms–0.94 ms | 133,282 | 5.79% | 46.79% |
| 0.94 ms–1.1 ms | 134,807 | 5.85% | 52.64% |
| 1.1 ms–1.29 ms | 131,737 | 5.72% | 58.36% |
| 1.29 ms–1.52 ms | 128,520 | 5.58% | 63.94% |
| 1.52 ms–1.78 ms | 126,023 | 5.47% | 69.41% |
| 1.78 ms–2.08 ms | 126,716 | 5.50% | 74.91% |
| 2.08 ms–2.44 ms | 124,564 | 5.41% | 80.32% |
| 2.44 ms–2.86 ms | 113,544 | 4.93% | 85.25% |
| 2.86 ms–3.36 ms | 94,796 | 4.12% | 89.36% |
| 3.36 ms–3.94 ms | 71,309 | 3.10% | 92.46% |
| 3.94 ms–4.62 ms | 48,856 | 2.12% | 94.58% |
| 4.62 ms–5.42 ms | 34,942 | 1.52% | 96.09% |
| 5.42 ms–6.35 ms | 24,020 | 1.04% | 97.14% |
| 6.35 ms–7.45 ms | 17,261 | 0.75% | 97.89% |
| 7.45 ms–8.73 ms | 12,704 | 0.55% | 98.44% |
| 8.73 ms–10.2 ms | 9,398 | 0.41% | 98.85% |
| 10.2 ms–12 ms | 6,880 | 0.30% | 99.14% |
| 12 ms–14.1 ms | 5,327 | 0.23% | 99.38% |
| 14.1 ms–16.5 ms | 4,083 | 0.18% | 99.55% |
| 16.5 ms–19.4 ms | 2,435 | 0.11% | 99.66% |
| 19.4 ms–22.7 ms | 1,565 | 0.07% | 99.73% |
| 22.7 ms–26.6 ms | 1,329 | 0.06% | 99.78% |
| 26.6 ms–31.2 ms | 1,106 | 0.05% | 99.83% |
| 31.2 ms–36.6 ms | 722 | 0.03% | 99.86% |
| 36.6 ms–42.9 ms | 484 | 0.02% | 99.88% |
| 42.9 ms–50.3 ms | 251 | 0.01% | 99.90% |
| 50.3 ms–59 ms | 98 | <0.01% | 99.90% |
| 59 ms–69.2 ms | 103 | <0.01% | 99.90% |
| 69.2 ms–81.1 ms | 1 | <0.01% | 99.90% |
| 81.1 ms–95.1 ms | 114 | <0.01% | 99.91% |
| 112 ms–131 ms | 28 | <0.01% | 99.91% |
| 131 ms–153 ms | 17 | <0.01% | 99.91% |
| 153 ms–180 ms | 1 | <0.01% | 99.91% |
| 180 ms–211 ms | 2 | <0.01% | 99.91% |
| 211 ms–247 ms | 2 | <0.01% | 99.91% |
| 247 ms–290 ms | 3 | <0.01% | 99.91% |
| 290 ms–340 ms | 3 | <0.01% | 99.91% |
| 340 ms–399 ms | 10 | <0.01% | 99.91% |
| 399 ms–467 ms | 32 | <0.01% | 99.91% |
| 467 ms–548 ms | 66 | <0.01% | 99.92% |
| 548 ms–643 ms | 129 | 0.01% | 99.92% |
| 643 ms–753 ms | 136 | 0.01% | 99.93% |
| 753 ms–884 ms | 290 | 0.01% | 99.94% |
| 884 ms–1.04 s | 885 | 0.04% | 99.98% |
| 1.04 s–1.21 s | 490 | 0.02% | 100% |
| Latency | Calls | Share | Cumulative |
|---|---|---|---|
| 13.9 ms–14.9 ms | 1 | <0.01% | 0.00% |
| 14.9 ms–16 ms | 1 | <0.01% | 0.00% |
| 16 ms–17.2 ms | 2 | <0.01% | 0.00% |
| 17.2 ms–18.4 ms | 1 | <0.01% | 0.00% |
| 18.4 ms–19.7 ms | 3 | <0.01% | 0.00% |
| 19.7 ms–21.2 ms | 2 | <0.01% | 0.00% |
| 21.2 ms–22.7 ms | 6 | <0.01% | 0.00% |
| 22.7 ms–24.3 ms | 14 | <0.01% | 0.00% |
| 24.3 ms–26.1 ms | 31 | <0.01% | 0.01% |
| 26.1 ms–28 ms | 36 | <0.01% | 0.01% |
| 28 ms–30 ms | 24 | <0.01% | 0.01% |
| 30 ms–32.2 ms | 49 | <0.01% | 0.02% |
| 32.2 ms–34.5 ms | 172 | 0.02% | 0.03% |
| 34.5 ms–37 ms | 299 | 0.03% | 0.06% |
| 37 ms–39.6 ms | 540 | 0.05% | 0.12% |
| 39.6 ms–42.5 ms | 1,187 | 0.12% | 0.24% |
| 42.5 ms–45.5 ms | 2,160 | 0.22% | 0.46% |
| 45.5 ms–48.8 ms | 3,886 | 0.39% | 0.85% |
| 48.8 ms–52.4 ms | 7,285 | 0.74% | 1.59% |
| 52.4 ms–56.1 ms | 13,883 | 1.40% | 2.99% |
| 56.1 ms–60.2 ms | 23,456 | 2.37% | 5.36% |
| 60.2 ms–64.5 ms | 33,042 | 3.34% | 8.70% |
| 64.5 ms–69.2 ms | 41,634 | 4.21% | 12.91% |
| 69.2 ms–74.2 ms | 45,875 | 4.64% | 17.55% |
| 74.2 ms–79.5 ms | 45,896 | 4.64% | 22.19% |
| 79.5 ms–85.2 ms | 41,098 | 4.16% | 26.35% |
| 85.2 ms–91.4 ms | 44,465 | 4.50% | 30.84% |
| 91.4 ms–98 ms | 64,991 | 6.57% | 37.42% |
| 98 ms–105 ms | 97,700 | 9.88% | 47.30% |
| 105 ms–113 ms | 116,900 | 11.82% | 59.12% |
| 113 ms–121 ms | 115,391 | 11.67% | 70.78% |
| 121 ms–129 ms | 99,025 | 10.01% | 80.80% |
| 129 ms–139 ms | 69,777 | 7.06% | 87.85% |
| 139 ms–149 ms | 43,765 | 4.43% | 92.28% |
| 149 ms–160 ms | 29,675 | 3.00% | 95.28% |
| 160 ms–171 ms | 15,449 | 1.56% | 96.84% |
| 171 ms–183 ms | 10,546 | 1.07% | 97.91% |
| 183 ms–197 ms | 6,243 | 0.63% | 98.54% |
| 197 ms–211 ms | 4,165 | 0.42% | 98.96% |
| 211 ms–226 ms | 2,914 | 0.29% | 99.25% |
| 226 ms–242 ms | 1,135 | 0.11% | 99.37% |
| 242 ms–260 ms | 1,072 | 0.11% | 99.48% |
| 260 ms–279 ms | 764 | 0.08% | 99.55% |
| 279 ms–299 ms | 657 | 0.07% | 99.62% |
| 299 ms–320 ms | 617 | 0.06% | 99.68% |
| 320 ms–343 ms | 188 | 0.02% | 99.70% |
| 343 ms–368 ms | 454 | 0.05% | 99.75% |
| 368 ms–395 ms | 160 | 0.02% | 99.76% |
| 395 ms–423 ms | 256 | 0.03% | 99.79% |
| 423 ms–454 ms | 78 | 0.01% | 99.80% |
| 454 ms–486 ms | 143 | 0.01% | 99.81% |
| 599 ms–643 ms | 3 | <0.01% | 99.81% |
| 643 ms–689 ms | 182 | 0.02% | 99.83% |
| 689 ms–739 ms | 220 | 0.02% | 99.85% |
| 739 ms–792 ms | 109 | 0.01% | 99.86% |
| 792 ms–849 ms | 376 | 0.04% | 99.90% |
| 849 ms–910 ms | 477 | 0.05% | 99.95% |
| 910 ms–976 ms | 339 | 0.03% | 99.98% |
| 976 ms–1.05 s | 159 | 0.02% | 100% |
Replica-local reads: lag is allowed. 70% reads / 30% mutations · HTTP/2 · 60.1 measured seconds.
Run passed. 8/8 group audits passed. 8 injected leader failures; quorum recovery 560–704 ms.
Apple M5 Pro; all replicas and load generators share one machine. Completed customer calls use the union measurement window; retries, worker traffic, and explicit replays do not inflate throughput. Reads may be stale; mutations and audits retain fresh checks.
Run it yourself
npm ci
npm run build
cargo build --release --bin flower
npm run bench
cargo build --release --bin flower --bin flower-bench-driver
npm run bench:stress
# Customize the base command, with one copy of each flag:
npm run bench -- --duration 30 --concurrency 16 --workers 4 \
--max-orders 64 --drain 180 --json bench/results/custom.json
npm run bench -- --help- What counts: successful customer calls (shop reads, orders, tips). A retried call counts once. Worker polls, deliveries, replays, setup and the audit don't count.
- The workload is closed-loop: each customer waits for one call before starting the next. The mix is 30% orders, 40% reads and 30% tips until the order cap, then orders turn into reads.
- Everything runs on one machine, servers and load generator together. Your production numbers depend on your workload and hardware.
- The run fails on unexpected errors, an unfinished drain, bad fencing or replays, or a failed business check.
- Reports go to
bench/results/as JSON and HTML, with the settings and binary used.
The base command starts one three-node group. bench:stress runs eight independent groups, each with its own customers, workers and leader crash. Its preset sets 256 customers per group, serial writer preparation and a 1,024-request writer queue, overriding the normal defaults. Environment variables override the preset. See the benchmark guide for all options and the measurement analysis for the current result.
To publish a new result to this page, run inside nix develop, build first and leave profiling off:
npm ci
cargo build --release --locked --bin flower --bin flower-bench-driver
npm run bench:stress
node scripts/publish-bench-results.mjs bench/results/latest.json
node scripts/publish-bench-results.mjs --checkWatch the live dashboard
The dashboard shows one tenant's stores, orders, oven timers, drone leases and rankings from a single watched pizza.dashboard({tenant}) value.
cargo build --release --bin flower
npm run demo:pizza
# Optional: fixed port, paused arrivals, stop automatically after two minutes
npm run demo:pizza -- --port 3030 --paused --duration 120Open the printed URL. It starts a fresh three-node cluster with three tenants. Place orders, tip goblins, pause arrivals or crash the leader. Ctrl+C stops everything and deletes its data.
The dashboard reads replica-local data. It can lag with no bound, and a reconnect may show an older revision. The benchmark's shop reads use pizza.shop.local the same way; pass --read-consistency fresh to measure fresh reads instead.
Compare HTTP/1.1 and HTTP/2
Method calls use HTTP/1.1 by default. Add --http2 to use pooled HTTP/2. Node-to-node traffic always uses HTTP/2.
npm run bench -- --http2 --duration 30 --concurrency 32 --workers 4 \
--max-orders 96 --chaos --json bench/results/http2.jsonKeep the workload and binary the same when you compare. The benchmark reuses request IDs when it retries; the HTTP/2 client itself doesn't retry uncertain calls.
Profile before tuning
Profile under the same mixed load before changing settings. Skip the leader crash for a first comparison.
CARGO_PROFILE_RELEASE_DEBUG=1 cargo build --release --bin flower
npm run bench -- --duration 30 --concurrency 32 --workers 4 \
--max-orders 96 --drain 180 \
--cpu-profile bench/results/profile-mixed.sample.txt \
--json bench/results/profile-mixed.json
# Batch timing for the same mixed workload (one group).
FLOWER_PROFILE_STORAGE=1 FLOWER_PROFILE_EVALUATOR=1 \
node bench/profile-mixed.mjs --http2 --duration 30 --concurrency 128 \
--workers 4 --max-orders 96 --json /tmp/flower-mixed-profile.json- The first command samples native stacks on the initial leader with macOS
/usr/bin/sample. It shows where threads spend time, not CPU percentages or TypeScript lines. - The second writes an extra
*-groups.jsonwith per-batch timings. Phases overlap, so don't add them up as wall time. node bench/profile-otel.mjstakes the same flags and writes an HTML report of request, writer, storage and Raft timing. See the OpenTelemetry guide.- Profiling slows things down. Rerun without it before quoting throughput.
More options are in the profiling commands.
Run the checks
Run the test suite from the checkout:
npm run check
cargo test
cargo build
node tests/e2e.mjs
node tests/e2e-http2.mjs
node tests/e2e-watch.mjsBenchmarks clean up their temporary data when they stop; pass --keep-data to keep it for inspection.