Benchmarks
How the benchmarks are run, what they compare, where m0serve loses and by how much, and the ways a benchmark of a server misleads.
Every number in the tables below is rendered from a dated,
environment-stamped JSON artifact in bench/results/,
and the figures quoted in the prose are recomputed from those same
artifacts and compared at the precision they are written to. Both are
CI-checked: a hand-edited table, or a sentence whose number has drifted,
fails the build naming the file.
Two figures are not covered by that, and each says so where it appears: the performance/efficiency core split and the cross-session variance number. Both are recorded observations from before the artifact system existed. The slow-view isolation result used to be the third and largest of them; it has an artifact now.
That is the only unusual claim this page makes. The performance claims themselves are mixed: m0serve is not the fastest server in this comparison on raw throughput, and the tables below say so. What it is faster at is a narrower thing — keeping a fast request fast while slow work is in flight — and the reason to read the rest is to see exactly where the line falls.
How to read this page
Five caveats, stated before the numbers rather than under them.
- Within-run ratios are the signal; absolute rows are not. Identical binaries move ~1.5x in absolute rps across sessions on this hardware, from thermal and load state alone (recorded observation, not an artifact). Compare rows inside one table, never a row here against a row in something else you read.
- Cores are measured, not configured. Each table's
corescolumn is sampled%cpuof the pids on the listen socket. The column exists because it caught a real error: a comparator invoked as--workers 1was running ~1.75 cores across its runtime's I/O threads, so every earlier raw-rps ratio had been comparing 1.75 cores against one. - The benchmark box has performance and efficiency cores (Apple M4, 4P
- 6E). An E-core serves this workload at 18.6k rps against a P-core's 81.7k — 4.4x slower (measured by pinning a worker to background QoS; not an artifact). So per-core rows are comparable only where the server plus the load generator fit in the P-cores: on this box, 1 and 2 workers. The 4-worker rows measure the scheduler, not the server, and are kept because removing them would hide that.
- The two ASGI artifacts predate the
--target-cpubaseline pin (commit30f044a), visible in theircommitlines; the WSGI and isolation artifacts were recorded after it. The pin's cost was measured separately and was zero — byte-identical code — so re-running the ASGI pair is provenance hygiene rather than a correction in waiting. Said here rather than discovered later. - One anomalous round per run is normal on this box. Three recorded runs of the layer split each had exactly one round land well off the other two, in a different position each time — which is why every bench here takes the median of three rather than a mean of one. The medians from those three runs agree to within 0.03 on the per-core ratio; the individual rounds do not.
The comparators are Granian 2.8.1 and uvicorn, both run with their own recommended settings, and every row is a single process unless the label says otherwise. Where a comparator wins, the row stays.
The short version
| question | answer |
|---|---|
| Fastest per core on bare WSGI? | No — Granian, by ~1.2x |
| Fastest per core on bare ASGI? | Against uvicorn --loop asyncio, yes — the executor is ahead by ~1.06x per core at 16 connections and 1.22x at 256. Against uvloop, no — uvicorn is ahead, and by ~1.36x with uvloop, which is what pip install uvicorn[standard] runs by default; 0.94x at 256 connections |
| Fastest fast-request tail under mixed load? | Yes — p99 ahead of uvicorn in every recorded run |
| Fastest HTTP layer, Python excluded? | Yes — but see the note on why that is not the interesting number |
The HTTP layer, and the bridge
This is the decomposition that matters, and it is why the "fastest HTTP layer" row above is marked as uninteresting. Splitting the server into the part that parses HTTP and the part that calls Python locates the gap instead of reporting one number for both:
Source: layer-split-20260826T135108Z.json — 2026-08-26T13:51:08+00:00, commit 476358b.
Environment: Python 3.14.7 free-threading build; granian 2.8.1; Apple M4 (10 cores); wrk -c16 -d10s, 3 rounds, medians.
| row | rps | cores | rps/core |
|---|---|---|---|
apps/hello — mojo-http HTTP layer, zero Python |
106,629 | 0.92 | 115,901 |
m0serve + bare WSGI, 1 worker |
82,629 | 0.97 | 85,185 |
granian + bare WSGI, 1 worker |
175,015 | 1.75 | 100,009 |
m0serve + bare WSGI, 4 workers |
158,338 | 3.12 | 50,750 |
granian + bare WSGI, 4 workers |
141,571 | 4.18 | 33,869 |
Cores are measured (sampled %cpu of the pids on the listen socket), not configured — the column exists because a "1 worker" comparator was found running 1.6 cores. Cross-session absolute rps on this hardware varies ~1.5x; within-run ratios are the signal.
Read down the rps/core column at one worker. Three numbers, and the
arithmetic between them is the finding:
- the HTTP layer with no Python at all (
apps/hello) runs at 115.9k rps/core — above Granian's end-to-end 100.0k - put the same bare WSGI application behind it and m0serve runs at 85.2k rps/core, so its own bridge costs 1.36x
- net, m0serve is 0.85x Granian per core; Granian is 1.17x ahead
So the deficit is the bridge — the per-request crossing into CPython — and not the parsing or the event loop, which is the whole reason to split the measurement rather than report one number.
What this table cannot tell you is how much Granian's own bridge costs, because there is no Granian-without-Python row to divide by. A cleaner decomposition — "1.0x HTTP layer × 1.35x bridge" — appeared in the working record and does not reconcile with this artifact: the measured gap is 1.17x, and for a 1.35x bridge term to hold, the HTTP layer would have to be 0.89x, which contradicts the row above it. Corrected rather than repeated; the per-side split needs a measurement nobody has taken yet.
Quoting the 115.9k hello row against Granian's 100.0k would be comparing a server that runs no Python to one that does. It is on this page because it locates the cost, not because it is a win.
ASGI throughput
apps/asgi_bare under wrk with browser-shaped headers, byte parity
asserted between the two responses, single process each:
Source: asgi-wrk-hello-20260827T195533Z.json — 2026-08-27T19:55:33+00:00, commit eee3878.
Environment: Python 3.13.6; Apple M4 (10 cores); wrk -t2 -c16 -d8s, browser headers; executor loop: uvloop.
| row | rps | cores | rps/core |
|---|---|---|---|
m0serve — zero-config executor (its loop is stamped above) |
60,268 | 0.99 | 60,876 |
uvicorn --loop asyncio |
57,014 | 0.99 | 57,590 |
uvicorn with uvloop — what pip install uvicorn[standard] runs by default |
81,670 | 0.99 | 82,494 |
Cores are measured (sampled %cpu of the pids on the listen socket), not configured — the column exists because a "1 worker" comparator was found running 1.6 cores. Cross-session absolute rps on this hardware varies ~1.5x; within-run ratios are the signal.
Against uvicorn --loop asyncio the executor wins this row, and the
cores column says what changed. It used to lose it at 0.83–0.90 cores
against uvicorn's 0.99 — wakeup-bound, not CPU-bound: every request
serialized through loop thread → submit datagram → executor thread →
completion datagram → loop thread, and both threads idled between
handoffs. On 2026-08-27 the pump was batched in both directions (+5%
here, +19% at 256 connections), and then inverted: Python calls into
Mojo for every event through a type built in-process, and the executor
thread never leaves run_forever — which is what finally let its
opportunistic uvloop pay. This row runs on uvloop at 0.99 cores and
reads 1.06x uvicorn --loop asyncio; at 256 connections, 1.22x. Against
uvicorn with uvloop it is still behind — 0.74x here, 0.94x at 256 — and
that row is on the page because it is the number a developer's own
uvicorn[standard] install produces: 0.74x. The concurrency tables and
the loop-by-loop comparison are in
WSGI_PERFORMANCE.md.
Worth recording because it inverted a conclusion: an earlier run of this
comparison used a stdlib http.client harness and reported 0.88–0.94x. The
assumption was that the stdlib client understated the Mojo layer's parsing
edge. Under wrk the ratio is 1.06x at 16 connections (0.72x before the
pump was batched and then inverted) — the stdlib client had been
flattering the executor as it stood, and the fix path derived from it
was aimed the wrong way.
Fast-request tail under mixed load
The measurement the executor exists for: how fast is a fast request while
slow ones are in flight. Four threads measure / while two hammer
/slow?ms=200.
Source: asgi-executor-20260825T172549Z.json — 2026-08-25T17:25:49+00:00, commit 58a35ed (dirty tree).
Environment: Python 3.13.6; Apple M4 (10 cores); seconds=8, threads=8.
| server | rps | fast p50 | fast p99 | errors |
|---|---|---|---|---|
m0serve — asyncio executor |
22,597 | 267 µs | 376 µs | 0 |
uvicorn |
23,642 | 173 µs | 407 µs | 0 |
Fast-request latency is measured while slow requests are in flight; rps and the two percentiles come from the same run, so they trade against each other rather than being separately optimised rows.
This is the row m0serve wins, and it is a narrow win stated narrowly. The fast-request p99 is ahead of uvicorn's — in this run and in every recorded run — because awaits overlap on the loop and the Mojo acceptor never runs application code. In the same run, uvicorn's p50 is better by about 1.5x and its throughput by a few percent. Both facts come from one run of one client, so they are a trade, not two independent scores.
The await-concurrency underneath is unambiguous in a way the percentiles are not: eight concurrent 1.5 s awaits complete in 1.51 s on one loop with zero threads, where the buffered bridge takes 12 s.
Slow-view isolation
The strongest claim this project makes, and now the one with an artifact
behind it. A synchronous view that blocks holds the connections pinned to
its event loop; --blocking-threads N puts a pool of handler threads
behind each loop so it stops doing that.
Read the first two rows across, then the next two. Without the pool the fast-route p99 climbs from ~1 ms to ~195 ms as slow views are added — most of the 200 ms hold, which is what "the connections pinned behind it" means arithmetically. With the pool it does not move. That is a ~100x change and the largest effect recorded anywhere in this repository, and it holds in both execution modes, which is the part that matters: prefork and threads fail identically and are fixed identically.
The control is the point. Both halves run in one pass, so the rows without the flag have to keep failing for the rows with it to mean anything.
Source: mixed-workload-20260826T141514Z.json — 2026-08-26T14:15:14+00:00, commit c182b62.
Environment: Python 3.14.7 free-threading build; granian 2.8.1; Apple M4 (10 cores); wrk -c16 -d10s, 2 rounds, medians.
| configuration | slow=0 | slow=1 | slow=2 |
|---|---|---|---|
--workers 4 |
1.0 ms | 190.7 ms | 195.9 ms |
--threads 4 |
1.7 ms | 195.2 ms | 200.5 ms |
--workers 4 +bt=4 |
2.4 ms | 7.4 ms | 7.4 ms |
--threads 4 +bt=4 |
2.2 ms | 2.5 ms | 1.6 ms |
granian bt=4 |
0.6 ms | 0.5 ms | 0.6 ms |
Fast-route p99, median across rounds, as concurrent slow requests are added. A row that stays flat isolated the slow work; a row that climbs toward the slow view's hold time had its connections stranded behind it. Both halves run in one pass, because a control that stops failing has stopped measuring anything.
And the comparator's row is better than ours. Granian's own
--blocking-threads is the architecture this feature copied, and its
fast-route p99 stays at ~0.6 ms across all three slow levels — flat, like
ours, but roughly 4x lower than our best row and 12x lower than
--workers 4 +bt=4. So the honest claim is the one about the stall: the
pool removes a hundredfold failure that is there without it. It does not
also win the tail that remains.
That row could not appear in this repository's earlier record of this
benchmark, which noted granian was absent because it is not in the lock
file's default groups. It is in the bench group, pinned at the version
every number here names, and installing it is one flag:
uv sync --group bench.
What this page does not measure
- TLS and HTTP/2. m0serve has neither; terminate at a proxy, which is gunicorn's answer too. A comparison against servers that do would be measuring the proxy.
- Real applications. Every row above serves a bare handler, because a
Django view's own work dominates and would hide the server difference
entirely. That cuts both ways: it makes these gaps visible, and it makes
them a smaller fraction of any real request than they look here. Two
framework rows — FastHTML and Django ASGI at
/— are measured all the same and kept in WSGI_PERFORMANCE.md ("Framework rows"), off this page for that reason. - Anything on Linux. These artifacts are macOS arm64. The CI matrix builds and smoke-tests Linux x86-64 and aarch64, but the benchmark box is one machine and the page says which.
Reproducing
The full procedure — the free-threaded interpreter swap, the rebuild inside
it, the granian install, and the three traps that each produced a
confident wrong answer — is in
WSGI_PERFORMANCE.md. It is deliberately
in the working record rather than here: this page is the result, that one
is the method and the dead ends.
After recording a new artifact, uv run poe render-bench-docs rewrites
every table on this page from it, and uv run poe check-docs fails if
anyone edits one by hand.