m0serve Realtime from a synchronous Python app

A mut argument is not a reference — 2026-09-29

B20 (PR #463) lost two writes the same way. _flush_inverted and dispatch_job took the asyncio executor's ExecutorState as a mut argument, code re-entered during the call wrote to that struct through its address, and when the call returned the caller stored the callee's copy over the struct. One lost write left an inverted m0serve that never exited on SIGTERM; the other left unanswered every request that an eager task answered in its first step. This is the rule behind it, measured on the pinned toolchain, and how far it reaches in this tree.

The measurement

scripts/probes/mut_copyback_matrix.py generates a matrix of small programs. In every case a call writes b = 42 into a struct through the struct's address, and the program reads b after the call. The script runs the optimized binary and reads each function's signature in the unoptimized LLVM IR (mojo build --emit llvm: KGEN's output, before LLVM's own passes). Three mechanisms lose the write.

M1: a mut argument of 256 bytes or less is passed by value and stored back. The callee takes the struct as an aggregate and returns it, and the caller brackets the call with a load of the whole struct and a store of the result:

%18 = load { i64, i64 }, ptr %15, align 8
%19 = tail call { i64, i64 } @"matrix::free_Word2(matrix::Word2,::SIMD[DType.int, 1])"({ i64, i64 } %18, i64 %14)
store { i64, i64 } %19, ptr %15, align 8

Any write made through the address during the call is erased, whichever field it touched. The matrix puts the struct on the heap behind an address that comes from an opaque call, so M2 below cannot be what loses it:

argument 256 B or less over 256 B
mut, mut self included by value and returned: lost ptr noalias: kept, but see M3
ref by value and returned: lost ptr: kept
read-only by value: the callee reads a copy taken before the write ptr: kept
Pointer[T, MutUntrackedOrigin] ptr: kept ptr: kept

The threshold is the struct's size in bytes: 256 bytes go by value and 257 by pointer, measured with byte arrays and with word fields. Nothing else in the matrix moves it. A free function and a mut self method behave alike, as do raising and non-raising, generic and concrete, TrivialRegisterPassable and not, and fields that are List, String or pointers. Two things avoid the copy: an @always_inline callee, and a callee that hands out the argument's own address (Pointer(to=s) escaping), which gets ptr noalias instead.

M2: a stack local reached through an Int is invisible across a tail call. KGEN emits a call as tail call when, as far as its own origin tracking can tell, no argument refers to the caller's frame, and an Int address refers to nothing it can track. An LLVM tail call promises the callee touches none of the caller's stack slots, so a local whose address the frame handed out as an Int is assumed unchanged: after the call the optimizer reuses the value the frame last stored, whether the local is read by name or back through the same Int:

case the call result
a local, read by name after the call tail call lost
a local, read back through the Int address tail call lost
the same, with one more argument pointing into the frame call kept
a heap object reached only by address tail call kept

M3: a mut argument over 256 bytes is noalias. Nothing is copied back, so a field the callee never touches keeps a write made through the address. The callee's own accesses are another matter:

case result
the callee reads b, a call writes it by address, the callee reads it again the first value, reused
the callee stores a = 7, a call reads it by address, the callee stores a = 9 the call saw 0
the call writes b, which the callee never touches kept

All three are optimizations. At -O0 every argument is a ptr noalias and every write is kept; from -O1 up the tables above hold. mojo run optimizes by default, so a unit test sees what the built binary does.

How far it reaches in m0serve

--census lists every function in a program's unoptimized IR that takes a struct by value and returns it. That is M1's shape, and an upper bound on it, since a function taking a value as var and returning the same type has the shape too. m0serve has 458 such arguments, most of them stdlib containers (List, String, Dict). The tree's own include:

  • every event-loop function that takes the backend as mut: 35 of them, from run_pass_once and _run_pass through the accept, request, response, stream, timer and shutdown modules. KqueueBackend is 24 bytes, and EpollBackend is four word-sized fields;
  • six mut self methods of ExecutorPort (56 bytes): _dispatch, _flush, _flush_inverted, _pass_with, _drain_step_with and _place_frame_with. B20's fix left _flush_inverted taking the port itself that way; see below for why that is harmless;
  • SSERegistry's five mutators, AcceptShare's four, WorkerSupervisor, ProvisionPool, WSState, ThreadSet.join_within, DetachingBackend.wait, and _deliver's PgListenSpec.

The struct types the tree reaches through a raw address fall on both sides of the threshold:

256 B or less: M1 over 256 B: M3
ExecutorState 168, ExecutorPort 56, ThreadedServer 200, HostContext 248, PgListenSpec 80, MojoPool 88, KqueueBackend 24, ThreadSet 24, AcceptShare 96 OffloadPool 584, LoopState 1136, WSGIHandler 808, ServeOptions 544, LoopShared 488, WSGIApp 296, PyBridge 288

Sizes are size_of on macOS arm64.

What was checked

A copy-back loses nothing when nothing writes the struct during the call. These were read against the source:

  • ExecutorPort: every field is set in __init__ and never again, so each copy-back stores what was already there.
  • The backend: its one inline field that changes is _n_ready, which wait sets and the same pass reads. The event buffer and the timer table live behind pointers.
  • PgListenSpec: read-only after construction.
  • _place_frame_with, which B20 left unchanged, holds the OffloadPool as ptr noalias (M3) across the passes it runs, and reads only stream_chunk_write from it, a descriptor fixed at setup. No write can be lost there, so it is unchanged.
  • _executor_serve reads state.pending_done by name after run_forever, which is M2's shape. In m0serve's IR that call is a plain call, not a tail call, so the read is fresh. That rests on how KGEN marks the call rather than on anything the source says.

The rest of the census, M3 for the large types included, was swept after this: the next section.

Who writes them: the sweep

A copy loses a write only when something writes the struct while the call holds it. The sweep (A2) took every type of the tree's own in two censuses, m0serve's and that of the Mojo host's gate application, apps/host_check, which adds the host's pool, producer and loop threads. For each it listed the inline fields that change after construction and looked for a writer of them that can run during the call: another thread, a signal handler, or code the call re-enters. A field behind a pointer is out of a copy's reach, since the store-back writes the same pointer back. A List's header is not: its data pointer, length and capacity are inline.

struct passed what changes inline, and who writes it during a call verdict
ExecutorState M1, when taken as mut stopping, drain_start and the list headers, by the port re-entered from inside its own calls live: B20's two sites, fixed in #463 and now guarded
the backend: KqueueBackend 24, EpollBackend 32, DetachingBackend M1, in 35 loop functions and DetachingBackend.wait _n_ready alone. Under M0_INVERTED a pass run inside a pass calls wait through the backend's address, and the outer frames store their own count back over it latent: the count restored is never read before the next wait rewrites it
ExecutorPort 56 M1, in six methods nothing: every field is set in __init__ safe
AcceptShare 96 M1, in five methods drained, handoffs_in, handoffs_out and left, only by its own methods, which call sendmsg, recvmsg and atomics. Siblings write the shared page, behind page latent
SSERegistry 104 M1, in five mutators four list headers and a count, only by the mutators, which call nothing latent
WorkerSupervisor 240 M1 in three methods, ptr noalias in the rest child_pids and the counters, only by the supervisor's own flow. Its signal handler writes global words (global_slot.mojo), never the struct latent
ThreadSet 24 M1, in spawn, join_all and join_within nothing: a count and two malloc'd addresses. Its threads write their blocks, behind them safe
PoolThreads, MojoPool 88, BlockingPool, ProducerThread, AsgiExecutor M1, in their deal, start and stop_and_join, and MojoPool into the host's _start_pool and _join_pool started, stragglers and the lanes, only by the owning thread. Pool, producer and executor threads write malloc'd blocks and atomics latent
ProvisionPool, WSState, HTTPChunkedDecoder, Socket, ServerMetrics M1 per-connection state, by callees that call nothing or one syscall latent
OffloadPool 584 M3 scalars, set before any thread starts or after the loop ends; the aborts and _drain_buf headers, by the loop thread alone. Pool and executor threads write list elements and ring words, behind pointers latent
LoopState 1136, WSGIHandler 808, PyBridge 288 M3 under M0_INVERTED, a nested pass writes them by address while outer frames hold them noalias: see below latent, not measured
HostContext 248, LoopShared 488, ThreadedServer 200, ServeOptions 544, PgListenSpec 80, Ring 16, the demo mount's Corpus read-only, M1 or M3 nothing after construction: other threads only read them safe

Latent means a mutable inline field that nothing writes during the call today, or whose lost value nothing reads. Sizes are bytes on macOS arm64.

The nested pass. The one path on which code a call re-enters writes a struct through its address, beyond the executor's own state, is under M0_INVERTED. An eager task's first step runs inside WSGIHandler.direct_job, and a frame it places through _place_frame_with can run a pass inside the pass. Review record B23 takes that path on for its own sake, since the inner pass overwrites the event buffer the outer one is reading, whatever the compiler does. The compiler's share of it:

  • M1: the outer frames hold the backend by value and store back only _n_ready, a dead value. A field added to the backend or to ExecutorPort for this path would be restored the same way: state a nested call must see belongs in ExecutorState, reached by address.
  • M3: _process_request holds LoopState as noalias across handler.direct_job(slot), which receives no pointer to it. It adds to st.offload.inflight before the call and returns after it, so nothing is read stale; a store the optimizer moved past the call would overwrite what the nested pass wrote there. Whether it moves one was not measured.
  • M3: spawn_asgi holds PyBridge as noalias across the Python call in which the first step runs. A nested spawn_asgi rewrites the bridge's two scratch lists, which the outer call uses only before that call.

Since B23 (#471, SPEC L32), no pass runs inside another. _place_frame_with is gone. A full chunk channel's frames are handed to the loop's handler instead of a pass making room. The port also refuses loop work while some is on the stack, through ExecutorState.in_pass, which is reached by address as M1 requires.

So the writers the three points above describe no longer run inside those calls. ExecutorState grew by two Bool fields and stays under 256 bytes, still resolved by address in every frame and still guarded by check-copyback.

M2. A scan of both programs' IR for M2's shape (a local whose address escapes into memory, then a tail call, then a load or store of the local) finds six sites, all in _executor_serve and serve_inverted, and every tail call among them is one of OffloadPool's descriptor getters, which write nothing. run_forever and run_forever_inverted, during which the port writes those locals, are called on a sub-object of a local, a pointer into the frame, so neither is a tail call.

C callbacks. m0-sqlite's virtual-table callbacks read and write only malloc'd words (the vtab, the cursor and the stored entry points), and m0-postgres polls PQnotifies, so libpq calls back into nothing.

On Linux. The matrix in the Linux container (aarch64, the same Mojo 1.1.0 build) prints the tables above cell for cell: the 256-byte threshold and all three mechanisms. The census of m0serve's Linux IR counts 459 arguments to macOS's 458, with the same types of the tree's own: EpollBackend, in six functions, in place of KqueueBackend, in five.

The guard

poe check-copyback (scripts/copyback_guard.py, SPEC L31, in test-all) fails when a listed type is passed by value anywhere in m0serve's IR, emitted with build-serve's flags. The list holds one type, ExecutorState, the only one the sweep found written by address while a call holds it. A type belongs on it once that is true of it.

  • It matches a type by its layout, read from the type's own constructor, not by the name the census prints. KGEN appends a raising function's error slot to its parameters, and the census printed the type of B20's _flush_inverted(mut self, mut st) as ?. The census now names such a function's parameters from the start, which leaves ? for a symbol one of whose zero-sized arguments KGEN dropped.
  • It flags a copy whether it is returned (M1) or not (a read-only copy, stale for the same reason), and a layout held inline by another argument's, whatever that argument is called.
  • It shows it can fail on every run. A control program takes the type in B20's two shapes and inside a copied struct, and all three copies must be seen. When they are not (the census is blind, or the type has grown past 256 bytes, onto M3's side, which this cannot see), it stops with 2. A --selftest over canned IR holds the judgement itself.

Sabotaged in a copy of m0-wsgi/src, precompiled beside a copy of the entry file: mut st put back in _flush_inverted fails it, naming that function, and put back in dispatch_job, naming the body it delegates to; an ExecutorState inside a struct taken mut fails it too; the unsabotaged copy passes. The copied entry file matters. Its own directory is searched before any -I, so a sabotaged .mojoc elsewhere on the path is ignored without a word, and the first attempt judged the real one and passed.

Both IRs compile in 7 s on an M4, and in 12 s in the Linux container.

The rule meanwhile

Never hold a struct as a mut, ref or read-only argument across a call during which something may write it by address. Resolve it by address in the frame that uses it (ref st = Pointer[T, MutUntrackedOrigin](unsafe_from_address=addr)[]), pass the address rather than the struct, and read it again after any call that can re-enter. ExecutorState's docstring states it for the executor, and poe check-copyback holds it there.

M2 adds one case the address does not cure: the frame that OWNS such a struct as a local, and computed its Int, reads a stale value after a tail call even through that Int. The owner either lets the struct live on the heap or reads it back only after calls that also pass a pointer into its frame.

Rerun the matrix after a toolchain move:

uv run --no-sync python scripts/probes/mut_copyback_matrix.py --evidence