m0serve Realtime from a synchronous Python app

A stream has no deadline — 2026-09-30

A design note from the engineering record. D57 is the decision; SPEC A24 names it.

The question

Pull request #412 gave a response a send deadline: --idle-timeout between two sends that move bytes, the way nginx's send_timeout works (SPEC A24). It left streams out on purpose. _arm_send_deadline skips a slot the loop holds as SSE or as a WebSocket, so a client that stops reading a stream keeps its slot until it leaves. The question was whether streams should get a timer too.

What was decided

No timer (the owner, 2026-09-30). Three reasons:

  • A stalled reader's slot costs what an idle reader's does. A stream lives as long as its client. A deadline would not stop abuse, because a client can read one byte a second and never trip it. How many slots one client may hold is a connection-count question.
  • Memory is bounded. A slot's queue stops at MAX_PENDING_BYTES (65536 bytes) in SSERegistry and in WSHub, and a frame that would pass it is dropped for that slot alone. SSERegistry holds Mojo SSE streams, DatastarStream's, and both kinds of m0serve hold. Beyond the queue, a slot holds what the kernel's socket buffers hold.
  • A client that vanishes without a FIN is reaped by TCP, because the heartbeat keeps bytes in flight. M0_SSE_HEARTBEAT_MS defaults to 15000, and on that timer an SSE slot gets a : heartbeat comment and a WebSocket slot a ping. Nothing acknowledges those bytes, so the kernel retransmits them until tcp_retries2 is spent and fails the socket with ETIMEDOUT. The loop then closes the slot, and the application's disconnect hook runs. That is the measurement below.

A stream the application writes through the chunk channel differs in two ways. That is an ASGI application's stream, or a WSGI iterable a pool thread streams. First, the loop puts no comment into its SSE, because an event may span two chunks and a comment between them corrupts the event (_heartbeat returns early for slot_channel_stream). Such a stream is reaped the same way only if the application writes something on a cadence of its own, as sse-starlette's ping does; a WebSocket on that path still gets the loop's ping. Second, its sends wait for drain credit (the WebSocket send window), so a stalled reader holds up the application's producer and no frame is dropped. None of this path was measured when the decision was written. It was on 2026-10-01, and a quiet stream of that kind held its slot after its client vanished; keepalive on the stream's socket is what reaps it now.

The measurement

The server was m0serve built from main at 13c7d40, in the Linux container m0lin (arm64, kernel 6.8.0), with --realtime and one worker. It served a WSGI view that holds an SSE stream (M0-Hold: stream) and one that holds a WebSocket (M0-Hold: websocket). A client connected and read the head. Then iptables dropped every packet to and from the client's port on lo, so the client vanished without a FIN. The probe polled the server's own count every 0.2 s from the moment the rule was in: subscribers or sockets in /health, and the server's socket in ss -tnoi.

tcp_retries2 was lowered to 5 in the container's network namespace, entered from the VM with nsenter -n. The setting is per namespace: after the write it read 5 inside the container and 15 in the VM's own namespace.

arm heartbeat tcp_retries2 the slot closed after the vanish
SSE hold 1 s 5 14.44 s, 14.23 s, 14.57 s
WebSocket hold 1 s 5 14.48 s
SSE hold 15 s, the default 15, the default 964.59 s
SSE hold, heartbeat off 0 5 held at 100 s
SSE hold, a live client with a zero window 1 s 5 held at 100 s
the same client, then vanished 1 s 5 39.94 s

With a 1 s heartbeat, the first unacknowledged heartbeat went out about 1.0 s after the vanish. The kernel retransmitted it with the timeout doubling from 201 ms (backoff:5, rto:6432 at the end) and gave up about 13 s after that first send. Linux derives the give-up time from tcp_retries2 and a 200 ms base: 12.6 s for 5, 924.6 s for 15, checked when a retransmission timer fires. In every reaped run the socket left ss and the count reached 0 in the same 0.2 s sample.

At the defaults, the first unacknowledged heartbeat went out 15.0 s after the vanish, and the slot closed 949.6 s after that, with the last retransmissions 120 s apart (backoff:15, rto:120000). So a client that vanishes holds its slot for up to about 16 minutes: the heartbeat period, then about 15.5 minutes of retransmission. The close fell between two heartbeats, 63.3 periods after the first, so the loop acted on the kernel's error report itself. An SSE slot keeps read interest while it idles, and the error wakes it.

With the heartbeat off, the slot stays. Nothing is in flight, so nothing times out. ss showed the socket established with an empty send queue and no timer for all 100 s, past the 60 s idle timeout, which a stream does not have. m0serve set no SO_KEEPALIVE on a connection at 13c7d40, so nothing else looked; it does now, on a stream's socket, and this arm is reaped (below). Once the rules were removed and the client closed, the count was 0 within 1.5 s.

A live client that stops reading keeps its slot, by design. It set a 4 KB receive buffer and read nothing after the head. Two publishes of 8 × 4000 bytes filled its window, and the server's socket went onto its persist timer with 58 KB unsent. For 100 s the kernel's window probes were answered: timer:(persist,…,0), no probe outstanding. TCP keeps a connection with a zero window open for as long as its probes are answered, and Linux does. Dropping that client's packets then left the probes unanswered. The count of unanswered probes rose to 5 and the slot closed 39.9 s after the drop. So a reader that stalls and then vanishes is reaped too, by the persist timer's probe limit, which is also tcp_retries2.

Not measured:

  • macOS. The packet filter and the TCP sysctls need root there.
  • A stalled reader that vanishes, at the default tcp_retries2. The probe interval grows to 120 s, so the bound is longer than the retransmission case's.
  • The chunk-channel streams above, until 2026-10-01 (the next section).

Keepalive on a stream's socket

Added 2026-10-01 (SPEC I32). Two of the arms above left a vanished client holding its slot for as long as the server ran: a stream with the heartbeat off, and, unmeasured then, the stream an application writes through the chunk channel, which the loop does not heartbeat. Measured the same way, at the same kernel, on a quiet ASGI stream (/stream-quiet in apps/asgi_bare: one event, then nothing): 30 s after its client's packets were dropped the application still counted the stream open.

A deadline is not the answer D57 gave, and this is not one. TCP keepalive is the kernel asking an idle peer whether it is still there: after M0_STREAM_KEEPALIVE_S seconds with nothing received (15) it sends an empty segment, again at that interval, and fails the socket when three in a row go unanswered. No byte enters the stream, so the reason the loop's comment is kept out of an application's stream does not apply, and a live client's kernel answers whatever its application is doing. The loop sets it where a slot becomes an SSE stream or a WebSocket (_keep_stream_alive in loop/response.mojo); 0 sets nothing.

arm heartbeat keepalive the slot after the vanish
ASGI stream, quiet none reaches it off held at 30 s
ASGI stream, quiet none reaches it 1 s closed at 4.18, 4.16, 4.14 s
ASGI stream, quiet none reaches it 15 s, the default closed at 61.52 s
SSE hold off 1 s closed at 4.16 s
WebSocket hold off 1 s closed at 4.17 s
SSE hold 1 s 1 s held at 60 s
ASGI stream, quiet, a live client none reaches it 1 s kept through 15 s
SSE hold, a live client with a zero window 1 s 1 s kept through 30 s

tcp_retries2 was the default, 15, throughout. The reap is the idle time plus three intervals: ss showed timer:(keepalive,997ms,0) as the rule went in and timer:(keepalive,124ms,3) in the last sample before the socket was gone, and the application's disconnect ran in the same 0.1 s.

Three things the table says:

  • A stream with nothing in flight is reaped in about a minute at the default, where it was held for good.
  • A heartbeated stream is not reaped sooner. Once a heartbeat is unacknowledged the kernel is retransmitting, and Linux sends no keepalive probe while data is outstanding: with both at 1 s the slot was still held at 60 s, on its retransmission timer (timer:(on,44sec,8)). The bound for that stream is the one measured above, about 16 minutes at the defaults. So a quiet stream is now reaped sooner than a heartbeated one. Shortening the second is a different change (TCP_USER_TIMEOUT on Linux). Whether that would also end a live client with a zero window is unmeasured, and it is D57's question.
  • A live client keeps its slot, idle or not reading: its kernel answers the probe, and with data queued the persist timer is what runs.

Not measured: macOS's reap, since dropping a client's packets there needs root. What is checked on macOS is that the option is on the server's socket with the idle time asked for (lsof -Tf prints SO=KEEPALIVE=15001 for a stream and nothing for a plain connection). The smoke's Linux leg repeats the first two rows on every pull request.

What it costs

A reader that stalls and then resumes gets later frames with silent gaps. A frame the cap refuses does not advance the slot's last-seen id, and a later frame that fits does. The client's Last-Event-ID moves past the gap, so a reconnect cannot replay the frames it lost. A WebSocket has no replay at all, so a message the cap refuses is missed. For a stream of whole states, such as DatastarStream(send_latest=True) and the blobs demo, that is the right degradation: the next frame supersedes what was lost. For a stream that is a log, it is not.

What would retire it

An application whose stream is a log rather than whole states, and whose clients stall. The answer then is to close a stream when its queue reaches the cap, so the client reconnects and replays from its Last-Event-ID, rather than to add a timer. Or a deployment that sees live stalled readers holding slots.