1.32.0 — Realtime keepalive and fewer per-request database writes
Editorial identity incomplete
2026-10-02
Realtime sockets now get a server keepalive and a 90-second idle window. Unset settings, chat delivery, API-key stamps, failed GeoIP lookups and routine 4xx incident events no longer cost SQL on every request or frame. This is a minor release because the account User no longer writes its realtime connection flag and routine 4xx incident events are now written by a job, which needs the jobs engine running.
Breaking
- A realtime connect or disconnect no longer writes the account
Userrow.User.metadata.realtime_connected,realtime_connected_atandrealtime_disconnected_atare no longer updated; existing values stay frozen at whatever was last written. - The account
Userno longer defineson_realtime_connected/on_realtime_disconnected. A subclass that callssuper()on either now raisesAttributeError. Custom identity models keep their hooks. - Incident events for routine 4xx responses from the REST dispatcher (permission denials,
mojo_rest_errorincluding 404s,api_denied,rest_value_error) are written by a job on theincident_handlerschannel a moment after the response, not before it. With no runner consuming that channel they wait in Redis and are dropped after one day. WS_UNAUTH_TIMEOUT,WS_CONNECT_RATE_LIMITandWS_MAX_CONNECTIONSare read once from Django settings when the realtime handler loads. A databaseSettingrow for any of them is now ignored, and changing one needs an ASGI restart.
Added
- Server keepalive: authenticated sockets receive
{"type": "ping", "ts": <epoch>}everyWS_SERVER_PING_SECONDS(default 20,0disables). A client{"type": "pong"}resets the idle timer and gets no reply. mojo.apps.realtime.signals.realtime_connection_changed(sender, user, connected, connection_id), sent once per authenticated socket on connect and on disconnect, after Redis presence has changed. A receiver that raises is logged and never reaches the socket.- Chat publishes
chat_member_removed,chat_member_bannedandchat_room_deletedon the room topic after the write commits, alongside the existingchat_member_left. WS_SUBSCRIPTION_RECHECK_SECONDS(default 300,<= 0re-checks every frame): how long a socket trusts a chat access decision.GEOIP_FAILURE_TTL(default 3600 s): how long a failed GeoIP lookup is cached before it is retried.report_event(..., defer=True)queues an incident event instead of writing it inline. Directreport_eventcallers are unchanged unless they opt in.INCIDENT_SYNC_CATEGORIES(settings file only) adds categories that always write inline, on top of the built-in security list.
Changed
- The authenticated idle cull is
WS_IDLE_TIMEOUT, default 90 s (was a hard-coded 30 s). Only frames from the client count as activity; server pings do not. pongis now a reserved client message type. It no longer reacheson_realtime_messageorREALTIME_MESSAGE_HANDLERS.settings.getcaches misses per scope: after the first lookup, an unset key costs one Redis read per scope and no SQL. ASettingsave invalidates at once; each cache hash expires after one hour as a backstop.- With Redis down,
settings.getreads the database for every scope instead of falling through to the settings file. - Chat delivery checks room access once per subscription instead of on every frame. Leave, remove, ban and room delete still cut delivery on the next frame; any other loss of access (a revoked
chat/manage_chatpermission, a deactivation outside the disable service) takes effect withinWS_SUBSCRIPTION_RECHECK_SECONDS. - A failed GeoIP lookup is stored with
provider: "failed"and served from cache untilGEOIP_FAILURE_TTLpasses, instead of re-running the provider chain on every event. Therefreshaction retries it immediately; a record that resolved before keeps its data on a failed refresh. - API-key
last_used(bothApiKeyand user API keys) is rewritten at most once everyAPI_KEY_TOUCH_SECONDS(default 300, settings file only) instead of on every request. - A deferred 4xx event's
createdis when the job wrote it. 5xx errors and security categories (auth failures, invalid or expired tokens, threat-intel hits) still write before the response.
Upgrade notes
- Clients that cannot tolerate an unknown
{"type": "ping"}frame: setWS_SERVER_PING_SECONDS = 0until they answer it with{"type": "pong"}. - Anything reading
User.metadata.realtime_connectedmust move to therealtime_connection_changedsignal orrealtime.is_online. The old key does not error; it just goes stale. - Deferred 4xx incident events need the jobs engine consuming
incident_handlers, which is a default channel. A deployment that narrowedJOBS_CHANNELSmust add it back, or those events are silently lost after a day. - Tests that count 4xx incident events right after a response must run the queued job first (
th.run_pending_jobs(channel="incident_handlers", ...)), including zero-count checks, which otherwise pass vacuously. GEOIP_FAILURE_TTLmay be lowered where geo signals feed risk scoring, so an address that failed once is retried sooner.- A setting created outside
Setting.save()(querysetupdate,bulk_create, raw SQL, a fixture load) can read as unset for up to an hour on a host that already cached the miss. Callpush_to_cache()on the row after such a write. - Move any database
Settingrows forWS_UNAUTH_TIMEOUT,WS_CONNECT_RATE_LIMITorWS_MAX_CONNECTIONSinto the settings file. EveryWS_*value needs an ASGI restart to change.