Collaboration Consistency
This document records the current collaboration consistency model after #865. It is a companion tocollaboration-layers.md (layering) and
collaboration-models.md (engine contracts).
It describes what the Nitro/Postgres collaboration host and browser session
guarantee today. The runtime-independent DocumentRoom semantics and host
capabilities are defined in
collaboration-runtime.md; the Nitro host supplies
their production adapters. This page does not describe a Durable Objects
runtime.
Durable write invariants
These invariants are load-bearing. Consistency tests treat them as the conformance baseline for future runtime hosts.1. One causally dependent write in flight per client session
A browser document session keeps at most one durableEvent write in flight.
- Later local edits stay pending until
DurableAck. - The outgoing queue preserves exact batch identity across disconnect.
- After reconnect, the session repairs first, then exact-resends pending batches.
2. Per-document durable serialization
Same-document durable appends serialize through a Postgres transaction-scoped advisory lock (pg_advisory_xact_lock). This works across
serverless instances that share the database.
3. Durable commit before DurableAck and fan-out
The server order is:- Validate and append batches in the locked transaction.
- Send
DurableAckto the writing peer. - Publish the
Eventon the realtime bus for local and remote peers.
4. Idempotent exact resend
Resending the samebatchId with the same payload is a no-op success.
Resending the same batchId with a different payload is a conflict.
5. Repair is explicit synchronization
RepairRequest / RepairResponse pages catch clients up from a cursor over
durable history. Repair is intentional synchronization, not an automatic
reaction to every conflict.
Live Event messages may arrive during an incomplete repair. Clients apply
both paths with batch-id deduplication and must still converge.
6. Presence is not durable convergence
Awareness/presence remains ephemeral. It must not be mixed into EG-walker document history or durable batch appends.EventConflictError semantics
On the realtime transport, conflicts are surfaced to clients as non-retryable protocolError messages with code conflict. The browser does
not enter a conflict-driven reconnect/retry loop. Missing history is
recovered on the next intentional repair/resync after a clean reconnect path,
not by blindly retrying the conflicting payload.
HTTP appends use a different surface. In document-events.post.ts,
EventConflictError maps to HTTP 409 with an { error } body (the conflict
message string only). Structured fields such as missingParentIds exist on the
server-side EventConflictError instance; they are not included in that HTTP
response. The HTTP handler logs conflictType, documentId, and
messageType via logDocumentMetric; it does not record bounded batch or
event ID samples. On the realtime route, structured conflict details are not
sent on the WebSocket Error payload.
Implementation: EVENT_CONFLICT_TYPE / EventConflictError in
apps/collab-nitro/server/utils/event-conflict.ts.
DuplicateIncomingEventId
MissingParentHistory
BatchPayloadConflict
StoredEventIdConflict
Multi-instance ownership model
The consistency harness models the intended ownership boundary:Conformance coverage
Phase 1 hardening tests live underapps/collab-nitro/test/:
repair-live-interleave.test.ts— repair ↔ live Event racesmulti-instance-collab.test.ts— write / repair / reconnect across instancescollaboration-convergence.property.test.ts— seeded randomized 2–3 client traces
DocumentRoom with
in-memory capabilities. The client queue and EG-walker/block-model replicas
remain test-owned, while authorization, repair, durable acknowledgement,
fan-out, and room lifecycle exercise the same runtime semantics as Nitro.
Optional soak: