mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-21 21:15:56 +08:00
* feat(realtime): add shared room identity + authorization spine (#5929) Introduces the foundation for a unified realtime "room" model spanning the Socket.IO presence server (apps/realtime), the durable SSE event log, and the ephemeral pub/sub fanout — all of which today reinvent their own room identity, naming, and authorization. - @sim/realtime-protocol/rooms: RoomRef { type, id }, ROOM_TYPES, and a roomName/parseRoomName codec. WORKFLOW deliberately maps to the bare id so the ~40 existing io.to(workflowId) callsites and presence state keys are unchanged; every other room type is namespaced so id spaces cannot collide. - @sim/platform-authz/rooms: authorizeRoom(userId, room, action) generalizing the exemplary authorizeWorkflowByWorkspacePermission — one resource->workspace resolver per room type, then the shared resolveEffectiveWorkspacePermission + permissionSatisfies gate. Pure foundation, no behavior change: nothing consumes these yet. Prune graph stays at 14/25 (platform-authz already depended transitively on realtime-protocol via apps/realtime). * refactor(realtime): generalize presence server to multi-room [2/N] (#5930) * refactor(realtime): generalize presence server to multi-room (RoomRef) Generalizes the Socket.IO presence layer from single-workflow-room-per-socket to a domain-neutral, multi-room-per-socket model keyed by RoomRef, so a second domain (workspace files, next PR) can reuse the same membership + presence engine. Behavior-preserving for workflow collaboration. IRoomManager is now domain-neutral (addUserToRoom/removeUserFromRoom/ getRoomForSocket/getRoomUsers/updateUserActivity/... all take a RoomRef). The workflow lifecycle broadcasts (deletion/revert/update/deploy) move out of the manager into WorkflowRoomService, composed over the generic manager. Backward-compat by design (no workflow migration, no regression): - Workflow Socket.IO room name stays the bare workflowId (roomName() maps workflow -> bare id), so the ~40 io.to(workflowId) callsites are untouched. - Workflow Redis presence keys stay workflow:{id}:users/:meta (the type prefix IS "workflow"). Multi-room correctness (from adversarial audit): - socket:{id}:workflow single-value key -> socket:{id}:rooms HASH (type->id). - The SHARED socket:{id}:session key is deleted only when the socket leaves its LAST room (refcount via HLEN) — a leave from one room no longer breaks the other room's handlers. - disconnect enumerates the socket's stored rooms and rebroadcasts presence per room, instead of picking an arbitrary socket.rooms entry. - presence broadcasts use a per-room-type event name (workflow keeps the bare presence-update; others are namespaced). Workflow handlers wrap manager calls with a shared workflowRoom(id) helper; UserPresence.workflowId -> room (the client never reads that field). Tests: existing 112 realtime tests pass unchanged (behavior gate) + 7 new multi-room tests (refcounted session, presence isolation, multi-room disconnect, per-type event names). tsc clean, boundaries + prune (14/25) green. * fix(realtime): harden multi-room disconnect + id-guard room removal Two fixes from an adversarial regression audit of the multi-room refactor: - Disconnect now handles `disconnecting` (where `socket.rooms` is still populated and authoritative) and falls back to the live Socket.IO room set for any room the manager's stored state no longer tracked. This restores reliable presence cleanup + departure broadcast even if the Redis `socket:{id}:rooms` key was evicted or TTL-expired — the one behavioral gap vs the pre-refactor disconnect. - REMOVE_ROOM_SCRIPT now only drops the socket's room mapping (and runs the last-room session cleanup) when the stored id matches the room being removed, matching the memory manager's existing id guard. Prevents a mismatched-room call from wiping a different room's mapping or the shared session. +1 test (id-guarded no-op removal). 120 realtime tests pass, tsc clean. * fix(realtime): only rebroadcast disconnect-fallback rooms whose removal succeeded Greptile 4/5 follow-up: the disconnecting-time fallback ignored removeUserFromRoom's boolean and rebroadcast presence even when the removal reported false. Now it only treats a room as removed (and rebroadcasts) when the manager confirms it — symmetric with removeSocketFromAllRooms, which already only returns rooms it actually removed. * fix(realtime): exclude the disconnecting socket from its farewell broadcast Greptile follow-up (transient-Redis-failure edge): if removeUserFromRoom fails on disconnect, the socket's presence entry can outlive it (room hashes have no TTL) and reappear as a ghost. Disconnect now broadcasts a correction to EVERY room the socket was in (union of the manager's removed rooms and the live Socket.IO membership) and passes the disconnecting socket id as excludeSocketId, so it is never shown as a collaborator regardless of whether the Redis delete succeeded. Any orphaned entry is still reclaimed by the next join's stale-presence sweep. broadcastPresenceUpdate gains an optional excludeSocketId; normal broadcasts are unchanged. +1 test. * fix(realtime): make presence broadcasts liveness-aware (root-cause ghost fix) Presence broadcasts now reconcile the stored list against the live Socket.IO membership (io.in(room).fetchSockets()) before emitting, via a shared filterVisiblePresence helper. This closes the residual behind the earlier disconnect fixes: an entry orphaned by a failed removal (room hashes have no TTL) could reappear in a LATER join's presence snapshot until the 75-min stale sweep. Now such an entry is never emitted, because a non-live socket is filtered out of every broadcast. Combined with excludeSocketId (which handles the disconnecting socket, still momentarily live). Fail-safe: on a fetchSockets throw or an empty result while entries remain, emit the unfiltered list rather than hide live collaborators. Also drops a dead guard in the disconnect union loop (rooms already removed are skipped by the wasInRooms check) and the now-unused isSameRoom import. +1 ghost-guard test. 122 realtime tests pass. * feat(files): live presence avatars + live file tree via realtime rooms (#5932) * refactor(tables): adopt shared durable event-log core (#5934) * fix(realtime): address post-merge review-comment findings (#5937) * fix(realtime): address post-merge review-comment findings A re-audit of every inline review comment on the merged stack surfaced real issues that the thread-resolutions and prior audits missed. Fixes: Presence server (#5930 comments): - connection.ts: snapshot `socket.rooms` SYNCHRONOUSLY before the first await. Socket.IO clears the room set once the synchronous part of a `disconnecting` handler returns, so reading it after `await removeSocketFromAllRooms` saw an empty set — the eviction fallback was dead. (Cursor: "Disconnect fallback misses live rooms".) - workflow-room-service: restore the original managers' final unconditional room-state wipe via a new `deleteRoom(room)` manager method, so a deleted workflow leaves no lingering presence/meta even if a per-socket removal failed or a socket joined mid-teardown. (Cursor: "Deletion skips final room wipe".) Files (#5932 comments): - workspace-file-manager.uploadWorkspaceFile now fans out the live-tree signal (all direct-upload paths: multipart fallback, copilot create, /api/files/upload, v1 files — the presigned path already notified). (Cursor: "Creates miss live tree fan-out".) - use-workspace-files-room: clear the pending retry timer on join success; and a module-scoped intended-room guard defers the unmount `leave` so a rapid remount re-claims the room and skips a stale leave — fixing presence flap + a leave-after-join race. (Cursor: "Retry timer survives join success" + "Remount churns files presence".) - workspace-files handler: roll back a partial join (leave room + remove presence) in the catch, mirroring the workflow join. (Cursor: "Join failure skips membership rollback".) +2 tests (deleteRoom). 127 realtime tests pass, both apps tsc clean, api-validation + boundaries green. * fix(files): scope workspace-files leave to a workspace (deferred-leave safety) Self-review of the deferred-leave guard found a real bug: leave-workspace-files was not workspace-scoped, so after a workspace switch (A->B) the deferred leave from A would evict the socket from its new room B. The leave now carries the workspaceId and the server no-ops if the socket's current files room differs. Also excludes the leaving socket from the leave broadcast (consistent with disconnect). * fix(realtime): close files-room presence leak + validate join payload Architecture-audit findings: - S1 (real Redis leak): the files room inherited the shared manager but not the workflow join's liveness sweep, so an UNGRACEFUL disconnect (pod crash — no `disconnecting` event) left its presence entry in the no-TTL room hash forever. Added a shared `sweepStalePresence(manager, room)` (fetchSockets liveness + remove not-live-AND-stale entries, matching the workflow 75min threshold) and run it on files join; also filter the join ack through `filterVisiblePresence` so a joiner never briefly sees an un-swept ghost. - S2: validate the client-supplied `workspaceId` on files join before it reaches the DB query (matches the /api/workspace-files-changed guard; fails closed). - N2: corrected the notify doc — it is awaited (guaranteed dispatch before a Node route returns) and hard-bounded to NOTIFY_TIMEOUT_MS, not "never block". +1 test (sweepStalePresence keeps live/fresh, reclaims not-live-stale). 128 realtime tests pass, both apps tsc clean, biome clean. * fix(realtime): workflow-deletion always notifies + cleans by socket.io membership Review-round findings on #5937: - Always emit `workflow-deleted` (was guarded by users.length>0), so a socket still in the Socket.IO room after a Redis presence eviction is told the workflow is gone before socketsLeave kicks it — the editor no longer keeps showing a deleted workflow. (Cursor: "Silent kick skips deletion event".) - Clean per-socket state for the UNION of live Socket.IO members and presence-tracked sockets, so an evicted/late-joined socket's room mapping + session are dropped too — not just presence-snapshot sockets. (Greptile: "Room deletion leaves reverse state".) - deleteRoom now logs AND rethrows on Redis failure (like addUserToRoom) so a failed wipe isn't reported as a clean deletion; the request surfaces it. (Greptile: "Room deletion failures are suppressed".) The two "deferred leave drops new membership" P1s were already fixed by the workspace-scoped leave in a prior commit (leave carries { workspaceId }; server no-ops on mismatch). 128 tests pass, tsc + biome clean. * refactor(files): drop module-scoped deferred-leave; rely on workspace-scoped leave Removes the one non-idiomatic construct (a module-level mutable `intendedFilesWorkspaceId` + queueMicrotask). It only guarded a same-workspace CONCURRENT remount, which doesn't occur in production (folder nav is shallow/no remount; list<->detail is sequential) — a dev-StrictMode-only case. The real cross-workspace race is already handled by the workspace-scoped leave: if B's join runs first (auto-leaving A), A's leave no-ops because the socket's current files room is B. Simpler, idiomatic, prod-correct. * feat(realtime): Yjs relay server for collaborative document editing [4/N] (#5941) Server-side Yjs relay for collaborative document editing (live carets + text selection) in the Files rich-markdown editor. Faithful y-websocket-style relay over the existing authenticated Socket.IO connection + shared room abstraction; in-memory Y.Doc + Awareness per file; awareness ownership binding, userId-keyed client-id uniqueness, seeder election with deadline re-election, concurrent-JOIN generation guard. 25 relay tests. Reviewed to Greptile 5/5 + Cursor pass across multiple rounds, plus an independent 4-lens audit (correctness/security/conventions/simplicity) and /simplify + /cleanup passes. * feat(files): collaborative document editing — client provider + editor (#5946) Client Yjs provider (FileDocProvider over the authenticated socket) + TipTap Collaboration/CollaborationCaret wiring for live carets + text-selection in the Files rich-markdown editor. Collaboration is a Files-page-only surface (explicit `collaborative` opt-in), disjoint from agent-streaming. Read-only + autosave-gated until synced+seeded. Merges into the realtime-rooms integration branch. * feat(tables): live collaboration — cell-selection presence + live mutation propagation (#5957) * feat(tables): live cell-selection presence — protocol + server + client hook The realtime spine for Google-Sheets-style table presence (mode A, socket): - @sim/realtime-protocol/table-presence: centralized wire protocol (events + TableCellSelection {anchor, focus, editing} + payloads) so server emits and client subscriptions can't drift. - ROOM_TYPES.TABLE + resolveTableWorkspace registered in ROOM_WORKSPACE_RESOLVERS (tableId -> workspace via userTableDefinitions, honoring archivedAt); roomName / presenceEventName / disconnect cleanup / authorizeRoom all derive automatically. - apps/realtime/src/handlers/tables.ts: join/leave (mirrors workspace-files) + a table-cell-selection relay (mirrors the workflow selection channel), broadcasting via roomName(room) since table rooms are namespaced. UserPresence gains a cell field threaded through the memory + Redis managers (Lua ARGV[7], null clears). - Extracted the duplicated resolveAvatarUrl into handlers/avatar.ts. - use-table-room.ts client hook: joins over the shared socket, tracks the roster (avatars) + patches per-socket cell deltas, exposes a throttled emitCellSelection. Grid UI (avatars + selection overlay) lands next; concurrent cell-value edits (last-write-wins via the durable log) are the follow-up PR. * feat(tables): render live cell-selection presence in the grid Wires the table presence room into the grid UI: - Page (table.tsx): useTableRoom (gated off in embedded/mothership mode) — renders <PresenceAvatars> in the header and passes remoteSelections + emitCellSelection down to the grid. - Grid emits its local selection: an effect resolves the index-based anchor/focus to stable (rowId, columnId) via refs and broadcasts it (with an editing flag for the active cell) through the throttled emitter. - RemoteSelectionOverlay: draws each remote viewer's selection in their color (getUserColor), a darker fill while editing, and name-on-hover — measured from live cell rects in the content wrapper's space (scrolls with the grid), hidden when rows are virtualized off-window, pointer-events-none so it never blocks cell clicks (hover via pointer hit-test). * test(tables): cover the table presence handler Mirrors workspace-files.test.ts: join auth/unavailable/denied/success, plus the cell-selection relay (asserts it persists via updateUserActivity and broadcasts on the namespaced roomName, not the bare id) and leave. * feat(tables): propagate manual cell edits live (last-write-wins) A manual row edit now appends a lightweight 'edit' event to the durable table stream; collaborators refetch the row (via the existing debounced rows-invalidate the job events use) so the winning value shows live. The event carries no value — peers refetch in their own wire format, so there's no auth-specific value translation on the wire, and last-write-wins falls out of the DB's committed order (the Google-Sheets model). Edits that also trigger a dispatch already emit dispatch/cell events; the debounce coalesces the two. * refactor(tables): apply /simplify findings - Drop the dead 'add unknown peer' upsert branch in use-table-room (Socket.IO ordering guarantees a peer is in the roster before their selection delta). - TableCellSelectionBroadcast = TablePresenceUser & { cell } (was a copy-paste). - Make TableGrid's presence props required + drop the unused empty-default/guard (only table.tsx mounts it, always passing both). - Drop the unused rowId from the 'edit' event (the handler invalidates all rows). - Overlay: subscribe scroll/resize/pointer listeners once per scroll element and cache the wrapper origin, so incoming deltas re-measure without re-subscribing and the pointer hit-test never forces a per-move layout read. - Server: cache the immutable socket session so a selection delta no longer reads it from Redis every time. * refactor(tables): apply /cleanup findings - Fix the remote-selection name label contrast: text-white is unreadable on the light-pastel user colors (same bug the Files caret fixed) → fixed dark #1a1a1a. - Re-measure via useLayoutEffect so a moving peer selection updates before paint (no one-frame position lag). - Drop 'mothership' from a comment (constitution copy rule). Six cleanup passes ran (effect, memo/callback, state, react-query, emcn, comment); the rest confirmed clean — all state/memos/callbacks/effects are load-bearing, presence correctly lives in useState (socket-pushed), and the edit→rows-invalidate granularity is right. * feat(tables): propagate every table mutation live (edit + schema signals) Comprehensive live collaboration for all user table mutations, via two value-less durable signals + named helpers (signalTableRowsChanged / signalTableSchemaChanged): - edit (rows refetch): single + batch row create, cell/row update, batch update, delete by id/filter, and upsert. - schema (definition + rows refetch): column add/update/delete, workflow-group add/update/delete, table rename, and CSV import (which can add columns). - Client handles 'schema' by invalidating the table detail (exact) + rows. Execution paths (column run, cancel-runs) and async jobs (delete/import-async, job-cancel) already propagate via cell/dispatch/job events — verified applyJob refetches on terminal. No reorder routes exist. Table archive (route DELETE) is a deliberate follow-up: it needs a table-deleted redirect event, not a refetch signal (which would 404). * refactor(tables): apply comprehensive /cleanup audit findings Holistic + react-query + comment audits over the whole PR: - Security/crash fix: a remote peer's rowId flowed unescaped into the overlay's querySelector — a hostile id ('x"]') threw SyntaxError inside a useLayoutEffect, crashing every other viewer's page. CSS.escape it, and validate + whitelist the untrusted cell payload server-side (shape + 200-char id bound) before it is stored/rebroadcast. - Simplify the CELL_SELECTION relay: the delta attached userId/userName/avatarUrl that the client discarded (identity comes from the roster). Drop them + the getUserSession lookup/cache entirely — the delta is now { socketId, cell }. - React Query: schema handler also invalidates lists() (parity with the local column-mutation set); document that the mutating client self-refetches by design. - Comment tightenings; biome fixed a stale import order in workspace-files.ts. * fix(tables): broadcast single-cell selections (focus falls back to anchor) Cursor High: a normal cell click leaves selectionFocus null (the grid treats it as a one-cell selection via focus ?? anchor), but the presence emit required BOTH anchor and focus to resolve — so the most common selection never broadcast and clicking even cleared a prior remote outline. Mirror the grid's focus ?? anchor semantics. * fix(tables): reviewer + regression + per-LOC audit findings Cursor review round (5 findings) + regression audit + per-LOC audit: - Presence roster snapshot now KEEPS the cell we already hold for a known socket, so a join/leave broadcast can't revert a fresher CELL_SELECTION delta. - Reset the selection throttle on table switch (was unmount-only), so a pending selection for table A can't flush into table B's room after a switch. - Metadata writes (column widths, display) use a new lightweight 'metadata' signal that refetches only the definition — a resize no longer forces peers to refetch rows. - Overlay re-measures on row add/remove/reorder via a tbody childList MutationObserver (a live refetch moves cells without a scroll/resize). - Document the actor self-refetch create caveat (scrolled multi-page insert) accurately. - isCellRef narrows to a partial instead of casting to the full type then re-checking; drop a redundant mount measure() (the layout effect covers it); text-[11px]→text-xs. * fix(tables): drop ineffective metadata propagation + re-measure overlay on column resize Cursor round onb8f28b04b: - Remove the 'metadata' signal entirely. The grid seeds columnWidths/pinnedColumns from metadata ONCE (metadataSeededRef) and deliberately never re-applies them (to avoid clobbering a local in-progress resize), so refetching the definition on a peer never surfaced their width/pin change — an ineffective path. Width/pin live-sync needs reconciliation that doesn't clobber a local resize; that's a deliberate follow-up, not a no-op refetch. Structural changes still propagate via 'schema'. - Overlay now also observes the content layer with the ResizeObserver, so a column resize (which grows the content, not the scroll container) re-measures remote outlines. - Presence-merge comment now states both sides of the trade-off. * fix(tables): re-broadcast local selection on (re)join Cursor Medium: a selection made before the room join completes (or held across a reconnect) was dropped server-side and never re-sent, so peers didn't see it until the local user moved it again. Track the current selection in a ref (set on every emit, cleared on table switch) and re-emit it from handleJoinSuccess once the room is joined. * fix(tables): re-broadcast selection when a peer's row change shifts it End-to-end lifecycle audit (Low-Med): the selection emit resolved the stable (rowId, columnId) only on selection/editing change, not when a live edit/schema refetch inserted/deleted/reordered rows. The index-based local selection then sat on a different logical row than the rowId peers held, so your outline showed on the old row until you moved. Re-run the emit on rows/displayColumns change and dedup an unchanged result (also drops the redundant null-on-open emit) so the broadcast stays consistent with the local highlight. * fix(tables): schema invalidates run-state/enrichment + guard stale join Cursor round oncdc8796b8(2 Medium): - schema handler used detail exact:true, so it skipped the activeDispatches + enrichmentDetails sibling queries the local invalidateTableSchema refreshes via a prefix match. After a peer deletes/restructures a workflow group, peers could keep a stale running badge or enrichment panel. Now invalidates both siblings too (rows stay on the debounce). - Guard against a stale join stealing the room: a fast table A->B switch could let A's async authorize finish after B, leave B, and strand the socket in A. Added a per-socket monotonic join generation checked after authorize (mirrors the file-doc relay's guard) + a test. * feat(tables): live column width/pin/order sync Collaborators now see each other's column resizes, pins, and reorders live — the last piece of Google-Sheets-style layout parity. - New lightweight `metadata` durable event kind (distinct from `schema`): only the table definition carries UI metadata, so peers refetch the definition alone — no rows/run-state refetch. The metadata PUT route now signals it. - The grid reconciles server metadata against its in-progress gesture: the column being actively resized keeps its live local width, and an in-flight column drag blocks a reorder apply — so a peer's change never reverts the local action. Each field is reference-guarded (React Query structural sharing keeps unchanged sub-objects stable), so an unrelated peer change doesn't re-apply the others. * fix(tables): escalate to schema signal when a reorder scrubs group deps Independent audit of the metadata-sync commit found a stale-run-state hole: a columnOrder PUT that moves a column left of a workflow group's leftmost column makes updateTableMetadata scrub that group's dependencies and write a new schema — a real structural change. But the route only fired the lightweight 'metadata' signal (detail-only refetch), so peers' and the actor's activeDispatches / enrichmentDetails queries stayed stale (a lingering running badge / enrichment panel) — exactly what the 'schema' handler exists to prevent. updateTableMetadata now reports whether it scrubbed the schema; the route emits signalTableSchemaChanged in that case and the light signalTableMetadataChanged otherwise. Width/pin/plain-reorder stay on the cheap detail-only path. * feat(realtime): accurate in-file presence + collaborative-caret polish (#5965) Per-session file-doc presence (avatars count other sessions like the canvas), StrictMode-safe stable Y.Doc (fixes blank-doc on join), flush caret cap + restored hover hit-slop, and three join-lifecycle race fixes unifying file-doc + workspace-files on one intent-tracked monotonic generation model. All findings root-caused with regression tests. * feat(files): smarter bullet delete/indent and untitled-file title sync (#5971) * improvement(files): smarter bullet delete/indent, fix empty-nested-bullet heading corruption Backspace at the start of a list item now outdents a nested item or clears a top-level item to a paragraph in place instead of deleting the row and jumping the caret to the previous block; Enter on an empty nested item outdents. Empty non-trailing top-level items still collapse cleanly since they cannot round-trip as a lifted paragraph. Also strips nested empty list-item marker lines on serialize: a nested empty bullet re-parsed as a Setext heading underline, silently turning its parent line into an H2 and dropping the bullet. Top-level empty items are preserved. * feat(files): sync an untitled file's name with its leading heading While a file is still named untitled(.md), typing a leading heading auto-renames the file after it (debounced), and renaming the file first seeds a leading H1 from the new name. One-shot: coupling stops once the file has a real name, and the heading seed always prepends so existing content is never clobbered. * fix(files): count inline atoms in list-item emptiness, keep multi-block items on Backspace Addresses review findings on the list Backspace logic: - Emptiness now uses the caret block's content.size (counts inline images/mentions), not textContent, so a bullet holding only a non-text atom is no longer treated as empty and deleted. - An empty first block whose item has sibling blocks removes only that block instead of lifting the whole item out of the list. * fix(files): preserve the untitled to named heading seed across a rename during editor load The parent captures the file name at mount (before content/session finish loading) and passes it as the transition baseline, so a rename that lands in the loading window is still seen as an untitled to named transition and the leading heading seed is not skipped. * fix(files): drop the name-to-heading seed, keep title sync one-way Removes the effect that inserted a leading H1 when an untitled file was renamed. On the collaborative Files page every open client observed the untitled-to-named transition and inserted into the shared doc, producing duplicate headings; it could also re-insert a heading a user had just deleted while a rename was in flight. Seeding document content from an async rename transition is the wrong model on a shared editor. The primary direction — typing a leading heading renames a still-untitled file — is unaffected (it never mutates the doc). * fix(files): keep empty lines between paragraphs on reload The chunked markdown parser (parseMarkdownToDoc) parses each block stripped of the blank lines between them, so it dropped the empty paragraphs @tiptap/markdown builds from runs of blank lines — a saved visual blank line silently vanished on the next load (the settle/reopen re-seed goes through the chunker). The whole-document parser preserves them, but whether a gap yields an empty paragraph is a global, block-type- dependent decision (kept between two paragraphs, dropped after a heading), so it can't be reconstructed block-locally. Route documents with empty-paragraph blank-line spacing to the whole-document parser for exact fidelity — the same tradeoff NON_CHUNKABLE makes; ordinary single-blank-line separation still takes the fast chunked path. Adds a suite asserting chunked output matches the whole-document parser for leading/trailing/between gaps and around lists/headings. * fix(files): only auto-name an untitled file when the user can edit The debounced untitled→filename hook ran on every onUpdate — including the mount-time seed and for view-only viewers — without checking edit permission, so a read-only user could schedule a rename they have no permission to make (a spurious, server-rejected write). Gate the derive-title on editor.isEditable (canEdit + settled + collab-ready, the same signal the autosave path uses), at both schedule and fire time. * fix(files): normalize line endings before the empty-paragraph guard; Enter/Backspace symmetry - markdown-parse: EMPTY_PARAGRAPH_SPACING/NON_CHUNKABLE tested the raw body, but a classic \r-only file (blank lines are \r) would miss the \n-anchored guard and still be chunked, dropping empties. Normalize line endings once up front so the routing guards, the chunker, and the parser all see the same \n. +CRLF/CR test cases. - keymap: Enter on an empty first block of a multi-block item now removes only that block (removeEmptyWrappedBlock) instead of exiting the list, mirroring the Backspace hasSiblingBlocks case — the trailing check no longer swallows multi-block items. +test. * fix(files): editor audit follow-ups (trailing-blank read-only, collab rename, over-strip) A 4-agent independent audit (UX vs inkeep + SOTA, cleanliness, adversarial correctness) surfaced these: - HIGH regression: files ending in a blank line opened READ-ONLY. The empty-paragraph routing preserved a TRAILING empty paragraph, but postProcess collapses trailing newlines → serialize/parse non-idempotent → isRoundTripSafe flipped the file read-only. A trailing empty paragraph can't be serialized stably, so parseMarkdownToDoc now strips trailing empty paragraphs and the guard no longer routes on trailing blanks. Interior/leading empties are unaffected. +regression tests. - Medium: the debounced untitled→filename rename fired on remote Yjs edits too, so every peer renamed and could rename from a not-yet-synced heading. Gate on isChangeOrigin (local edits only; false for non-collab surfaces). - Medium: stripEmptyListItemLines over-stripped a nested empty item that follows a same-indent sibling (a real placeholder the parser keeps). Narrowed to the actual Setext hazard — an empty item DIRECTLY under a shallower parent line — matching the function's own docstring intent. Probe-verified. +test. - Low: corrected untitled-title.ts docstring that described a reverse name→heading coupling removed during review. * fix(files): a remote edit must not cancel the local rename debounce The isChangeOrigin gate cleared the debounce timer BEFORE bailing on a remote update, so a peer's edit arriving within the 600ms window cancelled the local user's pending rename. Bail on isChangeOrigin first, before touching the timer; only local edits clear/reschedule it. * docs(files): correct EMPTY_PARAGRAPH_SPACING rationale after trailing-strip The stacked trailing-empty-paragraph strip made the older comment overstate a correctness necessity it no longer owns, mislabel trailing runs of 2+ blanks, and advertise dead CRLF handling. Reword to match what the code actually does. --------- Co-authored-by: Waleed Latif <walif6@gmail.com> * fix(realtime): access-revalidation multi-room safety + cleanup pass Fix a blocker surfaced by a full cleanup/simplify audit of the branch: the access-revalidation sweep (staging's workflow-only #5917) treated every entry in socket.rooms as a workflow id, but the generalized multi-room model puts namespaced files/tables/file-doc rooms on the same io. It would resolve those as bogus workflows, get null, and evict files/tables collaborators every ~30s. collectScanTargets now decodes each room name with parseRoomName and sweeps only workflow rooms; added a regression test and fixed the now-false TSDoc. Other audit fixes (all behavior-preserving): - workflow.ts reuses resolveAvatarUrl (drops db/user/eq imports duplicated from avatar.ts) - PresenceAvatars: mr-1 was baked into the shared component, silently adding a margin to the workflow sidebar stack; moved to an optional layout className, re-applied on the tables/file-doc header surfaces only - table DELETE routes only signal collaborators when rows were actually removed (matches PUT) - events.ts definition kind: drop the never-emitted reason:'schema', fix its doc - event-log: rename buildMemory -> buildEntry (it builds the entry on the Redis success path too, not just the memory fallback) - remove dead resolveWorkspaceIdForRoom export; parallelize per-socket removals in handleWorkflowDeletion; gate the table columnIndexById map on remote selections; move file-doc module TSDoc off the FileDocOwner interface; fix a stale @returns * fix(realtime,tables): close table-presence race + v1/copilot live-collab gaps Validated each issue with subagents before implementing the cleanest fix. - tables LEAVE in-flight-join race (B8): the table handler tracked no current-table intent, so an unscoped/same-table leave during an in-flight authorize left the socket stranded in the room (present in the roster, broadcasting a ghost until disconnect). Mirror workspace-files: a closure-local currentTableId + a leave that advances joinGeneration to cancel the racing join. + 3 regression tests. - v1 + copilot live-collab signal gap (D1): tables edited via the v1 public API or Sim/copilot emitted no edit/schema signal, so open collaborators didn't live-update. Add the signals at those call sites (add-only, matching the existing route seam) — never in the service, so execution writes can't double-emit. Sync-only for copilot bulk ops, guarded on affected/deleted count; async job branches stay covered by their kind:'job' events; create/delete/get untouched. - table join read consolidation (B5): sweepStalePresence returns its roster so the same-tab dedup reuses it instead of a second getRoomUsers. - shared authorize slice (B6): extract only the guard-safe authorize->allowed branch into resolveRoomJoinAuth, shared by the three room handlers (the full preamble stays inline — file-doc's generation capture sits mid-ladder and must not move). - resize-revert flicker (E3): a peer's value-less metadata event forces a refetch that could momentarily revert a just-finished local resize; a pendingWidthWriteRef keeps local widths leading until the width PUT settles. - embedded-mode stray emit (E5): gate emitCellSelection on a bound table id so the embedded surface stops broadcasting cell selections the server drops. * fix(tables): close two copilot live-collab signal gaps + harden presence sweep Follow-ups from a comprehensive review of the branch: - copilot batch_update_rows and import_file's inline append branch wrote rows but emitted no live-collab signal, so collaborators didn't see those edits live (the append's sibling replace branch already signalled). Add the guarded signal to both, matching the internal route. - sweepStalePresence now reads the roster before the fetchSockets liveness probe and returns it on a probe failure, so same-tab dedup still runs during a transient fetchSockets outage instead of being skipped. - reword an internal comment off the retired "mothership" term. * fix(realtime): guard table join commit + rollback against supersession; drop no-op eviction cleanup Review round on #5991: - Table join re-checked the generation only once after authorize, then awaited leave/sweep/avatar before joining + registering presence. A table switch or leave in that window stranded the socket in the wrong room, and the failure catch could tear down a newer successful join. Resolve the avatar up-front, re-check generation immediately before the membership commit (matching the file-doc join), and skip the rollback/error for a superseded join. + a post-authorize-window regression test. - access-revalidation cleanup treated removeUserFromRoom's no-op false as a transport failure and re-enqueued a still-connected socket forever. Only retry when the socket is still mapped to the room (a healthy null mapping means the entry is already gone). Repurposed the expired-mapping test to lock it. * fix(realtime): guard table join leave-prior against superseding join Round 2 on #5991: a superseded join's leave-prior could still run — during its getRoomForSocket await a newer join commits to its room, so currentRoom is that newer room and the superseded join would leave/remove/broadcast it before the final guard aborts. Re-check the generation immediately after the lookup await, before the leave mutation. Extended the post-authorize-window test to assert the superseded join never tears down the newer join's room. * fix(realtime): roll back a table join superseded during addUserToRoom Round 3 on #5991: after the final generation guard, A could join + register presence while a newer join B commits to its room during addUserToRoom's await — B's leave-prior can't observe A's half-written entry, so A's late write wins and strands the socket. Re-check after addUserToRoom and roll back A's own Socket.IO join + presence (scoped to A's room, never touching B). + a regression test hanging addUserToRoom mid-commit. * refactor(realtime): DRY table-join supersession guards; fix stale comment Cleanliness pass after the review rounds (no behavior change): - Extract the four identical `joinGeneration !== joinAttempt || socket.disconnected` checks into a named `superseded()` helper (the catch keeps its intentionally narrower check). - Remove a stale guard comment that was left stranded above the avatar resolve. - Document the best-effort rollback catch. * fix(realtime): file-doc rebind must not drop the current doc or leave a writable ghost Two Cursor findings on the file-doc client-id ownership rebind: - On a document switch, the prior room was left BEFORE the ownership check, so a CLIENT_ID_IN_USE rejection dropped the socket from the old doc without joining the new one (contradicting its own comment). Run the ownership check first, and leave the previous doc only once the rebind is guaranteed to succeed. - Reclaiming a client id removed the stale prior socket from owners + awareness only; its socketToRoomName + Socket.IO membership remained, and handleMessage's SYNC path gates on socketToRoomName (not owners), so it stayed able to write document frames until disconnect. Fully evict the reclaimed socket. + 2 tests. * refactor(realtime): serialize table join/leave to fix map-corruption at the root Round 4 on #5991 surfaced a race the generation guards structurally cannot fix: two concurrent joins for one socket race on the single-valued socket→room map — a stalled addUserToRoom for table A lands late, clobbers a newer join's map entry to A, and the rollback then wipes it, stranding the socket (map empty while it holds table B). Guards protect JS suspension points; they can't stop an in-flight Redis write from landing late. Fix per architecture review: serialize this socket's JOIN + LEAVE on a per-socket promise chain so their multi-step async Redis commits can never interleave — restoring the atomic-commit property the synchronous sibling handlers get for free. This DELETES the leave-prior guard and the post-commit rollback (the code that caused the bug); four generation guards collapse to two identical superseded() checks (skip a superseded queued op + one pre-commit check). Reworked the interleaving-specific tests into a fast-switch-skips-superseded test; the leave- cancels-join tests are unchanged. Local to the tables handler — no shared-infra change. * fix(realtime): always roll back a failed table join; re-elect file-doc seeder on reclaim Two review findings: - Table join: the catch skipped rollback when superseded, but a socket.join that landed before addUserToRoom threw leaves the socket in the Socket.IO room with no matching socket->room map entry — unreclaimable by any later op (cleanup keys off the map). Under serialization the skip is unnecessary (the newer op hasn't committed), so always roll back. Simpler + fixes the strand. - File-doc reclaim: fully evicting the prior socket didn't release the seeder role if it held it, so electSeederIfNeeded (which no-ops while seederSocketId is set) never re-elected and an unseeded doc stayed empty until the deadline. Clear the role on eviction so the join's election picks a new seeder. + 2 regression tests. * fix(realtime): close revoke-race ghost presence in workflow join An access-revalidation revoke landing between socket.join and addUserToRoom socketsLeaves the socket while its presence mapping does not yet exist, so cleanupEvictedSocket finds nothing to remove and the join then writes presence for a socket already out of the room — a ghost collaborator until the stale sweep. Hoist resolveAvatarUrl (the only await in that gap) above the re-auth check so the whole re-auth -> socket.join -> addUserToRoom section is await-free, matching the invariant the handler already relies on for the pre-join re-auth. + ordering regression test. * refactor(realtime): serialize workflow join/leave; drop dead room-authz limb Comprehensive independent audit follow-ups: - workflow.ts join/leave now use the same opChain + joinGeneration serialization as the sibling handlers (tables, file-doc, workspace-files). It was the only async presence path left unserialized, so a rapid workflow switch A->B (or a leave racing an in-flight join) could strand presence in room A — a ghost collaborator still receiving A's operation broadcasts until disconnect. The join now aborts a superseded op at start and again right before the membership commit, and the catch always rolls back a partial join. - leave-workflow drops the '&& session' gate: an idle user whose 1h session key expired (while the 24h room mapping is still live) can now leave cleanly instead of being stranded until disconnect. The room ref alone suffices. - authorizeRoom: remove the dead ROOM_TYPES.WORKFLOW resolver + its getActiveWorkflowContext import. Workflow authorizes through its own path and never flows through authorizeRoom; the map now honestly covers only the workspace-scoped types (files, file-doc, table). - Remove unused isSameRoom (zero callers) and a needless useMemo in PresenceAvatars (plain derivation, single copy). - Tests: 4 workflow serialization/leave regressions. All gates green: tsc (sim/realtime/packages) 0, 204 realtime + 11 protocol tests, biome, api-validation, boundaries, prune 14/25. * fix(realtime): align session TTL, harden committed joins from post-success rollback Per-module comprehensive audit follow-ups: - redis-manager: SESSION_TTL now tracks SOCKET_ROOMS_TTL (was 1h vs 24h). The room set outlived the session, and since getRoomForSocket reads the room set while the workflow handlers gate edits/presence on `room && session`, an active-but-idle collaborator got wedged into 'session expired' after 1h — sticky until reload (activity only EXPIREs the already-gone session; only addUserToRoom re-HSETs it). Both keys refresh together, so they now expire together (restores the pre-refactor consistency, where both shared one TTL). - workflow.ts + tables.ts: a 'committed' flag stops the join catch from rolling back a genuinely-joined user when a trailing ack/broadcast/metric step fails on a Redis blip (a pure getUniqueUserCount log-metric failure could otherwise kick a live collaborator after success was already acked). - workflow.ts: leave-prior now guards `currentRoom.id !== workflowId` (a same-workflow re-join no longer leave→re-adds and flickers peers' presence), and the join ack is liveness-filtered via filterVisiblePresence — both for parity with the tables handler. - platform-authz: honest docstring + 400 message for the workspace-scoped-only authorizeRoom map (workflow authorizes via its own path). - caret-presence: corrected an over-stated batching comment. - +1 workflow regression test (post-success failure keeps the user joined). Gates: tsc (sim/realtime/packages) 0, 205 realtime + 11 protocol tests, biome, api-validation, boundaries, prune. * fix(realtime): narrow join commit-guard to post-success; skip empty presence broadcast Two Cursor findings on the prior audit-fix commit: - The 'committed' guard in join-workflow/join-table was too broad: a failure BETWEEN the membership commit and the success ack (e.g. getWorkflowState) hit 'if (committed) return' and emitted neither success nor error, hanging the client while it sat in the room. Replaced with the narrower shape: only the purely-decorative post-success steps (peer broadcast + log-metric) are wrapped best-effort; anything before the success ack still rolls back and surfaces a retryable error, so the client retries instead of hanging — while the original goal (a benign broadcast/metric blip never kicking a live, acked user) holds. - broadcastPresenceUpdate read the roster via getRoomUsers, which swallows a Redis transport error to []. On a disconnect broadcast that emitted an empty roster and cleared every remaining collaborator's presence until the next healthy update. Split out a throwing readRoomUsers; broadcastPresenceUpdate now skips the broadcast on a read failure (getRoomUsers keeps its swallow contract). - Tests: pre-success failure rolls back + retryable error (no hang); post-success failure keeps the user joined. Gates: realtime tsc 0, 206 realtime tests, biome, boundaries, prune. * fix(realtime): harden seeder recovery, join-generation, and misc robustness Final line-by-line audit follow-ups (all LOW/MED, no P0/P1): - file-doc: a sole client whose seed FETCH fails was added to triedSeeders, re-election found nobody, and the document stayed permanently empty until reload. Re-offer seeding a bounded number of rounds (MAX_SEED_ROUNDS) before giving up. Also bound clientId to a non-negative integer (it is an ownership key). - tables + workspace-files: validate the room id BEFORE advancing joinGeneration, so a malformed/rejected join can't cancel a legitimate in-flight join. - workflow + tables + workspace-files: suppress the client-facing join error when the op was already superseded (a retryable error naming the abandoned room could make a client re-join and cancel its newer join). The rollback still runs. - redis-manager: set isConnected=true only after scriptLoad succeeds (and reset it on failure) so isReady() can't report ready while the Lua SHAs are null. - connection: apply the presence-bearing filter to the manager-removed set too (symmetry with the fallback path). - http.ts: validate workflowId on the four workflow endpoints (matching the files one). - client: clear a pending join-retry timer before rescheduling (reconnect churn no longer orphans a stray extra join); clear caret fade timers on plugin destroy; seed-effect cleanup reports NOT-ready (safe direction). - Tests: bounded seeder recovery, cell-selection strip-junk, TABLE round-trip, presenceEventName. Gates: tsc (sim/realtime/packages) 0, 208 realtime + 12 protocol tests, biome, api-validation, boundaries, prune. * fix(realtime): gate isReady() on loaded script SHAs, not just connection Follow-up to the prior isConnected change, which was incomplete: the redis client's 'ready' event flips isConnected=true on connect — before initialize() loads the Lua scripts — so a bare isConnected check reports ready while removeUserFromRoom / updateUserActivity would silently no-op on a null SHA. Gate isReady() on the SHAs too, so the POST endpoints return a retryable 503 during that startup window instead of proceeding against unloaded scripts. Standard readiness-probe discipline. * fix(files): offline read-only fallback when the realtime doc never syncs When the realtime server is unreachable (offline, server down, socket never connects), the collaborative editor would sit blank and read-only forever — content only arrives via provider sync events that never fire. Add a bounded connect-deadline to the Yjs provider: if no first sync lands within CONNECT_DEADLINE_MS, latch fatal and emit a synthetic non-retryable join-error — the exact path a real fatal rejection already uses, which seeds the file's stored content read-only. Latching fatal also stops a late reconnect from syncing server state in and merge-duplicating the locally-seeded content (the documented Yjs 'non-empty doc ignores initial value' gotcha). Deliberately NOT adding durable Yjs snapshot persistence / server-side seeding: TipTap can't run the markdown->Yjs conversion server-side (Collaboration extension errors under jsdom), and a durable binary snapshot would create a dual source of truth with the markdown file (edited by copilot / PUT / download). The client-seeder + bounded re-election is the correct architecture for a markdown-is-truth model; this closes its one real user-facing gap without persistence, a migration, or dual-truth. Timer cleared on first sync, on a real fatal rejection, and on destroy. +2 tests. Gates: sim tsc 0, 496 editor tests, biome. Needs live offline->reconnect verification. * fix(realtime): validate workflow join id before generation bump; scope file-doc join rollback to its target Two Cursor findings: - join-workflow bumped joinGeneration before validating workflowId (unlike tables / workspace-files, which I'd already fixed). A malformed/empty join could advance the counter and cancel a legitimate in-flight workflow switch. Validate the id first. +test. - The file-doc join catch called cleanupFileDocForSocket unconditionally, which keys off socketToRoomName. During a document SWITCH that fails before rebinding (e.g. a throw in client-id reclaim), that binding still points at the socket's PRIOR, valid document — so the rollback tore down a document the socket was validly in. Only run that cleanup when the binding already points at THIS join's target; otherwise the socket never registered as an owner here and the only leftover is a freshly-created empty room, dropped by destroyRoomIfIdle. Gates: realtime tsc 0, 209 tests, biome, boundaries, prune. * fix(files): drop late sync frames once fatal; file-doc join error suppression + retry-budget reset Final safety-audit findings: - CRITICAL: FileDocProvider.handleMessage had no fatal guard. After the connect deadline latched fatal and the editor fell back to a read-only local seed, a late SyncStep2 (slow server / flaky network / deploy) was still applied — merging server state into the seeded doc (content duplication) and flipping synced=true, which un-gated autosave and would persist the duplicate to the real file. fatal guarded (re)join but not inbound sync. Now handleMessage returns early when fatal. +test (late SyncStep2 after the deadline is ignored, doc stays empty + gated). - file-doc join catch now suppresses the client-facing error when superseded (matches workflow/tables/workspace-files) so a retryable error for an abandoned file can't make a client re-join and cancel the newer one. - table/workspace-files room hooks reset the retry budget on (re)connect so a prior full exhaustion doesn't block retries after a reconnect. - presence-visibility: corrected a stale TTL comment. Gates: tsc (sim/realtime) 0, 209 realtime + collab/hooks suites, biome, boundaries, prune. * feat(collab-doc): server-authoritative Yjs seeding (#6008) Server-authoritative Yjs seeding for collaborative file documents: the realtime relay fetches a Yjs seed built from the file's markdown (via the shared TipTap engine) and applies it once per room, replacing the client-seeder election/handshake entirely. - DOM-free markdown<->Yjs conversion core (markdownToYDoc / yDocToMarkdown / applyMarkdownToYDoc) reusing the client markdown engine for parity by construction - Internal x-api-key seed endpoint + realtime fetch; single attempt bounded under the client readiness deadline, guard-release for join-driven retry, read-only fallback on persistent failure - Client readiness gate = synced && server seed flag; jsdom wired for the Next standalone build (serverExternalPackages + outputFileTracingIncludes) Foundation only — copilot-into-doc + markdown projection + durable persistence are Stage C. * feat(collab-doc): Sim merge endpoint for copilot-into-doc (Stage C foundation) buildFileDocMergeUpdate(docState, markdown) computes the minimal Yjs diff that turns a live document into target markdown, via applyMarkdownToYDoc (a real updateYFragment diff, not a replace) — so a copilot rewrite merges with concurrent user edits instead of clobbering them. Exposed over the internal x-api-key /api/internal/file-doc/merge endpoint the realtime relay will call: the relay owns the doc, the app owns the conversion engine, so the relay ships the current state and applies the returned diff. Tested incl. concurrent-edit no-clobber. * feat(collab-doc): realtime apply-edit — merge copilot markdown into a live doc The relay can now stream a copilot edit into open editors: applyMarkdownToLiveFileDoc finds the seeded live room, ships its state to the app's /merge endpoint for a minimal CRDT diff, applies it (relaying to every editor, reconciled with concurrent user edits), and reports 'no-live-room' so the caller falls back to a direct file write when nothing is open. - Generalize the realtime->app request module (file-doc-seed.ts -> file-doc-app.ts) with a shared POST helper + fetchFileDocSeed/fetchFileDocMerge - POST /api/file-doc/apply-edit on the internal x-api-key HTTP surface, returning { applied } - Tests for the seeded-room merge relay and the no-live-room fallback * feat(collab-doc): stream copilot edits into open editors (Stage C) edit_content now, after its durable file write, best-effort merges the same markdown into the file's live collaborative document (markdown files only). If a collaborator has it open, the edit streams into their editor as a CRDT merge — reconciled with their concurrent typing — instead of the file silently changing under them; the editor's existing autosave mirrors the merged doc back to the file. No-op when nothing is open. Never blocks or fails the edit. * fix(collab-doc): strip frontmatter on merge; gate live-merge to markdown - buildFileDocMergeUpdate now strips YAML frontmatter (splitFrontmatter().body) exactly as the seed does. Copilot passes full-file content, so without this the frontmatter merged into the doc as editor content and autosave wrote it back over the file (corruption). - Gate the live-doc merge on isMarkdownFileName (new server-safe helper) instead of the over-broad !isDoc, so code/text edits don't pay the realtime round-trip for a format the collaborative editor never renders. * fix(collab-doc): make frontmatter collaborative so a merge can't revert it The editor re-attaches its open-time frontmatter on every autosave, so the Stage C merge (which triggers an autosave with no user action) could write stale YAML back over a copilot frontmatter change — silently dropping it. Carry the file's frontmatter in the doc's config map instead of locking it at open: the seed stores it, the merge updates it (only when it actually changed, preserving the no-op diff), and the editor re-attaches THAT value on save — falling back to the locked copy before the seed lands and for non-collaborative docs. A server-side frontmatter change is now reflected rather than reverted. New FILE_DOC_SEED.frontmatterKey; seed/merge tests cover it. * fix(collab-doc): re-sync draft on a frontmatter-only merge A server edit that changes only the frontmatter updates the config map but not the body fragment, so TipTap's onUpdate never fires — the autosave draft kept the stale open-time frontmatter, and an explicit save could revert the live change. Observe the config map and, on a frontmatter-only change, re-attach the new frontmatter to the current body and push a fresh draft (guarded on a synced body so it never races the seed's own onUpdate). * fix(collab-doc): order the apply-edit/merge timeouts (outer > inner) The Sim->realtime apply-edit timeout (4s) was shorter than the nested realtime->Sim merge timeout (8s), so the outer call could abort while the relay was still merging — the relay then applied the merge after edit_content had returned, racing a follow-on edit. Split the shared realtime->app timeout: the seed keeps 8s (it reads a cold blob), the merge gets a tight 3s (it is a pure conversion, no I/O), and the outer apply-edit is raised to 6s so it always outlives the inner merge. Cross-referenced in comments to prevent drift. * fix(collab-doc): address deep-audit findings (jsdom trace, timeouts, gate, races) A 4-agent LOC audit + precedent research (TipTap/Hocuspocus/Yjs docs confirm the core patterns are idiomatic) surfaced these real issues: - The /merge route also lazy-requires jsdom but was missing its outputFileTracingIncludes entry — a Docker/standalone build would 500 with MODULE_NOT_FOUND. Add it. - The four conversion timeouts encoded an ordering invariant (merge<applyEdit, seed<readiness) living only in prose across two apps. Hoist to a shared FILE_DOC_TIMEOUTS in @sim/realtime-protocol with a test asserting the ordering; both apps import it. - The copilot live-merge gate used an extension-only check, but the editor treats any text/markdown-MIME file as markdown. Replace isMarkdownFileName with a MIME-aware isMarkdownFile mirroring the client, so those files stream too. - applyMarkdownToLiveFileDoc had no per-file serialization; overlapping merges could each diff the same stale snapshot and apply out of order. Serialize per file via a promise chain. - Editor hardcoded 'config'/'initialContentLoaded' instead of FILE_DOC_SEED constants (drift). - Add .max() bounds to the merge contract body; document the best-effort-merge failure window honestly (open editor + merge failure can drop a copilot edit until reload — closed by the deferred durable-doc work). * feat(collab-doc): multi-replica shared Yjs backend + server-side markdown persistence Make collaborative file-doc editing correct across multiple ECS tasks (the per-process Y.Doc previously assumed one replica per file). - Shared Yjs backend over Redis Streams (apps/realtime file-doc-store): each file's stream is the ordered, replayable log of updates; a multiplexed XREAD tailer converges every task's in-memory doc. Coordinated single-seeder election (SET NX + empty-stream recheck) fixes split-brain seeding. - Doc-sync fans out to local clients + the stream; awareness stays on the Socket.IO adapter. Snapshot+XTRIM compaction trims only integrated entries. - Server-side persistence: project the live doc back to markdown via a new /api/internal/file-doc/persist endpoint, debounced during editing and flushed on last-disconnect, from the authoritative stream state. Collaborative editors no longer client-autosave, closing the copilot clobber-window. - Copilot merges apply through the stream (reach the live doc on any task) and serialize cross-task via a Redis merge lock. - Degrades to the original single-replica behavior when REDIS_URL is unset. * fix(collab-doc): harden distributed locks and durability from review Address Greptile + Cursor review of the multi-replica backend: - Merge lock: retry LONGER than the lock TTL (guaranteed acquisition, never merges against a shared base while a peer holds the lock) and AWAIT the stream write before releasing, so the next task never diffs a stale base. - Distributed locks (seed/merge/compact) now use ownership tokens + a compare-and-delete release (Lua), so a lock that expired and was re-acquired by another task is never stolen; acquisition fails CLOSED on Redis error. - Seed: publish the seed to the stream AWAITED under the lock before releasing, so a later seeder's empty-stream fence always sees it (closes the fence's publish-after-release gap); TTL kept at the readiness deadline. - Persist: persist the AUTHORITATIVE stream state even when this task's local doc was never seeded, and capture the local fallback synchronously so a last-disconnect flush never encodes an already-destroyed doc. - Publish: retry a transient xAdd failure so a Redis blip can't silently drop an edit from the shared log. * fix(collab-doc): only persist a doc a user actually edited Cursor review: a copilot durable write landing while a doc is being seeded could have the stale seed projected back over it on last-disconnect, clobbering the copilot edit even with no user changes. Gate server-side persistence on a genuine user edit (socket-origin update): a seed-only or merge-only doc is never projected back to the file (copilot writes the file durably itself), so it can't clobber a concurrent external write. * fix(collab-doc): close review-round race/durability gaps Address Greptile P1 + Cursor findings: - Seed publishes to the shared stream AWAITED *before* seeding the local doc, so a publish failure leaves the doc unseeded and the stream empty for a clean retry rather than serving an unpublished local seed a peer would re-seed over (split-brain). - flushPersist falls back to the synchronously-captured local snapshot when getStreamState throws (not only when it returns null), so a transient Redis read no longer drops the final durable write as the room is torn down. - streamHasContent fails CLOSED (returns true on xLen error): a Redis blip can no longer let the seed fence pass and double-seed. - Client collabReady initializes from the collaborative prop, so a collaborative editor never has a mount-window where client autosave could clobber the server write. * fix(collab-doc): unstick seed retry and persist tailed edits Cursor review: - ensureServerSeed clears serverSeedStarted when it aborts at the streamHasContent fence, so a fail-closed Redis xLen (or a genuine peer-seed) no longer strands the room unseeded with no retry. - Persistence dirty-tracking now marks a doc edited on any post-seed update, including a peer's edit relayed via the tailer (REDIS_ORIGIN), tracked via a seededObserved flag. The last task to leave persists real edits even if it only tailed them; the seed transition itself is still never counted, so a seeded-but-unedited doc is never projected back over the file. * fix(collab-doc): match client markdown post-processing on server persist Cursor review: yDocToFileMarkdown serialized the body with yDocToMarkdown only, but the editor save path runs postProcessSerializedMarkdown before applyFrontmatter. Server persist could therefore write markdown differing from a client save (empty list markers, callout un-escaping) — spurious blob churn / round-trip drift despite the byte-identical claim. Apply postProcessSerializedMarkdown in yDocToFileMarkdown so a server persist is byte-identical to a client save and the client's dirty-check baseline. Add a regression test guarding the composition. * fix(collab-doc): close durability gaps at deploy boundaries + audit polish From a comprehensive from-scratch audit (correctness, SOTA, cleanliness, feature-completeness): - Persist max-wait: a continuous edit burst kept resetting the 5s debounce and never persisted; cap it so a burst flushes at least every 20s, bounding unpersisted edits. - Graceful-shutdown flush: flushAllFileDocRooms awaited in shutdown so a rolling deploy / scale-in secures open edited rooms to durable markdown before exit, instead of relying on the stream + a surviving task. - Compacted-snapshot catch-up now marks the doc edited (REDIS_SNAPSHOT_ORIGIN): a snapshot folds seed+edits into one frame, so a task catching up purely from it no longer treats real edits as an unedited seed and skips persisting. - Polish: delete dead __setFileDocStoreForTest; bounded retry loop; rename acquirePersistSlot -> tryClaimPersistWindow with accurate docs; tailer object-identity guard; fix stale comments (seed route, edit-content autosave). * improvement(realtime): post-review fixes for collab dirty-state + relay lifecycle - collab editor no longer latches a spurious 'Unsaved changes' prompt: report dirty only when the client owns durability (canAutosave), since in a collaborative session the relay persists the doc server-side - guard shutdown against a double SIGINT/SIGTERM running teardown twice - disconnect local sockets before httpServer.close so shutdown exits gracefully instead of hitting the forced-exit timer (local-only, deploy-safe) - return 400 (not silent 200) on an invalid workspaceId in the files-changed fanout - clear a pending join-retry timer on reconnect so it can't fire a duplicate join * fix(realtime): make collab-doc seeding atomic to close split-brain window An adversarial concurrency audit found a split-brain double-seed vector: the seed used an advisory SET NX PX lock + a SEPARATE xLen fence + an unconditional xAdd (a non-atomic check-then-append). If the seed lock's TTL expired mid-seed (a >4s stall after the 8s seed fetch), a second task could acquire the freed lock over a still-empty stream, both fences read empty, and both append seeds with distinct Yjs client ids -> duplicated document content. - add an atomic SEED_IF_EMPTY_SCRIPT (append-iff-empty in one Redis step) + store.seedIfEmpty(); the emptiness check and append are now inseparable, so two tasks racing (even both past an expired lock) can never both seed - ensureServerSeed uses seedIfEmpty instead of streamHasContent-fence + publishAndWait; the seed lock is now purely an efficiency optimization (avoid a duplicate fetch), not a correctness dependency - fix the misleading comment that claimed a copilot merge is not counted as an edit in the multi-replica path (it round-trips as REDIS_ORIGIN and does count; a safe idempotent over-persist, never a lost edit) - add interleaving tests: seedIfEmpty atomicity/fence, the split-brain regression under an expired lock, a peer edit during attach catch-up, and concurrent two-task compaction * docs(realtime): align seed comments with the atomic-append correctness model Greptile flagged lingering doc drift: the module + shouldSeed comments still credited the seed lock + empty-stream check as the split-brain fix. Reframe them so the atomic seedIfEmpty is the exactly-once guarantee and shouldSeed is an efficiency gate only. * feat(realtime): live workspace tables list, sharing one invalidation-room impl (#6053) * feat(realtime): live workspace tables list, sharing one invalidation-room impl Bring the tables list to parity with the files list: a create/rename/move/delete/ restore now propagates to every viewer live instead of waiting out the 30s staleTime. Following the files pattern, but factoring the two into one shared implementation rather than copy-pasting. - add ROOM_TYPES.WORKSPACE_TABLES + its authz resolver (workspace-id-addressed, reuses the workspace resolver like workspace-files) - extract setupWorkspaceInvalidationRoom (server) and useWorkspaceInvalidationRoom (client) — the presence-free, workspace-scoped live-list room; files and tables now both bind to it, so they can never drift. Event/room names derive from the room type. Replaces the standalone workspace-files handler + hook - notifyWorkspaceTablesChanged fanout fired from the table service (createTable, renameTable, moveTableToFolder, deleteTable, restoreTable) so it covers both the HTTP routes AND copilot, which call the service directly - relay /api/workspace-tables-changed endpoint; wire the hook into the tables page - consolidate the handler test into one suite run against both room types * feat(realtime): live tables list also covers table-folder mutations Fold in the follow-up: a table folder create/rename/move/delete/restore now propagates to the tables list live too, so the browser is fully consistent. - generic notifyFolderResourceChanged(resourceType, workspaceId) dispatches the workspace live-list signal by resource type (a map, not a special-case if), so file/knowledge_base/workflow are no-ops today and gain liveness by adding a map entry when they adopt an invalidation room - fired from the shared folder lifecycle (createFolder/updateFolder/deleteFolder/ restoreFolder), covering routes AND copilot - the tables room hook now invalidates the table folders query too, not just the tables list, since the page renders both * fix(realtime): skip per-table live-list notify during a folder cascade A folder delete/restore already fires one folder-level notifyFolderResourceChanged for the whole subtree, but the cascade also calls deleteTable/restoreTable per table — each awaiting its own notifyWorkspaceTablesChanged. A folder with many tables would run N+1 sequential relay calls (each bounded by NOTIFY_TIMEOUT_MS), blocking the mutation. Add a skipNotify option the cascade passes so only the one folder-level notify fires. * feat(collab-doc): Hocuspocus binary persistence + Next 16 seed/persist fixes (#6059) * fix(collab-doc): make server-side seed conversion work under Next 16 / Turbopack Opening a file left both collaborators read-only and stalled ~12s: the server-side seed (markdown -> Yjs, run through the headless editor engine) was failing, so the doc never seeded and the editor never left its readiness gate. Two root causes, both latent until a real build/runtime (typecheck + unit tests don't exercise either), surfaced by the Next 16 upgrade: 1. Build boundary: the server seed route imported the shared editor schema (`createMarkdownContentExtensions`), which pulled in the React node-view components (`useEffect`) -> 'client component in a Server Component'. Split each node's React-free schema into its own `*-schema.ts` (code-block, image, raw-markdown-snippet); the client editor still injects the React node views via the existing `nodeViews` param, unchanged. 2. Runtime DOM: the converter installs a jsdom `window` on `globalThis`, but Turbopack's server bundle gives bundled `@tiptap/core` a `window` that does NOT read `globalThis`, so `elementFromString` threw 'no window object available'. Externalize the `@tiptap/*` packages the converter uses (native Node require, so their `window` reads the real global) and fix the converter's DOM guard to gate on `window` (what TipTap checks) with no sticky flag. Verified: seed route returns 200 with the Yjs update; 514 collab-doc + editor tests pass; schema byte-identical after the split. * feat(collab-doc): persist the Yjs binary and load it on cold-start (Hocuspocus pattern) Adopt the industry-standard Hocuspocus store/load-document pattern so a cold room open loads the file's last-persisted Yjs binary directly instead of re-converting markdown -> Yjs on every open. Rebuilding the CRDT from markdown on each connect is the exact anti-pattern Tiptap/Yjs warn against (fresh client ids -> duplicated content); it also forced the fragile server-side headless-editor conversion on every open. Now conversion runs only on a genuine first open or an external markdown edit. - new table workspace_file_collab_state(file_id PK->workspace_files cascade, doc_state bytea, source_hash, updated_at): the Yjs binary + a hash of the markdown it was derived from (bounded <=~1MB by the 256KB round-trip gate). Mirrors Hocuspocus's extension-database (binary in a DB column). Migration 0275. - persist upserts the binary (tagged with the exact markdown just written) - cold-start seed returns the cached binary when its source_hash matches the file's current markdown; otherwise converts (and the next persist refreshes the cache) - also externalize yjs / y-protocols / lib0 alongside @tiptap: bundling loaded a second yjs copy, so @tiptap/y-tiptap's 'instanceof Y.XmlElement' failed on app-created nodes ('Unexpected case') during Yjs -> markdown Verified end-to-end: seed -> persist -> seed returns the exact persisted binary (a cache hit, no re-conversion). 18 collab-doc + 51 realtime file-doc tests pass. * fix(collab-doc): best-effort cache read + drop dead barrel - seed: a cache-read failure (transient DB error, not-yet-migrated cache table) no longer aborts a cold room open — the durable markdown is already in hand, so fall through to conversion. Symmetric with persist's best-effort cache write. Addresses the Cursor Bugbot finding on the read/write asymmetry. - remove the collab-doc index.ts barrel: nothing imported it (every consumer uses direct ./seed / ./merge / ./converter imports), so it was dead re-export surface. De-export COLLAB_DOC_FIELD accordingly — it is used only inside converter.ts. * fix(collab-doc): stream every external file write into open editors, not just edit_content (#6070) * fix(collab-doc): stream every external file write into open editors, not just edit_content A copilot/mothership edit to an open markdown file did not appear live in another user's editor: the live-doc merge bridge (mergeEditIntoLiveFileDoc) was wired into the edit_content tool ONLY. Every other server-side write — the file tool (/api/tools/file/manage), function_execute (/api/function/execute via writeWorkspaceFileByPath), create_file overwrite, and the PUT /content route — went straight to updateWorkspaceFileContent and skipped the merge, so the durable file changed but the open editor never updated. Confirmed from live logs (the mothership 'Prepend sentence' ran read + file + function_execute — zero apply-edit calls) and Redis (the prepended text was absent from the doc stream). Centralize the merge at the one chokepoint every external writer shares: - updateWorkspaceFileContent gains an opt-out "syncLiveDoc" (default on) and, after the durable write, merges markdown writes into any open collaborative doc (best-effort; no-op when nobody has it open). Any current OR future writer is covered automatically. - persist.ts opts out (syncLiveDoc:false) — it IS the doc→markdown projection, so merging it back would be a persist→merge→persist self-loop. - create_file opts its empty shell out (real content arrives via a later write) so an open editor never flickers to empty on overwrite; threaded through writeWorkspaceFileByPath. - edit_content drops its now-redundant explicit merge call (the chokepoint handles it). - binary writers (image/video/audio/ffmpeg/download) are naturally excluded — the merge is gated to markdown, the only format the collaborative editor renders. Also bump the api-validation route baseline 994→996 to match the true route count already on this branch (pre-existing ratchet drift from an earlier merge; NOT added by this PR). * fix(collab-doc): defer setEditable out of the render phase (flushSync warning) The collab editability-reapply effect called editor.setEditable synchronously. In collab mode isEditable flips from readiness (synced + seeded), which is driven by a Yjs config.observe firing synchronously inside Y.applyUpdate — so the effect can run while React is mid-render. TipTap's React binding commits setEditable's transaction with flushSync, which throws "flushSync was called from inside a lifecycle method. React cannot flush when React is already rendering." Defer the setEditable to a microtask (runs right after the current commit, before paint), guarding against a destroyed editor or a stale value before it fires. Only the collab path (this effect) hit the warning; the streaming/settle effect's setEditable calls run on the non-collab path where isEditable isn't driven by a mid-render Yjs observer. * fix(rich-markdown-editor): defer non-collab settle/stream mutations off the render phase (flushSync) (#6073) * fix(rich-markdown-editor): defer non-collab settle/stream mutations off the render phase (flushSync) The non-collaborative streaming/settle effect called editor.setContent / setEditable / setTextSelection / focus directly in the effect body. setContent mounts the custom node views synchronously through the @tiptap/react flushSync path (tiptap#3764), so when this effect runs while React is mid-render it throws "flushSync was called from inside a lifecycle method." This is the second flushSync source (the collab editability effect was the first, fixed separately); it fires on the agent-streaming-into-a-non-collab-editor surface. Defer the effect-body view mutations to a microtask via a small runOffRender helper (runs right after the current commit, before paint; no-ops if the editor was torn down). The settle block is deferred as ONE microtask so setContent -> collapse selection -> setEditable -> focus keep their order. The streaming rAF tick is left untouched — it already runs off-render, so it keeps writing content directly. queueMicrotask is TipTap's own documented remedy for this warning. 497 rich-markdown-editor tests (incl. stream-settle-selection) pass; tsc + lint + api-validation + boundary + prune green. Needs a live check: stream an agent into a non-collab markdown file and confirm it still renders smoothly. * chore(rich-markdown-editor): trim verbose flushSync-defer comments * fix(rich-markdown-editor): drop superseded settle/stream microtasks via a run token runOffRender previously only guarded editor.isDestroyed, so if React ran the next reconcile pass (a newer stream or settle) before a queued microtask flushed, the stale microtask could still apply setContent/setEditable/setTextSelection over the newer state. Tag each effect run with an incrementing token; a deferred mutation applies only when its run is still the latest (and the editor is alive). A run token fits this effect's several early-return exits better than a per-exit cleanup flag. Addresses Greptile/Cursor review. * fix(rich-markdown-editor): never drop the settle selection-collapse under a superseded run The run token drops a superseded settle's microtask, but the settle had already flipped its state flags synchronously — so a pre-empting steady-sync run took the non-settle path and never collapsed the selection, leaving a post-stream select-all painting the leaf-in-selection decoration. Track the collapse as a debt (pendingCollapseRef): whichever deferred run ultimately applies — settle or the steady-sync path — clears it, so the collapse runs exactly once on the latest content. Addresses the Cursor review finding. * feat(tables): show live cell-selection carets in the embedded chat panel (#6081) The table cell-selection presence room was joined only on the dedicated /tables/[id] page (useTableRoom was passed an empty id in embedded mode). Join it in embedded too, so the mothership chat resource panel shows collaborators' live cell selections and broadcasts the local one. tableId is already resolved from props in embedded (the data event stream already uses it un-gated), and authz runs on join, so this is safe. Avatars are unaffected — they render only in the !embedded Resource.Header, so the panel gets carets without avatars. * feat(collab-doc): If-Match optimistic concurrency so persist never clobbers an out-of-band edit (#6085) * feat(collab-doc): optimistic-concurrency guard so persist never clobbers an out-of-band edit The relay projected the live Yjs doc back to durable markdown unconditionally (last-write-wins), so a persist already in flight when an external write landed could overwrite it. Add RFC 7232 If-Match optimistic concurrency end to end, reconciling through the CRDT (never rejecting user work): - updateWorkspaceFileContent gains an expectedUpdatedAt guard: the write commits only if the file is still at that version (checked against the SELECT ... FOR UPDATE-locked row, so it is atomic with the write), else it throws the new ContentVersionConflictError without clobbering. - persistFileDoc takes expectedVersion and returns a discriminated result (persisted | missing | conflict). On conflict it returns the current durable content + version instead of writing. - The relay tracks the durable version its live doc is synced to — set on seed, advanced when a durable write is merged in (apply-edit carries the version), and on each successful persist. It is held cluster-wide in Redis (filedoc:syncver:{name}) so whichever task persists reads the same version, with the per-room value as the single-pod fallback. - flushPersist sends that version as If-Match. On a conflict it merges the current durable content into the live doc (so the out-of-band edit AND the live edits converge) and retries (bounded), so even a last-leave flush racing an external write persists the reconciled result rather than losing the session's edits. Threads the version through the seed + persist contracts and the apply-edit payload. No schema change (reuses workspace_files.updatedAt as the version token). Tests: app-side CAS (match writes, mismatch throws + cleans up the orphan upload), relay conflict handled gracefully without clobber/loop; 236 realtime + 76 sim collab/uploads tests, tsc x2, lint, api-validation, boundaries, prune all green. * chore(collab-doc): heartbeat-refresh the synced-version key TTL alongside its stream Keep filedoc:syncver:{name} alive as long as the room's stream (it was only re-set on seed/merge/persist), so an open-but-idle doc's persist If-Match token can't expire and force a needless reconcile. * fix(collab-doc): stop persist-conflict retries when there is no live doc to reconcile On an If-Match conflict with no live doc to reconcile into (last collaborator gone, no shared stream), applyMarkdownToLiveFileDoc returns no-live-room; re-projecting the same pre-teardown snapshot would only re-conflict, so break the retry loop immediately and leave the out-of-band (durable) content authoritative — the intended conflict policy. Addresses Greptile review. * fix(collab-doc): close three optimistic-concurrency edge cases from review - Single-pod persist retry projected the pre-reconcile snapshot (captureState always returned the initial localState), while the synced version had been advanced by the reconcile — so the If-Match could pass and clobber the reconciled edit. captureState now re-reads the live doc on each attempt (falling back to the pre-teardown snapshot only once the room is gone). - The synced version was recorded from this task's own seed FETCH before knowing whether this task's seed actually won; a peer winning with a different version could leave a newer token than the stream content. Record it only inside the didSeed branch (the task whose seed won); peer-seeded tasks read the winner's cluster value. - Persist wrote UNCONDITIONALLY when no version was available (relay version momentarily missing), which could clobber non-empty durable content. It now returns conflict for a non-empty file with no version (reconcile/retry once the version is re-established); an empty file's first write stays unconditional. * fix(collab-doc): defer (not reconcile) on missing version, and use the freshest version token - Missing-version persist now returns 'deferred' instead of 'conflict'. A missing version token (a Redis blip on a peer-seeded task) is NOT a genuine out-of-band change, so triggering a reconcile would wipe live edits (incoming-wins) even though nothing changed durably. Deferred means: don't write, don't reconcile — leave the edits in the stream and let a later persist write them once the version is re-established. - currentVersion now takes the MAX of the cluster (Redis) and local room versions rather than always preferring Redis, so a lagged/failed fire-and-forget Redis set can't shadow a newer local value and cause spurious If-Match conflicts. Versions are monotonic epoch-ms, so the larger is the later sync. * fix(collab-doc): make persist If-Match teardown-race-immune and recover missing version on final flush Close two last-leave concurrency holes Cursor flagged: - Thread the reconciled version LOCALLY through the persist retry loop. After a conflict+reconcile the correct next If-Match is exactly result.version, so carry it in a local var instead of re-deriving from room.syncedVersion/Redis. On a last-leave flush destroyRoomIfIdle removes the room from the map before the async flush finishes, so mergeMarkdownIntoRoom's recordVersion can no longer update room.syncedVersion — threading makes each retry's precondition correct by construction, immune to that dropped mutation and to a best-effort Redis re-read. - Cache the resolved version back into room.syncedVersion in currentVersion() so a peer-seeded/tail-only task (which never sets it locally) or a later transient Redis read failure still resolves it from the last value seen (monotonic max, never regresses). - On a FINAL flush, briefly retry resolving the If-Match when the version read momentarily fails, rather than deferring and stranding the session's edits in the TTL'd stream — the version is cluster-wide and heartbeat-refreshed. * fix(collab-doc): stamp cluster sync version the moment the seed wins, before the liveness guard The winning seeder set the If-Match token (room + Redis filedoc:syncver) only after the liveness/seeded guard that follows seedIfEmpty. But the tailer can integrate the just-appended seed DURING the seedIfEmpty await, so isDocSeeded(room.doc) is already true when the guard runs and it returns early — leaving the stream holding seed content with no cluster version. Later persists then send no If-Match, the app returns `deferred`, and session edits stay only in the TTL'd stream (the exact stranding this PR prevents elsewhere). Move the version stamp to immediately after seedIfEmpty wins, before the guard. Recording it only once our seed won (not from the fetch) is preserved, so it still can't shadow a peer's winning seed. * fix(collab-doc): make the synced-version token monotonic at every write site The If-Match token is written fire-and-forget from the seed stamp, merges, and persists, both locally and to Redis. An out-of-order write (e.g. a seed's lagged setSyncedVersion landing after a later merge's) could regress it below the version the live doc already incorporates, causing spurious If-Match conflicts — and on a last-leave flush with no live room to reconcile into, a spurious conflict leaves durable authoritative and drops the session's edits. - setSyncedVersion now writes via SET_VERSION_IF_NEWER_SCRIPT (Redis-side compare-and-set): it overwrites only when the new value is greater, refreshing the TTL either way. - recordVersion / the persisted branch / the seed stamp all take Math.max instead of assigning room.syncedVersion directly. Versions are monotonic epoch-ms, so "newer" is a plain numeric compare, exact within a Lua double. * fix(collab-doc): close three last-leave persist edge cases from review - Stale snapshot after reconcile (High): the multi-task captureState fell back to the pre-await localState snapshot even after a reconcile advanced ifMatch, so a failed stream re-read could persist the pre-reconcile state against the new version and clobber the out-of-band edit the reconcile just incorporated. NULL localState after a reconcile so a failed read aborts instead. - Lock miss aborts reconcile (Medium): a merge-lock acquisition failure returned 'no-live-room', indistinguishable from an absent stream, so flushPersist treated transient contention as terminal. Return a distinct 'merge-unavailable' and handle it as retry-later (edits stay in the stream), never as "nothing to reconcile into". - Peer syncver never recovers (Medium): the winner's setSyncedVersion was fire-and-forget with swallowed errors — the only way a peer-seeded task learns the durable version — so a dropped write left that peer deferring forever. Make it retry (bounded) like appendUpdate/seedIfEmpty; the monotonic script keeps a racing retry a no-op. * fix(collab-doc): scope the persist If-Match to a content version so metadata bumps can't clobber edits The optimistic-concurrency validator was `updatedAt`, which rename/move/delete/restore also bump with no content change. A racing live-doc persist then saw a stale token, got `conflict`, reconciled the pre-edit durable body via updateYFragment (incoming-wins on overlap), and wiped the user's in-flight edits. Scope the validator to content (RFC 7232 semantics — validate the representation, not the row): - New `workspace_files.content_updated_at` (NOT NULL, `now()` fast-default — no table rewrite). Advances ONLY on content writes (upload / overwrite / create); metadata writes never touch it. - The FOR UPDATE CAS, the merge-notify version, and the seed version all use `content_updated_at`. A rename now leaves it unchanged, so the persist If-Match still matches -> no spurious conflict, no reconcile, no lost edits. Genuine out-of-band content writes still conflict and reconcile. - Consolidated the collab schema into one migration (the collab-state table + the new column) per request, rather than a separate follow-up migration. Relay/store/contracts unchanged (still a numeric monotonic version). * chore(collab-doc): condense the densest persist comments (no behavior change) Cleanup pass: tighten the three longest comment blocks added while hardening the persist path (currentVersion cache, ifMatch threading, final-flush version retry) without dropping any invariant. No dead code found (biome lint clean; all new symbols referenced). * fix(collab-doc): persist must return the content version, not updatedAt Follow-up to the content-scoped If-Match: persistFileDoc still returned `updatedAt` as the version in both the persisted and conflict results, while the CAS/seed/merge all guard on `content_updated_at`. A content write sets both to the same instant, so it was coincidentally correct — until they diverge: if a metadata write bumps `updatedAt` past `content_updated_at`, the conflict path returned the larger `updatedAt`, so the relay's re-persist sent an If-Match the CAS (which checks `content_updated_at`) could never match → perpetual conflict → dropped reconciled edits. Return `contentUpdatedAt` in both paths so the relay's token always matches what it's checked against. * fix(collab-doc): defer persist whenever the version is missing; guard the content-version test - Empty-file CAS race (Medium): the unconditional-write carve-out for size===0 read `record.size` outside the write transaction, so a concurrent first content write could land after the check and be clobbered. With content_updated_at NOT NULL every existing file always has a real version, so a missing expectedVersion is always transient — always defer, never write unconditionally. Removes the TOCTOU hole. - Content-version test (Low): the merge-chokepoint test kept updatedAt == contentUpdatedAt, so it passed even if wired to the wrong field. Mock distinct values and assert contentUpdatedAt, so a regression to updatedAt now fails the test. * fix(collab-doc): don't reconcile a conflict the live doc already reflects (would wipe newer edits) flushPersist reconciled the durable body into the live doc on every conflict. But when the conflict comes from a racing self-persist (or an apply-edit the chokepoint already merged), the durable body is a STALE SUBSET of the live stream, and the incoming-wins updateYFragment merge moves the doc backward — wiping newer in-flight edits, which the retry then persists. Before reconciling, re-check the freshest synced version. If it already covers the conflict version, the live doc has already incorporated that content (or is ahead), so skip the reconcile and just retry with the freshest version as If-Match — the re-projection captures the current live stream, preserving every edit. Only a genuine out-of-band change the live doc hasn't incorporated (freshest < conflict version) is reconciled in. freshest never exceeds the durable version, so this can't loop. * fix(collab-doc): make content_updated_at monotonic per file; skip-reconcile can't loop The If-Match token was stamped with app-local new Date() on each content write, so cross-instance clock skew could stamp a later write with an EARLIER content_updated_at — breaking the version ordering the whole optimistic-concurrency scheme (and the skip-reconcile branch's freshest>=version assumption) depends on. Under skew the relay's monotonic syncedVersion could exceed the durable version, sticking the If-Match: persist conflicts forever, exhausts retries, drops the session's edits. - Stamp content_updated_at strictly after the current committed value (we hold the row's FOR UPDATE lock): new Date(max(now, currentFile.contentUpdatedAt + 1ms)). Monotonic per file regardless of clocks; also removes same-millisecond collisions. updatedAt stays plain wall-clock (display/sort). - Skip-reconcile branch retries with result.version (the durable value the CAS will match), never freshest (which could exceed it and loop). Belt-and-suspenders now that the version is monotonic. * refactor(collab-doc): drop the destructive in-persist reconcile; adopt-version-and-retry on conflict The in-persist reconcile projected the durable body back over the live doc via updateYFragment ("make the doc match"). That is destructive: when the live stream is already ahead — the common case, because the write chokepoint (mergeEditIntoLiveFileDoc) already merged the out-of-band change into the stream — it moved the doc backward and wiped newer in-flight edits. This produced a run of races (stale snapshot, wipe-newer-edits, version-lag skip miss) that a full-document reconcile fundamentally can't avoid, since deciding when it's safe relies on a laggy cross-task version token. Remove it. On conflict, adopt the durable version as the new If-Match and retry: captureState re-reads the current stream (which holds the out-of-band change AND the live edits), so the re-projection persists the converged result. The durable change reaches the live doc via the chokepoint, never here. Trade-off: the only unmerged out-of-band write is one whose chokepoint merge itself failed (rare, logged), which we accept over the frequent reconcile-wipes-edits race. - flushPersist: conflict -> ifMatch = result.version, retry (bounded). No applyMarkdownToLiveFileDoc. - conflict response drops `markdown` (contract + relay type + persist) — no body needed, saves a blob fetch. applyMarkdownToLiveFileDoc stays (still used by the apply-edit route / the chokepoint). * fix(collab-doc): don't let a last-leave conflict retry clobber via the stale local snapshot Regression from dropping the reconcile: on conflict the retry adopts result.version and re-reads captureState. But after single-pod last-leave teardown the room is already destroyed, so captureState falls back to the pre-teardown localState (which lacks the out-of-band change); the retry then CAS-passes and overwrites the committed external write — undoing the external-wins last-leave policy. Null localState on the first conflict, so the retry can only use freshly-read authoritative state (stream / live doc). When none is available (single-pod room gone, or a transient stream-read failure) captureState returns null and the retry stops, leaving durable content authoritative. Covers both the single-pod and multi-task-stream-unavailable variants of the stale-snapshot clobber. * fix(collab-doc): stop (don't re-persist) on a persist conflict — closes the commit-window clobber The conflict retry adopted the durable version and immediately re-persisted the current stream, assuming the stream already held the out-of-band change. But an external write commits durable BEFORE its chokepoint merge (mergeEditIntoLiveFileDoc) reaches the stream, so a persist landing in that window CAS-passed with a stream that still lacked the external content and clobbered the committed write — not just the rare merge-failed path, but a race on every external write, worst at last-leave flushes. Make persist a single attempt: on conflict, STOP and leave durable authoritative. The chokepoint merges the change into the stream and — only once it is actually there — advances the synced version via its own recordVersion; a later flush (debounced or final) then projects the converged stream with a matching token. The session's edits stay in the stream meanwhile. The conflict handler deliberately does NOT advance the synced version, or the next flush would clobber with a still-behind stream. Removes the retry loop and PERSIST_CONFLICT_RETRIES. * improvement(tables): fire the live-rows signal on async delete, run cancel, and column run (#6094) * improvement(tables): fire the live-rows signal on async delete, run cancel, and column run These three table operations mutate row data but emitted no `rows` change signal, so open editors' grids stayed stale until a manual refresh (enrichment *results* already stream live via `cell` events; these are the bulk paths that don't emit per-cell events): - Async row delete (`runTableDelete`): signal as rows drop out (throttled with the existing progress event) and once more on completion — the `job` progress event only drives the delete meter, not the rows query. Covers the delete-async route and the copilot bulk-delete, since both share the runner. - Cancel runs (`cancel-runs` route): cancelling clears each affected row's exec state; the `dispatch: cancelled` events drop the run overlay but the client then renders authoritative DB state, so refetch. Only when something was actually cancelled. - Run column (`columns/run` route): starting a run bulk-clears the target group's cells to pending; refetch so the cleared cells show. Only when a dispatch was actually created. Guarded so no signal fires on a no-op/failure. Adds a delete-runner test asserting the completion signal. * fix(tables): guarantee the live-rows signal on every mutating path (review) - Delete runner (Greptile P1): a batch could commit and the job then cancel/supersede before the next throttled progress signal or `markJobReady`, bypassing both signals and leaving deleted rows on screen. Track `deletedAny` and fire the grid refetch in a `finally`, so it runs on EVERY exit — completion, cancel/supersede, mid-batch lock, or a rethrown error after a partial delete. - cancel-runs / columns/run routes (Cursor): the `cancelled > 0` / `if (dispatchId)` guards don't always reflect DB row changes — cancel tombstones exec state even when 0 dispatches were active, and a run bulk-clears cells then can return a null dispatchId. Signal unconditionally; a stale-but-harmless refetch beats a missed one. - Tests: assert the delete signal fires on the mid-run-cancel-after-delete path and NOT when nothing was deleted. * fix(tables): mark deletedAny before the page delete so a mid-page lock still refreshes the grid `deletePageByIds` commits in internal batches, so a delete lock landing mid-page can persist earlier batches and THEN throw TableLockedError — the catch returns without a count, so setting `deletedAny` from the return value missed it and the finally skipped the grid refetch. Set `deletedAny = true` before the call (any attempt may commit rows); an attempt that commits nothing only over-refetches (harmless). Adds a test asserting the signal fires when a page throws a mid-page lock. * fix(files): make embedded resource file view collaborative (#6095) * fix(files): make embedded resource file view collaborative The /chat resource panel rendered saved files through FileViewer without the collaborative opt-in, so a file open on the Files page and the same file open in the embedded panel never joined the same file-doc room — no live carets and no live content sync between the two surfaces. Pass collaborative on the EmbeddedFile FileViewer. Collaboration still self-gates on canEdit + non-streaming + workspace doc, so the agent token-stream preview (the dedicated streaming-file path, canEdit=false) is untouched. * fix(files): refcount file-doc room membership per shared socket Two collaborative surfaces in one tab (the Files editor and the embedded chat resource panel) share one Socket.IO connection, so both providers for the same file JOIN the same room over that socket. The server's LEAVE does socket.leave(name) with no membership refcount, so the first provider's destroy() would strand the second still-mounted one — no more live content or presence. Count live providers per file per socket (keyed by the stable Socket object, so it survives reconnects) and emit LEAVE only when the last provider for a file tears down. The single-provider path is unchanged (0->1->0). * feat(tables): propagate shared saved-view changes to collaborators live (#6100) Table views (named filter/sort/layout presets) are table-wide shared state — every reader sees every view — but view create/update/delete had no realtime signal, so a collaborator only saw another user's view changes on their own staleTime/focus refetch. Add a 'views' table event kind + signalTableViewsChanged, emitted from the views service (createTableView/updateTableView/deleteTableView, on real success only), and a client handler that invalidates the views query alone (no rows/definition refetch — a view is presentation state on the loaded table). Mirrors how row/schema/metadata changes already propagate. * test(tables): cover the views realtime signal (emit + on-success-only) (#6101) - events.test.ts: signalTableViewsChanged appends a single 'views' event carrying the tableId (through the real memory buffer). - views/service.test.ts: create/update/delete emit signalTableViewsChanged on real success, and DON'T on a no-op (a PATCH/DELETE targeting a missing view changes nothing, so it must not signal). Mirrors delete-runner's signal-path coverage; drives the DB via the shared dbChainMock. - Add tableViews to the comprehensive @sim/db/schema test mock so the service tests can queue the in-transaction existence row. * chore(ci): reconcile api-validation baselines after the staging merge The staging merge unioned realtime-rooms's own `as unknown as` cast (lib/collab-doc/converter.ts) with staging's zod-recursive-type cast (lib/api/contracts/tables.ts), so the non-test double-cast count is 9 — both casts pre-existed and were individually accepted on their branches. Also tighten rawJsonReads 6->5 to the true current count. Fixes the strict API contract boundary audit on realtime-rooms. * feat(copilot): stream file edits into the live collaborative Y.Doc (keep embedded view collaborative) (#6108) * feat(copilot): stream file edits into the live collaborative Y.Doc Copilot's file edits previously only reached the live doc once, at the final edit_content write, so a collaborative editor watching the file saw nothing until completion (streaming looked broken) and the client-side preview path was suppressed in collab mode. Make copilot a CRDT peer: as it streams append/update/patch content, merge the growing markdown into the file's live Y.Doc via the existing apply-edit path (a minimal updateYFragment diff, concurrent-edit-safe), throttled to ~250ms. version is omitted for these intermediate merges — they advance the live doc for viewers but are not durable checkpoints; the final edit_content write carries the real contentUpdatedAt and reconciles the durable file. Per the relay's persist gating, server-internal merges never schedule a persist, so a copilot-only stream produces zero intermediate file writes. - notify.ts: mergeEditIntoLiveFileDoc version is now optional (streaming omits it). - file-preview-adapter.ts: throttled live-doc merge at the edit_content stream hook. * fix(copilot): order + gate streaming live-doc merges; fast collab first render Harden the streaming merge (adversarial review): - Order + bound: dispatch through a per-file in-flight guard (drop-while-in-flight) so a stale out-of-order snapshot can never land after a newer one and regress the doc, and relay load is capped at one request per file regardless of rate. - No wipe: gate append/patch on the base file content having loaded — a base-less snapshot would diff to a delete-everything wipe of the seeded doc; update streams a full rewrite from scratch and needs no base. - Markdown-only gate: non-markdown files have no collaborative room, so skip the wasted relay round-trip. Fast collab first render (Issue 2): render the already-fetched markdown read-only via generateHTML while the collaborative doc seeds, with the editor mounted-but- hidden in the same layout box for a seamless swap on collabReady. Pure HTML — it never touches the Y.Doc (client seeding duplicates the doc), and generateHTML escapes text (raw-HTML snippets render escaped), so no XSS. * test(copilot): cover streaming file edits into the live collaborative Y.Doc Drives edit_content args_delta stream events through processFilePreviewStreamEvent and asserts the live-doc merge: fires with the growing FULL previewText and no version arg; is throttled (~250ms per file); is skipped for non-markdown files and for a base-less append (the delete-everything wipe guard); and runs at most one-in-flight per file. Verified to fail if any gate/guard is removed. * fix(collab-doc): coordinate live-doc merge ordering in one place; close durable-clobber race The second review found a residual: the durable edit_content write went through a different path than the adapter's in-flight guard, so a late straggler streaming merge could land after it and, via a persist, clobber the durable file's tail. Move the per-file coordination into mergeEditIntoLiveFileDoc (the one place both the streaming and durable paths call): a streaming (versionless) merge is dropped while one is in flight for the file; a durable (versioned) write instead WAITS for the in-flight streaming merge, so the final content is always the last merge applied and can't be regressed by a straggler. Simplifies the adapter (drops its Set + helper). Relocate the one-in-flight test to notify.test.ts (streaming-drops-while-busy + durable-waits-then-applies-last); the adapter test keeps throttle/gates/previewText. * fix(copilot): address review — order merges, exclude update, gate throttle, unhide stream Review round on #6108: - Greptile P1 (durable merges lose ordering): serialize ALL merges per file on one chain in mergeEditIntoLiveFileDoc (each chains after the current tail), so concurrent durable writes can't resume-and-fire out of order. notify now exposes isLiveDocMergeInFlight. - Cursor High (update stream blanks the doc): only append/patch stream — they build on the loaded base; update is a from-scratch rewrite whose partial snapshot would diff the full doc toward a fragment, so it applies atomically at the durable write. - Cursor Medium (throttle advances on a dropped merge): the adapter gates on !isLiveDocMergeInFlight, so the send throttle advances only on an actual dispatch — no lag, no backlog behind a slow relay. - Cursor Medium (placeholder hides a live stream): show the fast-render placeholder only when not streaming, so a stream that starts before the doc seeds shows through the editor. - Soften merge.ts/notify.ts comments per the lifecycle audit: only UNTOUCHED regions are preserved; a region the merge rewrites reconciles toward copilot's content. Tests updated: notify covers chain ordering + isLiveDocMergeInFlight; adapter covers append streaming, throttle, non-markdown/base-less/update skips, and the in-flight skip. * fix(collab-doc): reject stale durable merges at the relay (cross-process ordering) The in-process merge chain only orders merges within one apps/sim process. Two durable writes for the same file on DIFFERENT processes could reach the relay out of dispatch order; the relay recorded the version monotonically but still APPLIED the older markdown, regressing the live doc while the token stayed high (a later persist could then write the stale content back over the durable file). Enforce ordering at the relay — the single cross-process coordination point — using the existing Redis primitives: under the per-file Redis merge lock, read the cluster-wide synced version and SKIP a versioned merge that is not newer (a newer durable write already landed). Make recordVersion await setSyncedVersion so it is durable before the lock releases, so the next holder's staleness check reads a consistent value. Streaming (versionless) merges are unaffected — they carry no durable version and are ordered per-process by the caller. Adds a relay test asserting a stale/idempotent versioned merge returns 'stale' and never computes or publishes a diff. * fix(copilot): match durable path — detect markdown by MIME type + name at the stream gate The streaming gate checked isMarkdownFile with only the filename, while the durable merge uses type + name — so a text/markdown file without a .md extension was skipped mid-stream (it self-corrected at the durable write). Pass editIntent.contentType so streaming detects the same set of markdown files as the durable path. * test(copilot): assert throttle follow-through after an in-flight merge clears * fix(collab-doc): order streaming merges by streamedAt so a late snapshot can't regress a newer durable write * refactor(collab-doc): tidy merge-order docs + relay order object; cover multi-replica streaming stale-check * fix(collab-doc): order streaming merges by causal base version, not wall-clock A streaming snapshot now carries baseVersion (the durable contentUpdatedAt it was built from) instead of a wall-clock streamedAt. The relay drops the snapshot when a newer durable write landed since that base, so a concurrent human save can no longer be clobbered in the live doc and then persisted over the durable file. Skew-immune: both keys are DB-monotonic contentUpdatedAt values. * fix(collab-doc): derive streaming baseVersion as contentUpdatedAt ?? updatedAt Match the version line the seed/persist use so a legacy file with no content version still ships an ordered streaming snapshot instead of an unordered one. * fix(collab-doc): fail-closed on a streaming snapshot with no baseVersion The live-merge gate now requires a numeric baseVersion, not just loaded base content. A rare base with no file record (hence no version) would otherwise ship an unordered snapshot the relay can't stale-check, risking a clobber of a concurrent durable write. Skip the live merge instead; the durable write reconciles. * docs(collab-doc): document the accepted concurrent-independent-streams limitation * chore(ci): reconcile api-validation baseline after the staging merge The merge commit auto-merged the baseline at 1000; bump totalRoutes/zodRoutes to 1003 for staging's three new contract-bound routes (nonZodRoutes still 0). * test(files): update storage-accounting assertion to the mergeEditIntoLiveFileDoc options object * fix(collab-doc): trace the full yjs/tiptap external stack into the file-doc route bundles The seed/merge/persist internal routes run the collab-doc converter (markdown <-> Yjs via headless TipTap) server-side. Those deps are serverExternalPackages, and the standalone tracer only force-included jsdom — it does NOT follow yjs's ESM subpath imports of lib0 (lib0/logging, ...), so Docker/standalone builds shipped node_modules without them and the seed route 500'd (Cannot find module 'lib0/logging'). That left every collaborative document unseeded and permanently read-only on deployed envs. Force yjs, lib0, y-protocols, and @tiptap into the trace for all three routes. * fix(collab-doc): copy the full yjs/lib0 stack into the app image The seed/merge/persist routes run the converter (markdown <-> Yjs) server-side. yjs is a serverExternalPackage and the Next standalone tracer copies lib0 only partially — it drops the ESM subpath file lib0/logging.js that yjs.mjs imports via lib0's exports map, so the seed 500s ('Cannot find module lib0/logging') and every collaborative doc is stuck read-only. Verified in the running dev container: /app/node_modules/lib0 had 37/38 files, logging.js missing. outputFileTracingIncludes can't fix it — its globs resolve against apps/sim, but these deps hoist to the monorepo-root node_modules, so the glob matches nothing (my prior next.config attempt was a no-op; reverted). Instead COPY the complete lib0/yjs/y-protocols from the deps stage in the runner, overwriting the partial trace — the same pattern already used for isolated-vm. * feat(files): stream copilot edits into the collaborative doc smoothly (#6122) * feat(files): stream copilot edits into the collaborative doc smoothly - apply the agent stream client-side into the live Yjs binding as minimal updateYFragment diffs (like main's setContent, but incremental) so it renders smoothly AND broadcasts to every peer via CRDT — a collaborator on /files sees the stream for free - gate the apply on collabReady so diffs never land on an unseeded doc; keep the read-only placeholder visible until the seed swaps in - run streamed ops under a dedicated tx origin so they stay out of the user's undo stack - delete the throttled server-side streaming merge and the baseVersion ordering machinery it needed (relay + notify + session contract); the durable final write still reconciles open editors and seeds late joiners * fix(files): apply agent stream as a true CRDT peer + guard base-less snapshots Review round 1 (Greptile P1s): - apply the stream against a private shadow replica (seeded from the live doc at stream start) and relay only the agent's own delta into the shared doc, so a concurrent peer edit to a region the agent snapshot didn't include is no longer reverted (previously the whole-body reconcile deleted it) - gate append snapshots on "must extend the base": a base-less append fragment (emitted before the base loads) can no longer reconcile the seeded doc to a wipe; patch still legitimately replaces a mid-region - gate the apply on collabReady so diffs never land on an unseeded doc; keep the placeholder visible until the seed swaps in - plumb streamOperation through the preview surfaces to drive the append gate - add a peer-edit-preservation test (fails under whole-body reconcile) and refresh the undo-isolation + broadcast tests for the session API * fix(files): destroy the agent shadow deterministically on settle Cursor round 1 (Low): endAgentStream ran inside runOffRender, whose microtask is dropped when a rapid follow-up stream bumps the run token — leaking the shadow Y.Doc. Split it out into an unguarded microtask queued after the (droppable) final apply, so the shadow is always destroyed. * fix(files): agent stream frames skip the relay's durable persist Cursor round 1 (High): client-applied stream frames broadcast over the sync channel, so the relay stamped a socket origin and ran schedulePersist — durably writing partial agent content mid-stream, attributed to the watching user (the old server-merge applied with no origin and never did). Restore that behavior: - new FILE_DOC_MESSAGE_TYPE.SYNC_NO_PERSIST wire tag; the provider tags AGENT_STREAM_ORIGIN updates with it (normal user edits stay SYNC) - the relay applies it under an AgentSyncOrigin (carries the socket id for broadcast exclusion, but is not a plain string) so originSocketId() is null → no edited/schedulePersist/lastEditorUserId; excludeSocketId() still excludes the sender, and the update still publishes to the stream so peers converge - the copilot's final edit_content write remains the authoritative durable persist - tests: relay applies+fans-out but never persists a SYNC_NO_PERSIST frame (verified it fails if applied as a socket edit); provider tags agent edits * fix(files): open the stream shadow at start + private extend baseline Cursor round 2: - High (settle skips apply without session): the stream shadow is now opened on the first ready frame, BEFORE the extend gate — so an `update` rewrite (whose every frame is gated out until settle) and a stream that finishes before seed still get a session, and settle applies the final body via the reused-or-on-demand shadow instead of leaving the doc stale until the durable reconcile. - Medium (peer edits stall the stream): the extend gate now reads a private `lastStreamedBodyRef` (the agent's own last frame), snapshotted at stream start, not `lastSyncedBodyRef` which `onUpdate` clobbers on peer edits — so a collaborator typing can't make the growing snapshot stop prefixing the shown body and freeze it. - Medium (multi-replica over-persist): pre-existing, documented "safe over-persist" (a peer task tails the frame as REDIS_ORIGIN and marks edited) — refreshed the stale comment to describe the SYNC_NO_PERSIST source; copilot's edit_content write remains the authoritative durable persist. * fix(files): fail-close base-less previews + operation-based stream hold Cursor/Greptile round 3 (High + Medium) — remove the fragile string-prefix "extend gate", which was the root of both findings: - Server: `buildFilePreviewText` now fails closed for an `append` whose base content hasn't loaded (returns undefined, like patch/update), so a base-less fragment never reaches the client. This eliminates the base-less wipe at settle (Greptile P1) at the source; an empty file (existingContent === '') still previews normally. - Client: the collab streaming tick no longer string-prefixes the raw preview against the editor's canonical markdown (the '*' vs '-' / emphasis mismatch that froze every append frame — Cursor). The mid-stream hold is now purely operation-based: `update` waits for settle; append/patch/create apply each frame via the (peer-safe) shadow reconcile. lastStreamedBodyRef is now a plain dedup guard, not a prefix baseline. Keeps the shadow, durable write, and SYNC_NO_PERSIST unchanged. * fix(files): elect a single agent-stream writer across tabs Cursor round 4 (High): with the stream applied client-side, two tabs/windows on the same chat could each derive streamingContent (the reconnect/resume path re-consumes preview events) and each independently insert the stream under a different Yjs clientID, duplicating content until the durable reconcile. Fix — single-writer election via the file-doc awareness (new agent-stream-leader): - a client applying an agent stream announces `agentApplying` on its own awareness - only the leader (min clientID among announcers) applies mid-stream AND at settle; a non-leader renders the leader's ops via Yjs and does not apply (a non-leader applying the final body would re-insert the whole doc as a duplicate) - re-checked each frame, so it converges to one writer the moment awareness propagates; the sub-frame startup race is reconciled by the durable write - single-client (the common case) is unaffected: it is the only announcer, so it always leads * fix(files): gate the settle apply locally, not on a settle-time re-election Cursor round 5 (High): the settle recomputed leadership from live awareness and the leader cleared its announcement immediately, so a straggler peer that settled afterward became the sole announcer, self-elected, and applied finalBody through its base-seeded shadow — re-inserting the whole doc as a duplicate. Fix: gate the settle apply on a LOCAL didApplyStreamRef (set only when this client actually applied a mid-stream frame — i.e. it was the mid-stream leader whose shadow is up to date), not on a settle-time re-election. A client that never applied (non-leader, a held `update`, or a pre-seed stream) skips the final apply and converges via Yjs + the durable write. The mid-stream leader election (isAgentStreamLeader) is unchanged, so exactly one client's didApplyStreamRef is ever true. * fix(files): open the agent-stream shadow lazily on lead (no stale handoff) Greptile round 6 (P1): the leader race — (a) a mid-stream leadership handoff could apply from a stale pre-stream shadow, and (b) two tabs starting the same stream before awareness converges could both lead briefly. - (a) fixed: the shadow is now opened LAZILY in the tick, only when this client actually leads, seeded from the CURRENT doc — so a handoff successor diffs against the prior leader's ops (never a stale base) and a non-leader builds no shadow at all. Announce candidacy via a dedicated ref (decoupled from the shadow); settle still gates the final apply on didApplyStreamRef (leader-only). - (b) the pure startup race is inherent to eventually-consistent election. It is now the only residual: bounded to two tabs starting the SAME stream within the awareness-propagation window, transient (converges in a frame or two), and never persisted (SYNC_NO_PERSIST + the durable edit_content reconcile). Resumes are sequential, so the common multi-tab case elects cleanly. Documented inline; a server-granted lease would close it fully but at a round-trip cost on the common single-tab path, which isn't worth it. * fix(files): idempotent settle apply (update lands client-side; no straggler dup) Cursor round 6 (Medium): a lone client's `update` never applied client-side — held mid-stream, then skipped by the didApplyStreamRef settle gate — so the rewrite depended entirely on the durable merge (stale if delayed/failed). Root cause was over-correcting round 5. Now that the shadow is opened lazily in the tick (current-seeded), the round-5 base-shadow duplication is already gone, so didApplyStreamRef is unnecessary. Replaced it: settle applies the final body via `agentStreamSessionRef.current ?? beginAgentStream(editor)` — the leader reuses its up-to-date shadow (last throttled frame), while a client that never applied (non-leader, held `update`, pre-seed) opens a FRESH current-seeded shadow. Reconciling current->final is idempotent: a straggler that settles after another wrote the final reconciles to a noop. So a lone `update` applies at settle (no wait on the merge), and there's still no settle-time election or base-shadow dup. * fix(files): broadcast agent frames to the whole room (same-socket siblings) Cursor round 7 (Medium): SYNC_NO_PERSIST frames applied under an origin carrying the sender socket id, and excludeSocketId dropped that whole socket from the relay fan-out. A second FileDocProvider on the same socket (chat preview + Files editor) then missed all mid-stream ops and stayed stale until the durable reconcile — a regression from the old no-origin server merge, which reached both. Fix: the agent origin is now a plain AGENT_SYNC_ORIGIN symbol, and agent frames broadcast to the WHOLE room (no socket excluded), matching the old behavior — so a same-socket sibling provider stays live; the emitting provider no-ops on its own echo (the ops are already applied locally). originSocketId still returns null for the symbol, so it keeps skipping edited/schedulePersist. Removed excludeSocketId and the socket-carrying origin object. Updated the relay test to assert the whole-room broadcast (verified it fails if the sender is excluded). * fix(files): tag agent stream frames no-persist across replicas A peer task tailing an agent-streamed preview frame previously applied it as REDIS_ORIGIN, marking the seeded room edited and making a transient startup-race duplicate eligible for that task's last-disconnect flush. Mark agent frames with a stream field so peers apply them as REDIS_AGENT_ORIGIN, excluded from the edited/persist gate. The copilot's durable edit_content write stays the sole authority over file bytes. * fix(files): reseed agent shadow on lead regain + agent-only compaction Two multi-writer edge cases surfaced in review: - rich-markdown-editor: a client that led, lost leadership, then regained it reused its stale shadow (which never saw the interim leader's ops), re-emitting ops for content already present. Tear the shadow down when a client observes it is not the leader, so a regain rebuilds fresh from the current doc. - file-doc-store: compaction always stamped its snapshot REDIS_SNAPSHOT_ORIGIN (marks peers edited). A long agent-only stream crossing the threshold could fold preview content into a persist-eligible snapshot. Track whether a room integrated any real edit and stamp an agent-only snapshot REDIS_AGENT_ORIGIN so it stays no-persist. Both covered by falsification-verified tests. * fix(files): close realEdited data-loss race + elect a settle writer Independent audit surfaced two real gaps: - file-doc-store: realEdited was latched AFTER appendUpdate's awaits, but the edit already sits in room.doc synchronously. A concurrent agent-frame compaction could read realEdited=false, snapshot that real content, and stamp it a no-persist agent frame — a lost edit. Latch it synchronously (same tick as the doc mutation) before any await. Deterministic falsifiable test added. - rich-markdown-editor: at settle every tab applied the final body, and a non-leader's local microtask runs before the leader's final propagates, so both insert the tail (Yjs keeps both) -> duplicated tail. Elect a single settle writer (reliable — awareness is long converged by settle), reading leadership before clearing the announcement. Corrects the overclaiming idempotency comment and the handoff pick-up comment. Adds a y-tiptap internals upgrade-guardrail test. * fix(files): own presence per client id, not one-per-socket The shared workspace socket hosts one collaborative provider per mounted view, so the chat file preview and the standalone Files editor for the same file each bind their own Yjs client id over ONE socket. The relay owned a single client id per socket, so the later JOIN overwrote the earlier and dropped its awareness — which silently broke the single-writer agent-stream election (a peer stopped seeing the streaming provider's announcement and could self-elect, duplicating streamed text for the whole stream). Track ownership per (socket, client id): a socket owns a set of client ids; the awareness gate accepts a frame only if every id it carries is owned; cleanup drops all of a socket's ids; the roster stays one-entry-per-session. Reclaim and the same-user reconnect path evict just the reclaimed id, dropping the old socket only if it empties. Falsification-verified test added. * fix(files): make streamed file-preview accumulation replay-safe Guard deriveFilePreviewSession against re-delivered/replayed content events: apply a delta/snapshot only when previewVersion strictly advances, so a client re-render or stream replay can't double-append the tail (the duplicated-content bug) or regress on an older snapshot. * fix(files): fix new-file collab streaming latch and agent-edit duplication - Latch collab readiness so a new file's post-seed `synced` flap can no longer re-gate agent streaming (the stream previously showed only the seed and the rest appeared only on reload) - Relay defers the durable edit_content merge to an actively-streaming client: the client shadow stream and the server merge were both writing the same content into the live doc, duplicating it when the server ran ahead - Render the collaborator caret bar out of flow so a peer caret never nudges the surrounding text by ~1px - Remove dead code: unused FileDocMessageType alias, unnecessary LiveFileDocMergeOrder export Covered by tests: readiness latch (flap/offline/latch cases), relay merge deferral (single- and multi-replica), plus verified-failing guards. * test(copilot): fix loadWorkspaceFileTextForPreview mock to return { text } not a bare string The adapter reads previewBase.text to seed an append/patch base; the mock returned a bare '' so previewBase.text was undefined, making a base-less append fail closed (no file_preview_content). My PR's fail-close change exposed the wrong-shaped mock. --------- Co-authored-by: mzxchandra <129460234+mzxchandra@users.noreply.github.com>
53 lines
1.5 KiB
JSON
53 lines
1.5 KiB
JSON
{
|
|
"name": "@sim/realtime",
|
|
"version": "0.1.0",
|
|
"private": true,
|
|
"license": "Apache-2.0",
|
|
"type": "module",
|
|
"engines": {
|
|
"bun": ">=1.2.13",
|
|
"node": ">=20.0.0"
|
|
},
|
|
"scripts": {
|
|
"dev": "SIM_DB_ROLE=realtime DB_APP_NAME=sim-realtime bun --watch src/index.ts",
|
|
"start": "SIM_DB_ROLE=realtime DB_APP_NAME=sim-realtime bun src/index.ts",
|
|
"type-check": "tsc --noEmit",
|
|
"lint": "biome check --write --unsafe .",
|
|
"lint:check": "biome check .",
|
|
"format": "biome format --write .",
|
|
"format:check": "biome format .",
|
|
"test": "vitest run",
|
|
"test:watch": "vitest"
|
|
},
|
|
"dependencies": {
|
|
"@sim/audit": "workspace:*",
|
|
"@sim/auth": "workspace:*",
|
|
"@sim/db": "workspace:*",
|
|
"@sim/logger": "workspace:*",
|
|
"@sim/platform-authz": "workspace:*",
|
|
"@sim/realtime-protocol": "workspace:*",
|
|
"@sim/runtime-secrets": "workspace:*",
|
|
"@sim/security": "workspace:*",
|
|
"@sim/utils": "workspace:*",
|
|
"@sim/workflow-persistence": "workspace:*",
|
|
"@sim/workflow-types": "workspace:*",
|
|
"@socket.io/redis-adapter": "8.3.0",
|
|
"drizzle-orm": "^0.45.2",
|
|
"lib0": "0.2.117",
|
|
"postgres": "^3.4.5",
|
|
"redis": "5.10.0",
|
|
"socket.io": "^4.8.1",
|
|
"y-protocols": "1.0.7",
|
|
"yjs": "13.6.31",
|
|
"zod": "4.3.6"
|
|
},
|
|
"devDependencies": {
|
|
"@sim/testing": "workspace:*",
|
|
"@sim/tsconfig": "workspace:*",
|
|
"@types/node": "24.2.1",
|
|
"socket.io-client": "4.8.1",
|
|
"typescript": "^7.0.2",
|
|
"vitest": "^4.1.0"
|
|
}
|
|
}
|