Commit Graph
35 Commits
Author SHA1 Message Date
Waleed 06506bbd38 improvement(files): cache stream binding meta, tighten file cache-control, prune dead persist path (#6149)
* improvement(files): cache stream binding meta, tighten file cache-control, prune dead persist path

- Cache the agent-stream ProseMirror↔Yjs binding metadata per session and reuse it
  across streamed frames instead of rebuilding it (O(doc)) every frame; matches how
  y-tiptap's own binding maintains the mapping in place. Safe because the shadow doc
  only ever sees the agent's own reconciles.
- Default createFileResponse to a private, no-cache policy so auth-gated bytes are never
  stored in a shared cache; genuinely-public serve routes opt into public caching
  explicitly.
- Serve content-addressed (key=) embedded images with an immutable private cache to
  avoid re-downloading them on every doc re-open; fileId= embeds and public shares keep
  revalidating (the underlying key can change / a share can be revoked).
- Drop the dead conflict.version field from the collab-doc persist result and remove the
  wasted getWorkspaceFile re-read on the conflict path (the relay treats missing and
  conflict identically and never read version).
- Decorate-sort the files/folders lists (compute each sort key once) and index the
  move-menu subtree build (O(N) vs O(N^2)); ordering is preserved exactly.

* fix(files): serve inline images with no-cache so deletion/authorization is enforced per request

Embedded images are authenticated content whose backing file can be deleted or have its
access revoked at any time. Both key= and fileId= embeds serve
private, no-cache, must-revalidate, so every request re-runs the server-side
deletion/authorization check instead of serving a possibly-stale image from cache.
2026-07-31 20:09:45 -07:00
Waleedandmzxchandra 10bfb5d139 feat(realtime): shared room spine + live Files/Tables collaboration + Yjs document editing (#5991)
* feat(realtime): add shared room identity + authorization spine (#5929)

Introduces the foundation for a unified realtime "room" model spanning the
Socket.IO presence server (apps/realtime), the durable SSE event log, and the
ephemeral pub/sub fanout — all of which today reinvent their own room identity,
naming, and authorization.

- @sim/realtime-protocol/rooms: RoomRef { type, id }, ROOM_TYPES, and a
  roomName/parseRoomName codec. WORKFLOW deliberately maps to the bare id so the
  ~40 existing io.to(workflowId) callsites and presence state keys are unchanged;
  every other room type is namespaced so id spaces cannot collide.
- @sim/platform-authz/rooms: authorizeRoom(userId, room, action) generalizing the
  exemplary authorizeWorkflowByWorkspacePermission — one resource->workspace
  resolver per room type, then the shared resolveEffectiveWorkspacePermission +
  permissionSatisfies gate.

Pure foundation, no behavior change: nothing consumes these yet. Prune graph
stays at 14/25 (platform-authz already depended transitively on realtime-protocol
via apps/realtime).

* refactor(realtime): generalize presence server to multi-room [2/N] (#5930)

* refactor(realtime): generalize presence server to multi-room (RoomRef)

Generalizes the Socket.IO presence layer from single-workflow-room-per-socket
to a domain-neutral, multi-room-per-socket model keyed by RoomRef, so a second
domain (workspace files, next PR) can reuse the same membership + presence
engine. Behavior-preserving for workflow collaboration.

IRoomManager is now domain-neutral (addUserToRoom/removeUserFromRoom/
getRoomForSocket/getRoomUsers/updateUserActivity/... all take a RoomRef). The
workflow lifecycle broadcasts (deletion/revert/update/deploy) move out of the
manager into WorkflowRoomService, composed over the generic manager.

Backward-compat by design (no workflow migration, no regression):
- Workflow Socket.IO room name stays the bare workflowId (roomName() maps
  workflow -> bare id), so the ~40 io.to(workflowId) callsites are untouched.
- Workflow Redis presence keys stay workflow:{id}:users/:meta (the type prefix
  IS "workflow").

Multi-room correctness (from adversarial audit):
- socket:{id}:workflow single-value key -> socket:{id}:rooms HASH (type->id).
- The SHARED socket:{id}:session key is deleted only when the socket leaves its
  LAST room (refcount via HLEN) — a leave from one room no longer breaks the
  other room's handlers.
- disconnect enumerates the socket's stored rooms and rebroadcasts presence per
  room, instead of picking an arbitrary socket.rooms entry.
- presence broadcasts use a per-room-type event name (workflow keeps the bare
  presence-update; others are namespaced).

Workflow handlers wrap manager calls with a shared workflowRoom(id) helper;
UserPresence.workflowId -> room (the client never reads that field).

Tests: existing 112 realtime tests pass unchanged (behavior gate) + 7 new
multi-room tests (refcounted session, presence isolation, multi-room disconnect,
per-type event names). tsc clean, boundaries + prune (14/25) green.

* fix(realtime): harden multi-room disconnect + id-guard room removal

Two fixes from an adversarial regression audit of the multi-room refactor:

- Disconnect now handles `disconnecting` (where `socket.rooms` is still populated
  and authoritative) and falls back to the live Socket.IO room set for any room
  the manager's stored state no longer tracked. This restores reliable presence
  cleanup + departure broadcast even if the Redis `socket:{id}:rooms` key was
  evicted or TTL-expired — the one behavioral gap vs the pre-refactor disconnect.
- REMOVE_ROOM_SCRIPT now only drops the socket's room mapping (and runs the
  last-room session cleanup) when the stored id matches the room being removed,
  matching the memory manager's existing id guard. Prevents a mismatched-room
  call from wiping a different room's mapping or the shared session.

+1 test (id-guarded no-op removal). 120 realtime tests pass, tsc clean.

* fix(realtime): only rebroadcast disconnect-fallback rooms whose removal succeeded

Greptile 4/5 follow-up: the disconnecting-time fallback ignored
removeUserFromRoom's boolean and rebroadcast presence even when the removal
reported false. Now it only treats a room as removed (and rebroadcasts) when the
manager confirms it — symmetric with removeSocketFromAllRooms, which already only
returns rooms it actually removed.

* fix(realtime): exclude the disconnecting socket from its farewell broadcast

Greptile follow-up (transient-Redis-failure edge): if removeUserFromRoom fails on
disconnect, the socket's presence entry can outlive it (room hashes have no TTL)
and reappear as a ghost. Disconnect now broadcasts a correction to EVERY room the
socket was in (union of the manager's removed rooms and the live Socket.IO
membership) and passes the disconnecting socket id as excludeSocketId, so it is
never shown as a collaborator regardless of whether the Redis delete succeeded.
Any orphaned entry is still reclaimed by the next join's stale-presence sweep.

broadcastPresenceUpdate gains an optional excludeSocketId; normal broadcasts are
unchanged. +1 test.

* fix(realtime): make presence broadcasts liveness-aware (root-cause ghost fix)

Presence broadcasts now reconcile the stored list against the live Socket.IO
membership (io.in(room).fetchSockets()) before emitting, via a shared
filterVisiblePresence helper. This closes the residual behind the earlier
disconnect fixes: an entry orphaned by a failed removal (room hashes have no TTL)
could reappear in a LATER join's presence snapshot until the 75-min stale sweep.
Now such an entry is never emitted, because a non-live socket is filtered out of
every broadcast. Combined with excludeSocketId (which handles the disconnecting
socket, still momentarily live). Fail-safe: on a fetchSockets throw or an empty
result while entries remain, emit the unfiltered list rather than hide live
collaborators.

Also drops a dead guard in the disconnect union loop (rooms already removed are
skipped by the wasInRooms check) and the now-unused isSameRoom import.

+1 ghost-guard test. 122 realtime tests pass.

* feat(files): live presence avatars + live file tree via realtime rooms (#5932)

* refactor(tables): adopt shared durable event-log core (#5934)

* fix(realtime): address post-merge review-comment findings (#5937)

* fix(realtime): address post-merge review-comment findings

A re-audit of every inline review comment on the merged stack surfaced real
issues that the thread-resolutions and prior audits missed. Fixes:

Presence server (#5930 comments):
- connection.ts: snapshot `socket.rooms` SYNCHRONOUSLY before the first await.
  Socket.IO clears the room set once the synchronous part of a `disconnecting`
  handler returns, so reading it after `await removeSocketFromAllRooms` saw an
  empty set — the eviction fallback was dead. (Cursor: "Disconnect fallback
  misses live rooms".)
- workflow-room-service: restore the original managers' final unconditional
  room-state wipe via a new `deleteRoom(room)` manager method, so a deleted
  workflow leaves no lingering presence/meta even if a per-socket removal failed
  or a socket joined mid-teardown. (Cursor: "Deletion skips final room wipe".)

Files (#5932 comments):
- workspace-file-manager.uploadWorkspaceFile now fans out the live-tree signal
  (all direct-upload paths: multipart fallback, copilot create, /api/files/upload,
  v1 files — the presigned path already notified). (Cursor: "Creates miss live
  tree fan-out".)
- use-workspace-files-room: clear the pending retry timer on join success; and a
  module-scoped intended-room guard defers the unmount `leave` so a rapid remount
  re-claims the room and skips a stale leave — fixing presence flap + a
  leave-after-join race. (Cursor: "Retry timer survives join success" + "Remount
  churns files presence".)
- workspace-files handler: roll back a partial join (leave room + remove presence)
  in the catch, mirroring the workflow join. (Cursor: "Join failure skips
  membership rollback".)

+2 tests (deleteRoom). 127 realtime tests pass, both apps tsc clean,
api-validation + boundaries green.

* fix(files): scope workspace-files leave to a workspace (deferred-leave safety)

Self-review of the deferred-leave guard found a real bug: leave-workspace-files
was not workspace-scoped, so after a workspace switch (A->B) the deferred leave
from A would evict the socket from its new room B. The leave now carries the
workspaceId and the server no-ops if the socket's current files room differs.
Also excludes the leaving socket from the leave broadcast (consistent with
disconnect).

* fix(realtime): close files-room presence leak + validate join payload

Architecture-audit findings:

- S1 (real Redis leak): the files room inherited the shared manager but not the
  workflow join's liveness sweep, so an UNGRACEFUL disconnect (pod crash — no
  `disconnecting` event) left its presence entry in the no-TTL room hash forever.
  Added a shared `sweepStalePresence(manager, room)` (fetchSockets liveness +
  remove not-live-AND-stale entries, matching the workflow 75min threshold) and
  run it on files join; also filter the join ack through `filterVisiblePresence`
  so a joiner never briefly sees an un-swept ghost.
- S2: validate the client-supplied `workspaceId` on files join before it reaches
  the DB query (matches the /api/workspace-files-changed guard; fails closed).
- N2: corrected the notify doc — it is awaited (guaranteed dispatch before a Node
  route returns) and hard-bounded to NOTIFY_TIMEOUT_MS, not "never block".

+1 test (sweepStalePresence keeps live/fresh, reclaims not-live-stale). 128
realtime tests pass, both apps tsc clean, biome clean.

* fix(realtime): workflow-deletion always notifies + cleans by socket.io membership

Review-round findings on #5937:
- Always emit `workflow-deleted` (was guarded by users.length>0), so a socket
  still in the Socket.IO room after a Redis presence eviction is told the
  workflow is gone before socketsLeave kicks it — the editor no longer keeps
  showing a deleted workflow. (Cursor: "Silent kick skips deletion event".)
- Clean per-socket state for the UNION of live Socket.IO members and
  presence-tracked sockets, so an evicted/late-joined socket's room mapping +
  session are dropped too — not just presence-snapshot sockets. (Greptile: "Room
  deletion leaves reverse state".)
- deleteRoom now logs AND rethrows on Redis failure (like addUserToRoom) so a
  failed wipe isn't reported as a clean deletion; the request surfaces it.
  (Greptile: "Room deletion failures are suppressed".)

The two "deferred leave drops new membership" P1s were already fixed by the
workspace-scoped leave in a prior commit (leave carries { workspaceId }; server
no-ops on mismatch). 128 tests pass, tsc + biome clean.

* refactor(files): drop module-scoped deferred-leave; rely on workspace-scoped leave

Removes the one non-idiomatic construct (a module-level mutable
`intendedFilesWorkspaceId` + queueMicrotask). It only guarded a same-workspace
CONCURRENT remount, which doesn't occur in production (folder nav is shallow/no
remount; list<->detail is sequential) — a dev-StrictMode-only case. The real
cross-workspace race is already handled by the workspace-scoped leave: if B's
join runs first (auto-leaving A), A's leave no-ops because the socket's current
files room is B. Simpler, idiomatic, prod-correct.

* feat(realtime): Yjs relay server for collaborative document editing [4/N] (#5941)

Server-side Yjs relay for collaborative document editing (live carets + text selection) in the Files rich-markdown editor. Faithful y-websocket-style relay over the existing authenticated Socket.IO connection + shared room abstraction; in-memory Y.Doc + Awareness per file; awareness ownership binding, userId-keyed client-id uniqueness, seeder election with deadline re-election, concurrent-JOIN generation guard. 25 relay tests. Reviewed to Greptile 5/5 + Cursor pass across multiple rounds, plus an independent 4-lens audit (correctness/security/conventions/simplicity) and /simplify + /cleanup passes.

* feat(files): collaborative document editing — client provider + editor (#5946)

Client Yjs provider (FileDocProvider over the authenticated socket) + TipTap Collaboration/CollaborationCaret wiring for live carets + text-selection in the Files rich-markdown editor. Collaboration is a Files-page-only surface (explicit `collaborative` opt-in), disjoint from agent-streaming. Read-only + autosave-gated until synced+seeded. Merges into the realtime-rooms integration branch.

* feat(tables): live collaboration — cell-selection presence + live mutation propagation (#5957)

* feat(tables): live cell-selection presence — protocol + server + client hook

The realtime spine for Google-Sheets-style table presence (mode A, socket):

- @sim/realtime-protocol/table-presence: centralized wire protocol (events +
  TableCellSelection {anchor, focus, editing} + payloads) so server emits and
  client subscriptions can't drift.
- ROOM_TYPES.TABLE + resolveTableWorkspace registered in ROOM_WORKSPACE_RESOLVERS
  (tableId -> workspace via userTableDefinitions, honoring archivedAt); roomName /
  presenceEventName / disconnect cleanup / authorizeRoom all derive automatically.
- apps/realtime/src/handlers/tables.ts: join/leave (mirrors workspace-files) + a
  table-cell-selection relay (mirrors the workflow selection channel), broadcasting
  via roomName(room) since table rooms are namespaced. UserPresence gains a cell
  field threaded through the memory + Redis managers (Lua ARGV[7], null clears).
- Extracted the duplicated resolveAvatarUrl into handlers/avatar.ts.
- use-table-room.ts client hook: joins over the shared socket, tracks the roster
  (avatars) + patches per-socket cell deltas, exposes a throttled emitCellSelection.

Grid UI (avatars + selection overlay) lands next; concurrent cell-value edits
(last-write-wins via the durable log) are the follow-up PR.

* feat(tables): render live cell-selection presence in the grid

Wires the table presence room into the grid UI:
- Page (table.tsx): useTableRoom (gated off in embedded/mothership mode) —
  renders <PresenceAvatars> in the header and passes remoteSelections +
  emitCellSelection down to the grid.
- Grid emits its local selection: an effect resolves the index-based
  anchor/focus to stable (rowId, columnId) via refs and broadcasts it (with an
  editing flag for the active cell) through the throttled emitter.
- RemoteSelectionOverlay: draws each remote viewer's selection in their color
  (getUserColor), a darker fill while editing, and name-on-hover — measured from
  live cell rects in the content wrapper's space (scrolls with the grid),
  hidden when rows are virtualized off-window, pointer-events-none so it never
  blocks cell clicks (hover via pointer hit-test).

* test(tables): cover the table presence handler

Mirrors workspace-files.test.ts: join auth/unavailable/denied/success, plus the
cell-selection relay (asserts it persists via updateUserActivity and broadcasts
on the namespaced roomName, not the bare id) and leave.

* feat(tables): propagate manual cell edits live (last-write-wins)

A manual row edit now appends a lightweight 'edit' event to the durable table
stream; collaborators refetch the row (via the existing debounced rows-invalidate
the job events use) so the winning value shows live. The event carries no value —
peers refetch in their own wire format, so there's no auth-specific value
translation on the wire, and last-write-wins falls out of the DB's committed order
(the Google-Sheets model). Edits that also trigger a dispatch already emit
dispatch/cell events; the debounce coalesces the two.

* refactor(tables): apply /simplify findings

- Drop the dead 'add unknown peer' upsert branch in use-table-room (Socket.IO
  ordering guarantees a peer is in the roster before their selection delta).
- TableCellSelectionBroadcast = TablePresenceUser & { cell } (was a copy-paste).
- Make TableGrid's presence props required + drop the unused empty-default/guard
  (only table.tsx mounts it, always passing both).
- Drop the unused rowId from the 'edit' event (the handler invalidates all rows).
- Overlay: subscribe scroll/resize/pointer listeners once per scroll element and
  cache the wrapper origin, so incoming deltas re-measure without re-subscribing
  and the pointer hit-test never forces a per-move layout read.
- Server: cache the immutable socket session so a selection delta no longer reads
  it from Redis every time.

* refactor(tables): apply /cleanup findings

- Fix the remote-selection name label contrast: text-white is unreadable on the
  light-pastel user colors (same bug the Files caret fixed) → fixed dark #1a1a1a.
- Re-measure via useLayoutEffect so a moving peer selection updates before paint
  (no one-frame position lag).
- Drop 'mothership' from a comment (constitution copy rule).

Six cleanup passes ran (effect, memo/callback, state, react-query, emcn, comment);
the rest confirmed clean — all state/memos/callbacks/effects are load-bearing,
presence correctly lives in useState (socket-pushed), and the edit→rows-invalidate
granularity is right.

* feat(tables): propagate every table mutation live (edit + schema signals)

Comprehensive live collaboration for all user table mutations, via two value-less
durable signals + named helpers (signalTableRowsChanged / signalTableSchemaChanged):

- edit (rows refetch): single + batch row create, cell/row update, batch update,
  delete by id/filter, and upsert.
- schema (definition + rows refetch): column add/update/delete, workflow-group
  add/update/delete, table rename, and CSV import (which can add columns).
- Client handles 'schema' by invalidating the table detail (exact) + rows.

Execution paths (column run, cancel-runs) and async jobs (delete/import-async,
job-cancel) already propagate via cell/dispatch/job events — verified applyJob
refetches on terminal. No reorder routes exist. Table archive (route DELETE) is a
deliberate follow-up: it needs a table-deleted redirect event, not a refetch signal
(which would 404).

* refactor(tables): apply comprehensive /cleanup audit findings

Holistic + react-query + comment audits over the whole PR:

- Security/crash fix: a remote peer's rowId flowed unescaped into the overlay's
  querySelector — a hostile id ('x"]') threw SyntaxError inside a useLayoutEffect,
  crashing every other viewer's page. CSS.escape it, and validate + whitelist the
  untrusted cell payload server-side (shape + 200-char id bound) before it is
  stored/rebroadcast.
- Simplify the CELL_SELECTION relay: the delta attached userId/userName/avatarUrl
  that the client discarded (identity comes from the roster). Drop them + the
  getUserSession lookup/cache entirely — the delta is now { socketId, cell }.
- React Query: schema handler also invalidates lists() (parity with the local
  column-mutation set); document that the mutating client self-refetches by design.
- Comment tightenings; biome fixed a stale import order in workspace-files.ts.

* fix(tables): broadcast single-cell selections (focus falls back to anchor)

Cursor High: a normal cell click leaves selectionFocus null (the grid treats it as
a one-cell selection via focus ?? anchor), but the presence emit required BOTH anchor
and focus to resolve — so the most common selection never broadcast and clicking even
cleared a prior remote outline. Mirror the grid's focus ?? anchor semantics.

* fix(tables): reviewer + regression + per-LOC audit findings

Cursor review round (5 findings) + regression audit + per-LOC audit:
- Presence roster snapshot now KEEPS the cell we already hold for a known socket, so
  a join/leave broadcast can't revert a fresher CELL_SELECTION delta.
- Reset the selection throttle on table switch (was unmount-only), so a pending
  selection for table A can't flush into table B's room after a switch.
- Metadata writes (column widths, display) use a new lightweight 'metadata' signal
  that refetches only the definition — a resize no longer forces peers to refetch rows.
- Overlay re-measures on row add/remove/reorder via a tbody childList MutationObserver
  (a live refetch moves cells without a scroll/resize).
- Document the actor self-refetch create caveat (scrolled multi-page insert) accurately.
- isCellRef narrows to a partial instead of casting to the full type then re-checking;
  drop a redundant mount measure() (the layout effect covers it); text-[11px]→text-xs.

* fix(tables): drop ineffective metadata propagation + re-measure overlay on column resize

Cursor round on b8f28b04b:
- Remove the 'metadata' signal entirely. The grid seeds columnWidths/pinnedColumns
  from metadata ONCE (metadataSeededRef) and deliberately never re-applies them (to
  avoid clobbering a local in-progress resize), so refetching the definition on a peer
  never surfaced their width/pin change — an ineffective path. Width/pin live-sync needs
  reconciliation that doesn't clobber a local resize; that's a deliberate follow-up, not
  a no-op refetch. Structural changes still propagate via 'schema'.
- Overlay now also observes the content layer with the ResizeObserver, so a column
  resize (which grows the content, not the scroll container) re-measures remote outlines.
- Presence-merge comment now states both sides of the trade-off.

* fix(tables): re-broadcast local selection on (re)join

Cursor Medium: a selection made before the room join completes (or held across a
reconnect) was dropped server-side and never re-sent, so peers didn't see it until
the local user moved it again. Track the current selection in a ref (set on every
emit, cleared on table switch) and re-emit it from handleJoinSuccess once the room is
joined.

* fix(tables): re-broadcast selection when a peer's row change shifts it

End-to-end lifecycle audit (Low-Med): the selection emit resolved the stable
(rowId, columnId) only on selection/editing change, not when a live edit/schema
refetch inserted/deleted/reordered rows. The index-based local selection then sat
on a different logical row than the rowId peers held, so your outline showed on the
old row until you moved. Re-run the emit on rows/displayColumns change and dedup an
unchanged result (also drops the redundant null-on-open emit) so the broadcast stays
consistent with the local highlight.

* fix(tables): schema invalidates run-state/enrichment + guard stale join

Cursor round on cdc8796b8 (2 Medium):
- schema handler used detail exact:true, so it skipped the activeDispatches +
  enrichmentDetails sibling queries the local invalidateTableSchema refreshes via a
  prefix match. After a peer deletes/restructures a workflow group, peers could keep a
  stale running badge or enrichment panel. Now invalidates both siblings too (rows stay
  on the debounce).
- Guard against a stale join stealing the room: a fast table A->B switch could let A's
  async authorize finish after B, leave B, and strand the socket in A. Added a
  per-socket monotonic join generation checked after authorize (mirrors the file-doc
  relay's guard) + a test.

* feat(tables): live column width/pin/order sync

Collaborators now see each other's column resizes, pins, and reorders live —
the last piece of Google-Sheets-style layout parity.

- New lightweight `metadata` durable event kind (distinct from `schema`): only the
  table definition carries UI metadata, so peers refetch the definition alone — no
  rows/run-state refetch. The metadata PUT route now signals it.
- The grid reconciles server metadata against its in-progress gesture: the column
  being actively resized keeps its live local width, and an in-flight column drag
  blocks a reorder apply — so a peer's change never reverts the local action. Each
  field is reference-guarded (React Query structural sharing keeps unchanged
  sub-objects stable), so an unrelated peer change doesn't re-apply the others.

* fix(tables): escalate to schema signal when a reorder scrubs group deps

Independent audit of the metadata-sync commit found a stale-run-state hole: a
columnOrder PUT that moves a column left of a workflow group's leftmost column
makes updateTableMetadata scrub that group's dependencies and write a new schema —
a real structural change. But the route only fired the lightweight 'metadata'
signal (detail-only refetch), so peers' and the actor's activeDispatches /
enrichmentDetails queries stayed stale (a lingering running badge / enrichment
panel) — exactly what the 'schema' handler exists to prevent.

updateTableMetadata now reports whether it scrubbed the schema; the route emits
signalTableSchemaChanged in that case and the light signalTableMetadataChanged
otherwise. Width/pin/plain-reorder stay on the cheap detail-only path.

* feat(realtime): accurate in-file presence + collaborative-caret polish (#5965)

Per-session file-doc presence (avatars count other sessions like the canvas), StrictMode-safe stable Y.Doc (fixes blank-doc on join), flush caret cap + restored hover hit-slop, and three join-lifecycle race fixes unifying file-doc + workspace-files on one intent-tracked monotonic generation model. All findings root-caused with regression tests.

* feat(files): smarter bullet delete/indent and untitled-file title sync (#5971)

* improvement(files): smarter bullet delete/indent, fix empty-nested-bullet heading corruption

Backspace at the start of a list item now outdents a nested item or clears a
top-level item to a paragraph in place instead of deleting the row and jumping
the caret to the previous block; Enter on an empty nested item outdents. Empty
non-trailing top-level items still collapse cleanly since they cannot round-trip
as a lifted paragraph.

Also strips nested empty list-item marker lines on serialize: a nested empty
bullet re-parsed as a Setext heading underline, silently turning its parent line
into an H2 and dropping the bullet. Top-level empty items are preserved.

* feat(files): sync an untitled file's name with its leading heading

While a file is still named untitled(.md), typing a leading heading auto-renames
the file after it (debounced), and renaming the file first seeds a leading H1
from the new name. One-shot: coupling stops once the file has a real name, and
the heading seed always prepends so existing content is never clobbered.

* fix(files): count inline atoms in list-item emptiness, keep multi-block items on Backspace

Addresses review findings on the list Backspace logic:
- Emptiness now uses the caret block's content.size (counts inline images/mentions),
  not textContent, so a bullet holding only a non-text atom is no longer treated as
  empty and deleted.
- An empty first block whose item has sibling blocks removes only that block instead
  of lifting the whole item out of the list.

* fix(files): preserve the untitled to named heading seed across a rename during editor load

The parent captures the file name at mount (before content/session finish loading) and
passes it as the transition baseline, so a rename that lands in the loading window is still
seen as an untitled to named transition and the leading heading seed is not skipped.

* fix(files): drop the name-to-heading seed, keep title sync one-way

Removes the effect that inserted a leading H1 when an untitled file was renamed. On the
collaborative Files page every open client observed the untitled-to-named transition and
inserted into the shared doc, producing duplicate headings; it could also re-insert a heading
a user had just deleted while a rename was in flight. Seeding document content from an async
rename transition is the wrong model on a shared editor. The primary direction — typing a
leading heading renames a still-untitled file — is unaffected (it never mutates the doc).

* fix(files): keep empty lines between paragraphs on reload

The chunked markdown parser (parseMarkdownToDoc) parses each block stripped of the
blank lines between them, so it dropped the empty paragraphs @tiptap/markdown builds
from runs of blank lines — a saved visual blank line silently vanished on the next
load (the settle/reopen re-seed goes through the chunker). The whole-document parser
preserves them, but whether a gap yields an empty paragraph is a global, block-type-
dependent decision (kept between two paragraphs, dropped after a heading), so it can't
be reconstructed block-locally. Route documents with empty-paragraph blank-line spacing
to the whole-document parser for exact fidelity — the same tradeoff NON_CHUNKABLE makes;
ordinary single-blank-line separation still takes the fast chunked path. Adds a suite
asserting chunked output matches the whole-document parser for leading/trailing/between
gaps and around lists/headings.

* fix(files): only auto-name an untitled file when the user can edit

The debounced untitled→filename hook ran on every onUpdate — including the mount-time
seed and for view-only viewers — without checking edit permission, so a read-only user
could schedule a rename they have no permission to make (a spurious, server-rejected
write). Gate the derive-title on editor.isEditable (canEdit + settled + collab-ready,
the same signal the autosave path uses), at both schedule and fire time.

* fix(files): normalize line endings before the empty-paragraph guard; Enter/Backspace symmetry

- markdown-parse: EMPTY_PARAGRAPH_SPACING/NON_CHUNKABLE tested the raw body, but a classic
  \r-only file (blank lines are \r) would miss the \n-anchored guard and still be chunked,
  dropping empties. Normalize line endings once up front so the routing guards, the chunker,
  and the parser all see the same \n. +CRLF/CR test cases.
- keymap: Enter on an empty first block of a multi-block item now removes only that block
  (removeEmptyWrappedBlock) instead of exiting the list, mirroring the Backspace hasSiblingBlocks
  case — the trailing check no longer swallows multi-block items. +test.

* fix(files): editor audit follow-ups (trailing-blank read-only, collab rename, over-strip)

A 4-agent independent audit (UX vs inkeep + SOTA, cleanliness, adversarial correctness)
surfaced these:

- HIGH regression: files ending in a blank line opened READ-ONLY. The empty-paragraph
  routing preserved a TRAILING empty paragraph, but postProcess collapses trailing newlines
  → serialize/parse non-idempotent → isRoundTripSafe flipped the file read-only. A trailing
  empty paragraph can't be serialized stably, so parseMarkdownToDoc now strips trailing empty
  paragraphs and the guard no longer routes on trailing blanks. Interior/leading empties are
  unaffected. +regression tests.
- Medium: the debounced untitled→filename rename fired on remote Yjs edits too, so every peer
  renamed and could rename from a not-yet-synced heading. Gate on isChangeOrigin (local edits
  only; false for non-collab surfaces).
- Medium: stripEmptyListItemLines over-stripped a nested empty item that follows a same-indent
  sibling (a real placeholder the parser keeps). Narrowed to the actual Setext hazard — an empty
  item DIRECTLY under a shallower parent line — matching the function's own docstring intent.
  Probe-verified. +test.
- Low: corrected untitled-title.ts docstring that described a reverse name→heading coupling
  removed during review.

* fix(files): a remote edit must not cancel the local rename debounce

The isChangeOrigin gate cleared the debounce timer BEFORE bailing on a remote update, so
a peer's edit arriving within the 600ms window cancelled the local user's pending rename.
Bail on isChangeOrigin first, before touching the timer; only local edits clear/reschedule it.

* docs(files): correct EMPTY_PARAGRAPH_SPACING rationale after trailing-strip

The stacked trailing-empty-paragraph strip made the older comment overstate a
correctness necessity it no longer owns, mislabel trailing runs of 2+ blanks,
and advertise dead CRLF handling. Reword to match what the code actually does.

---------

Co-authored-by: Waleed Latif <walif6@gmail.com>

* fix(realtime): access-revalidation multi-room safety + cleanup pass

Fix a blocker surfaced by a full cleanup/simplify audit of the branch: the
access-revalidation sweep (staging's workflow-only #5917) treated every entry
in socket.rooms as a workflow id, but the generalized multi-room model puts
namespaced files/tables/file-doc rooms on the same io. It would resolve those
as bogus workflows, get null, and evict files/tables collaborators every ~30s.
collectScanTargets now decodes each room name with parseRoomName and sweeps
only workflow rooms; added a regression test and fixed the now-false TSDoc.

Other audit fixes (all behavior-preserving):
- workflow.ts reuses resolveAvatarUrl (drops db/user/eq imports duplicated
  from avatar.ts)
- PresenceAvatars: mr-1 was baked into the shared component, silently adding a
  margin to the workflow sidebar stack; moved to an optional layout className,
  re-applied on the tables/file-doc header surfaces only
- table DELETE routes only signal collaborators when rows were actually removed
  (matches PUT)
- events.ts definition kind: drop the never-emitted reason:'schema', fix its doc
- event-log: rename buildMemory -> buildEntry (it builds the entry on the Redis
  success path too, not just the memory fallback)
- remove dead resolveWorkspaceIdForRoom export; parallelize per-socket removals
  in handleWorkflowDeletion; gate the table columnIndexById map on remote
  selections; move file-doc module TSDoc off the FileDocOwner interface; fix a
  stale @returns

* fix(realtime,tables): close table-presence race + v1/copilot live-collab gaps

Validated each issue with subagents before implementing the cleanest fix.

- tables LEAVE in-flight-join race (B8): the table handler tracked no current-table
  intent, so an unscoped/same-table leave during an in-flight authorize left the
  socket stranded in the room (present in the roster, broadcasting a ghost until
  disconnect). Mirror workspace-files: a closure-local currentTableId + a leave that
  advances joinGeneration to cancel the racing join. + 3 regression tests.
- v1 + copilot live-collab signal gap (D1): tables edited via the v1 public API or
  Sim/copilot emitted no edit/schema signal, so open collaborators didn't live-update.
  Add the signals at those call sites (add-only, matching the existing route seam) —
  never in the service, so execution writes can't double-emit. Sync-only for copilot
  bulk ops, guarded on affected/deleted count; async job branches stay covered by
  their kind:'job' events; create/delete/get untouched.
- table join read consolidation (B5): sweepStalePresence returns its roster so the
  same-tab dedup reuses it instead of a second getRoomUsers.
- shared authorize slice (B6): extract only the guard-safe authorize->allowed branch
  into resolveRoomJoinAuth, shared by the three room handlers (the full preamble stays
  inline — file-doc's generation capture sits mid-ladder and must not move).
- resize-revert flicker (E3): a peer's value-less metadata event forces a refetch that
  could momentarily revert a just-finished local resize; a pendingWidthWriteRef keeps
  local widths leading until the width PUT settles.
- embedded-mode stray emit (E5): gate emitCellSelection on a bound table id so the
  embedded surface stops broadcasting cell selections the server drops.

* fix(tables): close two copilot live-collab signal gaps + harden presence sweep

Follow-ups from a comprehensive review of the branch:
- copilot batch_update_rows and import_file's inline append branch wrote rows
  but emitted no live-collab signal, so collaborators didn't see those edits
  live (the append's sibling replace branch already signalled). Add the guarded
  signal to both, matching the internal route.
- sweepStalePresence now reads the roster before the fetchSockets liveness probe
  and returns it on a probe failure, so same-tab dedup still runs during a
  transient fetchSockets outage instead of being skipped.
- reword an internal comment off the retired "mothership" term.

* fix(realtime): guard table join commit + rollback against supersession; drop no-op eviction cleanup

Review round on #5991:
- Table join re-checked the generation only once after authorize, then awaited
  leave/sweep/avatar before joining + registering presence. A table switch or
  leave in that window stranded the socket in the wrong room, and the failure
  catch could tear down a newer successful join. Resolve the avatar up-front,
  re-check generation immediately before the membership commit (matching the
  file-doc join), and skip the rollback/error for a superseded join. + a
  post-authorize-window regression test.
- access-revalidation cleanup treated removeUserFromRoom's no-op false as a
  transport failure and re-enqueued a still-connected socket forever. Only retry
  when the socket is still mapped to the room (a healthy null mapping means the
  entry is already gone). Repurposed the expired-mapping test to lock it.

* fix(realtime): guard table join leave-prior against superseding join

Round 2 on #5991: a superseded join's leave-prior could still run — during its
getRoomForSocket await a newer join commits to its room, so currentRoom is that
newer room and the superseded join would leave/remove/broadcast it before the
final guard aborts. Re-check the generation immediately after the lookup await,
before the leave mutation. Extended the post-authorize-window test to assert the
superseded join never tears down the newer join's room.

* fix(realtime): roll back a table join superseded during addUserToRoom

Round 3 on #5991: after the final generation guard, A could join + register
presence while a newer join B commits to its room during addUserToRoom's await
— B's leave-prior can't observe A's half-written entry, so A's late write wins
and strands the socket. Re-check after addUserToRoom and roll back A's own
Socket.IO join + presence (scoped to A's room, never touching B). + a regression
test hanging addUserToRoom mid-commit.

* refactor(realtime): DRY table-join supersession guards; fix stale comment

Cleanliness pass after the review rounds (no behavior change):
- Extract the four identical `joinGeneration !== joinAttempt || socket.disconnected`
  checks into a named `superseded()` helper (the catch keeps its intentionally
  narrower check).
- Remove a stale guard comment that was left stranded above the avatar resolve.
- Document the best-effort rollback catch.

* fix(realtime): file-doc rebind must not drop the current doc or leave a writable ghost

Two Cursor findings on the file-doc client-id ownership rebind:
- On a document switch, the prior room was left BEFORE the ownership check, so a
  CLIENT_ID_IN_USE rejection dropped the socket from the old doc without joining
  the new one (contradicting its own comment). Run the ownership check first, and
  leave the previous doc only once the rebind is guaranteed to succeed.
- Reclaiming a client id removed the stale prior socket from owners + awareness
  only; its socketToRoomName + Socket.IO membership remained, and handleMessage's
  SYNC path gates on socketToRoomName (not owners), so it stayed able to write
  document frames until disconnect. Fully evict the reclaimed socket. + 2 tests.

* refactor(realtime): serialize table join/leave to fix map-corruption at the root

Round 4 on #5991 surfaced a race the generation guards structurally cannot fix:
two concurrent joins for one socket race on the single-valued socket→room map —
a stalled addUserToRoom for table A lands late, clobbers a newer join's map entry
to A, and the rollback then wipes it, stranding the socket (map empty while it
holds table B). Guards protect JS suspension points; they can't stop an in-flight
Redis write from landing late.

Fix per architecture review: serialize this socket's JOIN + LEAVE on a per-socket
promise chain so their multi-step async Redis commits can never interleave —
restoring the atomic-commit property the synchronous sibling handlers get for free.
This DELETES the leave-prior guard and the post-commit rollback (the code that
caused the bug); four generation guards collapse to two identical superseded()
checks (skip a superseded queued op + one pre-commit check). Reworked the
interleaving-specific tests into a fast-switch-skips-superseded test; the leave-
cancels-join tests are unchanged. Local to the tables handler — no shared-infra change.

* fix(realtime): always roll back a failed table join; re-elect file-doc seeder on reclaim

Two review findings:
- Table join: the catch skipped rollback when superseded, but a socket.join that
  landed before addUserToRoom threw leaves the socket in the Socket.IO room with no
  matching socket->room map entry — unreclaimable by any later op (cleanup keys off
  the map). Under serialization the skip is unnecessary (the newer op hasn't
  committed), so always roll back. Simpler + fixes the strand.
- File-doc reclaim: fully evicting the prior socket didn't release the seeder role
  if it held it, so electSeederIfNeeded (which no-ops while seederSocketId is set)
  never re-elected and an unseeded doc stayed empty until the deadline. Clear the
  role on eviction so the join's election picks a new seeder. + 2 regression tests.

* fix(realtime): close revoke-race ghost presence in workflow join

An access-revalidation revoke landing between socket.join and addUserToRoom
socketsLeaves the socket while its presence mapping does not yet exist, so
cleanupEvictedSocket finds nothing to remove and the join then writes presence
for a socket already out of the room — a ghost collaborator until the stale
sweep. Hoist resolveAvatarUrl (the only await in that gap) above the re-auth
check so the whole re-auth -> socket.join -> addUserToRoom section is await-free,
matching the invariant the handler already relies on for the pre-join re-auth.
+ ordering regression test.

* refactor(realtime): serialize workflow join/leave; drop dead room-authz limb

Comprehensive independent audit follow-ups:

- workflow.ts join/leave now use the same opChain + joinGeneration serialization
  as the sibling handlers (tables, file-doc, workspace-files). It was the only
  async presence path left unserialized, so a rapid workflow switch A->B (or a
  leave racing an in-flight join) could strand presence in room A — a ghost
  collaborator still receiving A's operation broadcasts until disconnect. The
  join now aborts a superseded op at start and again right before the membership
  commit, and the catch always rolls back a partial join.
- leave-workflow drops the '&& session' gate: an idle user whose 1h session key
  expired (while the 24h room mapping is still live) can now leave cleanly
  instead of being stranded until disconnect. The room ref alone suffices.
- authorizeRoom: remove the dead ROOM_TYPES.WORKFLOW resolver + its
  getActiveWorkflowContext import. Workflow authorizes through its own path and
  never flows through authorizeRoom; the map now honestly covers only the
  workspace-scoped types (files, file-doc, table).
- Remove unused isSameRoom (zero callers) and a needless useMemo in
  PresenceAvatars (plain derivation, single copy).
- Tests: 4 workflow serialization/leave regressions.

All gates green: tsc (sim/realtime/packages) 0, 204 realtime + 11 protocol
tests, biome, api-validation, boundaries, prune 14/25.

* fix(realtime): align session TTL, harden committed joins from post-success rollback

Per-module comprehensive audit follow-ups:

- redis-manager: SESSION_TTL now tracks SOCKET_ROOMS_TTL (was 1h vs 24h). The room
  set outlived the session, and since getRoomForSocket reads the room set while the
  workflow handlers gate edits/presence on `room && session`, an active-but-idle
  collaborator got wedged into 'session expired' after 1h — sticky until reload
  (activity only EXPIREs the already-gone session; only addUserToRoom re-HSETs it).
  Both keys refresh together, so they now expire together (restores the pre-refactor
  consistency, where both shared one TTL).
- workflow.ts + tables.ts: a 'committed' flag stops the join catch from rolling back
  a genuinely-joined user when a trailing ack/broadcast/metric step fails on a Redis
  blip (a pure getUniqueUserCount log-metric failure could otherwise kick a live
  collaborator after success was already acked).
- workflow.ts: leave-prior now guards `currentRoom.id !== workflowId` (a same-workflow
  re-join no longer leave→re-adds and flickers peers' presence), and the join ack is
  liveness-filtered via filterVisiblePresence — both for parity with the tables handler.
- platform-authz: honest docstring + 400 message for the workspace-scoped-only
  authorizeRoom map (workflow authorizes via its own path).
- caret-presence: corrected an over-stated batching comment.
- +1 workflow regression test (post-success failure keeps the user joined).

Gates: tsc (sim/realtime/packages) 0, 205 realtime + 11 protocol tests, biome,
api-validation, boundaries, prune.

* fix(realtime): narrow join commit-guard to post-success; skip empty presence broadcast

Two Cursor findings on the prior audit-fix commit:

- The 'committed' guard in join-workflow/join-table was too broad: a failure
  BETWEEN the membership commit and the success ack (e.g. getWorkflowState) hit
  'if (committed) return' and emitted neither success nor error, hanging the
  client while it sat in the room. Replaced with the narrower shape: only the
  purely-decorative post-success steps (peer broadcast + log-metric) are wrapped
  best-effort; anything before the success ack still rolls back and surfaces a
  retryable error, so the client retries instead of hanging — while the original
  goal (a benign broadcast/metric blip never kicking a live, acked user) holds.
- broadcastPresenceUpdate read the roster via getRoomUsers, which swallows a Redis
  transport error to []. On a disconnect broadcast that emitted an empty roster and
  cleared every remaining collaborator's presence until the next healthy update.
  Split out a throwing readRoomUsers; broadcastPresenceUpdate now skips the
  broadcast on a read failure (getRoomUsers keeps its swallow contract).
- Tests: pre-success failure rolls back + retryable error (no hang); post-success
  failure keeps the user joined.

Gates: realtime tsc 0, 206 realtime tests, biome, boundaries, prune.

* fix(realtime): harden seeder recovery, join-generation, and misc robustness

Final line-by-line audit follow-ups (all LOW/MED, no P0/P1):

- file-doc: a sole client whose seed FETCH fails was added to triedSeeders,
  re-election found nobody, and the document stayed permanently empty until
  reload. Re-offer seeding a bounded number of rounds (MAX_SEED_ROUNDS) before
  giving up. Also bound clientId to a non-negative integer (it is an ownership key).
- tables + workspace-files: validate the room id BEFORE advancing joinGeneration,
  so a malformed/rejected join can't cancel a legitimate in-flight join.
- workflow + tables + workspace-files: suppress the client-facing join error when
  the op was already superseded (a retryable error naming the abandoned room could
  make a client re-join and cancel its newer join). The rollback still runs.
- redis-manager: set isConnected=true only after scriptLoad succeeds (and reset it
  on failure) so isReady() can't report ready while the Lua SHAs are null.
- connection: apply the presence-bearing filter to the manager-removed set too
  (symmetry with the fallback path).
- http.ts: validate workflowId on the four workflow endpoints (matching the files one).
- client: clear a pending join-retry timer before rescheduling (reconnect churn no
  longer orphans a stray extra join); clear caret fade timers on plugin destroy;
  seed-effect cleanup reports NOT-ready (safe direction).
- Tests: bounded seeder recovery, cell-selection strip-junk, TABLE round-trip,
  presenceEventName.

Gates: tsc (sim/realtime/packages) 0, 208 realtime + 12 protocol tests, biome,
api-validation, boundaries, prune.

* fix(realtime): gate isReady() on loaded script SHAs, not just connection

Follow-up to the prior isConnected change, which was incomplete: the redis client's
'ready' event flips isConnected=true on connect — before initialize() loads the Lua
scripts — so a bare isConnected check reports ready while removeUserFromRoom /
updateUserActivity would silently no-op on a null SHA. Gate isReady() on the SHAs
too, so the POST endpoints return a retryable 503 during that startup window instead
of proceeding against unloaded scripts. Standard readiness-probe discipline.

* fix(files): offline read-only fallback when the realtime doc never syncs

When the realtime server is unreachable (offline, server down, socket never connects),
the collaborative editor would sit blank and read-only forever — content only arrives
via provider sync events that never fire. Add a bounded connect-deadline to the Yjs
provider: if no first sync lands within CONNECT_DEADLINE_MS, latch fatal and emit a
synthetic non-retryable join-error — the exact path a real fatal rejection already uses,
which seeds the file's stored content read-only. Latching fatal also stops a late
reconnect from syncing server state in and merge-duplicating the locally-seeded content
(the documented Yjs 'non-empty doc ignores initial value' gotcha).

Deliberately NOT adding durable Yjs snapshot persistence / server-side seeding: TipTap
can't run the markdown->Yjs conversion server-side (Collaboration extension errors under
jsdom), and a durable binary snapshot would create a dual source of truth with the
markdown file (edited by copilot / PUT / download). The client-seeder + bounded
re-election is the correct architecture for a markdown-is-truth model; this closes its
one real user-facing gap without persistence, a migration, or dual-truth.

Timer cleared on first sync, on a real fatal rejection, and on destroy. +2 tests.

Gates: sim tsc 0, 496 editor tests, biome. Needs live offline->reconnect verification.

* fix(realtime): validate workflow join id before generation bump; scope file-doc join rollback to its target

Two Cursor findings:
- join-workflow bumped joinGeneration before validating workflowId (unlike tables /
  workspace-files, which I'd already fixed). A malformed/empty join could advance the
  counter and cancel a legitimate in-flight workflow switch. Validate the id first. +test.
- The file-doc join catch called cleanupFileDocForSocket unconditionally, which keys off
  socketToRoomName. During a document SWITCH that fails before rebinding (e.g. a throw in
  client-id reclaim), that binding still points at the socket's PRIOR, valid document —
  so the rollback tore down a document the socket was validly in. Only run that cleanup
  when the binding already points at THIS join's target; otherwise the socket never
  registered as an owner here and the only leftover is a freshly-created empty room,
  dropped by destroyRoomIfIdle.

Gates: realtime tsc 0, 209 tests, biome, boundaries, prune.

* fix(files): drop late sync frames once fatal; file-doc join error suppression + retry-budget reset

Final safety-audit findings:
- CRITICAL: FileDocProvider.handleMessage had no fatal guard. After the connect
  deadline latched fatal and the editor fell back to a read-only local seed, a
  late SyncStep2 (slow server / flaky network / deploy) was still applied — merging
  server state into the seeded doc (content duplication) and flipping synced=true,
  which un-gated autosave and would persist the duplicate to the real file. fatal
  guarded (re)join but not inbound sync. Now handleMessage returns early when fatal.
  +test (late SyncStep2 after the deadline is ignored, doc stays empty + gated).
- file-doc join catch now suppresses the client-facing error when superseded
  (matches workflow/tables/workspace-files) so a retryable error for an abandoned
  file can't make a client re-join and cancel the newer one.
- table/workspace-files room hooks reset the retry budget on (re)connect so a prior
  full exhaustion doesn't block retries after a reconnect.
- presence-visibility: corrected a stale TTL comment.

Gates: tsc (sim/realtime) 0, 209 realtime + collab/hooks suites, biome, boundaries, prune.

* feat(collab-doc): server-authoritative Yjs seeding (#6008)

Server-authoritative Yjs seeding for collaborative file documents: the realtime relay
fetches a Yjs seed built from the file's markdown (via the shared TipTap engine) and
applies it once per room, replacing the client-seeder election/handshake entirely.

- DOM-free markdown<->Yjs conversion core (markdownToYDoc / yDocToMarkdown /
  applyMarkdownToYDoc) reusing the client markdown engine for parity by construction
- Internal x-api-key seed endpoint + realtime fetch; single attempt bounded under the
  client readiness deadline, guard-release for join-driven retry, read-only fallback on
  persistent failure
- Client readiness gate = synced && server seed flag; jsdom wired for the Next standalone
  build (serverExternalPackages + outputFileTracingIncludes)

Foundation only — copilot-into-doc + markdown projection + durable persistence are Stage C.

* feat(collab-doc): Sim merge endpoint for copilot-into-doc (Stage C foundation)

buildFileDocMergeUpdate(docState, markdown) computes the minimal Yjs diff that turns a live
document into target markdown, via applyMarkdownToYDoc (a real updateYFragment diff, not a
replace) — so a copilot rewrite merges with concurrent user edits instead of clobbering them.
Exposed over the internal x-api-key /api/internal/file-doc/merge endpoint the realtime relay
will call: the relay owns the doc, the app owns the conversion engine, so the relay ships the
current state and applies the returned diff. Tested incl. concurrent-edit no-clobber.

* feat(collab-doc): realtime apply-edit — merge copilot markdown into a live doc

The relay can now stream a copilot edit into open editors: applyMarkdownToLiveFileDoc finds
the seeded live room, ships its state to the app's /merge endpoint for a minimal CRDT diff,
applies it (relaying to every editor, reconciled with concurrent user edits), and reports
'no-live-room' so the caller falls back to a direct file write when nothing is open.

- Generalize the realtime->app request module (file-doc-seed.ts -> file-doc-app.ts) with a
  shared POST helper + fetchFileDocSeed/fetchFileDocMerge
- POST /api/file-doc/apply-edit on the internal x-api-key HTTP surface, returning { applied }
- Tests for the seeded-room merge relay and the no-live-room fallback

* feat(collab-doc): stream copilot edits into open editors (Stage C)

edit_content now, after its durable file write, best-effort merges the same markdown into the
file's live collaborative document (markdown files only). If a collaborator has it open, the
edit streams into their editor as a CRDT merge — reconciled with their concurrent typing —
instead of the file silently changing under them; the editor's existing autosave mirrors the
merged doc back to the file. No-op when nothing is open. Never blocks or fails the edit.

* fix(collab-doc): strip frontmatter on merge; gate live-merge to markdown

- buildFileDocMergeUpdate now strips YAML frontmatter (splitFrontmatter().body) exactly as
  the seed does. Copilot passes full-file content, so without this the frontmatter merged
  into the doc as editor content and autosave wrote it back over the file (corruption).
- Gate the live-doc merge on isMarkdownFileName (new server-safe helper) instead of the
  over-broad !isDoc, so code/text edits don't pay the realtime round-trip for a format the
  collaborative editor never renders.

* fix(collab-doc): make frontmatter collaborative so a merge can't revert it

The editor re-attaches its open-time frontmatter on every autosave, so the Stage C merge
(which triggers an autosave with no user action) could write stale YAML back over a copilot
frontmatter change — silently dropping it.

Carry the file's frontmatter in the doc's config map instead of locking it at open: the seed
stores it, the merge updates it (only when it actually changed, preserving the no-op diff),
and the editor re-attaches THAT value on save — falling back to the locked copy before the
seed lands and for non-collaborative docs. A server-side frontmatter change is now reflected
rather than reverted. New FILE_DOC_SEED.frontmatterKey; seed/merge tests cover it.

* fix(collab-doc): re-sync draft on a frontmatter-only merge

A server edit that changes only the frontmatter updates the config map but not the body
fragment, so TipTap's onUpdate never fires — the autosave draft kept the stale open-time
frontmatter, and an explicit save could revert the live change. Observe the config map and,
on a frontmatter-only change, re-attach the new frontmatter to the current body and push a
fresh draft (guarded on a synced body so it never races the seed's own onUpdate).

* fix(collab-doc): order the apply-edit/merge timeouts (outer > inner)

The Sim->realtime apply-edit timeout (4s) was shorter than the nested realtime->Sim merge
timeout (8s), so the outer call could abort while the relay was still merging — the relay
then applied the merge after edit_content had returned, racing a follow-on edit.

Split the shared realtime->app timeout: the seed keeps 8s (it reads a cold blob), the merge
gets a tight 3s (it is a pure conversion, no I/O), and the outer apply-edit is raised to 6s
so it always outlives the inner merge. Cross-referenced in comments to prevent drift.

* fix(collab-doc): address deep-audit findings (jsdom trace, timeouts, gate, races)

A 4-agent LOC audit + precedent research (TipTap/Hocuspocus/Yjs docs confirm the core
patterns are idiomatic) surfaced these real issues:

- The /merge route also lazy-requires jsdom but was missing its outputFileTracingIncludes
  entry — a Docker/standalone build would 500 with MODULE_NOT_FOUND. Add it.
- The four conversion timeouts encoded an ordering invariant (merge<applyEdit, seed<readiness)
  living only in prose across two apps. Hoist to a shared FILE_DOC_TIMEOUTS in
  @sim/realtime-protocol with a test asserting the ordering; both apps import it.
- The copilot live-merge gate used an extension-only check, but the editor treats any
  text/markdown-MIME file as markdown. Replace isMarkdownFileName with a MIME-aware
  isMarkdownFile mirroring the client, so those files stream too.
- applyMarkdownToLiveFileDoc had no per-file serialization; overlapping merges could each
  diff the same stale snapshot and apply out of order. Serialize per file via a promise chain.
- Editor hardcoded 'config'/'initialContentLoaded' instead of FILE_DOC_SEED constants (drift).
- Add .max() bounds to the merge contract body; document the best-effort-merge failure window
  honestly (open editor + merge failure can drop a copilot edit until reload — closed by the
  deferred durable-doc work).

* feat(collab-doc): multi-replica shared Yjs backend + server-side markdown persistence

Make collaborative file-doc editing correct across multiple ECS tasks (the
per-process Y.Doc previously assumed one replica per file).

- Shared Yjs backend over Redis Streams (apps/realtime file-doc-store): each
  file's stream is the ordered, replayable log of updates; a multiplexed XREAD
  tailer converges every task's in-memory doc. Coordinated single-seeder
  election (SET NX + empty-stream recheck) fixes split-brain seeding.
- Doc-sync fans out to local clients + the stream; awareness stays on the
  Socket.IO adapter. Snapshot+XTRIM compaction trims only integrated entries.
- Server-side persistence: project the live doc back to markdown via a new
  /api/internal/file-doc/persist endpoint, debounced during editing and flushed
  on last-disconnect, from the authoritative stream state. Collaborative editors
  no longer client-autosave, closing the copilot clobber-window.
- Copilot merges apply through the stream (reach the live doc on any task) and
  serialize cross-task via a Redis merge lock.
- Degrades to the original single-replica behavior when REDIS_URL is unset.

* fix(collab-doc): harden distributed locks and durability from review

Address Greptile + Cursor review of the multi-replica backend:

- Merge lock: retry LONGER than the lock TTL (guaranteed acquisition, never
  merges against a shared base while a peer holds the lock) and AWAIT the stream
  write before releasing, so the next task never diffs a stale base.
- Distributed locks (seed/merge/compact) now use ownership tokens + a
  compare-and-delete release (Lua), so a lock that expired and was re-acquired by
  another task is never stolen; acquisition fails CLOSED on Redis error.
- Seed: publish the seed to the stream AWAITED under the lock before releasing,
  so a later seeder's empty-stream fence always sees it (closes the fence's
  publish-after-release gap); TTL kept at the readiness deadline.
- Persist: persist the AUTHORITATIVE stream state even when this task's local doc
  was never seeded, and capture the local fallback synchronously so a
  last-disconnect flush never encodes an already-destroyed doc.
- Publish: retry a transient xAdd failure so a Redis blip can't silently drop an
  edit from the shared log.

* fix(collab-doc): only persist a doc a user actually edited

Cursor review: a copilot durable write landing while a doc is being seeded could
have the stale seed projected back over it on last-disconnect, clobbering the
copilot edit even with no user changes.

Gate server-side persistence on a genuine user edit (socket-origin update): a
seed-only or merge-only doc is never projected back to the file (copilot writes
the file durably itself), so it can't clobber a concurrent external write.

* fix(collab-doc): close review-round race/durability gaps

Address Greptile P1 + Cursor findings:

- Seed publishes to the shared stream AWAITED *before* seeding the local doc, so
  a publish failure leaves the doc unseeded and the stream empty for a clean
  retry rather than serving an unpublished local seed a peer would re-seed over
  (split-brain).
- flushPersist falls back to the synchronously-captured local snapshot when
  getStreamState throws (not only when it returns null), so a transient Redis
  read no longer drops the final durable write as the room is torn down.
- streamHasContent fails CLOSED (returns true on xLen error): a Redis blip can no
  longer let the seed fence pass and double-seed.
- Client collabReady initializes from the collaborative prop, so a collaborative
  editor never has a mount-window where client autosave could clobber the server
  write.

* fix(collab-doc): unstick seed retry and persist tailed edits

Cursor review:

- ensureServerSeed clears serverSeedStarted when it aborts at the streamHasContent
  fence, so a fail-closed Redis xLen (or a genuine peer-seed) no longer strands the
  room unseeded with no retry.
- Persistence dirty-tracking now marks a doc edited on any post-seed update,
  including a peer's edit relayed via the tailer (REDIS_ORIGIN), tracked via a
  seededObserved flag. The last task to leave persists real edits even if it only
  tailed them; the seed transition itself is still never counted, so a
  seeded-but-unedited doc is never projected back over the file.

* fix(collab-doc): match client markdown post-processing on server persist

Cursor review: yDocToFileMarkdown serialized the body with yDocToMarkdown only,
but the editor save path runs postProcessSerializedMarkdown before applyFrontmatter.
Server persist could therefore write markdown differing from a client save (empty
list markers, callout un-escaping) — spurious blob churn / round-trip drift despite
the byte-identical claim.

Apply postProcessSerializedMarkdown in yDocToFileMarkdown so a server persist is
byte-identical to a client save and the client's dirty-check baseline. Add a
regression test guarding the composition.

* fix(collab-doc): close durability gaps at deploy boundaries + audit polish

From a comprehensive from-scratch audit (correctness, SOTA, cleanliness,
feature-completeness):

- Persist max-wait: a continuous edit burst kept resetting the 5s debounce and
  never persisted; cap it so a burst flushes at least every 20s, bounding
  unpersisted edits.
- Graceful-shutdown flush: flushAllFileDocRooms awaited in shutdown so a rolling
  deploy / scale-in secures open edited rooms to durable markdown before exit,
  instead of relying on the stream + a surviving task.
- Compacted-snapshot catch-up now marks the doc edited (REDIS_SNAPSHOT_ORIGIN): a
  snapshot folds seed+edits into one frame, so a task catching up purely from it
  no longer treats real edits as an unedited seed and skips persisting.
- Polish: delete dead __setFileDocStoreForTest; bounded retry loop; rename
  acquirePersistSlot -> tryClaimPersistWindow with accurate docs; tailer
  object-identity guard; fix stale comments (seed route, edit-content autosave).

* improvement(realtime): post-review fixes for collab dirty-state + relay lifecycle

- collab editor no longer latches a spurious 'Unsaved changes' prompt: report
  dirty only when the client owns durability (canAutosave), since in a
  collaborative session the relay persists the doc server-side
- guard shutdown against a double SIGINT/SIGTERM running teardown twice
- disconnect local sockets before httpServer.close so shutdown exits gracefully
  instead of hitting the forced-exit timer (local-only, deploy-safe)
- return 400 (not silent 200) on an invalid workspaceId in the files-changed fanout
- clear a pending join-retry timer on reconnect so it can't fire a duplicate join

* fix(realtime): make collab-doc seeding atomic to close split-brain window

An adversarial concurrency audit found a split-brain double-seed vector: the
seed used an advisory SET NX PX lock + a SEPARATE xLen fence + an unconditional
xAdd (a non-atomic check-then-append). If the seed lock's TTL expired mid-seed
(a >4s stall after the 8s seed fetch), a second task could acquire the freed
lock over a still-empty stream, both fences read empty, and both append seeds
with distinct Yjs client ids -> duplicated document content.

- add an atomic SEED_IF_EMPTY_SCRIPT (append-iff-empty in one Redis step) +
  store.seedIfEmpty(); the emptiness check and append are now inseparable, so
  two tasks racing (even both past an expired lock) can never both seed
- ensureServerSeed uses seedIfEmpty instead of streamHasContent-fence +
  publishAndWait; the seed lock is now purely an efficiency optimization
  (avoid a duplicate fetch), not a correctness dependency
- fix the misleading comment that claimed a copilot merge is not counted as an
  edit in the multi-replica path (it round-trips as REDIS_ORIGIN and does count;
  a safe idempotent over-persist, never a lost edit)
- add interleaving tests: seedIfEmpty atomicity/fence, the split-brain
  regression under an expired lock, a peer edit during attach catch-up, and
  concurrent two-task compaction

* docs(realtime): align seed comments with the atomic-append correctness model

Greptile flagged lingering doc drift: the module + shouldSeed comments still
credited the seed lock + empty-stream check as the split-brain fix. Reframe them
so the atomic seedIfEmpty is the exactly-once guarantee and shouldSeed is an
efficiency gate only.

* feat(realtime): live workspace tables list, sharing one invalidation-room impl (#6053)

* feat(realtime): live workspace tables list, sharing one invalidation-room impl

Bring the tables list to parity with the files list: a create/rename/move/delete/
restore now propagates to every viewer live instead of waiting out the 30s
staleTime. Following the files pattern, but factoring the two into one shared
implementation rather than copy-pasting.

- add ROOM_TYPES.WORKSPACE_TABLES + its authz resolver (workspace-id-addressed,
  reuses the workspace resolver like workspace-files)
- extract setupWorkspaceInvalidationRoom (server) and useWorkspaceInvalidationRoom
  (client) — the presence-free, workspace-scoped live-list room; files and tables
  now both bind to it, so they can never drift. Event/room names derive from the
  room type. Replaces the standalone workspace-files handler + hook
- notifyWorkspaceTablesChanged fanout fired from the table service (createTable,
  renameTable, moveTableToFolder, deleteTable, restoreTable) so it covers both the
  HTTP routes AND copilot, which call the service directly
- relay /api/workspace-tables-changed endpoint; wire the hook into the tables page
- consolidate the handler test into one suite run against both room types

* feat(realtime): live tables list also covers table-folder mutations

Fold in the follow-up: a table folder create/rename/move/delete/restore now
propagates to the tables list live too, so the browser is fully consistent.

- generic notifyFolderResourceChanged(resourceType, workspaceId) dispatches the
  workspace live-list signal by resource type (a map, not a special-case if), so
  file/knowledge_base/workflow are no-ops today and gain liveness by adding a map
  entry when they adopt an invalidation room
- fired from the shared folder lifecycle (createFolder/updateFolder/deleteFolder/
  restoreFolder), covering routes AND copilot
- the tables room hook now invalidates the table folders query too, not just the
  tables list, since the page renders both

* fix(realtime): skip per-table live-list notify during a folder cascade

A folder delete/restore already fires one folder-level notifyFolderResourceChanged
for the whole subtree, but the cascade also calls deleteTable/restoreTable per
table — each awaiting its own notifyWorkspaceTablesChanged. A folder with many
tables would run N+1 sequential relay calls (each bounded by NOTIFY_TIMEOUT_MS),
blocking the mutation. Add a skipNotify option the cascade passes so only the one
folder-level notify fires.

* feat(collab-doc): Hocuspocus binary persistence + Next 16 seed/persist fixes (#6059)

* fix(collab-doc): make server-side seed conversion work under Next 16 / Turbopack

Opening a file left both collaborators read-only and stalled ~12s: the server-side
seed (markdown -> Yjs, run through the headless editor engine) was failing, so the
doc never seeded and the editor never left its readiness gate. Two root causes,
both latent until a real build/runtime (typecheck + unit tests don't exercise
either), surfaced by the Next 16 upgrade:

1. Build boundary: the server seed route imported the shared editor schema
   (`createMarkdownContentExtensions`), which pulled in the React node-view
   components (`useEffect`) -> 'client component in a Server Component'. Split each
   node's React-free schema into its own `*-schema.ts` (code-block, image,
   raw-markdown-snippet); the client editor still injects the React node views via
   the existing `nodeViews` param, unchanged.

2. Runtime DOM: the converter installs a jsdom `window` on `globalThis`, but
   Turbopack's server bundle gives bundled `@tiptap/core` a `window` that does NOT
   read `globalThis`, so `elementFromString` threw 'no window object available'.
   Externalize the `@tiptap/*` packages the converter uses (native Node require, so
   their `window` reads the real global) and fix the converter's DOM guard to gate
   on `window` (what TipTap checks) with no sticky flag.

Verified: seed route returns 200 with the Yjs update; 514 collab-doc + editor tests
pass; schema byte-identical after the split.

* feat(collab-doc): persist the Yjs binary and load it on cold-start (Hocuspocus pattern)

Adopt the industry-standard Hocuspocus store/load-document pattern so a cold room
open loads the file's last-persisted Yjs binary directly instead of re-converting
markdown -> Yjs on every open. Rebuilding the CRDT from markdown on each connect is
the exact anti-pattern Tiptap/Yjs warn against (fresh client ids -> duplicated
content); it also forced the fragile server-side headless-editor conversion on every
open. Now conversion runs only on a genuine first open or an external markdown edit.

- new table workspace_file_collab_state(file_id PK->workspace_files cascade,
  doc_state bytea, source_hash, updated_at): the Yjs binary + a hash of the markdown
  it was derived from (bounded <=~1MB by the 256KB round-trip gate). Mirrors
  Hocuspocus's extension-database (binary in a DB column). Migration 0275.
- persist upserts the binary (tagged with the exact markdown just written)
- cold-start seed returns the cached binary when its source_hash matches the file's
  current markdown; otherwise converts (and the next persist refreshes the cache)
- also externalize yjs / y-protocols / lib0 alongside @tiptap: bundling loaded a
  second yjs copy, so @tiptap/y-tiptap's 'instanceof Y.XmlElement' failed on
  app-created nodes ('Unexpected case') during Yjs -> markdown

Verified end-to-end: seed -> persist -> seed returns the exact persisted binary (a
cache hit, no re-conversion). 18 collab-doc + 51 realtime file-doc tests pass.

* fix(collab-doc): best-effort cache read + drop dead barrel

- seed: a cache-read failure (transient DB error, not-yet-migrated cache table)
  no longer aborts a cold room open — the durable markdown is already in hand, so
  fall through to conversion. Symmetric with persist's best-effort cache write.
  Addresses the Cursor Bugbot finding on the read/write asymmetry.
- remove the collab-doc index.ts barrel: nothing imported it (every consumer uses
  direct ./seed / ./merge / ./converter imports), so it was dead re-export surface.
  De-export COLLAB_DOC_FIELD accordingly — it is used only inside converter.ts.

* fix(collab-doc): stream every external file write into open editors, not just edit_content (#6070)

* fix(collab-doc): stream every external file write into open editors, not just edit_content

A copilot/mothership edit to an open markdown file did not appear live in another
user's editor: the live-doc merge bridge (mergeEditIntoLiveFileDoc) was wired into the
edit_content tool ONLY. Every other server-side write — the file tool
(/api/tools/file/manage), function_execute (/api/function/execute via
writeWorkspaceFileByPath), create_file overwrite, and the PUT /content route — went
straight to updateWorkspaceFileContent and skipped the merge, so the durable file
changed but the open editor never updated. Confirmed from live logs (the mothership
'Prepend sentence' ran read + file + function_execute — zero apply-edit calls) and
Redis (the prepended text was absent from the doc stream).

Centralize the merge at the one chokepoint every external writer shares:
- updateWorkspaceFileContent gains an opt-out "syncLiveDoc" (default on) and, after
  the durable write, merges markdown writes into any open collaborative doc (best-effort;
  no-op when nobody has it open). Any current OR future writer is covered automatically.
- persist.ts opts out (syncLiveDoc:false) — it IS the doc→markdown projection, so merging
  it back would be a persist→merge→persist self-loop.
- create_file opts its empty shell out (real content arrives via a later write) so an open
  editor never flickers to empty on overwrite; threaded through writeWorkspaceFileByPath.
- edit_content drops its now-redundant explicit merge call (the chokepoint handles it).
- binary writers (image/video/audio/ffmpeg/download) are naturally excluded — the merge is
  gated to markdown, the only format the collaborative editor renders.

Also bump the api-validation route baseline 994→996 to match the true route count already
on this branch (pre-existing ratchet drift from an earlier merge; NOT added by this PR).

* fix(collab-doc): defer setEditable out of the render phase (flushSync warning)

The collab editability-reapply effect called editor.setEditable synchronously. In collab
mode isEditable flips from readiness (synced + seeded), which is driven by a Yjs
config.observe firing synchronously inside Y.applyUpdate — so the effect can run while React
is mid-render. TipTap's React binding commits setEditable's transaction with flushSync, which
throws "flushSync was called from inside a lifecycle method. React cannot flush when React is
already rendering." Defer the setEditable to a microtask (runs right after the current commit,
before paint), guarding against a destroyed editor or a stale value before it fires.

Only the collab path (this effect) hit the warning; the streaming/settle effect's setEditable
calls run on the non-collab path where isEditable isn't driven by a mid-render Yjs observer.

* fix(rich-markdown-editor): defer non-collab settle/stream mutations off the render phase (flushSync) (#6073)

* fix(rich-markdown-editor): defer non-collab settle/stream mutations off the render phase (flushSync)

The non-collaborative streaming/settle effect called editor.setContent / setEditable /
setTextSelection / focus directly in the effect body. setContent mounts the custom node views
synchronously through the @tiptap/react flushSync path (tiptap#3764), so when this effect runs
while React is mid-render it throws "flushSync was called from inside a lifecycle method." This
is the second flushSync source (the collab editability effect was the first, fixed separately);
it fires on the agent-streaming-into-a-non-collab-editor surface.

Defer the effect-body view mutations to a microtask via a small runOffRender helper (runs right
after the current commit, before paint; no-ops if the editor was torn down). The settle block is
deferred as ONE microtask so setContent -> collapse selection -> setEditable -> focus keep their
order. The streaming rAF tick is left untouched — it already runs off-render, so it keeps writing
content directly. queueMicrotask is TipTap's own documented remedy for this warning.

497 rich-markdown-editor tests (incl. stream-settle-selection) pass; tsc + lint + api-validation
+ boundary + prune green. Needs a live check: stream an agent into a non-collab markdown file and
confirm it still renders smoothly.

* chore(rich-markdown-editor): trim verbose flushSync-defer comments

* fix(rich-markdown-editor): drop superseded settle/stream microtasks via a run token

runOffRender previously only guarded editor.isDestroyed, so if React ran the next reconcile
pass (a newer stream or settle) before a queued microtask flushed, the stale microtask could
still apply setContent/setEditable/setTextSelection over the newer state. Tag each effect run
with an incrementing token; a deferred mutation applies only when its run is still the latest
(and the editor is alive). A run token fits this effect's several early-return exits better
than a per-exit cleanup flag. Addresses Greptile/Cursor review.

* fix(rich-markdown-editor): never drop the settle selection-collapse under a superseded run

The run token drops a superseded settle's microtask, but the settle had already flipped its state
flags synchronously — so a pre-empting steady-sync run took the non-settle path and never collapsed
the selection, leaving a post-stream select-all painting the leaf-in-selection decoration. Track the
collapse as a debt (pendingCollapseRef): whichever deferred run ultimately applies — settle or the
steady-sync path — clears it, so the collapse runs exactly once on the latest content. Addresses the
Cursor review finding.

* feat(tables): show live cell-selection carets in the embedded chat panel (#6081)

The table cell-selection presence room was joined only on the dedicated /tables/[id] page
(useTableRoom was passed an empty id in embedded mode). Join it in embedded too, so the
mothership chat resource panel shows collaborators' live cell selections and broadcasts the
local one. tableId is already resolved from props in embedded (the data event stream already
uses it un-gated), and authz runs on join, so this is safe. Avatars are unaffected — they
render only in the !embedded Resource.Header, so the panel gets carets without avatars.

* feat(collab-doc): If-Match optimistic concurrency so persist never clobbers an out-of-band edit (#6085)

* feat(collab-doc): optimistic-concurrency guard so persist never clobbers an out-of-band edit

The relay projected the live Yjs doc back to durable markdown unconditionally (last-write-wins), so
a persist already in flight when an external write landed could overwrite it. Add RFC 7232 If-Match
optimistic concurrency end to end, reconciling through the CRDT (never rejecting user work):

- updateWorkspaceFileContent gains an expectedUpdatedAt guard: the write commits only if the file is
  still at that version (checked against the SELECT ... FOR UPDATE-locked row, so it is atomic with
  the write), else it throws the new ContentVersionConflictError without clobbering.
- persistFileDoc takes expectedVersion and returns a discriminated result (persisted | missing |
  conflict). On conflict it returns the current durable content + version instead of writing.
- The relay tracks the durable version its live doc is synced to — set on seed, advanced when a
  durable write is merged in (apply-edit carries the version), and on each successful persist. It is
  held cluster-wide in Redis (filedoc:syncver:{name}) so whichever task persists reads the same
  version, with the per-room value as the single-pod fallback.
- flushPersist sends that version as If-Match. On a conflict it merges the current durable content
  into the live doc (so the out-of-band edit AND the live edits converge) and retries (bounded), so
  even a last-leave flush racing an external write persists the reconciled result rather than losing
  the session's edits.

Threads the version through the seed + persist contracts and the apply-edit payload. No schema change
(reuses workspace_files.updatedAt as the version token). Tests: app-side CAS (match writes, mismatch
throws + cleans up the orphan upload), relay conflict handled gracefully without clobber/loop; 236
realtime + 76 sim collab/uploads tests, tsc x2, lint, api-validation, boundaries, prune all green.

* chore(collab-doc): heartbeat-refresh the synced-version key TTL alongside its stream

Keep filedoc:syncver:{name} alive as long as the room's stream (it was only re-set on
seed/merge/persist), so an open-but-idle doc's persist If-Match token can't expire and force a
needless reconcile.

* fix(collab-doc): stop persist-conflict retries when there is no live doc to reconcile

On an If-Match conflict with no live doc to reconcile into (last collaborator gone, no shared
stream), applyMarkdownToLiveFileDoc returns no-live-room; re-projecting the same pre-teardown
snapshot would only re-conflict, so break the retry loop immediately and leave the out-of-band
(durable) content authoritative — the intended conflict policy. Addresses Greptile review.

* fix(collab-doc): close three optimistic-concurrency edge cases from review

- Single-pod persist retry projected the pre-reconcile snapshot (captureState always returned the
  initial localState), while the synced version had been advanced by the reconcile — so the If-Match
  could pass and clobber the reconciled edit. captureState now re-reads the live doc on each attempt
  (falling back to the pre-teardown snapshot only once the room is gone).
- The synced version was recorded from this task's own seed FETCH before knowing whether this task's
  seed actually won; a peer winning with a different version could leave a newer token than the stream
  content. Record it only inside the didSeed branch (the task whose seed won); peer-seeded tasks read
  the winner's cluster value.
- Persist wrote UNCONDITIONALLY when no version was available (relay version momentarily missing), which
  could clobber non-empty durable content. It now returns conflict for a non-empty file with no version
  (reconcile/retry once the version is re-established); an empty file's first write stays unconditional.

* fix(collab-doc): defer (not reconcile) on missing version, and use the freshest version token

- Missing-version persist now returns 'deferred' instead of 'conflict'. A missing version token (a
  Redis blip on a peer-seeded task) is NOT a genuine out-of-band change, so triggering a reconcile
  would wipe live edits (incoming-wins) even though nothing changed durably. Deferred means: don't
  write, don't reconcile — leave the edits in the stream and let a later persist write them once the
  version is re-established.
- currentVersion now takes the MAX of the cluster (Redis) and local room versions rather than always
  preferring Redis, so a lagged/failed fire-and-forget Redis set can't shadow a newer local value and
  cause spurious If-Match conflicts. Versions are monotonic epoch-ms, so the larger is the later sync.

* fix(collab-doc): make persist If-Match teardown-race-immune and recover missing version on final flush

Close two last-leave concurrency holes Cursor flagged:

- Thread the reconciled version LOCALLY through the persist retry loop. After a
  conflict+reconcile the correct next If-Match is exactly result.version, so carry
  it in a local var instead of re-deriving from room.syncedVersion/Redis. On a
  last-leave flush destroyRoomIfIdle removes the room from the map before the async
  flush finishes, so mergeMarkdownIntoRoom's recordVersion can no longer update
  room.syncedVersion — threading makes each retry's precondition correct by
  construction, immune to that dropped mutation and to a best-effort Redis re-read.

- Cache the resolved version back into room.syncedVersion in currentVersion() so a
  peer-seeded/tail-only task (which never sets it locally) or a later transient
  Redis read failure still resolves it from the last value seen (monotonic max,
  never regresses).

- On a FINAL flush, briefly retry resolving the If-Match when the version read
  momentarily fails, rather than deferring and stranding the session's edits in the
  TTL'd stream — the version is cluster-wide and heartbeat-refreshed.

* fix(collab-doc): stamp cluster sync version the moment the seed wins, before the liveness guard

The winning seeder set the If-Match token (room + Redis filedoc:syncver) only after the
liveness/seeded guard that follows seedIfEmpty. But the tailer can integrate the just-appended
seed DURING the seedIfEmpty await, so isDocSeeded(room.doc) is already true when the guard runs
and it returns early — leaving the stream holding seed content with no cluster version. Later
persists then send no If-Match, the app returns `deferred`, and session edits stay only in the
TTL'd stream (the exact stranding this PR prevents elsewhere).

Move the version stamp to immediately after seedIfEmpty wins, before the guard. Recording it only
once our seed won (not from the fetch) is preserved, so it still can't shadow a peer's winning
seed.

* fix(collab-doc): make the synced-version token monotonic at every write site

The If-Match token is written fire-and-forget from the seed stamp, merges, and persists, both
locally and to Redis. An out-of-order write (e.g. a seed's lagged setSyncedVersion landing after a
later merge's) could regress it below the version the live doc already incorporates, causing
spurious If-Match conflicts — and on a last-leave flush with no live room to reconcile into, a
spurious conflict leaves durable authoritative and drops the session's edits.

- setSyncedVersion now writes via SET_VERSION_IF_NEWER_SCRIPT (Redis-side compare-and-set): it
  overwrites only when the new value is greater, refreshing the TTL either way.
- recordVersion / the persisted branch / the seed stamp all take Math.max instead of assigning
  room.syncedVersion directly.

Versions are monotonic epoch-ms, so "newer" is a plain numeric compare, exact within a Lua double.

* fix(collab-doc): close three last-leave persist edge cases from review

- Stale snapshot after reconcile (High): the multi-task captureState fell back to the pre-await
  localState snapshot even after a reconcile advanced ifMatch, so a failed stream re-read could
  persist the pre-reconcile state against the new version and clobber the out-of-band edit the
  reconcile just incorporated. NULL localState after a reconcile so a failed read aborts instead.

- Lock miss aborts reconcile (Medium): a merge-lock acquisition failure returned 'no-live-room',
  indistinguishable from an absent stream, so flushPersist treated transient contention as
  terminal. Return a distinct 'merge-unavailable' and handle it as retry-later (edits stay in the
  stream), never as "nothing to reconcile into".

- Peer syncver never recovers (Medium): the winner's setSyncedVersion was fire-and-forget with
  swallowed errors — the only way a peer-seeded task learns the durable version — so a dropped
  write left that peer deferring forever. Make it retry (bounded) like appendUpdate/seedIfEmpty;
  the monotonic script keeps a racing retry a no-op.

* fix(collab-doc): scope the persist If-Match to a content version so metadata bumps can't clobber edits

The optimistic-concurrency validator was `updatedAt`, which rename/move/delete/restore also bump
with no content change. A racing live-doc persist then saw a stale token, got `conflict`,
reconciled the pre-edit durable body via updateYFragment (incoming-wins on overlap), and wiped the
user's in-flight edits.

Scope the validator to content (RFC 7232 semantics — validate the representation, not the row):
- New `workspace_files.content_updated_at` (NOT NULL, `now()` fast-default — no table rewrite).
  Advances ONLY on content writes (upload / overwrite / create); metadata writes never touch it.
- The FOR UPDATE CAS, the merge-notify version, and the seed version all use `content_updated_at`.
  A rename now leaves it unchanged, so the persist If-Match still matches -> no spurious conflict,
  no reconcile, no lost edits. Genuine out-of-band content writes still conflict and reconcile.
- Consolidated the collab schema into one migration (the collab-state table + the new column) per
  request, rather than a separate follow-up migration.

Relay/store/contracts unchanged (still a numeric monotonic version).

* chore(collab-doc): condense the densest persist comments (no behavior change)

Cleanup pass: tighten the three longest comment blocks added while hardening the persist path
(currentVersion cache, ifMatch threading, final-flush version retry) without dropping any invariant.
No dead code found (biome lint clean; all new symbols referenced).

* fix(collab-doc): persist must return the content version, not updatedAt

Follow-up to the content-scoped If-Match: persistFileDoc still returned `updatedAt` as the version
in both the persisted and conflict results, while the CAS/seed/merge all guard on
`content_updated_at`. A content write sets both to the same instant, so it was coincidentally
correct — until they diverge: if a metadata write bumps `updatedAt` past `content_updated_at`, the
conflict path returned the larger `updatedAt`, so the relay's re-persist sent an If-Match the CAS
(which checks `content_updated_at`) could never match → perpetual conflict → dropped reconciled
edits. Return `contentUpdatedAt` in both paths so the relay's token always matches what it's checked
against.

* fix(collab-doc): defer persist whenever the version is missing; guard the content-version test

- Empty-file CAS race (Medium): the unconditional-write carve-out for size===0 read `record.size`
  outside the write transaction, so a concurrent first content write could land after the check and
  be clobbered. With content_updated_at NOT NULL every existing file always has a real version, so a
  missing expectedVersion is always transient — always defer, never write unconditionally. Removes
  the TOCTOU hole.
- Content-version test (Low): the merge-chokepoint test kept updatedAt == contentUpdatedAt, so it
  passed even if wired to the wrong field. Mock distinct values and assert contentUpdatedAt, so a
  regression to updatedAt now fails the test.

* fix(collab-doc): don't reconcile a conflict the live doc already reflects (would wipe newer edits)

flushPersist reconciled the durable body into the live doc on every conflict. But when the conflict
comes from a racing self-persist (or an apply-edit the chokepoint already merged), the durable body
is a STALE SUBSET of the live stream, and the incoming-wins updateYFragment merge moves the doc
backward — wiping newer in-flight edits, which the retry then persists.

Before reconciling, re-check the freshest synced version. If it already covers the conflict version,
the live doc has already incorporated that content (or is ahead), so skip the reconcile and just retry
with the freshest version as If-Match — the re-projection captures the current live stream, preserving
every edit. Only a genuine out-of-band change the live doc hasn't incorporated (freshest < conflict
version) is reconciled in. freshest never exceeds the durable version, so this can't loop.

* fix(collab-doc): make content_updated_at monotonic per file; skip-reconcile can't loop

The If-Match token was stamped with app-local new Date() on each content write, so cross-instance
clock skew could stamp a later write with an EARLIER content_updated_at — breaking the version
ordering the whole optimistic-concurrency scheme (and the skip-reconcile branch's freshest>=version
assumption) depends on. Under skew the relay's monotonic syncedVersion could exceed the durable
version, sticking the If-Match: persist conflicts forever, exhausts retries, drops the session's edits.

- Stamp content_updated_at strictly after the current committed value (we hold the row's FOR UPDATE
  lock): new Date(max(now, currentFile.contentUpdatedAt + 1ms)). Monotonic per file regardless of
  clocks; also removes same-millisecond collisions. updatedAt stays plain wall-clock (display/sort).
- Skip-reconcile branch retries with result.version (the durable value the CAS will match), never
  freshest (which could exceed it and loop). Belt-and-suspenders now that the version is monotonic.

* refactor(collab-doc): drop the destructive in-persist reconcile; adopt-version-and-retry on conflict

The in-persist reconcile projected the durable body back over the live doc via updateYFragment
("make the doc match"). That is destructive: when the live stream is already ahead — the common case,
because the write chokepoint (mergeEditIntoLiveFileDoc) already merged the out-of-band change into the
stream — it moved the doc backward and wiped newer in-flight edits. This produced a run of races
(stale snapshot, wipe-newer-edits, version-lag skip miss) that a full-document reconcile fundamentally
can't avoid, since deciding when it's safe relies on a laggy cross-task version token.

Remove it. On conflict, adopt the durable version as the new If-Match and retry: captureState re-reads
the current stream (which holds the out-of-band change AND the live edits), so the re-projection
persists the converged result. The durable change reaches the live doc via the chokepoint, never here.
Trade-off: the only unmerged out-of-band write is one whose chokepoint merge itself failed (rare,
logged), which we accept over the frequent reconcile-wipes-edits race.

- flushPersist: conflict -> ifMatch = result.version, retry (bounded). No applyMarkdownToLiveFileDoc.
- conflict response drops `markdown` (contract + relay type + persist) — no body needed, saves a blob
  fetch. applyMarkdownToLiveFileDoc stays (still used by the apply-edit route / the chokepoint).

* fix(collab-doc): don't let a last-leave conflict retry clobber via the stale local snapshot

Regression from dropping the reconcile: on conflict the retry adopts result.version and re-reads
captureState. But after single-pod last-leave teardown the room is already destroyed, so captureState
falls back to the pre-teardown localState (which lacks the out-of-band change); the retry then CAS-passes
and overwrites the committed external write — undoing the external-wins last-leave policy.

Null localState on the first conflict, so the retry can only use freshly-read authoritative state
(stream / live doc). When none is available (single-pod room gone, or a transient stream-read failure)
captureState returns null and the retry stops, leaving durable content authoritative. Covers both the
single-pod and multi-task-stream-unavailable variants of the stale-snapshot clobber.

* fix(collab-doc): stop (don't re-persist) on a persist conflict — closes the commit-window clobber

The conflict retry adopted the durable version and immediately re-persisted the current stream,
assuming the stream already held the out-of-band change. But an external write commits durable BEFORE
its chokepoint merge (mergeEditIntoLiveFileDoc) reaches the stream, so a persist landing in that window
CAS-passed with a stream that still lacked the external content and clobbered the committed write — not
just the rare merge-failed path, but a race on every external write, worst at last-leave flushes.

Make persist a single attempt: on conflict, STOP and leave durable authoritative. The chokepoint merges
the change into the stream and — only once it is actually there — advances the synced version via its own
recordVersion; a later flush (debounced or final) then projects the converged stream with a matching
token. The session's edits stay in the stream meanwhile. The conflict handler deliberately does NOT
advance the synced version, or the next flush would clobber with a still-behind stream. Removes the retry
loop and PERSIST_CONFLICT_RETRIES.

* improvement(tables): fire the live-rows signal on async delete, run cancel, and column run (#6094)

* improvement(tables): fire the live-rows signal on async delete, run cancel, and column run

These three table operations mutate row data but emitted no `rows` change signal, so open editors'
grids stayed stale until a manual refresh (enrichment *results* already stream live via `cell` events;
these are the bulk paths that don't emit per-cell events):

- Async row delete (`runTableDelete`): signal as rows drop out (throttled with the existing progress
  event) and once more on completion — the `job` progress event only drives the delete meter, not the
  rows query. Covers the delete-async route and the copilot bulk-delete, since both share the runner.
- Cancel runs (`cancel-runs` route): cancelling clears each affected row's exec state; the
  `dispatch: cancelled` events drop the run overlay but the client then renders authoritative DB state,
  so refetch. Only when something was actually cancelled.
- Run column (`columns/run` route): starting a run bulk-clears the target group's cells to pending;
  refetch so the cleared cells show. Only when a dispatch was actually created.

Guarded so no signal fires on a no-op/failure. Adds a delete-runner test asserting the completion signal.

* fix(tables): guarantee the live-rows signal on every mutating path (review)

- Delete runner (Greptile P1): a batch could commit and the job then cancel/supersede before the next
  throttled progress signal or `markJobReady`, bypassing both signals and leaving deleted rows on
  screen. Track `deletedAny` and fire the grid refetch in a `finally`, so it runs on EVERY exit —
  completion, cancel/supersede, mid-batch lock, or a rethrown error after a partial delete.
- cancel-runs / columns/run routes (Cursor): the `cancelled > 0` / `if (dispatchId)` guards don't
  always reflect DB row changes — cancel tombstones exec state even when 0 dispatches were active, and
  a run bulk-clears cells then can return a null dispatchId. Signal unconditionally; a stale-but-harmless
  refetch beats a missed one.
- Tests: assert the delete signal fires on the mid-run-cancel-after-delete path and NOT when nothing
  was deleted.

* fix(tables): mark deletedAny before the page delete so a mid-page lock still refreshes the grid

`deletePageByIds` commits in internal batches, so a delete lock landing mid-page can persist earlier
batches and THEN throw TableLockedError — the catch returns without a count, so setting `deletedAny`
from the return value missed it and the finally skipped the grid refetch. Set `deletedAny = true` before
the call (any attempt may commit rows); an attempt that commits nothing only over-refetches (harmless).
Adds a test asserting the signal fires when a page throws a mid-page lock.

* fix(files): make embedded resource file view collaborative (#6095)

* fix(files): make embedded resource file view collaborative

The /chat resource panel rendered saved files through FileViewer without
the collaborative opt-in, so a file open on the Files page and the same
file open in the embedded panel never joined the same file-doc room —
no live carets and no live content sync between the two surfaces.

Pass collaborative on the EmbeddedFile FileViewer. Collaboration still
self-gates on canEdit + non-streaming + workspace doc, so the agent
token-stream preview (the dedicated streaming-file path, canEdit=false)
is untouched.

* fix(files): refcount file-doc room membership per shared socket

Two collaborative surfaces in one tab (the Files editor and the embedded
chat resource panel) share one Socket.IO connection, so both providers for
the same file JOIN the same room over that socket. The server's LEAVE does
socket.leave(name) with no membership refcount, so the first provider's
destroy() would strand the second still-mounted one — no more live content
or presence.

Count live providers per file per socket (keyed by the stable Socket object,
so it survives reconnects) and emit LEAVE only when the last provider for a
file tears down. The single-provider path is unchanged (0->1->0).

* feat(tables): propagate shared saved-view changes to collaborators live (#6100)

Table views (named filter/sort/layout presets) are table-wide shared state —
every reader sees every view — but view create/update/delete had no realtime
signal, so a collaborator only saw another user's view changes on their own
staleTime/focus refetch.

Add a 'views' table event kind + signalTableViewsChanged, emitted from the
views service (createTableView/updateTableView/deleteTableView, on real
success only), and a client handler that invalidates the views query alone
(no rows/definition refetch — a view is presentation state on the loaded
table). Mirrors how row/schema/metadata changes already propagate.

* test(tables): cover the views realtime signal (emit + on-success-only) (#6101)

- events.test.ts: signalTableViewsChanged appends a single 'views' event
  carrying the tableId (through the real memory buffer).
- views/service.test.ts: create/update/delete emit signalTableViewsChanged
  on real success, and DON'T on a no-op (a PATCH/DELETE targeting a missing
  view changes nothing, so it must not signal). Mirrors delete-runner's
  signal-path coverage; drives the DB via the shared dbChainMock.
- Add tableViews to the comprehensive @sim/db/schema test mock so the
  service tests can queue the in-transaction existence row.

* chore(ci): reconcile api-validation baselines after the staging merge

The staging merge unioned realtime-rooms's own `as unknown as` cast
(lib/collab-doc/converter.ts) with staging's zod-recursive-type cast
(lib/api/contracts/tables.ts), so the non-test double-cast count is 9 —
both casts pre-existed and were individually accepted on their branches.
Also tighten rawJsonReads 6->5 to the true current count. Fixes the strict
API contract boundary audit on realtime-rooms.

* feat(copilot): stream file edits into the live collaborative Y.Doc (keep embedded view collaborative) (#6108)

* feat(copilot): stream file edits into the live collaborative Y.Doc

Copilot's file edits previously only reached the live doc once, at the final
edit_content write, so a collaborative editor watching the file saw nothing
until completion (streaming looked broken) and the client-side preview path
was suppressed in collab mode.

Make copilot a CRDT peer: as it streams append/update/patch content, merge the
growing markdown into the file's live Y.Doc via the existing apply-edit path
(a minimal updateYFragment diff, concurrent-edit-safe), throttled to ~250ms.
version is omitted for these intermediate merges — they advance the live doc
for viewers but are not durable checkpoints; the final edit_content write
carries the real contentUpdatedAt and reconciles the durable file. Per the
relay's persist gating, server-internal merges never schedule a persist, so a
copilot-only stream produces zero intermediate file writes.

- notify.ts: mergeEditIntoLiveFileDoc version is now optional (streaming omits it).
- file-preview-adapter.ts: throttled live-doc merge at the edit_content stream hook.

* fix(copilot): order + gate streaming live-doc merges; fast collab first render

Harden the streaming merge (adversarial review):
- Order + bound: dispatch through a per-file in-flight guard (drop-while-in-flight)
  so a stale out-of-order snapshot can never land after a newer one and regress
  the doc, and relay load is capped at one request per file regardless of rate.
- No wipe: gate append/patch on the base file content having loaded — a base-less
  snapshot would diff to a delete-everything wipe of the seeded doc; update streams
  a full rewrite from scratch and needs no base.
- Markdown-only gate: non-markdown files have no collaborative room, so skip the
  wasted relay round-trip.

Fast collab first render (Issue 2): render the already-fetched markdown read-only
via generateHTML while the collaborative doc seeds, with the editor mounted-but-
hidden in the same layout box for a seamless swap on collabReady. Pure HTML — it
never touches the Y.Doc (client seeding duplicates the doc), and generateHTML
escapes text (raw-HTML snippets render escaped), so no XSS.

* test(copilot): cover streaming file edits into the live collaborative Y.Doc

Drives edit_content args_delta stream events through processFilePreviewStreamEvent
and asserts the live-doc merge: fires with the growing FULL previewText and no
version arg; is throttled (~250ms per file); is skipped for non-markdown files
and for a base-less append (the delete-everything wipe guard); and runs at most
one-in-flight per file. Verified to fail if any gate/guard is removed.

* fix(collab-doc): coordinate live-doc merge ordering in one place; close durable-clobber race

The second review found a residual: the durable edit_content write went through a
different path than the adapter's in-flight guard, so a late straggler streaming
merge could land after it and, via a persist, clobber the durable file's tail.

Move the per-file coordination into mergeEditIntoLiveFileDoc (the one place both the
streaming and durable paths call): a streaming (versionless) merge is dropped while
one is in flight for the file; a durable (versioned) write instead WAITS for the
in-flight streaming merge, so the final content is always the last merge applied and
can't be regressed by a straggler. Simplifies the adapter (drops its Set + helper).

Relocate the one-in-flight test to notify.test.ts (streaming-drops-while-busy +
durable-waits-then-applies-last); the adapter test keeps throttle/gates/previewText.

* fix(copilot): address review — order merges, exclude update, gate throttle, unhide stream

Review round on #6108:
- Greptile P1 (durable merges lose ordering): serialize ALL merges per file on one chain in
  mergeEditIntoLiveFileDoc (each chains after the current tail), so concurrent durable writes
  can't resume-and-fire out of order. notify now exposes isLiveDocMergeInFlight.
- Cursor High (update stream blanks the doc): only append/patch stream — they build on the loaded
  base; update is a from-scratch rewrite whose partial snapshot would diff the full doc toward a
  fragment, so it applies atomically at the durable write.
- Cursor Medium (throttle advances on a dropped merge): the adapter gates on !isLiveDocMergeInFlight,
  so the send throttle advances only on an actual dispatch — no lag, no backlog behind a slow relay.
- Cursor Medium (placeholder hides a live stream): show the fast-render placeholder only when not
  streaming, so a stream that starts before the doc seeds shows through the editor.
- Soften merge.ts/notify.ts comments per the lifecycle audit: only UNTOUCHED regions are preserved;
  a region the merge rewrites reconciles toward copilot's content.

Tests updated: notify covers chain ordering + isLiveDocMergeInFlight; adapter covers append streaming,
throttle, non-markdown/base-less/update skips, and the in-flight skip.

* fix(collab-doc): reject stale durable merges at the relay (cross-process ordering)

The in-process merge chain only orders merges within one apps/sim process. Two durable
writes for the same file on DIFFERENT processes could reach the relay out of dispatch
order; the relay recorded the version monotonically but still APPLIED the older markdown,
regressing the live doc while the token stayed high (a later persist could then write the
stale content back over the durable file).

Enforce ordering at the relay — the single cross-process coordination point — using the
existing Redis primitives: under the per-file Redis merge lock, read the cluster-wide
synced version and SKIP a versioned merge that is not newer (a newer durable write already
landed). Make recordVersion await setSyncedVersion so it is durable before the lock
releases, so the next holder's staleness check reads a consistent value. Streaming
(versionless) merges are unaffected — they carry no durable version and are ordered
per-process by the caller.

Adds a relay test asserting a stale/idempotent versioned merge returns 'stale' and never
computes or publishes a diff.

* fix(copilot): match durable path — detect markdown by MIME type + name at the stream gate

The streaming gate checked isMarkdownFile with only the filename, while the durable merge
uses type + name — so a text/markdown file without a .md extension was skipped mid-stream
(it self-corrected at the durable write). Pass editIntent.contentType so streaming detects
the same set of markdown files as the durable path.

* test(copilot): assert throttle follow-through after an in-flight merge clears

* fix(collab-doc): order streaming merges by streamedAt so a late snapshot can't regress a newer durable write

* refactor(collab-doc): tidy merge-order docs + relay order object; cover multi-replica streaming stale-check

* fix(collab-doc): order streaming merges by causal base version, not wall-clock

A streaming snapshot now carries baseVersion (the durable contentUpdatedAt it was
built from) instead of a wall-clock streamedAt. The relay drops the snapshot when a
newer durable write landed since that base, so a concurrent human save can no longer
be clobbered in the live doc and then persisted over the durable file. Skew-immune:
both keys are DB-monotonic contentUpdatedAt values.

* fix(collab-doc): derive streaming baseVersion as contentUpdatedAt ?? updatedAt

Match the version line the seed/persist use so a legacy file with no content
version still ships an ordered streaming snapshot instead of an unordered one.

* fix(collab-doc): fail-closed on a streaming snapshot with no baseVersion

The live-merge gate now requires a numeric baseVersion, not just loaded base
content. A rare base with no file record (hence no version) would otherwise ship
an unordered snapshot the relay can't stale-check, risking a clobber of a
concurrent durable write. Skip the live merge instead; the durable write reconciles.

* docs(collab-doc): document the accepted concurrent-independent-streams limitation

* chore(ci): reconcile api-validation baseline after the staging merge

The merge commit auto-merged the baseline at 1000; bump totalRoutes/zodRoutes to
1003 for staging's three new contract-bound routes (nonZodRoutes still 0).

* test(files): update storage-accounting assertion to the mergeEditIntoLiveFileDoc options object

* fix(collab-doc): trace the full yjs/tiptap external stack into the file-doc route bundles

The seed/merge/persist internal routes run the collab-doc converter (markdown <-> Yjs
via headless TipTap) server-side. Those deps are serverExternalPackages, and the
standalone tracer only force-included jsdom — it does NOT follow yjs's ESM subpath
imports of lib0 (lib0/logging, ...), so Docker/standalone builds shipped node_modules
without them and the seed route 500'd (Cannot find module 'lib0/logging'). That left
every collaborative document unseeded and permanently read-only on deployed envs.
Force yjs, lib0, y-protocols, and @tiptap into the trace for all three routes.

* fix(collab-doc): copy the full yjs/lib0 stack into the app image

The seed/merge/persist routes run the converter (markdown <-> Yjs) server-side. yjs is a
serverExternalPackage and the Next standalone tracer copies lib0 only partially — it drops the
ESM subpath file lib0/logging.js that yjs.mjs imports via lib0's exports map, so the seed 500s
('Cannot find module lib0/logging') and every collaborative doc is stuck read-only. Verified in
the running dev container: /app/node_modules/lib0 had 37/38 files, logging.js missing.

outputFileTracingIncludes can't fix it — its globs resolve against apps/sim, but these deps hoist
to the monorepo-root node_modules, so the glob matches nothing (my prior next.config attempt was a
no-op; reverted). Instead COPY the complete lib0/yjs/y-protocols from the deps stage in the runner,
overwriting the partial trace — the same pattern already used for isolated-vm.

* feat(files): stream copilot edits into the collaborative doc smoothly (#6122)

* feat(files): stream copilot edits into the collaborative doc smoothly

- apply the agent stream client-side into the live Yjs binding as minimal
  updateYFragment diffs (like main's setContent, but incremental) so it renders
  smoothly AND broadcasts to every peer via CRDT — a collaborator on /files sees
  the stream for free
- gate the apply on collabReady so diffs never land on an unseeded doc; keep the
  read-only placeholder visible until the seed swaps in
- run streamed ops under a dedicated tx origin so they stay out of the user's
  undo stack
- delete the throttled server-side streaming merge and the baseVersion ordering
  machinery it needed (relay + notify + session contract); the durable final
  write still reconciles open editors and seeds late joiners

* fix(files): apply agent stream as a true CRDT peer + guard base-less snapshots

Review round 1 (Greptile P1s):
- apply the stream against a private shadow replica (seeded from the live doc at
  stream start) and relay only the agent's own delta into the shared doc, so a
  concurrent peer edit to a region the agent snapshot didn't include is no longer
  reverted (previously the whole-body reconcile deleted it)
- gate append snapshots on "must extend the base": a base-less append fragment
  (emitted before the base loads) can no longer reconcile the seeded doc to a wipe;
  patch still legitimately replaces a mid-region
- gate the apply on collabReady so diffs never land on an unseeded doc; keep the
  placeholder visible until the seed swaps in
- plumb streamOperation through the preview surfaces to drive the append gate
- add a peer-edit-preservation test (fails under whole-body reconcile) and refresh
  the undo-isolation + broadcast tests for the session API

* fix(files): destroy the agent shadow deterministically on settle

Cursor round 1 (Low): endAgentStream ran inside runOffRender, whose microtask is
dropped when a rapid follow-up stream bumps the run token — leaking the shadow
Y.Doc. Split it out into an unguarded microtask queued after the (droppable) final
apply, so the shadow is always destroyed.

* fix(files): agent stream frames skip the relay's durable persist

Cursor round 1 (High): client-applied stream frames broadcast over the sync
channel, so the relay stamped a socket origin and ran schedulePersist — durably
writing partial agent content mid-stream, attributed to the watching user (the old
server-merge applied with no origin and never did). Restore that behavior:

- new FILE_DOC_MESSAGE_TYPE.SYNC_NO_PERSIST wire tag; the provider tags
  AGENT_STREAM_ORIGIN updates with it (normal user edits stay SYNC)
- the relay applies it under an AgentSyncOrigin (carries the socket id for
  broadcast exclusion, but is not a plain string) so originSocketId() is null →
  no edited/schedulePersist/lastEditorUserId; excludeSocketId() still excludes the
  sender, and the update still publishes to the stream so peers converge
- the copilot's final edit_content write remains the authoritative durable persist
- tests: relay applies+fans-out but never persists a SYNC_NO_PERSIST frame
  (verified it fails if applied as a socket edit); provider tags agent edits

* fix(files): open the stream shadow at start + private extend baseline

Cursor round 2:
- High (settle skips apply without session): the stream shadow is now opened on
  the first ready frame, BEFORE the extend gate — so an `update` rewrite (whose
  every frame is gated out until settle) and a stream that finishes before seed
  still get a session, and settle applies the final body via the reused-or-on-demand
  shadow instead of leaving the doc stale until the durable reconcile.
- Medium (peer edits stall the stream): the extend gate now reads a private
  `lastStreamedBodyRef` (the agent's own last frame), snapshotted at stream start,
  not `lastSyncedBodyRef` which `onUpdate` clobbers on peer edits — so a collaborator
  typing can't make the growing snapshot stop prefixing the shown body and freeze it.
- Medium (multi-replica over-persist): pre-existing, documented "safe over-persist"
  (a peer task tails the frame as REDIS_ORIGIN and marks edited) — refreshed the
  stale comment to describe the SYNC_NO_PERSIST source; copilot's edit_content write
  remains the authoritative durable persist.

* fix(files): fail-close base-less previews + operation-based stream hold

Cursor/Greptile round 3 (High + Medium) — remove the fragile string-prefix
"extend gate", which was the root of both findings:

- Server: `buildFilePreviewText` now fails closed for an `append` whose base
  content hasn't loaded (returns undefined, like patch/update), so a base-less
  fragment never reaches the client. This eliminates the base-less wipe at
  settle (Greptile P1) at the source; an empty file (existingContent === '')
  still previews normally.
- Client: the collab streaming tick no longer string-prefixes the raw preview
  against the editor's canonical markdown (the '*' vs '-' / emphasis mismatch
  that froze every append frame — Cursor). The mid-stream hold is now purely
  operation-based: `update` waits for settle; append/patch/create apply each
  frame via the (peer-safe) shadow reconcile. lastStreamedBodyRef is now a plain
  dedup guard, not a prefix baseline.

Keeps the shadow, durable write, and SYNC_NO_PERSIST unchanged.

* fix(files): elect a single agent-stream writer across tabs

Cursor round 4 (High): with the stream applied client-side, two tabs/windows on
the same chat could each derive streamingContent (the reconnect/resume path
re-consumes preview events) and each independently insert the stream under a
different Yjs clientID, duplicating content until the durable reconcile.

Fix — single-writer election via the file-doc awareness (new agent-stream-leader):
- a client applying an agent stream announces `agentApplying` on its own awareness
- only the leader (min clientID among announcers) applies mid-stream AND at settle;
  a non-leader renders the leader's ops via Yjs and does not apply (a non-leader
  applying the final body would re-insert the whole doc as a duplicate)
- re-checked each frame, so it converges to one writer the moment awareness
  propagates; the sub-frame startup race is reconciled by the durable write
- single-client (the common case) is unaffected: it is the only announcer, so it
  always leads

* fix(files): gate the settle apply locally, not on a settle-time re-election

Cursor round 5 (High): the settle recomputed leadership from live awareness and
the leader cleared its announcement immediately, so a straggler peer that settled
afterward became the sole announcer, self-elected, and applied finalBody through
its base-seeded shadow — re-inserting the whole doc as a duplicate.

Fix: gate the settle apply on a LOCAL didApplyStreamRef (set only when this client
actually applied a mid-stream frame — i.e. it was the mid-stream leader whose
shadow is up to date), not on a settle-time re-election. A client that never
applied (non-leader, a held `update`, or a pre-seed stream) skips the final apply
and converges via Yjs + the durable write. The mid-stream leader election
(isAgentStreamLeader) is unchanged, so exactly one client's didApplyStreamRef is
ever true.

* fix(files): open the agent-stream shadow lazily on lead (no stale handoff)

Greptile round 6 (P1): the leader race — (a) a mid-stream leadership handoff
could apply from a stale pre-stream shadow, and (b) two tabs starting the same
stream before awareness converges could both lead briefly.

- (a) fixed: the shadow is now opened LAZILY in the tick, only when this client
  actually leads, seeded from the CURRENT doc — so a handoff successor diffs
  against the prior leader's ops (never a stale base) and a non-leader builds no
  shadow at all. Announce candidacy via a dedicated ref (decoupled from the
  shadow); settle still gates the final apply on didApplyStreamRef (leader-only).
- (b) the pure startup race is inherent to eventually-consistent election. It is
  now the only residual: bounded to two tabs starting the SAME stream within the
  awareness-propagation window, transient (converges in a frame or two), and
  never persisted (SYNC_NO_PERSIST + the durable edit_content reconcile). Resumes
  are sequential, so the common multi-tab case elects cleanly. Documented inline;
  a server-granted lease would close it fully but at a round-trip cost on the
  common single-tab path, which isn't worth it.

* fix(files): idempotent settle apply (update lands client-side; no straggler dup)

Cursor round 6 (Medium): a lone client's `update` never applied client-side —
held mid-stream, then skipped by the didApplyStreamRef settle gate — so the
rewrite depended entirely on the durable merge (stale if delayed/failed).

Root cause was over-correcting round 5. Now that the shadow is opened lazily in
the tick (current-seeded), the round-5 base-shadow duplication is already gone,
so didApplyStreamRef is unnecessary. Replaced it: settle applies the final body
via `agentStreamSessionRef.current ?? beginAgentStream(editor)` — the leader
reuses its up-to-date shadow (last throttled frame), while a client that never
applied (non-leader, held `update`, pre-seed) opens a FRESH current-seeded shadow.
Reconciling current->final is idempotent: a straggler that settles after another
wrote the final reconciles to a noop. So a lone `update` applies at settle (no
wait on the merge), and there's still no settle-time election or base-shadow dup.

* fix(files): broadcast agent frames to the whole room (same-socket siblings)

Cursor round 7 (Medium): SYNC_NO_PERSIST frames applied under an origin carrying
the sender socket id, and excludeSocketId dropped that whole socket from the
relay fan-out. A second FileDocProvider on the same socket (chat preview + Files
editor) then missed all mid-stream ops and stayed stale until the durable
reconcile — a regression from the old no-origin server merge, which reached both.

Fix: the agent origin is now a plain AGENT_SYNC_ORIGIN symbol, and agent frames
broadcast to the WHOLE room (no socket excluded), matching the old behavior — so a
same-socket sibling provider stays live; the emitting provider no-ops on its own
echo (the ops are already applied locally). originSocketId still returns null for
the symbol, so it keeps skipping edited/schedulePersist. Removed excludeSocketId
and the socket-carrying origin object. Updated the relay test to assert the
whole-room broadcast (verified it fails if the sender is excluded).

* fix(files): tag agent stream frames no-persist across replicas

A peer task tailing an agent-streamed preview frame previously applied it
as REDIS_ORIGIN, marking the seeded room edited and making a transient
startup-race duplicate eligible for that task's last-disconnect flush. Mark
agent frames with a stream field so peers apply them as REDIS_AGENT_ORIGIN,
excluded from the edited/persist gate. The copilot's durable edit_content
write stays the sole authority over file bytes.

* fix(files): reseed agent shadow on lead regain + agent-only compaction

Two multi-writer edge cases surfaced in review:

- rich-markdown-editor: a client that led, lost leadership, then regained it
  reused its stale shadow (which never saw the interim leader's ops), re-emitting
  ops for content already present. Tear the shadow down when a client observes it
  is not the leader, so a regain rebuilds fresh from the current doc.

- file-doc-store: compaction always stamped its snapshot REDIS_SNAPSHOT_ORIGIN
  (marks peers edited). A long agent-only stream crossing the threshold could
  fold preview content into a persist-eligible snapshot. Track whether a room
  integrated any real edit and stamp an agent-only snapshot REDIS_AGENT_ORIGIN
  so it stays no-persist.

Both covered by falsification-verified tests.

* fix(files): close realEdited data-loss race + elect a settle writer

Independent audit surfaced two real gaps:

- file-doc-store: realEdited was latched AFTER appendUpdate's awaits, but the
  edit already sits in room.doc synchronously. A concurrent agent-frame
  compaction could read realEdited=false, snapshot that real content, and stamp
  it a no-persist agent frame — a lost edit. Latch it synchronously (same tick
  as the doc mutation) before any await. Deterministic falsifiable test added.

- rich-markdown-editor: at settle every tab applied the final body, and a
  non-leader's local microtask runs before the leader's final propagates, so
  both insert the tail (Yjs keeps both) -> duplicated tail. Elect a single
  settle writer (reliable — awareness is long converged by settle), reading
  leadership before clearing the announcement. Corrects the overclaiming
  idempotency comment and the handoff pick-up comment.

Adds a y-tiptap internals upgrade-guardrail test.

* fix(files): own presence per client id, not one-per-socket

The shared workspace socket hosts one collaborative provider per mounted view,
so the chat file preview and the standalone Files editor for the same file each
bind their own Yjs client id over ONE socket. The relay owned a single client id
per socket, so the later JOIN overwrote the earlier and dropped its awareness —
which silently broke the single-writer agent-stream election (a peer stopped
seeing the streaming provider's announcement and could self-elect, duplicating
streamed text for the whole stream).

Track ownership per (socket, client id): a socket owns a set of client ids; the
awareness gate accepts a frame only if every id it carries is owned; cleanup
drops all of a socket's ids; the roster stays one-entry-per-session. Reclaim and
the same-user reconnect path evict just the reclaimed id, dropping the old socket
only if it empties. Falsification-verified test added.

* fix(files): make streamed file-preview accumulation replay-safe

Guard deriveFilePreviewSession against re-delivered/replayed content events:
apply a delta/snapshot only when previewVersion strictly advances, so a client
re-render or stream replay can't double-append the tail (the duplicated-content
bug) or regress on an older snapshot.

* fix(files): fix new-file collab streaming latch and agent-edit duplication

- Latch collab readiness so a new file's post-seed `synced` flap can no longer
  re-gate agent streaming (the stream previously showed only the seed and the
  rest appeared only on reload)
- Relay defers the durable edit_content merge to an actively-streaming client:
  the client shadow stream and the server merge were both writing the same
  content into the live doc, duplicating it when the server ran ahead
- Render the collaborator caret bar out of flow so a peer caret never nudges
  the surrounding text by ~1px
- Remove dead code: unused FileDocMessageType alias, unnecessary
  LiveFileDocMergeOrder export

Covered by tests: readiness latch (flap/offline/latch cases), relay merge
deferral (single- and multi-replica), plus verified-failing guards.

* test(copilot): fix loadWorkspaceFileTextForPreview mock to return { text } not a bare string

The adapter reads previewBase.text to seed an append/patch base; the mock returned
a bare '' so previewBase.text was undefined, making a base-less append fail closed
(no file_preview_content). My PR's fail-close change exposed the wrong-shaped mock.

---------

Co-authored-by: mzxchandra <129460234+mzxchandra@users.noreply.github.com>
2026-07-31 18:48:12 -07:00
Waleed e98715da01 fix(realtime): noindex the socket server's 404 responses (#6129)
* fix(realtime): noindex the socket server's 404 responses

The sockets.* hostnames are served by this server and return a plain JSON
404, which Google Search Console reports as crawl errors. Mark unmatched
routes noindex so crawlers drop the hostnames instead of retrying them.

* fix(realtime): noindex every response, not just the 404

Setting the header only on the 404 fallback covered the one response that
crawlers already drop on status code alone, while /health — the sole route
returning 200 with a body, and so the only indexable surface on the socket
hostnames — stayed uncovered, with a test pinning it that way.

Set it once on the handler instead. Node merges setHeader values into
writeHead and no branch sets X-Robots-Tag, so it reaches every response.
2026-07-31 12:55:48 -07:00
Waleed 02311cafb9 perf(db): stop scanning every workspace file on workflow archive, cover the billing period aggregate (#6015)
* perf(files): stop loading every workspace file to archive a workflow's backing files

`cleanupWorkflowAliasBacking` runs on every workflow delete. It loaded every
`context='workspace'` row for the workspace — active and soft-deleted, all
columns — then discarded >99% of them in JS to find the handful of
`.changelogs/<workflowId>.md` and `.plans/<workflowId>/**` rows it owns.

In production this was the single worst query at 37.83% of total database
runtime: p50 6s, p95 9s, max 13.1s, 641 calls/day, 34.1M rows read. It is
index-served, so the cost is not a missing index — it is 59,061 rows and
~630MB of cold buffers read to archive a few files.

A file's `folderPath` is derived solely from its `folderId`, so matching on
folder membership is equivalent to the path comparison it replaces. The load
and JS filter become one targeted UPDATE keyed on a handful of folder ids.

Folders that are soft-deleted are still included when resolving which files a
workflow owns: path resolution ignores `deletedAt`, so a live file parented to
an archived folder previously matched and must continue to.

Also:
- Add `getWorkspaceShares`, replacing an id-list `IN` clause that grew with the
  file count (59,061 elements on the worst workspace) with one indexed lookup
  on `workspace_id`. Callers read the map by id, so a superset is equivalent.
- Drop `all` from the client-reachable file scope enum. It drops the
  `deleted_at` predicate and so cannot use the partial index serving the other
  two. No client requests it; server callers reach that scope directly.
- Set `fetch_types: false` on the app and realtime pools. postgres.js otherwise
  runs a blocking `pg_catalog.pg_type` roundtrip before each new connection's
  first query — 95,722 of them per day. It builds array parsers only; Drizzle
  already parses this schema's two `text[]` columns itself.

* perf(nav): add route loading boundaries and cover the billing period aggregate

Dynamic routes prefetch only down to the nearest loading boundary, and the
server stops prefetching at the first one. With no loading.tsx anywhere in the
workspace tree and a dynamic layout, Link prefetch was yielding almost nothing
and every navigation waited on a full server round trip with no feedback.

Add loading.tsx only where the fallback is provably what renders today, so
perceived speed improves without changing what users see:

- home, chat/[chatId]: reuse HomeFallback, already each page's own Suspense
  fallback.
- integrations, skills: reuse the tab-header chrome, byte-identical to each
  page's own Suspense fallback.

Deliberately not added to workspace root, settings, or w, whose pages are
redirect-only or already self-fallbacking — a boundary there would paint a
skeleton that does not match the destination.

Billing:

- Add usage_log_billing_period_cost_idx, trailing the remaining predicate
  columns and `cost` so the period aggregates resolve index-only. Confirmed
  against production: the aggregate currently runs as an Index Scan touching
  478,559 buffers because `cost` is absent from the existing index. That index
  is superseded but left in place; dropping it is a separate migration so a
  planner regression costs nothing to revert.
- Collapse the two aggregates behind /api/billing into one scan using
  SUM(...) FILTER. Besides halving the work on a nav-path query, it removes a
  latent inconsistency: as separate statements the two sums could observe
  different snapshots, making the copilot subset exceed the total.

Caching was considered and rejected for the usage aggregate. Its callers
include usage enforcement, threshold billing, and overage calculation, where a
stale-low read permits overspend and a stale-high read double-charges.

* test(files): cover cleanupWorkflowAliasBacking and correct a stale share mock

cleanupWorkflowAliasBacking had no test coverage, and the rewrite that replaced
its load-everything-then-filter body rests on a subtle equivalence: a file's
folderPath is derived solely from its folderId, and path resolution ignores
deletedAt. The second half is the easy part to get wrong — a live file parented
to an archived folder still resolves to a backing path and must still be
archived, so folders are filtered by deletedAt only when choosing which folders
to archive, never when deciding which files the workflow owns.

These tests pin that distinction: they fail if the archived-folder case is
dropped from file ownership, and separately assert that archived folders stay
out of the folder update and that unrelated workflows are never touched.

Also point the workspace files route test at getWorkspaceShares. Its mock still
named getSharesForResources, which the route no longer imports. The suite passed
regardless because all three cases exercise the upload path, so the GET listing
has no coverage — but the stale name would have handed the first GET test an
undefined function.

* fix(perf): drop the loading boundaries and correct the usage_log index order

Adversarial review found real regressions in both.

Route loading boundaries — all four removed:

The premise in their TSDoc was wrong. Each page's existing `<Suspense>` exists so
nuqs can prerender; `useSearchParams` only suspends during SSR, so on a client
navigation those fallbacks never painted. Hoisting them to loading.tsx did not
"show the same frame" — it made a previously invisible blank frame visible.

For chat, Next keys the Suspense boundary by cache key, so a chatId change mounts
a fresh suspended boundary and always commits its fallback. Switching chats would
have gone from "previous chat stays on screen" to a blank surface for the whole
RSC round trip, on the highest-frequency navigation in the product. Partial
prefetch only warms the fallback, so no latency was saved to offset it.

The integrations and skills fallbacks were correct for their own pages but also
became the fallback for four detail routes that render different chrome, inserting
a wrong intermediate frame. Fixing that needs per-child boundaries and new skeleton
UI that cannot be verified without a browser, for pages whose only server work is
`await params`. Not worth it.

Every remaining loading.tsx in this app paints real chrome. A blank one was against
the grain, and the measurable win was zero.

usage_log index:

Column order was wrong. The daily-refresh rollup filters entity type, id and
period_start but NOT period_end, so putting period_end fourth ended the usable
prefix at column three and left user_id and created_at as in-index filters rather
than scan boundaries — turning a ~2.2k-entry bitmap scan into a ~266k-entry scan.
Verified in production that period_start functionally determines period_end
(13,612 groups, zero with more than one end), so the slot bought no selectivity.
user_id and created_at now follow the shared prefix; period_end rides as payload.

Also corrected the claim that this supersedes usage_log_billing_entity_period_idx.
That index deduplicates to 17 MB across 1.42M entries because it has no
high-cardinality key column, which is what keeps prefix-only bitmap scans cheap.
It is retained deliberately, not pending a drop.

Harden cleanupWorkflowAliasBacking: gate the UPDATE on the ownership filter list
itself. `and()` and `or()` both drop undefined, so a clause that resolved to
nothing would have left a WHERE of workspace + context + not-deleted and archived
every file in the workspace. Tests now assert no UPDATE is issued when the
workflow owns nothing, plus the filename and context predicates.

Note fetch_types also disables array serializers, not just parsers; documented.

* docs(db): correct the fetch_types constraint note after differential testing

Ran both settings against a real Postgres with the schema's actual text[] columns.
Drizzle-typed selects and .returning() are byte-identical either way; only a raw
db.execute projecting an array column differs, yielding the wire form.

The previous note also claimed a raw JS-array bind fails because the serializers
come from the same catalog fetch. It does fail — but under both settings, because
Drizzle expands an array into a row constructor before postgres.js ever sees it.
That is unrelated to this flag, so the claim is removed rather than left implying
a constraint this change introduces.

* chore(db): generate the drizzle snapshot for the usage_log index migration

0271 was hand-written, so drizzle-kit's state never learned about the new index —
the next `generate` would have re-emitted it as a fresh migration against an
already-migrated database.

Ran `drizzle-kit generate` to produce the snapshot, then restored the
CONCURRENTLY form: drizzle emits a plain CREATE INDEX, which takes an ACCESS
EXCLUSIVE lock and would block writes on a 4M-row table for the duration of the
build. The generated column order matched the hand-written SQL exactly, which
also confirms schema.ts and the migration agree.

`generate` is now a no-op, and check:migrations still passes.
2026-07-28 12:39:10 -07:00
Vikhyath MondretiandClaude Fable 5 f43b52c569 fix(realtime): evict revoked collaborators from live workflow rooms (#5917)
* fix(realtime): evict revoked collaborators from live workflow rooms via periodic read-access re-validation

* fix

* fix(realtime): close join/eviction race and make sweep cleanups independent

Re-authorize immediately before socket.join so an in-flight join cannot
reverse a sweep eviction, and run the sweep's best-effort cleanups
independently so a room-state failure cannot skip the presence broadcast.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): single-flight role resolution and retry failed eviction cleanup

Coalesce concurrent role resolutions per (user, workflow) so a slow stale
read can never overwrite a recorded revocation, and defer failed eviction
room-state cleanups into a per-sweep retry queue so collaborators are not
left with a stale presence entry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): scope room removal to the target workflow and detect swallowed cleanup failures

Honor the workflowIdHint as the target room in both room managers so
removing a stale room cannot clobber the mapping of a room the socket has
since moved to, and confirm eviction cleanup via the returned workflowId
plus an unswallowed mapping read so Redis failures actually defer into the
retry queue.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): treat unconfirmed removals as failures and refresh role cache on fresh verify

Treat any null removal result as a failed cleanup (the sweep always passes
the target room, so null only means failure — including with expired
mapping keys), move the same-room rejoin guard to a synchronous check
immediately before the removal, and record verifyWorkflowAccess's fresh
decision into the role cache so a re-granted user is not blocked by a
stale cached revocation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): prefer a mid-flight recorded role decision over the in-flight query result

If a fresh authoritative read (join-time verify) records a decision while
a single-flighted resolution's query is in flight, keep the recorded
decision instead of overwriting it with the potentially stale result.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): isolate the security scan from the Redis cleanup lane

Run the revocation scan (local sockets + DB only) and the best-effort
room-state cleanup as independently-guarded lanes so a hanging Redis
command can stall only presence cleanup, never revocation enforcement.
Evictions now enqueue cleanup instead of awaiting it, and the scan no
longer reads presence for a fallback role.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): bound authorization waits in the revocation scan

Race each socket's authorization check against a per-socket timeout and
cap the whole pass with a budget below the sweep interval, so a hanging
DB query skips that socket for the pass (never evicting on uncertainty)
instead of wedging the scan lane and starving subsequent ticks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): round-robin the revocation scan so hung checks cannot starve later sockets

Resume each scan pass after the last target the previous pass processed,
so a fixed prefix of hanging authorization checks can never repeatedly
consume the pass budget and leave sockets behind it unexamined.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 16:45:40 -07:00
d24bc7eccb feat(agent-stream): thinking and tool streaming (#5671)
* feat(agent-stream): add agent-events thinking/tool streaming for chat and canvas

Ship the agent-events-v1 protocol with provider tool loops, dual-gated chat thinking, DeepSeek/Groq/OpenAI reasoning wiring, and ChatGPT-like thinking chrome.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): clear stuck streaming UI and format db snapshot

Biome was failing CI on migrations/meta/0261_snapshot.json. Also settle
assistant streaming/tool flags when SSE ends without a terminal frame,
without clobbering Stop's finalized content.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): satisfy biome format and import order

Auto-format the sim package for CI lint:check, and repair the Anthropic
streaming tool-loop payload after an unsafe delete-to-undefined rewrite.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): keep drained answer on abort and update migration journal test

Treat AbortError from reader.cancel as a cancelled pump result so soft-complete
retains answerText. Point the workspace storage migration journal assertion at
0261_chat_include_thinking.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): keep Stop notice when server emits cancel error

Ignore terminal SSE error frames after the user aborts so
"Client cancelled request" cannot overwrite "Response stopped by user".

Co-authored-by: Cursor <cursoragent@cursor.com>

* improvement(chat): ChatGPT-style thinking shimmer and stick-to-bottom scroll

Add left-to-right shimmer on live thinking label/body, keep scroll working by
shimmering an inner node, and follow the answer only while near the bottom.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): stop pump on client disconnect; soft-complete agents only

Abort the agent stream pump when the projected HTTP body is cancelled so
provider work does not continue after disconnect. Limit AbortError soft-success
to Agent blocks so Function/HTTP cancels still fail in logs.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): persist includeThinking across pause snapshots

Paused chat runs with Include thinking enabled were dropping the flag when
serializing the pause snapshot, so resume always rebuilt streams without
thinking/tool SSE frames.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): keep drained answer text when stream times out

Persist pump answerText onto the streaming execution before throwing on
timeout, and carry that partial content into the failed block output so
logs match what the client already saw.

Co-authored-by: Cursor <cursoragent@cursor.com>

* improvement(chat): auto-collapse tools chrome when tool streaming ends

Match thinking UX: open while tools run, collapse when finished, and keep
the panel open only if the user manually reopens it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): settle canvas stream chrome on failure paths

Clear agentStreamActive and settle running tool chips when blocks error,
timeouts cancel runs, or execution ends without stream:done so the output
panel does not stay on live Thinking/Using tools chrome.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(lint): organize imports in terminal console store

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): mark open tools cancelled on HITL pause

Pause can interrupt a tool loop without tool end events; settling those
chips as success incorrectly showed unfinished tools as complete.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(db): drop branch-local 0261 migration ahead of staging merge

* chore(db): regenerate include_thinking migration as 0266 post staging merge

* fix(providers): resolve type errors in streaming tool loop call sites

* fix(agent-stream): gate agent events opt-in and correct provider loop behavior

- streamToolCalls and provider thinking requests now require run-level
  agentEvents opt-in (canvas on, chat dual-gated, API off) so existing
  runs keep pre-agent-events behavior exactly
- OpenAI reasoning summaries opt-in + strip-and-retry on unverified-org 400
- streaming loops run tool postProcess again (firecrawl/exa async results)
- bedrock live loop falls back to silent path for responseFormat
- deepseek: reasoning_content pass-back unconditional, 'none' sends disabled
- groq: x_groq.usage fallback, reasoning params gated, qwen none disables
- gemini: functionCall parts echoed verbatim, local ids only for events
- truncated turns (max_tokens/length) no longer execute partial tool calls
- MAX_TOOL_ITERATIONS exit flushes last turn text as final answer
- iterations reports actual model calls; shared loop plumbing extracted

* refactor(agent-stream): consolidate protocol, dedupe client/server plumbing, hygiene

- canonical ChatStreamFrame union + type guards consumed by server emitters
  and the chat client; stream_error restored to legacy log-only handling
- strip thinking/tool args from providerTiming on public final envelopes
- shared tool-chip lifecycle module for chat, canvas, and console store
- shared sink-to-execution-events forwarder replaces the copy-pasted
  adapter in the execute route and HITL manager; LIVE_ONLY event set shared
- stream:thinking payload field renamed data->text; canvas thinking batched
- abort reasons carried as AbortError DOMExceptions so raw fetch consumers
  classify correctly; thinking cap renamed to chars and scope-documented
- kimi wired for agent events like the other compat providers
- deleted dead exports/step-N comments; fixtures match real wire shapes;
  loop tests use explicit mocks instead of importOriginal

* test(agent-stream): cover the dual-gated execution path and typed abort reasons

- chat route tests assert agentEvents reaches executeWorkflow only when
  policy and protocol header agree
- execution-limits tests assert AbortError-typed reasons
- executor metadata type carries agentEvents

* fix(deploy-modal): align include-thinking spacing with the modal's 6.5px rhythm

* docs(agent-stream): autogenerate per-model thinking/tool stream support on the Agent block page

- capabilities.thinking.streamed ('full' | 'summary' | 'none') on models.ts,
  explicit for the Anthropic family where visibility varies per generation;
  getThinkingStreamVisibility exposes the derivation for docs and UI alike
- scripts/sync-agent-stream-docs.ts regenerates the support tables between
  markers in workflows/blocks/agent.mdx from the model registry and
  STREAMING_TOOL_CALL_PROVIDERS; --check fails on drift or missing metadata
- wired agent-stream-docs:check into CI next to the other sync gates

* feat(anthropic): request summarized thinking display for omitted-default Claude models

The newest Claude generations (Fable 5, Sonnet 5, Opus 4.8/4.7) default
thinking.display to omitted — empty thinking blocks, no deltas. On
agent-events runs Sim now opts back in with display: 'summarized', driven
by the registry's streamed metadata; legacy runs keep the exact
pre-agent-events request shape. Registry, generated docs, and the family
capability table updated accordingly.

* docs(skills): cover thinking.streamed and agent-stream docs sync in model skills

* chore(deps): upgrade @anthropic-ai/sdk to 0.114.0 and adopt official types

- adaptive thinking, display, and output_config are now SDK-typed; the only
  remaining custom payload field is output_format (beta-header structured
  outputs, which the SDK models as output_config.format instead)
- anthropic stream events narrow on the SDK's discriminated unions instead
  of anonymous casts; compat deltas type content/tool_calls from the OpenAI
  SDK with vendor reasoning fields as an explicit optional extension
- @sim/auth exposes an explicit VerifyAuth contract so its declarations no
  longer reference better-auth's nested zod instance (TS2883 under fresh
  install layouts); realtime consumer aligned
- docs app zod pinned to the repo's exact 4.3.6 so ai SDK types bind the
  same zod instance (docs type-check was latently broken)
- knowledge embedding tests made hermetic against local .env keys and
  hosted rotation fallback

* refactor(providers): replace legacy as-any stream casts with annotated typed casts

* refactor(providers): finish provider audit — remove dead byte-stream helper, annotate remaining legacy casts

Audit of all 26 providers for the agent-events feature confirmed every
streaming execution declares agent-events-v1 and every adapter emits
AgentStreamEvent objects. Cleanup from the audit: the unconsumed legacy
createOpenAICompatibleStream byte helper is deleted, and the remaining
streamResponse-as-any casts (xai, nvidia, kimi, meta, zai, sakana) are
annotated typed casts matching the groq/deepseek fix.

* feat(streaming): stream answer text live during tool loops via turn_end protocol

The live tool loops buffered all answer text per model turn (classification
of intermediate vs final is only known at turn end), so gated surfaces saw
thinking stream, then dead air with the thinking chrome stuck open, then the
whole answer at once.

Loops now emit text deltas live as `turn: 'pending'` plus a `turn_end`
event per turn. The pump buffers pending text and projects it to the byte
path (answerText/logs/memory/legacy clients) only on a final turn_end, so
all settled semantics are unchanged. Gated surfaces render the pending text
as it streams and reconcile with a reset when a turn resolves to tools:

- public chat: live `chunk` frames from the sink + dual-gated `chunk_reset`;
  byte-path frame emission is suppressed to avoid duplicates (kept for
  response-format transformed streams via clientStreamTransformed)
- canvas: forwarder emits live `stream:chunk` + `stream:chunk_reset`; the
  execute route and HITL resume readers stop re-emitting byte chunks; panel
  chat tracks per-block segments and replaces content on flush
- chat client: per-block text segments, chunk_reset handling, and thinking
  chrome now settles on tool start as well as first answer chunk

* fix(streaming): address validated review findings across provider gating and reset reconciliation

Three-reviewer pass over the branch, findings validated against staging:

- agent-handler forwards agentEvents to executeProviderRequest — the flag was
  computed but dropped in the field-by-field copy, so provider-side thinking
  requests (OpenAI summaries, Gemini includeThoughts, Anthropic summarized
  display) never activated on opted-in runs
- openai: restore summary:'auto' alongside explicit reasoning effort — staging
  always paired them; gating summary purely on agentEvents changed legacy
  payloads
- gemini: Gemini 2 + tools + responseFormat falls back to the silent path;
  the live loop never applied the deferred responseSchema for AUTO tools
- openai-compat loop: malformed tool-argument JSON fails the call instead of
  executing with defaulted {} args (staging parsed inside the execution try)
- openai-compat parser: a vendor id arriving after a synthesized start no
  longer renames the call (start/end ids stayed consistent)
- stream-pump: abort closes the byte projection so a drain blocked on
  backpressure cannot deadlock teardown
- chunk_reset removes the block from the client text order (deployed chat +
  panel chat) so a reset block re-registers at arrival position — fixes
  separator/order corruption when parallel blocks stream around a reset
- resume route echoes the negotiated X-Sim-Stream-Protocol response header
  (parity with the chat route); docs: [DONE] wire shape + final-vs-error
  terminal semantics corrected

* chore(deps): exempt pinned @anthropic-ai/sdk 0.114.0 from the release-age gate

CI's bun install --frozen-lockfile blocks 0.114.0 (published 2026-07-23,
younger than the 7-day supply-chain gate). The pin is exact and was vetted
for the agent-events streaming work; following the existing bunfig pattern,
the exclusion ages out on 2026-07-30 and should be dropped then.

* chore(providers): fix double-cast-allowed annotation placement for the strict boundary audit

The audit only recognizes the annotation on the line directly above the cast;
two annotations had drifted behind intervening code lines (groq stream params,
deepseek loop messages) and the OpenAI reasoning-summary widening cast was
never annotated. No behavior change.

* fix(chat): settle straggler tool chips as error when final reports failure

A failed run can still terminate with a `final` frame carrying success: false;
running chips previously settled green regardless of the outcome.

* fix(canvas): wire agent stream chrome into run-from-block

Run-from-block executions emit the same live stream:thinking/stream:tool
events as full runs but registered none of the handlers, so the terminal
never showed thinking or tool chips on that path. The per-run chrome
(batched thinking writes + tool chip lifecycle + settlement on stream done,
block error, and every terminal execution state) is extracted into a shared
createAgentStreamChrome factory consumed by both paths.

---------

Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai>
2026-07-23 19:39:03 -07:00
Waleed fac8ec0c51 fix(workflow-edges): enforce edge/block validation server-side, not just client-side (#5571)
* fix(workflow-edges): enforce edge/block validation server-side, not just client-side

Dragging a connection that creates a cycle correctly refused to render
client-side, but the cyclic edge was still queued for realtime persistence
and written to the DB, so it reappeared after refresh. Root cause: cycle
detection (and several other edge/block rules) only lived in the client
Zustand store and was never enforced by the realtime persistence layer,
which is the actual source of truth on reload.

- Move wouldCreateCycle, edge scope-boundary, annotation-only-block,
  duplicate-edge, and block-name-conflict checks into @sim/workflow-types
  so the client store, the collaborative queueing layer, and
  apps/realtime's DB write path all share one implementation
- Wire these into apps/realtime/database/operations.ts's edge-add and
  block-rename handlers as the authoritative gate
- Client-side behavior is unchanged (same call sites, same error messages,
  same rule ordering) — verified via existing + new test coverage

* fix(workflow-edges): address review findings on realtime edge validation

- Select triggerMode when fetching blocks for edge-add validation —
  isKnownWorkflowTriggerBlock checked block.triggerMode but the column
  was never fetched from the DB, so trigger-mode blocks could still
  receive an incoming edge (Cursor Bugbot)
- Make filterUniqueWorkflowEdges incremental, so two duplicate edges
  within the same BATCH_ADD_EDGES payload are also deduped instead of
  both surviving (Greptile)

* fix(workflow-edges): normalize empty-string handles in duplicate-edge check

filterUniqueWorkflowEdges compared handles with ??, so a `sourceHandle: ''`
edge wasn't recognized as a duplicate of an existing null-handle edge —
even though both get persisted as the same null value at insert time
(edge.sourceHandle || null). Falsy-coalesce in the comparison so '' and
null/undefined are treated as the same "no handle" state everywhere.
(Greptile)

* improvement(workflow-edges): dedup realtime edge-add validation, reuse block-name-conflict helper

/simplify pass on the already-merged-quality PR before sign-off:

- Extract filterEdgesForPersist in apps/realtime/src/database/operations.ts:
  the single-edge ADD and batch BATCH_ADD_EDGES handlers hand-inlined the
  same six-step validation pipeline (missing block, protected target,
  annotation-only, trigger-target, scope boundary, duplicate, cycle) and
  each independently re-fetched blocksById/existingEdgesForCycleCheck. Two
  copies of one rule in the same file was exactly the drift risk this PR
  otherwise closes across client/server. One shared helper now backs both.
  Net -63 lines despite the new shared function.
- Fix a real bug this surfaced: droppedCounts was keyed by the free-text,
  block-id-bearing scope-boundary message, so it could never aggregate
  across edges/runs. Now keyed by a stable 'scope boundary' reason.
- use-collaborative-workflow.ts's collaborativeUpdateBlockName still
  hand-rolled the empty/reserved/duplicate block-name-conflict checks this
  PR centralized as getWorkflowBlockNameConflict (already adopted by
  store.ts). Switched it to the shared helper, which also fixes a latent
  check-order mismatch between the two (this pre-check ran reserved before
  duplicate; the store's real gate — after this PR's own store.ts change —
  runs duplicate before reserved, so they could disagree on which toast a
  name that was both reserved and duplicate would surface).

* docs(workflow-edges): document the per-workflow write-serialization invariant

No behavior change. Address Greptile P1 on filterEdgesForPersist ('concurrent
duplicate writes can persist without a per-workflow write guard') by
documenting, at the actual mechanism, why the concern doesn't apply here:
persistWorkflowOperation's leading 'UPDATE workflow SET updatedAt ... WHERE
id = workflowId' already takes a row lock that serializes every operation
(including edge adds) for a given workflowId — a second concurrent call
blocks on that UPDATE until the first transaction commits or rolls back, so
the validate-then-insert sequence in filterEdgesForPersist can never
interleave across two writers on the same workflow.

Verified empirically, not just by reading: ran two concurrent transactions
against a throwaway local Postgres against the exact statement shape used
here (UPDATE the parent row, sleep to simulate the read/validate window,
insert, commit). The second transaction's UPDATE blocked for the full
duration of the first's transaction and only proceeded once the first
committed — confirming the row lock, not any application-level guard,
already provides the serialization Greptile flagged as missing.

Added a comment at the lock site (not a second, redundant advisory lock)
so a future change can't silently break this invariant by making the
UPDATE conditional/skippable as a perceived no-op optimization.

* fix(workflow-edges): validate edges in BATCH_ADD_BLOCKS before persisting

Real gap Cursor's PR summary flagged ('BATCH_ADD_BLOCKS edge inserts... not
fully covered by the new server pipeline'), confirmed by reading the code:
this handler bulk-inserted the edges from a block-paste/duplicate/import
payload directly into workflowEdges with zero validation — no missing-block,
protected-target, annotation-only, trigger-target, scope-boundary,
duplicate, or cycle check. A client sending edges through this operation
instead of BATCH_ADD_EDGES could bypass every rule this PR otherwise
enforces server-side, exactly the class of gap the PR exists to close.

Routes it through the same filterEdgesForPersist used by the other two
edge-add handlers. Runs after the block insert in this same handler, so the
shared helper's blocksById lookup also sees the blocks this same batch just
inserted (a transaction observes its own prior writes).
2026-07-10 15:25:25 -07:00
Waleed b117758eaa fix(agent): scope nested tool canonical-mode overrides by instance, not type (#5534)
Two tool entries of the same type inside an Agent block's tool-input array
(e.g. two Table tools) shared a single canonical-mode override keyed by
${toolType}:${canonicalId}, so switching basic/advanced mode on one field
silently switched it on every other instance of the same tool type -
including at execution time, where the wrong basic/advanced value could be
resolved for the second tool.

Rescope the override key to the tool's position in its tool-input array
(${toolIndex}:${canonicalId}) instead of its type, and thread that index
through every consumer: the editor (read + write), execution
(agent-handler/providers), search-index, and fork/promote remapping.
2026-07-09 11:21:05 -07:00
Waleed 5db62b8bdc feat(observability): extend audit log, PostHog, and storage metering coverage (#5269)
* feat(observability): extend audit log, PostHog, and storage metering coverage

Instruments previously-uncaptured resources and actions across the three
observability layers (audit log, PostHog product analytics, usage metering):

- Exfiltration audit: file downloads (workspace/public-share/v1), table,
  workflow, and workspace exports.
- Credential access audited at the token-issuance boundary (success-only).
- Full revenue trail: invoice paid/failed, overage billed, disputes, credit
  fulfillment, subscription lifecycle, plan/seat changes.
- Session/login lifecycle audit via Better Auth hooks (login, blocked sign-in,
  logout, session revoke, account delete).
- v1/admin programmatic surface and copilot tool handlers instrumented.
- Dead audit/PostHog constants wired (lock/unlock, table, custom tool, etc.).
- Storage metering extended to KB documents and copilot files.
- PostHog person identify + workspace/organization group hygiene.

Audit package hardened to null the actor FK for system actors (admin-api) and
adds an awaitable recordAuditNow for pre-delete hooks. All instrumentation is
fire-and-forget and never blocks or breaks the primary operation.

* fix(observability): audit file/credential egress only on success

Address Cursor Bugbot review:
- GET OAuth token: emit CREDENTIAL_ACCESSED/credential_used after
  refreshTokenIfNeeded succeeds (matches POST), not before.
- File export: audit on each actual success exit via a shared helper —
  including the non-markdown serve redirect (previously unaudited) and
  after the zip is generated (previously before asset fetch).

* fix(audit): record null actor for anonymous public-share downloads

Address Greptile P1: setting actorId to the file owner made every anonymous
external download read as a self-download, undermining the exfiltration trail.
recordAudit now accepts a null actor; the public content/inline routes record
actorId=null with the owner in metadata.sharedByUserId and rely on ip/user-agent
for the forensic trail. The misleading owner-attributed file_downloaded PostHog
event is dropped on these anonymous paths.

* fix(observability): audit admin exports after zip; tidy comments

- Admin workflow/workspace ZIP exports now audit only after the archive is
  built (via a local helper), so a zip/build failure no longer logs an export.
- Remove redundant 'success-only placement' inline comments and tighten the
  remaining design-rationale notes (export helper now TSDoc).

* fix(billing): complete the chargeback + overage financial trail

Address Cursor review:
- handleDisputeClosed now records CHARGE_DISPUTE_CLOSED for every closed
  dispute (won/lost/warning_closed), unblocking only on favorable outcomes.
  dispute.status in the metadata distinguishes the outcome, so lost
  chargebacks are no longer missing from the trail.
- Threshold overage now emits OVERAGE_BILLED + overage_billed even when credits
  fully cover the overage (settledVia: 'credits' vs 'stripe'), so credit-settled
  overages are audited instead of silently returning null.

* fix(storage): decrement copilot quota before deleting metadata

Address Greptile P1: deleting the metadata row before the decrement meant a
decrement failure left the quota permanently inflated with no record to retry
from. Decrement first; only remove the metadata row once it succeeds.

* fix(observability): atomic copilot storage release; signup-blocked action

- Copilot file delete now releases storage via a single transaction
  (releaseDeletedFileStorage): the soft-delete is the idempotency claim and the
  decrement shares the transaction, so neither a partial failure (inflated
  counter) nor a retry (double-decrement) can desync the quota. Resolves the
  conflicting Greptile/Cursor ordering findings.
- Policy-blocked sign-ups now record USER_SIGNUP_BLOCKED instead of
  USER_SIGNIN_BLOCKED, so account-lifecycle events aren't mislabeled.

* fix(storage): meter copilot ingest centrally for path symmetry

Address Cursor review: only uploadCopilotFile incremented storage, but copilot
files also enter via the generic upload route and presigned uploads — all of
which persist metadata through insertFileMetadata. Move the increment into
insertFileMetadata (scoped to context='copilot', on genuine insert/restore) so
every ingest path is symmetric with the delete-time decrement, and drop the now
redundant per-path increment in uploadCopilotFile. Other contexts are metered by
their own managers and remain unaffected.

* fix(audit): skip WORKFLOW_EXPORTED when admin export is empty

Address Cursor review: when every requested workflow fails to load, the admin
export still returned 200 but recorded a successful export of zero workflows.
Guard auditExport so an empty result records nothing.

* fix(storage): settle copilot accounting before deleting the blob

Address Cursor review: the blob was removed before releaseDeletedFileStorage, so
a release failure left the counter inflated and the metadata active with the blob
gone. Now the atomic soft-delete + decrement runs first and the blob is deleted
only if it succeeds, so a failure leaves the file fully intact and retryable.

* fix(storage): decrement KB document storage atomically with deletion

Address Cursor review: hardDeleteDocuments deleted the rows then decremented
best-effort, so a decrement failure left billed storage inflated with no row to
reconcile. Resolve each owner's subscription up front, then decrement inside the
same transaction that deletes the embeddings/documents (decrementStorageUsageInTx,
also now shared by releaseDeletedFileStorage), so the counter and the content
commit or roll back together. Connector docs remain excluded.

* fix(audit): don't treat email verification as a login

Address Cursor review: /verify-email could emit USER_LOGIN when verification
mints or refreshes a session, mislabeling (or double-counting) a non-sign-in as
a login. Restrict isLoginPath to genuine sign-in entrypoints.

* fix(audit): null actor FK when the user lookup throws

Address Greptile P1: the catch branch left the original actorId, so a system
actor like 'admin-api' (or a since-deleted user) would FK-violate the insert and
lose the row when the existence lookup errored. Mirror the not-found branch —
null the FK with a readable label — so the audit row always persists.

* revert(audit): drop session/account-lifecycle auth instrumentation

Remove the Better Auth login/logout/session-revoke/account-delete/blocked-signin
audit hooks from auth.ts — they touch sensitive auth paths and are noisy. Also
removes the now-unused taxonomy (USER_LOGIN/_FAILED, USER_SIGNIN_BLOCKED,
USER_SIGNUP_BLOCKED, USER_LOGOUT, SESSION_REVOKED, ACCOUNT_DELETED,
ACCOUNT_EMAIL_CHANGED actions; SESSION/USER resource types) and the awaitable
recordAuditNow helper that only the pre-delete hook used. auth.ts is back to the
staging baseline.

* fix(storage): harden copilot+KB delete accounting against read errors and concurrency

Final-audit follow-ups:
- deleteCopilotFile: a failed metadata *read* (vs a genuine missing row) now
  blocks the blob delete too, so a transient read error can't leave an active
  row un-decremented with the blob gone.
- hardDeleteDocuments: drive the per-user decrement from the delete's
  returning() (the rows this tx actually removed), so two concurrent deletes of
  the same ids can't both decrement.

* chore(observability): drop two unused definitions

Final-audit cleanup: remove the orphaned knowledge_base_searched PostHog event
(never wired; KB search analytics already flow through the OpenTelemetry channel)
and the redundant AuditAction.SUBSCRIPTION_UPDATED (subscription/plan changes are
audited via ORG_PLAN_CONVERTED). Every remaining new action/event has a real
emit site.

* fix(knowledge): key hard-delete result off rows actually deleted

Address Cursor review: hardDeleteDocuments returned existingIds.length (the
requested count) and cleaned storage for the full pre-tx set, even though the
decrement was driven by the rows the transaction actually deleted. Under a
concurrent delete that claimed some ids first, that overstated the result and
re-touched storage for rows this call didn't delete. Drive the storage cleanup,
log, and return value from deletedDocs (the returning() rows) so all four are
consistent.

* fix(storage): gate copilot quota on all ingest paths; tidy inline comments

- Copilot uploads via the generic /api/files/upload route and presigned URLs now
  run checkStorageQuota before writing (matching uploadCopilotFile), so no copilot
  ingest path can grow usage past the plan limit. The central increment stays in
  insertFileMetadata.
- Trim/remove redundant inline comments across the diff; keep only concise notes
  on non-obvious decisions (left pre-existing comments untouched).

* fix(analytics): omit empty workspace_id on workspace-less file downloads

Address Greptile P1: the generic key-based /api/files/download route and the
no-workspace markdown-export path emitted file_downloaded with workspace_id:''
creating a phantom '' bucket in PostHog. Make workspace_id optional on the event
and omit it when there is no workspace (workspace-scoped routes still pass it).

* fix(analytics): omit empty workspace_id/workflow_id on copilot_chat_sent

Final-sweep nit: copilot_chat_sent sent '' for workspace_id/workflow_id in the
agent (workspace-less) branch, same phantom-bucket issue as file_downloaded.
Make both properties optional and omit them when absent.

* fix(analytics): clear stale org PostHog group on personal workspaces

Address Cursor review: switching from a team workspace to a personal one left
the previous organization group set, so later events kept rolling up under it.
When the active workspace has no organizationId, resetGroups() to drop the stale
org group, then re-apply the workspace group.

* fix: address review — export audit timing, copilot delete signal, group reset

- Table export (Cursor MED): audit fires before streaming begins, not after
  controller.close(), so a mid-stream failure still records the partial export.
- deleteCopilotFile (Greptile P1): throw when storage accounting can't be settled
  instead of silently returning, so callers can detect the file was not deleted.
- workspace-scope-sync (Cursor MED): only resetGroups() once workspace metadata is
  loaded (activeWorkspace present), so a team workspace doesn't transiently lose
  its org group while organizationId is still null during load.

* refactor(observability): final line-justification cleanup

Address final audit flags:
- Table import: move the 'columns added' audit OUT of the import transaction
  into a post-commit auditTableColumnsAdded() helper called by the three
  tx-owning callers, so a mid-import row-batch rollback no longer logs a false
  'added N columns' (matches the PR's success-only discipline).
- Omit empty-string analytics dimensions on workflow_lock_toggled (workspace_id)
  and organization_created (name), consistent with file_downloaded/copilot_chat_sent.
- Restore an unrelated capacity-check comment removed incidentally.
- Move copilot_chat_sent emit above the traceparent comment so the comment sits
  with the code it documents.

* fix(table): attribute import column audit to the importing user

Address Cursor review: addImportColumns (async createColumns import path) now
threads the importing userId into auditTableColumnsAdded instead of falling back
to table.createdBy, so column additions are attributed to the actual member who
ran the import rather than the table creator.

* feat(observability): drop copilot from storage accounting

Per design decision: copilot files are working/conversational artifacts, so
gating their materialization on storage quota would fail an agent operation
mid-flow when a user is over limit, and metering them would inflate usage enough
to block KB/workspace uploads indirectly. Remove copilot quota gates + copilot
ingest metering entirely (revert copilot-file-manager, metadata, and the upload
route's copilot branch to baseline) and drop the now-unused releaseDeletedFileStorage.
KB document metering + quota enforcement (the deliberate, persistent storage path)
is unchanged.

* fix(analytics): switch workspace + org PostHog groups together

Address Cursor review: gate the group-sync effect on workspace metadata being
loaded (activeWorkspace present) so the workspace and organization groups always
update atomically. Acting during the load window paired the new workspace group
with the previous workspace's org group; until metadata loads, events stay
consistently attributed to the previous workspace.

* fix(table): audit async export at authorization, not job completion

Address Cursor review: async exports only emitted TABLE_EXPORTED when the
background job reached 'ready', so an authorized export whose job later failed or
was abandoned left no audit trail — inconsistent with the sync export route,
which audits before streaming. Move the audit + analytics to the async route's
authorization point (after the job is claimed/dispatched) and remove it from the
runner. Drop the now-unused userId from TableExportPayload.

* fix(billing): never let payment_failed instrumentation skip user blocking

Address Greptile P1: the payment_failed audit hoisted an unguarded
isSubscriptionOrgScoped DB read (for entity_type) directly before the
attempt-count user-blocking block. A transient failure of that read would throw
out of the handler and skip blocking. Wrap the whole audit/analytics block in
try/catch (best-effort), and let the blocking compute its own isSubscriptionOrgScoped
as it did originally — instrumentation can no longer abort payment processing.

* chore(observability): trim verbose inline comments to concise notes

* fix(billing): wrap new dispute/enterprise/subscription instrumentation in Stripe webhook idempotency

recordAudit/captureServerEvent calls added for charge disputes, enterprise
subscription provisioning, and free->paid subscription creation ran
unconditionally with no idempotency guard, unlike their sibling handlers
in the same files. Stripe redelivers webhooks at-least-once, so a retry
would double-record the audit row and PostHog event even though the
underlying DB writes were already idempotent.

* fix(observability): tag org-scoped audit metadata with organizationId, guard concurrent table delete

The org-scoped self-service audit-log endpoint matches org-level rows via
metadata.organizationId (or resourceType=organization). Six recordAudit
call sites (subscription create/cancel, admin credit issuance, threshold
overage billing, charge disputes, credit purchase, invoice payment
succeeded/failed) tagged org-scoped events with a differently-named key
(referenceId/entityId/targetOrgId), making them invisible to org admins
querying their own audit trail despite being stored in the DB. Add the
missing organizationId key everywhere the pattern was missed.

Also guard deleteTable's archive UPDATE with isNull(archivedAt) so a
concurrent duplicate delete request is a no-op instead of re-archiving
and re-firing a duplicate TABLE_DELETED audit row.

Adds test coverage for the ORG_MEMBER_ADDED audit/analytics emission in
acceptInvitation, which previously had none.

* fix(test): queue resolveBillingActorId's owner lookup in payment-failure email test

handleInvoicePaymentFailed's new payment_failed audit instrumentation
resolves the billing actor via an extra db.select before
sendPaymentFailureEmails runs. The test's fixed select-response queue
didn't account for it, so the org-admin lookup consumed the wrong
queued row and the assertion saw zero email sends. Production behavior
is unaffected — each query is independent; this was a mock-queue
ordering issue only.

* fix(billing): don't let a post-commit usage-limit sync failure suppress the seat audit

Cursor Bugbot flagged: reconcileOrganizationSeats committed the seat
change and Stripe outbox enqueue in a transaction, then called
syncSubscriptionUsageLimits outside it before recording the audit/
PostHog events. A thrown sync left a genuinely-changed seat count with
no ORG_SEAT_PROVISIONED/DEPROVISIONED trail. Wrap the sync in its own
try/catch (log + continue) so the events always fire for a committed
change, matching the fire-and-forget instrumentation pattern used
elsewhere in this PR.

* fix(billing): guard the new subscription-created instrumentation block

Greptile P1: the org-scope check I added for the audit-metadata fix
(isSubscriptionOrgScoped) was a raw DB call with no error guard, unlike
resolveSubscriptionActorId next to it. Since it ran inside the
idempotency lambda with an unconditional rethrow, a transient DB error
would abort and retry the whole webhook after the free -> paid usage
reset had already committed. Wrap the actor/org-scope resolution +
audit + analytics block in try/catch, matching the same guarded
instrumentation pattern already used in handleInvoicePaymentFailed.
2026-07-08 23:23:59 -07:00
Waleed d078c84ee6 chore(typescript): upgrade to TypeScript 7 (native Go compiler) (#5521)
* chore(typescript): upgrade to TypeScript 7 (native Go compiler)

Bumps typescript to ^7.0.2 across every workspace package. Full
bun run type-check/lint/build/test all pass; apps/sim's type-check
(the one needing an 8GB heap bump) drops from ~55s to ~7s wall time.

Migration fixes required by TS7's stricter defaults:
- baseUrl removed: drop it from 5 tsconfigs (paths already resolved
  relative to tsconfig dir, so behavior is unchanged) and prefix the
  one bare (non-relative) paths entry each in apps/sim and
  apps/realtime with './'
- moduleResolution=node10 removed: switch packages/cli and
  packages/ts-sdk to "bundler", matching the rest of the monorepo
- types now defaults to [] instead of auto-including every @types/*
  package: add "types": ["node"] to the shared base tsconfig (this
  is fundamentally a Node monorepo, so this restores prior behavior
  in one place instead of duplicating it per-package), add explicit
  @types/node deps to packages that now rely on it transitively via
  @sim/db/@sim/logger, and add "declare module '*.css'" to the two
  packages with plain (non-module) CSS side-effect imports that
  TS7's stricter checker now flags
- packages/logger's isomorphic `typeof window` check no longer needs
  DOM lib in every consumer: replaced with `'window' in globalThis`
- packages/testing and apps/realtime's fetch/DOM mocks need DOM lib
  where they're compiled, since they model the browser Fetch API
- the `typescript` npm package no longer exports the classic
  Compiler API from its main entry (moved to unstable/ast subpaths);
  apps/sim's Function-block route used it at runtime to strip
  import statements from user code, so that one call site now uses
  Microsoft's official transition package, @typescript/typescript6
- Next.js 16.2.6's own TypeScript-detection heuristic hardcodes a
  path TS7 no longer ships, and its auto-install fallback assumes
  npm/pnpm; added @typescript/native-preview as a devDependency to
  apps/sim and apps/docs so Next detects a valid native compiler
  instead of trying (and failing) to auto-install one

Not merging yet: TS 7.0.2 published today and is still inside this
repo's bunfig.toml minimumReleaseAge (7-day) supply-chain gate, so
`bun install` will fail for everyone until 2026-07-15. Opening this
now to get it through review; hold the actual merge until then.

* fix(typescript): address Greptile review findings on TS7 upgrade

- packages/logger: 'window' in globalThis treats a shim that leaves
  globalThis.window explicitly undefined as browser-only, silently
  dropping production server logs. Restore the original
  typeof !== 'undefined' semantics via an inline cast instead, so it
  stays correct without requiring DOM lib in every consumer.
- packages/ts-sdk, packages/cli: both are tsc-built, published as
  Node ESM (package.json "type": "module" with an "exports" map).
  "moduleResolution": "bundler" is too permissive for that target -
  it accepts import patterns (e.g. extensionless relative imports)
  that Node's actual ESM resolver rejects at runtime. Switch both to
  "module"/"moduleResolution": "nodenext", the correct pairing for a
  published Node ESM package. Verified real tsc builds (not just
  --noEmit) still succeed for both.

* chore(bunfig): temporarily disable minimumReleaseAge gate for TS7 install

TS 7.0.2 published today, still inside the 7-day gate. Lowering to 0
to unblock this merge; will restore to 604800 in an immediate follow-up
commit right after merging.
2026-07-08 16:41:23 -07:00
Theodore Li 0613cebfab feat(db): resolve DATABASE_URL per role (DATABASE_URL_<ROLE> with fallback) (#5276)
* feat(db): resolve DATABASE_URL per role (DATABASE_URL_<ROLE> with fallback)

* fix(db): pin realtime process to SIM_DB_ROLE=realtime so both pools share the role

Without it, the realtime process left SIM_DB_ROLE unset: the shared @sim/db
client defaulted role to 'web' (web pool profile + DATABASE_URL_WEB) while
socketDb used 'realtime', so the two pools diverged after cutover. Set it at the
process level (bootstrap + dev/start scripts), mirroring DB_APP_NAME, so the
shared client and socketDb both resolve the realtime profile and URL.
2026-06-30 15:17:55 -04:00
Theodore Li 845a6276d9 perf(db): per-role Postgres connection-pool profiles (#5232)
* perf(db): drive Postgres pool size + application_name from per-role profiles

Replace ad-hoc DB_APP_NAME sizing with a per-role profile map keyed by
SIM_DB_ROLE (web/trigger/realtime), defaulting to web. Trigger machines
open a small pool instead of 15 to avoid PgBouncer connection exhaustion.
Also size realtime's separate socketDb pool down to 10.

* fix(db): throw on invalid SIM_DB_ROLE instead of silently using web pools

* fix(db): use Object.hasOwn for SIM_DB_ROLE validation to avoid prototype keys
2026-06-27 15:15:05 -04:00
Theodore Li a68d38ae75 feat(db): attribute Postgres connections by runtime via application_name (#5211)
* feat(db): attribute Postgres connections by runtime via application_name

* improvement(db): label migration-runner connection sim-migrate; trim DB_APP_NAME comment

* fix(db): label realtime's shared @sim/db connections sim-realtime too

The realtime process uses both its own socketDb pool and the shared @sim/db
client (handlers, preflight, permissions). Only socketDb was labeled, so the
shared client defaulted to sim-app, mislabeling much of realtime's DB traffic.
Set DB_APP_NAME=sim-realtime at the process level (bootstrap before the dynamic
@/index import for prod; dev/start scripts for local) so both clients report it.
2026-06-25 17:11:23 -04:00
Theodore Li e1c3c7f6c9 feat(secrets): ingest env secrets at container runtime instead of fanning into ECS taskdef (#5189)
* feat(secrets): ingest env secrets at container runtime instead of fanning into ECS taskdef

The app/socket ECS taskdefs were ~42KB, ~93% of which was the secrets[] array:
268 pointer entries each restating the full ~78-char secret ARN, marching toward
the 64KB taskdef limit and growing ~150 bytes per hosted key added. The secret
blob itself is only ~18KB/268 keys.

Move secret delivery to container boot: new @sim/runtime-secrets loadRuntimeSecrets()
reads SIM_ENV_SECRET_ID, fetches the combined secret once, and hydrates process.env
(no-clobber, no-op when unset, fail-fast). Bootstrap entrypoints for app + realtime
await it before importing the real server (env-flags reads env at module load). The
app bootstrap is bun-bundled in the Dockerfile builder stage since it runs outside
the Next standalone bundle; realtime keeps full node_modules and runs the TS entry.

Backward-compatible: with the current fan-out taskdef the loader no-ops and the app
reads the injected env vars unchanged. The matching infra change (empty secrets[] +
SIM_ENV_SECRET_ID) ships separately, after this image is live.

* fix(runtime-secrets): address review feedback

- Move the binary-secret guard outside the retry loop (sendWithRetry) so a
  missing SecretString throws immediately instead of burning 3 attempts + backoff.
- Bound each Secrets Manager request with AbortSignal.timeout(5s) so a stalled
  response can't hang boot indefinitely.
- Drop the redundant @aws-sdk/client-secrets-manager pin from apps/realtime; it
  resolves transitively via @sim/runtime-secrets.
- Add a test for the non-retriable binary-secret path.
2026-06-24 16:23:37 -04:00
Vikhyath Mondreti 91f9dfdaec improvement(governance): derived access (#5134)
* improvement(governance): org-ws-credential roles clarity

* revert isHosted

* improvement(credentials): code cleanup

* address comments

* make kb cascade delete on user hard delete

* revert env flags

* chore(db): drop local 0242 migration to regenerate after merging staging

Our 0242 collides with staging's 0242. Remove it (and its snapshot +
journal entry) so the KB-cascade migration can be regenerated with the
correct number on top of the merged staging migrations.

* chore(db): regenerate kb→workspace cascade migration as 0243

Regenerated via drizzle-kit generate on top of the merged staging
migrations (staging took 0242). Re-applied the safety edits: NOT VALID
+ separate VALIDATE on the FK re-add, and the -- migration-safe note on
the DROP. check:migrations passes.

* improve copy

* update docs
2026-06-19 12:47:09 -07:00
Vikhyath Mondreti d14bc78d03 fix(realtime): re-check workspace role on mutating socket events (#5080)
* fix(realtime): re-check workspace role on mutating socket events

* address comments
2026-06-16 10:50:36 -07:00
Waleed ae075f871f Revert "fix(realtime): re-validate socket role and evict revoked collaborator…" (#5051)
This reverts commit 4ab8760ae0.
2026-06-14 22:20:12 -07:00
Waleed 4ab8760ae0 fix(realtime): re-validate socket role and evict revoked collaborators (#5050)
Socket.io authorized workflow access only at join and cached the workspace
role in presence, so a removed or downgraded collaborator kept live read/
write access until they disconnected.

- Re-validate the cached role against the permissions table on mutating
  events, bounded by a short TTL; refresh or evict on change
- Add /api/permissions-updated so the app reconciles active rooms, evicting
  revoked users (cross-pod) and refreshing downgraded roles
- Notify realtime on workspace member removal and permission changes
2026-06-14 21:19:32 -07:00
Vikhyath Mondreti f7b40fe4a4 fix(db-part-1): eliminate pool self-deadlock from nested checkouts inside transactions (#4975)
* fix(db-part-1): eliminate pool self-deadlock from nested checkouts inside transactions

* update docs
2026-06-11 17:29:13 -07:00
Vikhyath Mondreti d9f78c0f56 improvement(sockets): make offline mode recoverable and stop transient races tripping it (#4980)
* improvement(sockets): make offline mode recoverable and stop transient races tripping it

* data persistence issues should trigger offline mode and force refresh

* code cleanup
2026-06-11 16:52:59 -07:00
Waleed 272bad9718 feat(realtime): preflight schema-compatibility check on startup (#4940)
* feat(realtime): preflight schema-compatibility check on startup

The socket service authorizes every connection with a full-row query against
the workflow table. When a deploy ships a realtime image whose compiled schema
is ahead of/behind the live DB (e.g. a column dropped by a migration the image
predates), that query fails on every request and silently breaks persistence —
yet the process stays up and the shallow /health probe keeps returning 200, so
the deploy looks healthy while serving nothing.

Run one representative workflow query before listen(): a schema mismatch throws,
propagates to the entrypoint, and the task exits non-zero and never goes healthy,
so CodeDeploy auto-rolls-back instead of shifting traffic onto broken tasks.

Schema-class errors (undefined column/table/function) fail fast; connection-class
errors retry with backoff so a cold DB at boot does not flap. Runs once at
startup, never on the per-probe LB health check, to avoid a DB blip mass-
terminating the fleet (cascading failure).

* fix(realtime): unwrap cause for schema codes, drop sleep after final attempt

- isSchemaMismatch now walks the error.cause chain — drizzle wraps the driver
  error, so the SQLSTATE often lives on the inner cause, not the outer throw.
  Without this a wrapped 42703/42P01 was retried 5x and mis-reported as
  "database unreachable" instead of failing fast.
- No longer sleeps after the final failed attempt (~6-10s of dead wait that
  undermined the fail-fast contract); sleep now only happens between attempts.
- Tests: assert sleep is called exactly 4 times on exhaustion, and add a
  wrapped-cause fail-fast case.
2026-06-09 22:03:03 -07:00
+1 39418f9615 improvement(mothership): v0.2 (#4923)
* CONTRACTS

* updates

* Prompt caching

* Fix regression

* updates to prompt caching

* Prompt caching trace

* VFS updates

* Changelog and plan

* VFS update

* Improvement/mothership (#4775)

* ff

* System role in cache

* Fixes

* Add dynamic user info

* Improve tool search

* improvement(platform): workspace UI/UX overhaul + integrations catalog

Rework the workspace around the AI-workspace model: a Mothership home, a
top-level Skills route, connected-credential and integration-detail pages,
and a polished sidebar/settings surface. Replace the notifications store
with a unified toast system (provider-level dismiss/pause, countdown ring).

Integrations & catalog:
- Add a BlockMeta layer (tags + catalog templates) scoped to catalog-visible
  integrations; every catalog integration carries >=7 grounded templates.
- Rework the taxonomy: each block declares category tools|blocks|triggers.
  3rd-party services are 'tools'; first-party primitives (postgres, mysql,
  knowledge, file, search, stt/tts, image/video generators, thinking, etc.)
  are 'blocks'. Versioned blocks follow the upgrade paradigm (old hidden,
  latest in toolbar/docs).
- Generate integrations.json + tool docs canonically from block configs.

Architecture & cleanup:
- Consolidate block data extraction behind a single latest-version strategy
  (getCanonicalBlocksByCategory; version-consistent getBlockMeta).
- Unify version-suffix handling in @sim/utils/string (stripVersionSuffix /
  isVersionedType, with tests); registry, generate-docs, tools/utils, and
  integrations all route through it.
- Repair latent broken barrels, remove dead code, fix BlockMeta-related type
  errors and 5 broken docs links.

Behavior-preserving for block execution and the toolbar's tool/block listing.

* refactor(platform): remove forms, templates, and creators features

Remove three standalone features and their supporting code:
- Forms: form-deployment pages, API routes, execution path, and docs.
- Templates: the template gallery (landing + workspace) and template APIs.
- Creators: creator-profile routes and contracts.

Add a super-user permissions module (lib/permissions/super-user) and an
organizations API contract; update the audit/db/testing packages, billing,
and the session/theme providers accordingly.

* New doc setup

* Fixes

* Fixes

* Logs tools

* test(workflows): update archiveWorkflow update count after forms removal

The forms feature was removed, dropping the form-table update from
archiveWorkflow. Update the stale assertion from 8 to 7 tx.update calls.

* Add finalization on error

* Fixes

* change dev CI to bun run db:push

* Updates

Please enter the commit message for your changes. Lines starting

* Contrat update

* improvement(nested): subagents

* updates

* upgrade

* improvement(knowledge): polish tag filter dropdowns (#4816)

* improvement(logs): object storage backed tracespans (#4787)

* improvement(logs): obj storage backed tracespans

* fix storage write context

* fix tests

* address comments

* address comments

* chore(db): remove migration 0219 to regenerate after staging merge

Drops the 0219_robust_shard SQL, its snapshot, and the journal entry so the
trace-spans/cost schema migration can be regenerated on top of the latest
staging migration chain (avoids a number collision with staging's migrations).

Co-authored-by: Cursor <cursoragent@cursor.com>

* improvement(billing): accurate per-member usage via shared ledger helper

Per-member/per-user usage in the org-member routes now adds the usage_log
ledger to the currentPeriodCost baseline (which is no longer incremented),
via a shared getOrgMemberLedgerByUser helper to avoid repeating the
subscription→period→ledger lookup across the admin and member-facing routes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* regen migrations

* update migration

* address comments

* more code cleanup

* incorrect type cast

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* improvement(providers): harden OpenAI-compatible providers + add tests (#4796)

* improvement(providers): harden OpenAI-compatible providers + add tests

* fix(vllm): let tool-loop errors propagate instead of returning silent partial success

* fix(litellm): force tool_choice 'none' on final structured-output call

The deferred final call used tool_choice 'auto', so the model could emit
another tool_calls round instead of the structured answer, leaving content
stale. Use 'none' (matching vLLM/Fireworks) on both the streaming and
non-streaming final calls so the model must return the structured response.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(providers/ollama): drop tools from post-tool streaming call

Ollama ignores tool_choice (not in its supported fields), so vLLM/Fireworks'
tool_choice:'none' guard is a no-op here. Omit tools from the final streaming
payload instead so the summarization turn can't emit dropped tool calls.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(litellm): spread payload into deferred final call so reasoning_effort carries over

The non-streaming deferred finalPayload hand-picked fields and dropped
reasoning_effort (and any future payload field), diverging from the streaming
path which spreads ...payload. Spread payload here too for consistency.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(providers/ollama): restore enrichment TSDoc block

Keeps parity with sibling Chat Completions providers (cerebras/mistral/xai).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(fireworks): restore TSDoc on utils helpers

Restore the TSDoc blocks on supportsNativeStructuredOutputs,
createReadableStreamFromOpenAIStream, and checkForForcedToolUsage —
TSDoc is the codebase documentation standard and should not have been
stripped.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(litellm): remove inline rationale comments (codebase uses TSDoc)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(providers/ollama): drop orphaned enrichment TSDoc

The block documented a function that now lives in trace-enrichment.ts, so it
documents nothing in this file.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* chore(copilot): deprecate mcp server (#4797)

* chore(copilot): deprecate mcp

* update error codes

* deprecate copilot api v1 route

* feat(integrations): hosted API keys for Findymail, Prospeo, and Wiza (#4777)

* feat(integrations): hosted API keys for Findymail, Prospeo, and Wiza

Add hosted-key support across all credit-consuming Findymail, Prospeo, and Wiza operations so Sim provides the key when a workspace has not brought its own. Register the three BYOK providers, consolidate Wiza's two-step reveal into a single polling wiza_individual_reveal op, and hide the API key field on hosted Sim for hosted operations.

* fix(integrations): harden Wiza reveal polling, soften enrichment getCost guards

Address Greptile + Cursor Bugbot review on #4777: return explicit failures from the Wiza individual_reveal poller instead of throwing (thrown errors were swallowed into a false queued success), short-circuit when the initial reveal is already terminal, tolerate transient 5xx/429 during polling, and return 0 (not throw) from Findymail getCost when the contacts/employees array is absent.

* chore(integrations): biome formatting after wiza merge resolution

* fix(wiza): type isTerminalReveal param structurally for next build typecheck

* feat(enrichments): add Findymail, Prospeo, Wiza to work-email waterfall

* feat(enrichments): add Wiza + Prospeo phone reveal to phone-number waterfall

* feat(enrichments): opportunistic identifiers + LinkedIn URL input across work-email & phone cascades

* fix(tables): reduce column header chevron size and fix sidebar shadow bleed (#4800)

* feat(slack): add install + privacy section to integration landing page (#4799)

* feat(slack): add install + privacy section to integration landing page

Adds a hand-authored, slug-keyed landing-content module (separate from the generated integrations.json so it survives regeneration) and renders an install walkthrough + privacy-policy link on integration pages when present. Also refreshes generated docs (data-enrichment entry, icon mappings, tool mdx).

* fix(landing): render privacy section independently, align CTA analytics label

* docs(landing): clarify the Slack install button is behind sign-in

* refactor(landing): bake integration landing content into generated json via docs-gen

Moves landing content (install walkthrough + privacy) out of a render-time augment and into the generation pipeline: generate-docs reads the pure-data content map and writes landingContent into integrations.json, so the page reads a single source (integration.landingContent). Canonical types live in integrations/data/types.ts.

* improvement(enrichments): align enrichments sidebar with design system (#4801)

* improvement(enrichments): align enrichments sidebar with design system

* fix(enrichments): consistent close button pattern and fix url link hover

* fix(misc): upgrade path change for new better-auth version, billing issue for workflow block agent usage (#4803)

* fix(misc): upgrade path change for new better-auth version, double-billing for workflow block agent usage

* fail loudly if stripe sub id missing

* fix(copilot): seq migration (#4804)

* chore(db): drop redundant idx_webhook_on_workflow_id_block_id index (#4809)

Removed because (workflow_id, block_id) is a left-prefix of idx_webhook_on_workflow_id_block_id_updated_at_desc, which fully covers it. The dropped index was non-unique and enforced no constraint.

* perf(copilot): read chat transcripts from copilot_messages (R+1 cutover) (#4808)

* perf(copilot): read chat transcripts from copilot_messages, not JSONB

Flip user-facing chat reads from the legacy copilot_chats.messages JSONB
array (5.7GB, 99% TOAST) to the normalized copilot_messages table via a
new loadCopilotChatMessages helper ordered by seq NULLS LAST, created_at,
id — the verified canonical order. Both chat-detail getters
(getAccessibleCopilotChat, getAccessibleCopilotChatWithMessages) now drop
the messages column from their metadata select (no more whole-array
detoast on every load) and assemble the transcript from the table after
authorization. This cascades to the copilot + mothership GET endpoints
and to resolveOrCreateChat's conversationHistory (the LLM payload).

The normalize/effective-transcript pipeline is source-agnostic
(copilot_messages.content == a JSONB array element), so transcripts are
byte-identical. Dual-write and the JSONB column stay in place as the
internal-logic source and fallback; removing JSONB writes is a later step.

Prod integrity verified before cutover: 0 messages missing, 0 NULL-seq,
0 dup keys/seq, 0 orphans, order-parity vs JSONB = 0 mismatches.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(copilot): cover auth-deny on a found row skips the messages query

Address PR review: exercise the `if (!authorized) return null` contract —
when the chat row exists but authorization fails, the getter returns null
and never issues the copilot_messages read.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tables): right-align run/stop in embedded toolbar; workflow cells format like normal cells (#4806)

* fix(tables): right-align run/stop in the embedded table toolbar

Add a right-aligned `trailing` slot to ResourceOptionsBar and move the embedded
mothership table's run/stop control into it, so Filter + Sort stay left-aligned
and run/stop sits opposite on the right. No-op for the search-bearing consumers
(logs, resource list), which don't pass `trailing`.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(tables): workflow-output cells format values like normal cells

Workflow-output columns short-circuited in resolveCellRender and rendered their
value as plain text, so a sim-resource URL / external URL / JSON / date produced
by a workflow never got the chip, favicon link, or typed formatting a normal
cell gets. Factor value formatting into a shared `resolveValueKind` helper used
by both the workflow-value branch and the plain-cell branch; the workflow branch
keeps the typewriter reveal for plain streaming text via a `typewriter` flag.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(tables): detect resource/URL links on workflow output regardless of column type

Workflow output columns default to `json` (columnTypeForLeaf), so routing their
values through the type-based formatter (a) gated chip/URL promotion behind
`column.type === 'string'` — a URL produced by a json-typed output never became
a chip — and (b) JSON.stringify'd plain string values, adding quotes and losing
the typewriter reveal. Detect links (sim-resource chip / favicon URL) on the
value string directly for workflow outputs, falling back to the plain `value`
kind; plain cells keep the type-based formatting. Addresses Greptile P2 on #4806.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(icons): repair broken integration icon rendering (#4810)

* fix(icons): repair broken integration icon rendering

Two distinct bugs left integration icons broken on the /integrations page
(visible at 32-40px, hidden at the toolbar's 16px):

1. Corrupted SVG paths (Notion, Greptile, Granola, Calendly, Grafana, Bedrock):
   over-minified data dropped elliptical-arc flag digits (e.g. `A1 1 0 5.9 7`
   instead of `A1 1 0 0 0 5.9 7`); Granola's cubic stream was truncated. Browsers
   abort path parsing at the first invalid arc flag, so each rendered as a fragment
   or blank. Replaced with correct path data from canonical sources, preserving each
   icon's existing fill/gradient and bgColor.

2. Invisible glyph (Bright Data): its icon uses fill='currentColor' but bgColor was
   '#FFFFFF', and every surface forces text-white on the glyph - white-on-white.
   Changed bgColor to Bright Data's brand blue (#3d7ffc) so the white glyph reads,
   matching the white-glyph-on-brand-chip convention.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(icons): restore Calendly dual-tone brand colors

Addresses review feedback: the previous fix replaced the broken Calendly icon
with a monochrome #006BFF path, dropping the cyan #0ae8f0 accent from the
original dual-tone mark. Restored the two-tone logo (blue + cyan) using clean,
valid path data, cropped to a tight square viewBox so it fills the chip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* improvement(icons): enlarge icons, fix Zoom contrast and Quiver chip

- Zoom: glyph was blue-on-blue (#0B5CFF on #2D8CFF chip); switched to
  currentColor so it renders as a white glyph on the blue chip.
- Quiver: chip bgColor #000000 -> #FFFFFF to match the icon's near-white box,
  and enlarged the mark slightly (viewBox crop).
- Enlarged (tightened viewBox, verified no clipping): RevenueCat, Prospeo,
  Granola, Firecrawl, Enrich.so, and the AWS icons (RDS, DynamoDB, SQS,
  CloudFormation, Athena, CloudWatch, SES, Bedrock, S3).
- ZoomInfo left unchanged: it is a full red rounded-square logo that already
  fills its frame, so a crop would clip it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(icons): use Bright Data wordmark on white chip; repair Circleback

- Bright Data: replaced the flame glyph with the official two-tone 'bright data'
  wordmark (provided asset), centered in a symmetric viewBox. Reverted the chip
  bgColor from #3d7ffc to #FFFFFF since the blue wordmark is invisible on a blue
  chip (the wordmark is designed for a light background).
- Circleback: a minifier had rounded the pattern's image scale to scale(0),
  collapsing the embedded logo to zero size (invisible). Restored the correct
  scale (1/280 = 0.00357142857) so the C. mark renders.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(docs): sync Quiver block color card to white chip

Reflects the Quiver bgColor change (#000000 -> #FFFFFF) in the docs block info card.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* improvement(icons): enlarge AWS/Cloudflare/Dagster icons, fully white Zoom

- Enlarged (tighter viewBox, render-verified, no clipping): Cloudflare, Dagster,
  and the red AWS icons AWS IAM, Identity Center, Secrets Manager, SES, STS.
  Identity Center was anomalously small (filled ~32% of its frame); the group is
  now sized consistently (~80% fill).
- Zoom: the camera lens triangle was still #0B5CFF (blue-on-blue); switched it to
  currentColor so the whole camera renders white on the blue chip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* docs(wiza): consolidate individual reveal into a single operation

Merges the separate Start/Get Individual Reveal operations into one Individual
Reveal operation in the Wiza docs and integrations data (operationCount 5 -> 4).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* improvement(icons): size remaining AWS icons to match the set (~80% fill)

Bring RDS, DynamoDB, SQS, CloudFormation, Athena, CloudWatch and S3 up to the
same ~80% fill as the AWS IAM/Identity Center/Secrets Manager/SES/STS group, so
all AWS icons are visually consistent. Bedrock left as-is (already ~92% fill).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(icons): use Bright Data flame mark, enlarge ZoomInfo

- Bright Data: the full 'bright data' wordmark was illegible at chip size.
  Replaced with just the flame-'i' brand mark (blue #4280f6 on the white chip),
  centered.
- ZoomInfo: cropped the viewBox toward the white 'Zi' so it's larger; the red
  rounded-square background still fills the chip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* improvement(icons): enlarge CrowdStrike icon

The falcon mark sat small in its chip because the icon used a wide 768x500
viewBox (letterboxed in the square chip). Switched to a square viewBox centered
on the mark so it fills ~80%, consistent with the other icons.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tables): serialize schema mutations to prevent parallel column clobber (#4812)

* Make workflow description nullable

* fix(tables): serialize schema mutations to prevent parallel column clobber

* fix(tables): load workflow outside schema lock; use DbOrTx for getTableById

* fix(tables): scale idle timeout in updateColumnType to avoid aborting large type changes

* fix(tables): skip stale remap types when workflowId changes concurrently

* fix(tables): scale idle timeout in updateColumnConstraints for large tables

* fix(wait): resume live/draft async waits and preserve cell context on chained waits (#4814)

* Make workflow description nullable

* fix(wait): resume live/draft async waits and preserve cell context on chained waits

* improvement(knowledge): polish tag filter dropdowns

* improvement(knowledge): soften filter section labels

* improvement(knowledge): soften list filter labels

* fix(security): harden SSO domain registration, webhook path isolation, and CSV export (#4813)

* fix(security): harden KB file access, SSO domain registration, webhook path isolation, env secrets, and CSV export

* fix(sso): scope domain conflict query with indexed lower(domain) filter

Address PR review: avoid a full-table scan on every SSO provider
registration by filtering candidate rows in SQL with
lower(domain) = <normalized>, keeping the in-memory ownership check.
Also tighten the normalizeSSODomain TSDoc.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: condense env route security comments

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* icons update

* chore(security): tighten inline comments in CSV export and KB file authorization

Condense verbose comment blocks to concise TSDoc/single-line form; no behavior change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(security): validate internal serve origin in KB file authorization

Replace the bypassable isInternalFileUrl substring check in resolveInternalKbKey
with an origin allow-list (base URL, internal API base URL, TRUSTED_ORIGINS).
A crafted external host whose path is /api/files/serve/<victim-key> no longer
resolves to the victim key. Relative same-origin URLs are unaffected.

* style(sso): use idiomatic sql lower() comparison for domain conflict query

Match the repo's prevailing `sql`lower(col) = value`` idiom for the
case-insensitive SSO domain conflict lookup.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(security): align workspace env admin gate with hasWorkspaceAdminAccess

Use the same admin check the secrets UI uses (owner, admin permission, or
org-admin) so owners and org-admins are not wrongly denied their own decrypted
workspace secrets, while read-only members remain restricted to names only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(sso): rely on lower(domain) match for conflict detection, drop dead in-memory recheck

Address PR review: the SQL `lower(domain) = <normalized>` predicate already
excludes rows that the in-memory `normalizeSSODomain(...) === domain` recheck
claimed to catch, making that recheck dead/misleading code. Match on the
canonical lower-cased domain and filter purely by ownership. Malformed legacy
values (wildcards, schemes, ports) never match an email domain at sign-in, so
excluding them is not a gap. Test DB mock now applies the lower() predicate so
the casing-variant case is genuinely exercised.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(security): scope webhook deploy path conflict to active webhooks

findConflictingWebhookPathOwner omitted the isActive filter that the
runtime dispatcher (findAllWebhooksForPath) applies, so an inactive but
non-archived webhook from another workflow (e.g. after undeploy or
failure auto-disable) would permanently block any new deployment on that
path even though it never receives deliveries. Align the guard with the
runtime isActive + archivedAt filter; the earliest-owner runtime check
remains the authoritative cross-tenant protection. Also trims verbose
TSDoc on the webhook path-isolation helpers.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(security): exclude archived workflows from webhook deploy path conflict

findConflictingWebhookPathOwner now joins workflow and filters
isNull(workflow.archivedAt), matching the runtime dispatcher
(findAllWebhooksForPath). A webhook on an archived workflow can never
receive deliveries at runtime, so it must not block legitimate path reuse
with a 409.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(security): anchor KB file ownership to earliest document in any state

A KB file's owner is now the earliest document referencing its key regardless of
state (active/archived/deleted/excluded); access is granted only when that owning
document is still active. Closes the residual where an attacker could plant an
active document to claim a file whose original document was archived or deleted.

* updated greptile icon

* revert(security): drop KB file authorization changes

Reverts the knowledge-base file-access work (origin-pinning / owner-pinning /
origin allow-list in verifyKBFileAccess) and its test. The other hardening fixes
(SSO domain registration, webhook path isolation, workspace env secrets, CSV
export) are unchanged. apps/sim/app/api/files/authorization.ts is restored to its
origin/staging baseline.

* fix(sso): treat caller's own user-scoped provider as owned during conflict check

Self-hosters often register SSO user-scoped via the CLI script (no
SSO_ORGANIZATION_ID). If they later enable organizations and reconfigure the
same domain org-scoped through the UI, the conflict check previously treated
their own user-scoped row as another tenant's and returned a misleading 409.
Recognize the caller's own user-scoped provider as owned so that migration is
allowed, while still blocking another user's or another org's domain.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* revert(security): remove workspace-env admin gate

Defer to a credential-based access model (separate change). Restores
GET /api/workspaces/[id]/environment to main behavior and removes the test.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(security): consolidate webhook path-collision check into one helper

Extract findConflictingWebhookPathOwner to lib/webhooks/utils.server.ts as
the single source of truth for cross-tenant path-collision detection, used by
both webhook creation paths (deploy sync and the manual /api/webhooks route).

This also repairs two latent issues in the manual route's previous inline
check, which queried with limit(1) and only webhook.archivedAt:
- limit(1) inspected one arbitrary row, so a same-workflow row could mask a
  foreign collision (false negative). The shared helper scans all matching
  rows.
- It omitted isActive/workflow.archivedAt, so inactive or archived-workflow
  webhooks (which never receive deliveries) permanently blocked path reuse.
  The helper mirrors the runtime dispatcher's filter.

Same-workflow webhook reuse for upsert is now a separate, explicit lookup.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* fix(security): block private/reserved IPs for hosted 1Password Connect SSRF (#4818)

* fix(security): block private/reserved IPs for hosted 1Password Connect SSRF

* test(security): use real isPrivateOrReservedIP and cover IPv6 edge cases

* improvement(integrations): validate and expand devin, cursor, and greptile (#4820)

* improvement(integrations): validate and expand devin, cursor, and greptile

- devin: fix missing org_id path segment on all session endpoints, add 7 session sub-resource tools (list messages/attachments, get/append/replace tags, archive, terminate), pagination, and is_archived output
- cursor: add get_api_key_info, list_models, list_repositories tools
- greptile: align block and docs
- normalize array outputs to default [] and tighten types

* refactor(cursor): simplify list_repositories v2 array normalization

Collapse the redundant `?? []` + `Array.isArray` double-guard into a
single Array.isArray check, per PR review feedback.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(devin): scope session-tag mapping to tag ops and normalize array tag inputs

- Only map sessionTags into the tools tags param for append/replace operations,
  preventing stale sessionTags state from clobbering create_session tags
- Fall back to a wired tags value when sessionTags is empty for tag operations
- Normalize tag inputs (string or wired string[]) via normalizeTags so array
  values from other blocks no longer throw on .split

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(cursor): restore base64 file data in legacy download_artifact metadata

The legacy CursorBlock exposes only content + metadata (no v2 file
output), so metadata.data was the only way legacy-block workflows could
access downloaded artifact bytes. Restore the base64 data field and
document it in the outputs/type instead of dropping it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(devin): coerce terminateArchive to archive flag for boolean-wired input

* docs(integrations): regenerate tool docs for new devin and cursor operations

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* fix(search-replace): don't auto-navigate when content edits invalidate the active match (#4819)

* fix(search-replace): don't auto-navigate when content edits invalidate the active match

* fix(search-replace): clear afterReplaceIndexRef on apply failure and zero matches

* fix(search-replace): remove duplicate setActiveSearchTarget(null) on close

* fix(search-replace): move afterReplaceIndexRef write inside handleApply past the guard

* fix(search-replace): auto-navigate when hydration resolves with no prior active match

* chore(search-replace): remove inline comments

* fix(search-replace): revert !activeMatchId guard that caused immediate re-navigation after deselect

* improvement(enrichments): limit company-info to fields both providers return (#4817)

Hunter's company dataset returns null industry/foundedYear for many large companies (verified against the live API for Microsoft, Amazon, Google), so under the first-non-empty-wins cascade those columns appeared inconsistently across rows. Limit company-info outputs to employee count and description — the fields Hunter and PDL both reliably return — so every row is consistent. employeeCount is a string so Hunter's range bucket and PDL's exact count share the column.

* fix(files): don't reject external URLs containing '..' in file parse validation (#4821)

* fix(files): don't reject external URLs containing '..' in file parse validation

The file block's file_fetch operation rejected any external URL whose path
contained '..' (e.g. Slack files-pri slugs with a literal '...') with
'Access denied: path traversal detected'. Traversal checks only apply to
local paths — external http(s) URLs are fetched with SSRF protection
downstream and are never resolved against the filesystem, so they now
short-circuit as valid. Internal /api/files/serve/ URLs keep full traversal
protection.

* test(files): fix external-URL assertion to handle undefined error

* test(files): assert success explicitly in external-URL traversal test

* fix(files): keep traversal protection for https URLs matching internal serve paths

* feat(google-sheets): add row filtering to read with numeric operators (#4822)

* feat(google-sheets): add row filtering to read with numeric operators

Adds client-side row filtering to the Google Sheets read (v2) operation.
Filter the returned rows by a header column using text operators
(contains, not_contains, exact, not_equals, starts_with, ends_with) and
numeric/ordering operators (gt, gte, lt, lte). Filtering lives in a pure,
unit-tested helper (filterSheetRows) and runs over the fetched read range;
an optional `filter` output reports whether the column was found and how
many rows matched.

Also hardens the surrounding tools:
- trim spreadsheetId in write/update/append URL builders (matches read)
- URL-encode the v1 read default range
- expose valueInputOption for the update operation in the block

Backwards compatible: with no filter requested, read output is byte-
identical and the `filter` field is omitted. The filterMatchType union is
widened additively (4 -> 10 values).

* fix(google-sheets): correct filter metadata for missing column and header-only sheets

- matchedRows is now 0 (not totalRows) when the filter column is not found,
  so it no longer contradicts applied=false / columnFound=false
- columnFound now reflects an actual header lookup for empty/header-only
  sheets instead of being hardcoded true
- add tests covering header-only and empty sheets with present/absent columns

* fix(selectors): fetch all pages for paginated dropdown list routes (#4823)

* fix(selectors): fetch all pages for paginated dropdown list routes

Dropdown selectors fetched only the first page of paginated provider
APIs, silently hiding results past page one. Add bounded server-side
draining to the list routes across Microsoft Graph, Google, Notion,
Atlassian, Linear, AWS CloudWatch, and offset/token REST APIs, plus a
shared client-side drain cap in the selector hook. Response shapes,
stored values, and tool execution are unchanged; CloudWatch list tools
still honor a caller-supplied limit. Also fixes the Word file picker
that was searching for .xlsx files.

* fix(selectors): harden JSM and Monday pagination draining

- JSM service-desk/request-type drains advance `start` by the actual row
  count returned (not the fixed page size) and stop on an empty page, so a
  short non-final page can't skip items.
- Monday boards drain now checks `response.ok` per page, surfacing a
  mid-drain HTTP failure instead of treating it as an empty final page and
  returning a partial 200.

* docs(selectors): clarify JSM drain advances start by actual row count

The offset-advancement fix (advance `start` by the rows returned, not the
fixed page size) landed in 7b19788a8; update the TSDoc to match so it no
longer reads as advancing by `limit`.

* fix(selectors): drain fetchPage in direct fetchList callers

Making `fetchList` optional left three direct callers (outside the
useSelectorOptions hook) calling it unguarded, which broke the build's
type check. Route them through a shared `loadAllSelectorOptions` helper
that uses `fetchList` when present and otherwise drains `fetchPage`.
This also prevents a regression: `confluence.spaces` / `knowledge.documents`
now paginate via `fetchPage` only, and these callers (search/replace,
value resolution) would otherwise have silently returned no options.

* chore(selectors): rename MAX_PAGE_PAGES to MAX_NOTION_PAGES for readability

* fix(sso): re-check domain conflict before write and reject IP-address domains (#4825)

* improvement(copilot): make copilot_messages the sole transcript store, remove JSONB dual-write (#4826)

Stop writing/reading the legacy copilot_chats.messages JSONB column now that
reads are cut over to copilot_messages. Make appendCopilotChatMessages the
primary write (throws on failure instead of swallowing), repoint peripheral
readers (workspace VFS, chat cleanup, data drains, fork, superuser import) to
copilot_messages, and persist the assistant turn inside finalizeAssistantTurn's
transaction so it commits atomically with the stream-marker clear. The column
itself is dropped in a follow-up migration after this bakes.

* feat(tables): expand filter operators (not-contains, starts/ends-with, not-in, empty) (#4827)

Add does-not-contain ($ncontains), starts-with ($startsWith), ends-with
($endsWith), not-in-array ($nin, previously executed server-side but unexposed
in the UI), and is-empty/is-not-empty ($empty) filter operators end-to-end —
SQL builder, condition types, query-builder converters/constants, the filter
UI, the Table tools/block descriptions, and docs.

Also fix correctness bugs in the filter builder surfaced by the wider operator
set:
- Same-column AND rules (e.g. age > 18 AND age < 65, or name startsWith 'A'
  AND name endsWith 'Z') silently overwrote each other because the AND group
  was keyed by column name. They now merge into one operator object, which
  also makes Filter -> rules -> Filter round-trip losslessly for multi-operator
  columns.
- $nin values were not split into an array like $in, and textual-match values
  like "123" were numeric-coerced (breaking the ILIKE path).
- A non-boolean $empty operand from the raw API silently inverted the check; it
  now coerces 'true'/'false' strings and otherwise returns a 400.

* improvement(copilot): stop persisting tool-call result outputs in transcripts (#4829)

Opening a Mothership task could take many seconds because a single persisted
assistant message in copilot_messages.content can reach hundreds of MB, almost
entirely inside contentBlocks[].toolCall.result.output (e.g. a get_workflow_logs
or run_workflow result). The DB query is ~2ms; the cost is detoasting that
payload, shipping it to the browser, and parsing it.

These outputs are dead weight on the Sim side: they are never rendered (the
thread shows only tool name/title/status) and never replayed to the model (the
upstream copilot service owns conversation memory). So drop result.output before
it is persisted, keeping result.success/error plus the tool metadata.

- add stripToolResultOutput() in persisted-message.ts
- apply it in messages-store toRow (covers every write path) and in
  loadCopilotChatMessages (existing rows render fast on read)

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* feat(providers): add Together AI, Baseten, and Ollama Cloud model providers (#4830)

* feat(providers): add Together AI, Baseten, and Ollama Cloud model providers

* fix(providers): guard Ollama streaming fast-path with hasActiveTools

Match Together/Baseten/Fireworks: when tools are supplied but all are
filtered out (usageControl 'none'), take the single streaming call instead
of an extra non-streaming round-trip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(providers): filter non-chat model types from Together model list

* refactor(providers): dedupe Ollama Cloud upstream schema

ollamaCloudUpstreamResponseSchema was byte-for-byte identical to
ollamaUpstreamResponseSchema (both /api/tags endpoints return the same
{ models: [{ name }] } shape). Drop the duplicate and reuse the shared schema.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* fix(knowledge): calendar view sync, deduplicate popover animation classes, type-safe filter cast

* cleanup(knowledge): remove TRIGGER_BORDER_CLASS duplication, inline displayLabel, drop enabledFilterParam alias

---------

Co-authored-by: Vikhyath Mondreti <vikhyathvikku@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Waleed <walif6@gmail.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Theodore Li <theo@sim.ai>
Co-authored-by: andresdjasso <andresdjasso@users.noreply.github.com>

* feat(blocks): add BlockMeta to Quiver and Linq; fix invalid block config fields; update skills

Block fixes:
- Add QuiverBlockMeta (tags + 3 templates: icon generator, diagram creator, vectorizer)
- Fix QuiverBlock: remove invalid tags field from BlockConfig, IntegrationType.Design →
  IntegrationType.AI (Design doesn't exist in the enum)
- Fix GreptileBlock: remove invalid tags field from BlockConfig,
  IntegrationType.DeveloperTools → IntegrationType.DevOps
- Fix LinqBlock: remove invalid tags field from BlockConfig (tags belong only in BlockMeta)

Skills:
- add-block: add dedicated BlockMeta section with structure, rules, and registration
  pattern; add BlockMeta checklist items
- add-integration: add BlockMeta to block structure template, add rules clarifying
  that tags must NOT appear on BlockConfig and integrationType must be a valid enum
  value; update registry snippet to include blocksMeta; add checklist items

* fix(integrations): fix category dropdown by defining missing LANDING_INTEGRATIONS_DATA_PATH and regenerating integrations.json

The staging merge introduced landing-content.ts but forgot to define
LANDING_INTEGRATIONS_DATA_PATH in generate-docs.ts, causing the script
to crash before writing integrations.json.

The stale JSON had integrationTypes (plural array) from an older script
version, while the Integration type and workspace UI both read
integrationType (singular string) — so ALL_CATEGORY_SECTIONS bucketed
to undefined and the category filters never appeared in the dropdown.

Fixed by adding the missing path constant and re-running the generator.
integrations.json now has 192 entries with the correct integrationType field.

* fix(sidebar): restore resize handle on all pages

commit 3109104582 wrapped the resize handle in {(isCollapsed ||
isOnWorkflowPage) && ...} and added a useEffect that resets sidebar
width to SIDEBAR_WIDTH.MIN whenever the user navigates away from a
workflow page. Together these made the sidebar non-resizable on Tasks,
Tables, Knowledge Base, and every other non-workflow page.

Restore the staging behavior: always render the resize handle and
remove the effect that forced the width reset on page transitions.

* fix(sidebar): match staging onKeyDown and tabIndex on resize handle

The resize handle was still conditionalizing onKeyDown and tabIndex
on isCollapsed, blocking keyboard accessibility of the separator role
when expanded. Staging always attaches both unconditionally.

onKeyDown={isCollapsed ? handleEdgeKeyDown : undefined} → onKeyDown={handleEdgeKeyDown}
tabIndex={isCollapsed ? 0 : undefined}                 → tabIndex={0}

* feat(integrations): show connected credentials on integration detail page

When navigating to /integrations/google-docs (or any integration), a
Connected section now appears above the templates listing all workspace
credentials tied to that provider. Each row links back to the credential
detail page (/integrations/connected/${id}) for management actions.

Pairs with the earlier change that routes connected items from the
integrations list to the provider detail page instead of directly to
the credential detail page.

* fix(integrations): rename Add in chat to Add to Sim

* fix(skills): rename Add button to Add to Sim

* fix(platform): restore M1/M2/M3 regressions and LazyMotion on landing page

M1 — Invitation guard: re-introduce usePermissionConfig().isInvitationsDisabled
alongside the workspace inviteDisabledReason check. The flag now also
respects NEXT_PUBLIC_DISABLE_INVITATIONS and EE permission-group
disableInvitations, not just billing policy.

M2 — Settings redirects: /settings/integrations and /settings/skills
now server-redirect to /integrations and /skills respectively so old
bookmarks and emails don't silently land on General.

M3 — Starter block search exclusion: restore block.type !== 'starter'
guard in the search store so users cannot add a duplicate Starter block
via the command palette.

LazyMotion: restore LazyMotion + domMax/domAnimation wrappers and m.*
components in landing-preview-panel and landing-preview-home. The
removal was accidental (the full motion bundle was left after an
import cleanup), which caused the entire framer-motion feature set to
load eagerly on the landing page.

* fix(integrations): revert connected list to credential detail; remove settings redirects

* feat(sidebar): restore workspace switcher search with updated styling

Shows a search input in the workspace dropdown when the user has more
than 3 workspaces (WORKSPACE_SEARCH_THRESHOLD). Keyboard navigation:
ArrowDown/Up to move through results, Enter to switch, resets on close.

Styled to match the current branch (border-1/surface-5 tokens, sm text,
11px Search icon) rather than the old staging styles. Highlight state
is wired through chipVariants active prop so it follows the same active
appearance as clicked/hovered items.

* fix(sidebar): clean up workspace search — layout, memo, and effect guard

* fix(sidebar): align workspace rename input selection style with workflow rename

* perf(sidebar): eliminate React re-renders during sidebar drag resize

Previously, every mousemove during resize called setSidebarWidth(), which
both updated the --sidebar-width CSS variable (sync) and set Zustand state
(async). This caused:
  1. A 1-frame transition flash on mousedown — isResizing state had to
     round-trip through React before the is-resizing CSS class was applied,
     so the width transition fired for the first pixel of movement.
  2. A React re-render per pixel dragged — components reading sidebarWidth
     from the store (avatars, usage-indicator) lagged one frame behind the
     container, making the + and ... buttons appear to jump ahead.

New approach:
  - handleMouseDown adds is-resizing directly to the sidebar DOM node before
    any React involvement (synchronous, no frame lag).
  - mousemove writes only to the CSS custom property (zero React renders).
  - mouseup persists the final width to Zustand exactly once.
  - isResizing / setIsResizing state removed from the store and hook — they
    are no longer needed since the class is managed via direct DOM mutation.

* perf(sidebar): add requestAnimationFrame throttle to resize mousemove handler

* fix(sidebar): fix drag-right lag caused by WorkspaceChrome overflow-hidden transition

The sidebar-container's is-resizing class correctly suppressed its own width
transition, but the two wrapper divs in WorkspaceChrome both have
transition-[width]/transition-transform with a 175ms ease. The outer wrapper
also has overflow-hidden, so while the sidebar content was at the correct width
instantly, it was visually clipped by the outer wrapper which was still
animating — causing the + and ... buttons to appear to lag behind the resize
line on drag-right (not on drag-left, since shrinking doesn't clip content).

Fix: add sidebar-shell-outer/sidebar-shell-inner class names to both chrome
wrappers, and suppress their transitions via html.sidebar-resizing rule when
a drag is active. The html.sidebar-resizing class is toggled directly in the
resize hook alongside is-resizing, so it takes effect synchronously on mousedown.

* fix(icons): redesign Download icon to match Upload style; fix Upload/Download confusion

Download icon was missing the tray/shelf line at the bottom that Upload has,
making it look like a plain arrow rather than a matched pair. Updated Download
to use the same viewBox, stroke weight, and three-path structure as Upload
(tray + stem + arrowhead), just pointing down.

Also fix 5 places where Upload (↑) was incorrectly used for download/export
actions:
  - files.tsx: two Download action rows in toolbar and context menu
  - tables/table.tsx: Export CSV toolbar button
  - table-context-menu.tsx: Export CSV context menu item
  - logs.tsx: Export toolbar button
  - landing-preview-logs.tsx: decorative Export button

Import CSV and actual upload actions correctly keep the Upload icon.

* fix(icons): replace Upload with Download on all remaining export/download actions

- panel.tsx: Export workflow dropdown item
- context-menu.tsx: Export in sidebar workflow context menu
- chat.tsx: Export chat button
- output-panel.tsx: Export console CSV button
- terminal.tsx: Export console CSV button
- resource-content.tsx: Export table as CSV + Download file buttons

* fix(icons): fix remaining Upload→Download on download actions in files and logs

- action-bar.tsx: download button in files toolbar
- file-row-context-menu.tsx: Download item in file context menu
- file-download.tsx: both download buttons in log details file viewer

* updated block skills, settings pages, modals, buttons -> chips, blocks missing metadata

* updated skill modal

* improvement(resource-header): refine breadcrumb truncation ux

* Fixes

* improvement(resource): add floating overflow text tooltips

* wire up credits counter

* improvement(resource-header): mute path dropdown title

* refactor(resource-header): share floating-tooltip engine, prune dead overlay tooltips (#4844)

Clean up the breadcrumb truncation feature for reuse and correctness:

- Extract useFloatingTooltip / useIsOverflowing / FloatingTooltip into a shared
  floating-tooltip module. BreadcrumbSegment and FloatingOverflowText now consume
  one implementation instead of duplicating ~150 lines of positioning, velocity,
  overflow-detection, and portal logic.
- Replace the hardcoded terminal-label regex in ResourceHeader with a typed
  `terminal` flag on BreadcrumbItem (set by the document chunk/loading crumbs),
  decoupling the generic header from knowledge-base copy.
- Clear the path-popover close timeout on unmount and reuse the shared
  POPOVER_ANIMATION_CLASSES constant.
- Drop the redundant manual overflow-state writes (fixes a sticky fade mask).
- Revert FloatingOverflowText inside Combobox `overlayContent` back to plain
  truncating spans across files/logs/tables/scheduled-tasks/document: the combobox
  overlay is pointer-events-none, so the tooltip handlers never fired there.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(emcn,resource-header): address PR #4844 review feedback

- useIsOverflowing now uses a callback ref so the ResizeObserver follows the
  element across mount/unmount/reassignment instead of capturing it once at mount.
  Safe for conditionally rendered consumers of the shared hook. (greptile P2)
- Move POPOVER_ANIMATION_CLASSES out of chip-date-picker implementation internals
  into emcn/components/popover/popover-animation.ts, exported from the
  @/components/emcn barrel. Consumers now import from the module boundary.
  (greptile P2)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Prompt surface

* Regenerated contracts

* upgrade table and styling upgrade

* Literalness gaps

* Updates

* Remove prefix for running code

* Display names for reads

* Tool catalog

* fix schema to include integration

* fix(files): align delete icon with tables view (Trash → Trash2)

Co-Authored-By: waleed <waleed@simstudio.ai>

* fix(mothership): preserve blockType for integration contexts in sent messages

Integration mention chips were missing their provider icons in sent messages
because blockType was dropped when mapping ChatContext to messageContexts.
renderIntegrationTile returns null without blockType, silently hiding the icon.

* fix(mothership): allow 'integration' resource type in chat resources API

The VALID_RESOURCE_TYPES allowlist was missing 'integration', causing a
400 error when adding integrations to the Mothership resource tab — so
they never persisted and disappeared on refresh.

* fix(ui): Add "File" title next to file resource header

* fix(ui): fix resource header columns being bolded

* fix(resource): keep the floating tooltip from jumping on click

Gate the focus-driven show behind :focus-visible so a mouse click (which
focuses the trigger) no longer re-shows the tooltip anchored to the element's
bottom edge. On click the tooltip now hides cleanly instead of jumping down;
keyboard focus still shows it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* perf(sidebar): eliminate unnecessary re-renders in workspace switcher for non-search users

- onMouseEnter: only set highlightedIndex when showSearch is true, preventing
  a state update + re-render on every workspace row hover for users with ≤ 3
  workspaces where the search is never shown
- onOpenChange: only reset workspaceSearch and highlightedIndex when showSearch
  is true, since both values are always already at their defaults for non-search
  users and setting them triggers a pointless re-render during dropdown close
- data-workspace-row-idx: only set when showSearch is true since the scroll
  effect that reads this attribute is already gated on showSearch

* feat(search): context-aware cmd-k results on the integrations page

When cmd-k is opened on the integrations page, show two new result
groups: connected accounts (visible even with empty input) and catalog
integrations (appear once the user types). Selecting an OAuth integration
deep-links to its detail page with ?connect=oauth so the connect modal
auto-opens. Non-OAuth integrations navigate to the plain detail page.

Both groups are gated to the integrations page only and respect the
hideIntegrationsTab permission. The credentials fetch shares the same
React Query cache key as the integrations page itself (no double fetch).

* refactor(emcn): make the floating tooltip the one canonical Tooltip

Replace the Radix-based emcn Tooltip with the cursor-following floating tooltip so
every tooltip in the app uses one consistent style. Built on the shared
floating-tooltip engine (relocated into emcn), not a parallel implementation.

- Move the floating-tooltip engine into emcn/components/tooltip and export it from
  the barrel; re-point its consumers (FloatingOverflowText, resource-header)
- Extend the FloatingTooltip bubble to render arbitrary children (+ role/id for
  a11y) so it can back general tooltips, not just overflow text
- Rebuild emcn Tooltip (Root/Trigger/Content/Provider/Shortcut/Preview) on
  useFloatingTooltip — compound API preserved, ~350 call sites unchanged, legacy
  side/align props accepted and ignored (the tooltip follows the cursor). Removes
  @radix-ui/react-tooltip usage (package kept for a later cleanup; react-slot
  retained for asChild)

Note: general tooltips now show instantly (no hover delay) and follow the cursor.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* style(emcn): put tooltip text on the design scale (text-caption)

Replace the tooltip's ad-hoc `text-xs` + `leading-[18px]` with the semantic
`text-caption` (12px) font-size token so the text styling is fully on the design
scale and self-documenting, matching how the rest of the system is set up. The
color already used the global `--text-body` token. No visual change (still 12px
with a ~18px line height).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Mship byok

* feat(sidebar): add empty task state and inline task creation

- Show "No tasks yet" in the Tasks section (expanded and collapsed) when the list is empty
- Clicking + now creates a task via the API and navigates directly to it, rather than navigating to home
- Add isCreatingTaskRef guard to prevent double-click from spawning multiple tasks
- Disable + button while creation is pending
- Fall back to home navigation on creation error

* invite, billing, home

* Sandbox warning

* improvement(seats): auto purchase seats on invitations into workspace (#4857)

* improvement(seats): auto purchase seats on invitations into workspace

* improve sampling for seat drift reconciler

* address comments

* feat(knowledge): align connector UI with integrations page styling

- ConnectorTypeCard now matches integration rows: brand-colored rounded-xl tile, ArrowRight, title/subtitle hierarchy
- ConnectorCard icon upgraded from flat surface-4 to branded tile (white icon on brand bg, graceful fallback)
- Connector header badges use chipVariants instead of custom Button classes
- Add-connector search input aligned to integrations style (h-[30px], rounded-lg, border-1)

* fix(icons): trim Folder SVG viewBox to remove right-side whitespace

The folder path only extends to x≈14.33 in a 15-unit viewBox, leaving
~0.5 units of empty space on the right. At 12px rendered size this
produces ~0.4px extra gap (visible as ~1px on retina displays) compared
to solid icons like the workflow color square. Trimming the viewBox to
14.5 units makes the folder fill its chip slot evenly.

* fix(user-input): restore draft text synchronously to preserve contexts on nav

The SSR-safe approach (empty useState + effect restore) created a timing
window where the sync effect in useContextManagement fired with message=''
before the value was set, clearing any restored contexts. Folder and workflow
contexts (not re-added by applyAutoMentions) were lost on every nav-back.

Revert to the staging approach: initialize value synchronously from the
draft store so message is already populated when effects run, matching
the behavior on staging.

* fix(queue): render context chips in queued messages

- Remove plainMentions from queued message rows so context chips render
  with icons, consistent with sent messages
- Fix computeMentionRanges to use '/' prefix for skill contexts (content
  has the slash trigger restored at submit time, not '@')

* fix(mothership): remove integrations from add-resource dropdown

* fix(mothership): comment out integrations from add-resource dropdown

* File serializer

* Skills

* chore(db): drop form, templates, template_creators, template_stars tables

These tables backed the Forms and Templates platform features which were
intentionally removed from this branch. Clean up the DB schema to match.

* chore(db): add migration metadata for 0224 drop tables

* block icons, sidebar, toolbar

* chore: remove remaining dead code for template-profile feature

- Remove 'template-profile' from SettingsSection union type
- Remove 'template-profile' entry from SECTION_TITLES
- Remove now-redundant template-profile guard in settings sidebar
- Remove commented-out template-profile nav item

* fix(multi-select): preserve anchor on range selection for tasks and folders

After a shift+click range, the anchor (lastSelectedTaskId / lastSelectedFolderId)
was being updated to the end of the range (toId). This caused subsequent
shift+clicks to extend from the wrong point instead of the original click.

Standard behavior: anchor stays at the initial click (fromId) so repeated
shift+clicks always expand/contract relative to where you started.

* Update skills

* Load skill tool

* Fix vfs dynamic context encoding

* Media subagent

* feat(emcn): add SearchInput component and unify search bars platform-wide

- Add SearchInput to emcn: 30px chip-family filled search input matching the
  integrations page pattern (border-1, surface-5, leading Search icon)
- Migrate all 22 search bars across settings, EE pages, and integrations to
  SearchInput (only layout classes allowed at callsites)
- Rename Sim Keys -> Sim API Keys in nav/title; page copy now says API key
- Remove components/ui input, label, and verified-badge; migrate consumers
  to emcn equivalents or raw inputs (table cell editor, wand prompt bar)
- Delete dead EE skeleton files (data-drains, data-retention)
- General settings: Home Page chip moves to header left as navigation

* fix(files,tables): restore new-file editor autofocus and CSV import error toasts

Both were dropped on the staging line and regressed vs production (main):

- Files: the new-file editor autofocus chain (files.tsx -> file-viewer ->
  text-editor) was stripped by the react-doctor dead-code pass in #4544,
  which misread the prop-drilled `autoFocus` (consumed by an imperative
  `editor.focus()` effect) as unused. Restored the prop through all three
  layers and the one-shot focus effect so creating a new file focuses the
  editor immediately.
- Tables: CSV import failures were silently logged with no user feedback.
  Restored the per-file and generic `toast.error` surfacing.

* Move superagnet back into superagent

* feat(home): score suggested actions by workspace signals

- Derive the suggestion pool from the curated block template catalog
  (1,343 prompts across 172 blocks) instead of 15 hardcoded entries
- Fix inverted relevance: prompts for connected providers are now boosted
  4x (instantly runnable) instead of excluded; unconnected discounted 0.4x
- Weight by featured (3x), popular category (1.5x), and resource gaps
  (no tables -> boost table starters; has KBs -> dampen KB-creation prompts)
- Weighted sampling without replacement, max one suggestion per block
- Connect rows weighted by catalog template count; 2 for fresh workspaces,
  1 once something is connected
- Key the catalog map by both versioned and base block types so gmail_v2
  templates resolve (gmail, github, notion, linear were silently dropped)
- Replace derive-in-effect state with a useMemo keyed by a shuffle nonce
- Add suggested_action_clicked / suggested_actions_shuffled /
  suggested_actions_toggled PostHog events

* billing, teammates

* improvement(credentials): credentials invites, secrets tab wiring up (#4874)

* improvement(credentials): move away from invite notion

* wire up secrets ui/ux

* address comments

* get consistent styling by removing emcninput + text area

* styling consistency

* remove fallback

* address comment:

* refactor(ui): migrate settings & workspace UI to chip design system

Migrate modals to ChipModal (showDivider, hint, resizable, size, leading,
ChipModalTabs), standardize Chip variants, add ChipCombobox wrapper, and apply
chip inputs/dropdowns across settings, knowledge, logs, tables, inbox, EE tabs.
Render ChipModalTabs as a ChipSwitch segmented control. Align /settings/secrets
detail with /integrations, refresh whitelabeling, and restore file-editor
autofocus and CSV import error toasts.

* Update media

* fix(mothership): restore integrations to useAvailableResources for @ mention

Integrations were fully removed from useAvailableResources which broke
the @ mention menu since user-input shares the same hook. Now integrations
are always included in the hook but excluded at the AddResourceDropdown
component level, keeping them out of the sidebar + menu while remaining
available for @ mention autocomplete.

* refactor(settings): chip design-system consistency pass across all tabs

Extract a shared chip-field shell (CHIP_FIELD_SHELL/CHIP_FIELD_INPUT) mirroring
Input variant='chip' and route secrets, credential detail, and integrations
credential detail through it (30px height, font-medium, focus ring). Add a
Discard action to the secrets header when dirty. Group BYOK providers into
Models/Search & web/Enrichment sections and align its tiles to the integrations
tile.

Normalize list-row typography to text-[14px]/text-[12px] and icon tiles to
rounded-xl + border across api-keys, copilot, custom-tools, mcp,
workflow-mcp-servers, credential-sets, and access-control. Tone the secrets
Details chip and per-row affordances to ghost. Fix token correctness: raw
tailwind colors to design tokens (data-drains), missing chip variant on the
Snowflake role input, ColorInput chip-field reuse (whitelabeling), error token
and icon sizes (workflow-mcp-servers), border token (mcp), no-results sizing
(secrets), row chrome (recently-deleted), hover token and Button to Chip
(access-control), and deduped textarea chrome (sso).

* feat(settings): unify filter dropdowns on ChipSelect (integrations style)

Add a ChipSelect emcn component — a filled chip trigger + chevron opening a
DropdownMenu, matching the integrations category filter — supporting single
select, multi-select (checkbox rows), grouped options, and optional in-menu
search. Migrate every settings/EE filter dropdown off ChipCombobox to it:
audit-logs (resource-type multi + time-range), data-retention, data-drains,
general, admin (grouped tool picker), inbox status filter, and the workflow
MCP-server pickers. The SSO provider-id field stays an editable combobox since
it accepts free-text slugs.

Also fix audit findings: chip the MCP client-secret input, normalize an MCP
error-text size, and drop now-dead destination-icon code in data-drains.

* fix(mentions): require explicit @ for integration mentions; decorate sent messages robustly

- Bare integration names in prose (Monday, Notion, Clay) are no longer
  auto-converted to mentions or chipped — mention treatment is strictly
  opt-in via a token-starting @ (fixes the scunthorpe problem)
- @-prefixed mentions still canonicalize casing (@slack -> @Slack) on both
  the keystroke fast-path and bulk paths (paste, template, draft, STT)
- Sent/queued messages now self-sufficiently decorate @IntegrationName
  tokens via a text scan, covering messages sent before the input pass
  ran or authored outside the chat input
- Integration contexts missing a resolvable blockType (messages persisted
  before blockType was saved) are backfilled by label lookup so their
  mention pills render the brand icon again

* refactor(settings): section the API keys page like secrets

Wrap Workspace, Personal, and the allow-personal-keys toggle in SettingsSection
(muted label + divider) instead of bare bold headers, matching the secrets and
BYOK pages.

* fix(settings): ChipSelect renders above modals + full-width form mode

Raise the ChipSelect menu to --z-popover so it layers above modal surfaces
(--z-modal) instead of opening behind them. Add a fullWidth prop that stretches
the trigger and right-aligns the chevron for form-field use, and apply it to the
workflow MCP-server pickers.

* fix(emcn): ChipSelect uses the emcn flat chevron, not lucide's square one

The lucide ChevronDown is square; rendering it at the chip's 9x7 footprint
stretched it. Switch to the custom emcn ChevronDown (built for that wide aspect),
matching the integrations filter and ChipDropdown.

* fix(emcn): ChipSelect trigger hugs its content (w-fit)

In a stacked form layout the trigger was stretched by align-items: stretch,
leaving an empty gap to the right of the value. Add w-fit so the chip sizes to
its content (a compact pill) everywhere; fullWidth form selects are unaffected.

* fix(emcn): ChipSelect uses a square lucide chevron

Revert to lucide's ChevronDown sized square (size-[14px]) so it renders crisp,
matching the standard select chevron used by Combobox.

* renamed tasks to chats

* rename and file change

* Wspace resource refs

* improvement(billing): wire up billing, org, teammates tabs + remove deprecated subscription tab (#4887)

* improvement(billing): wire up billing, org, teammates tabs + remove depr subscription tab

* pass exec timeout to tool routes

* reuse helper

* address comments

* address disable comment

* Vfs updates

* chore(db): remove migration 0224 to regenerate on top of staging

Co-authored-by: Cursor <cursoragent@cursor.com>

* VFS updates and linter

* fix type errors and regen migration?

* File block v5

* Remove docs

* Fix

* Doccer image

* Nuke the dag

* chore(db): drop branch migration 0226 ahead of staging merge; will regenerate

* chore(db): regenerate migration 0226 after staging merge

* Fix lint

* externalize before compaction in fallback'

* Run workflow improvements

* Loosen run options validation

* Lint

* fix save/discard chips to be consistent

* fix(ui): remove smodal tabs in favor of chip modal tabs

* update byok manager component

* fix(ui): skip auto-scrolling on mouse highlight of workspace

* Add creds to wspace context

* Wscontext

* fix(platform): restore settings redirects, forgot-password Enter submit, and tag tooltip visibility

- Re-add SETTINGS_REDIRECTS so /settings/integrations and /settings/skills
  deep links redirect to their top-level routes instead of rendering an
  empty settings panel (accidentally removed in 86da193cc3 one minute
  after cca5054cf6 added it)
- Add opt-in onSubmit to ChipModalField input/email variants and wire it
  in the forgot-password modal so Enter submits again (lost in the
  ChipModal conversion)
- Knowledge tag tooltip: drop the max-h/overflow-y-auto clamp that the
  pointer-events-none floating tooltip made unreachable; truncate each
  tag row instead so all tags stay visible with bounded height

* feat(usage-limits): org member specific limits (#4893)

* feat(usage-limits): org member specific limits

* revert feature flags

* fix type issues

* address comments

* address comments and address other billing gapes'

* fix

* improvements

* fix remaining credits display

* fix

* pass ws id to route validations

* chore(db): remove migration 0226_third_spot before staging merge

* chore(db): regenerate migration as 0227 after staging merge

* Block

* feat(telemetry): add posthog + audit coverage for new platform actions

Audit log (compliance/permission-relevant only):
- org_seat.provisioned — seat auto-purchased when an invite acceptance
  grows the org (actor = accepting user, includes seat delta)
- org_plan.converted — Pro→Team conversion triggered by invite acceptance
- org_seat.drift_reconciled — hourly cron healed a drifted seat count
- credential_member.added/removed/role_changed — credential sharing
  surface was previously fully unaudited
- table.created — parity with existing table.updated/deleted
- skill updates now record skill.updated instead of mislabeled
  skill.created

PostHog:
- seats_provisioned, credential_shared/unshared,
  environment_updated/deleted (key counts only, never names/values)
- table_import_started/completed — background CSV imports previously had
  zero failure observability
- table_exported, file_downloaded, skill_updated
- credential_connected now fires for OAuth completions (draft-hooks),
  credential_deleted for OAuth disconnects; previously only manual
  credentials were tracked
- table_workflow_run gains deployment_mode (live/deployed/mixed)

Also: logger.warn on all new credential-admin 403 denials (members +
environment routes); invite-created audit enriched with
enforcedFixedSeats/plan.

Deliberately excluded as noise: per-hour dead-letter audit rows (would
re-record the same stuck event every cron run) and a duplicate
system-actor org_plan.converted in the Stripe outbox handler.

* feat(home): fill textarea on suggested prompt click instead of sending

Clicking a prompt action in the Suggested Actions panel now populates
the Mothership user-input textarea (via applyAutoMentions) and focuses
it with the caret at end, rather than immediately submitting. The user
can review, edit, and send manually.

* fix(templates): name owning integration in featured block template prompts

Four featured prompts omitted their owning integration name, making them
unbranded and disconnected from their title. Each prompt now explicitly
names GitHub or Google Sheets so mentionifyIntegrations renders the chip
and the copy reads as a self-contained agent-building instruction.

* fix(templates): rewrite fragment prompts so each names its integration and reads as a complete instruction

Featured and non-featured block template prompts that either omitted
the owning integration's canonical name or were phrased as marketing
fragments rather than natural user instructions have been rewritten.
Each updated prompt now starts with an imperative verb ("Build a
workflow that…"), names the owning integration explicitly so the
@-mention chip renders correctly, and aligns with the entry's title.

Files changed: salesforce, hubspot (×2), github, slack, airtable,
firecrawl, iam.

* fix(templates): name integration in remaining block template prompts

- google_docs.ts: replace 'Google Doc' with 'Google Docs document' in 4 prompts; update title 'Meeting notes to Google Doc' to 'Meeting notes to Google Docs'
- google_sheets.ts: replace 'Google Sheet' with 'Google Sheets spreadsheet/Google Sheets' in 3 non-featured prompts
- slack.ts: replace 'Google Doc' with 'Google Docs document' in 'Daily standup summary'
- stripe.ts: replace 'Google Sheet'/'Slacks' with 'Google Sheets'/'Slack' in 'Weekly metrics report'
- reddit.ts: add 'Reddit' to 3 prompts that only referenced subreddits
- notion.ts: rewrite featured prompt to start with a verb and name Notion
- jira.ts: rewrite featured marketing-fragment prompt to start with a verb
- linear.ts: rewrite featured marketing-fragment prompt to start with a verb
- gmail.ts: rewrite featured marketing-fragment prompt to start with a verb and name Gmail

* fix(icons): convert monochrome dark brand icons to currentColor for dark mode

LinkupIcon, InfisicalIcon, IntercomIcon, LumaIcon, GranolaIcon, OnePasswordIcon,
and RailwayIcon were hardcoded to black/near-black fills/strokes, making them
invisible in bare DARK mode. Convert all to currentColor so they follow the
theme-aware text color.

Add iconColor: '#286efa' to IntercomBlock (Intercom brand blue is a confident
mid-tone, safe on both themes).

Two-tone icons (StagehandIcon, AgentPhoneIcon, QuiverIcon) are left unchanged
because their white details are structural — converting to single-tone would
destroy logo legibility.

* removed color, user input, sidebar, suggested actions, chips/emcn

* removed color migration

* Table bump

* improve(blocks): audit block catalog metadata for accuracy and fill gaps

- Fix template prompts that claimed capabilities blocks don't have
  (fabricated triggers, non-existent tools) across ~98 integrations
- Normalize integration tags to family conventions and valid union values
- Align alsoIntegrations and modules with template prompt content
- Add BlockMeta for circleback, imap, and rss trigger integrations
- Add templates to clickhouse and greptile metas
- Remove duplicate PageSpeed deploy-gate template

* chore(telemetry): drop org_seat.drift_reconciled audit

System self-heal bookkeeping doesn't belong in the user-facing audit
trail — membership and seat-purchase changes are already audited, and
the cron's logger output covers ops visibility.

* chore(db): consolidate branch migrations into single 0227

* fix(emcn): make ChipModal scroll internally when content exceeds viewport

The Modal→ChipModal migration dropped the old ModalBody scroll container:
ChipModal renders ModalContent bare (no overflow-hidden) and its wrappers
had no min-h-0 chain, so tall modals (e.g. New data drain with an S3
destination) overflowed max-h-[84vh] off-screen with no way to scroll.

Complete the flex min-h-0 chain through ChipModal's frame and give
ChipModalBody flex-1 min-h-0 overflow-y-auto — header/footer stay pinned,
body scrolls only when constrained. Short modals are unaffected (max-h
caps, it doesn't stretch), and dropdowns inside the body are Radix-portaled
so the new scroll container cannot clip them. Five modals that had locally
patched this with max-h-[Nvh] overrides keep working unchanged.

* fix(emcn): don't close modal when dismissing a dropdown via outside click

Radix dispatches pointer-down-outside to every open dismissable layer at
once, so clicking outside an open dropdown/select inside a modal closed
both the dropdown and the modal in one jarring step. ModalContent now
prevents its own dismissal while a portaled popper layer is open — the
first outside click closes just the popper, the next one closes the
modal.

* docs(skills): de-duplicate and correct agent docs; canonical styling tokens; mirror sim-sandbox rule

* docs(skills): broaden boundary-raw-fetch scope note, sync cursor rule mirrors

* fix(emcn): harden modal popper guard and exempt caret-anchored dropdowns from body scroll

Adversarial review follow-ups to the ChipModal scroll + dismissal fixes:

- Popper guard now requires data-state="open" inside the popper wrapper,
  so a dropdown that is merely animating closed no longer swallows the
  next outside click on the modal (DropdownMenu has an exit animation
  that keeps its wrapper mounted briefly)
- Port the same guard to SModalContent for consistency
- custom-tool-modal: opt its body out of the chrome scroll container
  (flex-none + overflow-visible); the caret-anchored EnvVar/Tag
  autocomplete dropdowns are absolute-positioned inside the body and
  must spill past its bounds rather than clip against a scroll boundary

* refactor(billing): drop unnecessary useCallback wrappers from event handlers

* fix(data-drains): complete chip migration of destination forms and fill missing placeholders

Finishes the in-flight FormField → ChipModalField conversion for all
destination form specs and adds the placeholders that several inputs
never had (S3 bucket/region/access keys, Azure account key, Datadog API
key, webhook signing secret/bearer token) — the cause of the New Data
Drain modal showing placeholder-less inputs inconsistently.

* fix(emcn): broaden modal popper guard to onInteractOutside

Covers the focusOutside dismissal path too: when a popper's focus scope
unwinds on close, the transient focus shift could still dismiss the
modal (and simultaneous body pointer-events lock teardown could freeze
the page). Same data-state="open" scoping as the pointer guard.

* fix

* fix(emcn): harden modal outside interactions

* Remove migration 0227 to regenerate after merge

Drop the 0227_org_member_usage_limit migration, its snapshot, and journal
entry so it can be regenerated cleanly after merging improvement/platform.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Regenerate migration as 0228 after merging improvement/platform

drizzle-kit generate re-emits the org_member_usage_limit table, document
uploaded_by, and table_run_dispatches triggered_by_user_id changes on top of
improvement/platform's 0227 snapshot. Also collapse the admin.tsx emcn import
left multi-line after merge conflict resolution.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix pre-existing type errors surfaced by post-merge type-check

These were present on dev before the merge; the post-merge type-check made
them visible.

- workspace-vfs.ts: materializeEnvironment now emits full oauth integration
  objects (id, providerId, displayName, role) so the summary matches
  WorkspaceMdData['oauthIntegrations'] (consumer reads credentialId from id)
- workflows/utils.ts: narrow resolved workflow rows to non-null workspaceId
  (the inArray filter already guarantees it) so the resolved result type holds

Co-authored-by: Cursor <cursoragent@cursor.com>

* Slack trigger

* fix tool mark complete bugs

* change thinking icon to just be blimp

* fix tests

* MCP fixes

* Fix files

* refactor use chat

* multiple fixes

* Fix deploys

* fix billing, tracing, retry lifecycle issues

* remove stale generated migration

Drop migration 0228 and its metadata so it can be regenerated after merging staging.

Co-authored-by: Cursor <cursoragent@cursor.com>

* merge staging in

* fix sockets err classification

* sockets invite flow fix

* remove stale generated migration 0229 before staging merge

Co-authored-by: Cursor <cursoragent@cursor.com>

* regen migration

* Update bun lock

* Fix schema issues

* Fix lint

* Fix build

---------

Co-authored-by: Vikhyath Mondreti <vikhyathvikku@gmail.com>
Co-authored-by: Theodore Li <theo@sim.ai>
Co-authored-by: Waleed <walif6@gmail.com>
Co-authored-by: Emir Karabeg <emirkarabeg@berkeley.edu>
Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai>
Co-authored-by: andres <k62hc5kjst@privaterelay.appleid.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: andresdjasso <andresdjasso@users.noreply.github.com>
Co-authored-by: waleed <waleed@simstudio.ai>
2026-06-09 13:39:51 -07:00
Vikhyath Mondreti 1933fc4f11 fix(env): schema treatment of empty string (#4862) 2026-06-03 10:21:47 -07:00
Waleed a7b0bd311d fix(deps): upgrade vitest to ^4.1.0 to patch critical Vitest UI advisory (GHSA-5xrq-8626-4rwp) (#4837)
* fix(deps): upgrade vitest to ^4.1.0 to patch critical Vitest UI advisory (GHSA-5xrq-8626-4rwp)

- Bump vitest and @vitest/coverage-v8 to ^4.1.0 across all workspaces (only patched release for the critical 'Vitest UI server arbitrary file read/execute' advisory; no 3.x backport exists)
- Widen @sim/testing peer range to ^3.0.0 || ^4.0.0
- Migrate constructor mocks to class expressions: vitest 4 uses Reflect.construct for mocks invoked with new, and arrow/function implementations are not constructable (function expressions also get reverted to arrows by biome's useArrowFunction)
- Remove deprecated test.poolOptions from apps/sim/vitest.config.ts (options are now top-level in vitest 4)

* fix(deps): exclude vulnerable vitest 4.0.x from @sim/testing peer range

Tighten the v4 arm of the peer range to >=4.1.0 <5.0.0 so the peer
requirement cannot be satisfied by the unpatched 4.0.x builds that
GHSA-5xrq-8626-4rwp affects.

* fix(testing): make vitest 4 constructor mocks type-check cleanly

- logging-session & mcp-oauth mocks: a class passed to mockImplementation has
  a construct signature that isn't assignable to its (...args) => any parameter,
  failing tsc. Use named function declarations instead (constructable via
  Reflect.construct, assignable to mockImplementation, and not rewritten to
  arrows by biome's useArrowFunction).
- database.mock.ts: vitest 4's generic vi.fn typings no longer break the
  self-referential cycle on the transaction callback's tx param; loosen tx and
  annotate the callback's return type to resolve the implicit-any errors.

* test(isolated-vm): de-flake queue-capacity scheduler tests

The 'queue is full' and 'per-owner queued limit' tests relied on
'await sleep(1)' to assume the first request had reached the queue before
submitting the overflow request. The first request only enqueues after an
async spawn-failure chain (acquireWorker -> spawn exit -> resolve null ->
enqueue), which isn't guaranteed within 1ms under CI load — the overflow
request then found an empty queue and hit the 200ms queue-wait timeout
instead of the capacity rejection.

Replace the wall-clock barrier with a deterministic, event-driven one: hold
the single global concurrency slot (IVM_MAX_CONCURRENT=1) with an active
worker and await an explicit 'dispatched' signal (fired when the worker
receives its execute message, after the scheduler counts it active). The
follow-up requests then deterministically hit the synchronous enqueue path.
Also drops the queue-wait timeout from 200ms to 50ms, so the tests run faster.
2026-06-01 16:11:35 -07:00
Waleed 8d7bbbc670 chore(utils): migrate to shared random/ID utilities and add enforcement linting (#4623)
* chore(utils): migrate to shared random/ID utilities and add enforcement linting

- Replace all Math.random(), crypto.randomUUID(), crypto.randomBytes(), nanoid, and uuid usages with shared @sim/utils/random and @sim/utils/id helpers across 72 files
- Add new @sim/utils exports: deepClone, omit, filterUndefined (object), truncate (string), backoffWithJitter, parseRetryAfter (retry), getErrorMessage (errors)
- Sweep all getErrorMessage, sleep, deepClone callsites across 500+ files to use shared utilities
- Add Biome noRestrictedImports rule to catch nanoid, uuid, and crypto named imports at lint time
- Add scripts/check-utils-enforcement.ts to catch Math.random and crypto.* global property access
- Add check:utils script to package.json

* chore(utils): replace deepClone wrapper with structuredClone built-in

deepClone() was a one-line wrapper around structuredClone(), which is
universally available in Node 17+ and all modern browsers. Removing the
abstraction reduces indirection and means contributors don't need to
learn a project-specific name for a well-known built-in.

- Remove deepClone from packages/utils/src/object.ts and index.ts
- Replace all 17 call sites with structuredClone() directly
- Update check:utils script suggestion text
- Update CLAUDE.md and global.md docs

* fix(utils): add missing biome noRestrictedImports rule and correct truncate docs

- Add noRestrictedImports to biome.json under style — bans nanoid and uuid
  package imports at lint time (crypto.randomUUID/randomBytes are caught by
  the check:utils grep script which handles global property access)
- Correct truncate() TSDoc and parameter name: sliceLength makes it clear
  that total output length is sliceLength + suffix.length, matching the
  behavior all callers were already written to expect

* fix(utils): add missing getErrorMessage imports at 4 call sites

The sweep agents added getErrorMessage calls without the corresponding
import in 4 files, causing test failures. Added the missing imports.

* fix(utils): fix build errors from getErrorMessage sweep and retry.ts Turbopack issue

- Fix retry.ts cross-file import: Turbopack cannot resolve './random.js' for
  internal package imports; inline the jitter crypto call directly
- Add missing getErrorMessage imports to 32 files where the sweep added calls
  without the corresponding import (caught by type-check and test runs)
- Remove accidental getErrorMessage import from crowdstrike/query/route.ts
  which has its own domain-specific getErrorMessage for parsing CrowdStrike's
  JSON error format
- Fix use-sub-block-value.ts type error from structuredClone narrowing:
  add 'as T' cast at emitValue callsite (safe — valueCopy is always a
  structural copy of newValue)

* fix(tools): use toError in crowdstrike catch block instead of local getErrorMessage

The catch block was calling the local getErrorMessage function which
parses CrowdStrike API JSON responses, not JavaScript Error objects.
Use toError(error).message to correctly extract the message from a
caught value in this context.
2026-05-15 17:31:27 -07:00
Waleed b5dba82ac9 improvement(db): reduce connection saturation and egress hotspots (#4594)
* improvement(db): reduce connection saturation and egress hotspots

* fix(vfs): preserve native content type in copilot SQL projection

* fix(vfs): guard jsonb_array_elements against non-array contentBlocks
2026-05-13 23:39:59 -07:00
Vikhyath Mondreti 43f53bb7d3 feat(execution): payload size bottlenecks with lazy execution value hydration, safer materialization, and batched parallel execution (#4560)
* improvement(resolver): lazy resolution for underlying fields greater than 10MB

* progress

* feat(parallel): batching

* codegen to allow inline substitution

* address comments

* ui inconsistencies

* cleanup redundant code

* address more comments

* address comments

* replace helper

* fix tests
2026-05-12 14:57:32 -07:00
Vikhyath Mondreti 6cb779601a feat(search-replace): search & replace, cut, deploy modal ui flicker (#4507)
* feat(search): workflow search and replace

* fix alignment

* fix hidden fields bug

* fix loops/parallel badge case

* resource resolver

* add cut

* update docs

* address comments

* make source code for func blocks dispay resolved code instead

* fix match issue

* fix padding
2026-05-07 20:34:53 -07:00
Theodore LiandClaude Opus 4.7 98f8e854eb improvement(tables): extract TablesDetail wrapper, ship trigger followups (#4476)
* ui improvements

* Update status pils, make checkbox column sticky

* add Run workflow to context menu

* Refactor dispatching logic

* fix checkbox width to be smaller if csv is small

* Add drag behavior for workflows, stop workflow on multi select

* fix z index of checkbox to left, add view workflow button

* Switch to emcn buttons for Add inputs

* Split up workflow sidebar from column sidebar, refactor cells

* Lint and add auto run toggle

* fix column reordering, add action bar

* Create and use emcn square

* Reconcile post-merge: drop positionMap, use rowId-based selection

Staging refactored Tables UI to decouple from DB position (gutter from
array index, checkedRows keyed by rowId, no PositionGapRows). Bring
HEAD's action-bar / context-menu helpers in line: contextMenuRowIds,
selectedRowIds, actionBarRowIds now key off row.id and walk `rows`
directly. Drop the maxPosition / positionMap derived state. Collapse
COLUMN_SIDEBAR_WIDTH_CSS to a numeric COLUMN_SIDEBAR_WIDTH used by both
the sidebar shell and the table's reserved padding-right.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(table): backfill remapped workflow outputs from execution logs

When a workflow column is re-pointed to a different (blockId, path),
populate its existing rows with the new output's value pulled from saved
execution logs instead of leaving them empty until the next run. Rows
where the new mapping has no logged value clear (matching the previous
behavior for those rows), but rows where the workflow already has the
new output's value surface immediately.

Refactor backfillAddedGroupOutputs into a generalized
backfillGroupOutputsFromLogs helper with an `overwrite` flag — used in
both the added-outputs path (preserves hand-edited values) and the new
remapped path (overwrites since the new mapping is the source of truth).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): update column type when remapping workflow output

A remap that changes the output's leaf type (string → number, json →
boolean, etc.) was leaving the column's declared type stale. The clear-
then-backfill flow then failed schema validation on every row, so the
backfill silently aborted and the column stayed empty.

Resolve the new leaf type via flattenWorkflowOutputs +
columnTypeForLeaf for each mappingUpdate, and patch
schema.columns[i].type before the schema write. The clear-tx then
backfill ordering now works end-to-end across type changes. If the
workflow or its target output can't be resolved (workflow deleted,
block removed), fall back to leaving the column type alone — the
backfill will skip rows whose picked value doesn't match, same as
before.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): stringify objects instead of "[object Object]" in cells

If a column's declared type lags its row data (e.g. a workflow column
mid-remap, where the schema cache hasn't refetched yet but the row data
already has the new mapping's value), formatValueForInput and the
cell-render text variant fell through to String(value) and rendered
"[object Object]". JSON-stringify objects in both spots so the transient
skew shows the actual data.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): drop extra left border on workflow group meta header

The meta cell had border-r/b/l while regular headers have only border-r/b.
With border-separate tables, that extra 1px left border shifted the
meta cell's content one pixel right of the columns below it.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): align workflow meta header without dropping its left border

Restore border-l and pull the cell back -1px with -ml-px so the visible
left border overlaps the previous cell's right border instead of adding
1px to the meta cell's box. Content lines up with the columns below.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): draw meta header left border via ::before pseudo

Adding border-l to the meta cell shifted its content right by 1px
because table-fixed + border-separate honors the border inside the
colspan'd cell's width budget. -ml-px doesn't work on <th>. Render the
visible left edge via a ::before at left: -1px instead — paints over
the prior cell's right border without consuming any of the meta cell's
content area. Content lines up with the columns below.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(tables): add missing barrels, drop doubled-path imports

Match the convention used by logs/components: every component folder
exposes its public API via index.ts so consumers import from the folder
name, not from its internal filenames.

- New barrels: column-config-sidebar/, workflow-sidebar/,
  table-action-bar/, table/cells/, table/headers/.
- Rename table-filter/index.tsx → index.ts (barrel is not a component).
- Top-level components/index.ts re-exports every sibling folder so
  external consumers have one import path.
- Replace `from '../foo/foo'` doubled paths in table.tsx with the
  shorter barrel-anchored form.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(tables): introduce TablesDetail wrapper as thin passthrough

Phase 1 step 0 of the wrapper extraction (see plan
okay-lets-make-a-shimmying-trinket.md). page.tsx now renders
TablesDetail, which today is a passthrough to <Table>. Subsequent
commits lift surface state out of <Table> into this wrapper one piece
at a time.

The mothership chat path (<Table embedded>) is untouched — <Table>
stays exportable as a lower-level component for embedded contexts.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(tables): lift slideout panel state into TablesDetail wrapper

The three right-edge slideout panels (column config, workflow config,
execution details) move out of <Table> into the wrapper. The wrapper
owns a single useReducer that encodes the at-most-one-open invariant
as a discriminated union — opening any one panel automatically closes
the others. <Table> emits open requests via three new callback props.

Also extract <ExecutionDetailsSidebar> from inline-in-table.tsx to its
own folder so the wrapper can compose it cleanly. Update the embedded
mothership callsite (resource-content.tsx) to render <TablesDetail
embedded> instead of <Table embedded>.

Phase 1 step 1 of the wrapper extraction. <Table> shrinks from 3849 →
3787 lines; <TablesDetail> grows from 19 → 145 lines.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(tables): lift delete-table modal + mutation into wrapper

The delete-table confirmation modal and `useDeleteTable` mutation move
out of <Table> into TablesDetail. <Table> exposes a new
`onRequestDeleteTable` callback fired by the page-header Delete action.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(tables): lift CSV import dialog into wrapper

ImportCsvDialog moves out of <Table>. Grid exposes
`onRequestImportCsv` fired by the page-header menu item; wrapper owns
the open state and renders the dialog.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(tables): lift RowModal (edit + delete) into wrapper

Both RowModal instances move out of <Table> into the wrapper. Grid
emits `onOpenRowModal(row)` (Space key) and
`onRequestDeleteRows(snapshots)` (context menu).

Post-delete cleanup (push undo, clear selection) needs grid-internal
state, so the grid populates an `afterDeleteRowsSinkRef` callback that
the wrapper's modal `onSuccess` invokes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(tables): lift delete-columns modal into wrapper

The destructive delete-columns confirmation modal moves into the
wrapper. Grid emits `onRequestDeleteColumns(names)`; the cascade itself
(per-column mutation, undo push, columnOrder + columnWidths cleanup)
stays in the grid as a sink the wrapper invokes on confirm — too
grid-internal to lift cleanly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(tables): lift run/stop mutations + TableActionBar to wrapper

useRunGroup and useCancelTableRuns move out of <Table> into the
wrapper, along with the <TableActionBar> render. Grid receives
onRunGroup, onRunRows, onStopRow, onStopRows, onStopAll, and
cancelRunsPending as props — used by the per-row gutter Play/Stop, the
workflow-group meta-cell run menu, and the right-click context menu's
Run/Stop on selection items.

Action-bar selection state (actionBarRowIds, runningInActionBar,
hasWorkflowColumns) is derived from grid-internal state, so the grid
emits a `SelectionSnapshot` via `onSelectionChange` from a useEffect.
Wrapper uses the snapshot to drive the floating <TableActionBar>.

Phase 2 step 1 of the wrapper extraction.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(tables): lift queryOptions to wrapper

queryOptions (filter + sort) moves out of <Table> into the wrapper,
making it a single source of truth that drives one useTable call. The
wrapper passes the bundle down to the grid; sort/filter handlers in
the grid call onQueryOptionsChange.

Eliminates the previous double-useTable pattern (one for the grid's
filtered/sorted view, one in the wrapper's hardcoded null/null query
for sidebar metadata).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(tables): lift page header (breadcrumbs/options/filter) to wrapper

Phase 3 of the wrapper extraction. The full page-header surface moves
out of <Table>:

- ResourceHeader (breadcrumbs, table-rename UI, headerActions, createTrigger)
- ResourceOptionsBar (sort + filter toggle)
- TableFilter (filter panel — wrapper owns filterOpen state)
- RunStatusControl (in the leading actions when runs are active)

useRenameTable + useInlineRename for the breadcrumb name move to the
wrapper. The grid populates pushTableRenameUndoSinkRef so the rename is
still part of the grid's undo stack.

Extract NewColumnDropdown and RunStatusControl from inline-in-table.tsx
to their own folders so the wrapper composes them cleanly without
reaching into the grid's internals.

Hoist generateColumnName from grid-internal useCallback to a shared util
so both the page-header and inline-header NewColumnDropdowns use the
same logic.

After this lift <Table> is the data grid only — no page surface, no
modals, no slideouts, no breadcrumbs. The selection snapshot now
includes totalRunning so the wrapper can render the page-header
RunStatusControl from outside the grid.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(tables): cleanup pass on TablesDetail wrapper extraction

Six-pass cleanup against the wrapper extraction diff:

- Effects: add content-compare bailout to onSelectionChange emit so
  unchanged snapshots don't churn wrapper re-renders.
- Memos: drop unnecessary activeSortState memo, fold into sortConfig.
- Callbacks: remove ~10 useCallbacks with no observed reference (sidebars
  not memoized, modals not memoized, inline arrows on non-memoized
  children); keep the ones that feed into <DataRow>/<RunStatusControl>/
  <ResourceHeader> (memoized) or grid-side useCallback deps.
- Dead props: drop onQueryOptionsChange/onRequestDeleteTable/
  onRequestImportCsv from <Table> — the page-header lift made them
  unused but the props weren't removed.
- React Query: drop redundant tableWorkflowGroupsRef (created when
  onRunRows was useCallback-wrapped; after callback cleanup it can read
  the query data directly).
- emcn: normalize Loader sizing to h-[14px] w-[14px] to match the
  codebase convention.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): re-seed columnOrder when columns change server-side

The metadata-seed effect short-circuited after the first seed, so any
later schema change (e.g. adding a workflow output column) couldn't
push the new column into local columnOrder. The new column would then
fall into the "remaining" bucket of `displayColumns` and render at the
end of the table — until the user refreshed and the grid re-mounted
with the now-current metadata.

Drop the `metadataSeededRef.current` short-circuit from the early
return so the effect can also reach the after-first-load re-seed
branch, which already does the right thing (only re-seeds when the set
of columns changes, leaves pure-reorder cases alone).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(tables): rename wrapper to <Table>, grid to <TableGrid>

Match the naming convention used elsewhere in the workspace
(workflow.tsx → <Workflow>, base.tsx → <Base>, logs.tsx → <Logs>).

- tables-detail.tsx → table.tsx (exports <Table>)
- components/table/ → components/table-grid/ (exports <TableGrid>)
- components/table-grid/table.tsx → table-grid.tsx
- Drop <ExecutionDetailsSidebar> — was a 3-line passthrough
  (executionId → useLogByExecutionId → <LogDetails>); inline directly
  into table.tsx where it's used.
- Flatten components/run-status-control/ folder to a single
  components/run-status-control.tsx file. 25-line single-use component
  with no internal subdirs — folder was overhead. Matches knowledge's
  max-badge.tsx precedent.

Net: 1 wrapper rename + grid rename + 2 folder collapses, all imports
updated. The mothership chat callsite updates from <TablesDetail
embedded> to <Table embedded>.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): don't show "Waiting" for autoRun=false workflow groups

A workflow group with autoRun=false never fires from the scheduler —
the cell stays empty until the user clicks Run manually. Treating
empty cells as "Waiting" misleads the user into thinking the group
will auto-fire once deps are filled, which it won't.

Skip autoRun=false groups when computing the per-row waiting labels
so their cells render the empty-dash instead of the Waiting pill.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(copilot): regenerate tool catalog from copilot dev (#247)

Pulls in the workflow_group operations on user_table:
add_workflow_group / update_workflow_group / delete_workflow_group /
add_workflow_group_output / delete_workflow_group_output /
run_workflow_group, plus the autoRun / blockId / dependencies / groupId
parameters and a tightened mapping description for import_file.

Also picks up biome import-order fixes from `bun run lint`.

* improvement(table): action bar in mothership + per-execution mode

Three related improvements to the table action bar:

1. Reposition from `position: fixed` to `position: absolute` inside
   the table's container. Fixed-positioning anchored to the viewport,
   which centered the bar across the whole window instead of the table
   panel — wrong in mothership embedded view, where the table sits in
   the right half. Absolute scopes the bar to the table's bounds.

2. Show the bar for single-execution highlights — when the user
   selects one workflow-output cell, or 1 row × N cols all within the
   same workflow group. The bar enters per-execution mode with Run /
   Stop / View execution buttons targeting that one cell or group.

3. Skip View execution for cancelled cells. A cancelled cell may have
   been cancelled before the worker ever picked the job up, so its
   executionId can't be relied on. Tighten the gate everywhere
   (context menu + action bar) to only `completed` / `error` / `running`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): backfill on add_workflow_group_output, don't re-run

addWorkflowGroupOutput (the one-shot single-output add path used by
the copilot user_table tool) was calling triggerWorkflowGroupRun({
mode: 'all' }) after appending the output — that re-fired the workflow
on every row. Trace a307ed8fd5fe2d931aa84dedab5a60f0 shows ~75
workflow-group-cell jobs enqueued in the seconds after a single
add_workflow_group_output call.

Replace with backfillGroupOutputsFromLogs (overwrite: false), the same
flow updateWorkflowGroup uses when receiving newOutputColumns. Reads
each row's saved trace spans and writes the new output's value back —
no compute beyond a JSONB write per row, no double-billing the user
for runs they didn't ask for.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): drop sql.raw quote-escaping in column-name interpolation

Six call sites in lib/table/service.ts built JSON-key string literals
at runtime via `sql.raw(\`'\${name.replace(/'/g, "''")}'\`)` for use
with PostgreSQL's `data->'key'` / `data->>'key'` operators. Practically
safe (NAME_PATTERN gates column names to alphanumeric+underscore at
insert time) but a smelly pattern that breaks the moment validation
loosens.

Both `data->` and `data->>` accept a parameterized text value as the
key, so the `sql.raw` is unnecessary. Replace each with a normal
`${name}::text` binding. No behavior change; eliminates the manual
quote-escaping surface.

Affected sites: renameColumn (the data-rewrite UPDATE), upsertRow's
match filter, updateColumnType's IS-NOT-NULL gate, updateColumnConstraints'
required-check + unique-duplicate-check.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(copilot-tool): forward autoRun + mappingUpdates on update_workflow_group

The sim-side service and contracts already accept both fields, but the
copilot tool's update_workflow_group handler was dropping them on the
floor. Now `args.autoRun` (toggle the persisted auto-fire flag) and
`args.mappingUpdates` (per-output (blockId, path) swap) get forwarded
through to updateWorkflowGroup.

Pairs with the upcoming copilot-side change that exposes these in the
tool catalog JSON / Go handler / prompting (see copilot branch
redo-workflow-tools).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): keep gutter border visible when hovering Run-row button

The per-row Run button sat flush against the row-gutter cell's right
border. Its hover background (rounded-rect surface-2) painted over the
border line for the 20px height of the button, making the gutter
divider appear to disappear at the hovered row.

Add mr-px to the button so the hover bg stops 1px short of the cell's
right edge, leaving the divider intact.

* fix(table): unify auto-fire and manual run paths in scheduler

scheduleWorkflowGroupRuns now owns eligibility, autoRun semantics,
dep evaluation, and enqueue for both paths. Auto-fire callers omit
opts; manual callers (triggerWorkflowGroupRun) pass { groupId,
isManualRun: true } to bypass the autoRun=false skip and (for
autoRun=false groups) the dep check.

Per-row /run-workflow-group route delegates to triggerWorkflowGroupRun
with rowIds=[rowId]. Single server-side path for both manual entry
points.

Also: optimisticallyScheduleNewlyEligibleGroups skips autoRun=false
groups so editing a row's data doesn't phantom-mark autoRun=false
output cells as Queued.

* fix(table): render empty cells as blank, not em-dash

Empty cells (any column type) showed an em-dash placeholder. Drop it
so empty cells render blank — matches what the user expects when
nothing's there.

* fix(table): per-row Run fires autoRun=false groups regardless of deps

handleRunRow filtered out every group whose deps weren't satisfied,
which silently dropped autoRun=false groups (since their deps usually
aren't satisfied — that's the whole point of autoRun=false). Click
Run row, the autoRun=false group's cells stayed empty.

Mirror the scheduler's semantics: autoRun=false bypasses the dep check,
autoRun=true still requires deps.

* fix ui shape

* improvement(table): collapse run ops into run_column, derive action-bar buttons from selection

The action bar now reflects what's actually selected:
- Selection-driven scope (cells the user highlighted, not their full rows)
- Play visible when there's anything empty/failed; Refresh when there's anything completed; both for mixed
- run_cell / run_row deleted; everything funnels through run_column
- Per-row gutter Play, right-click "Run workflows on N rows", and column-header menu all share the canonical run path
- Shared RunMode type from the contract; cleanup pass via /simplify (readExecution / isExecInFlight reuse, runScope helper, flat onViewExecution prop)

* chore(copilot): regen tool catalog after dropping run_cell / run_row + dependencies.workflowGroups

Mirror the copilot-side catalog change so the generated TS catalog matches the deployed copilot tool surface.

* fix(table): atomic per-key writes for executions, plus run-op race fixes

The executions blob on user_table_rows was read-modify-written wholesale on every
update. Concurrent writers (a column edit and a manual-retry stamp, two pickup
calls, a cancel and a cascade) each computed a merge from their own snapshot,
and the last writer clobbered keys it never touched — producing stuck "queued"
cells, vanished stamps, and stale completed exec records reappearing after
retries.

Fixes:
- updateRow / batchUpdateRows now apply executionsPatch via a SQL jsonb merge
  expression. Each writer only mutates the keys it explicitly patches; other
  keys are preserved. Eliminates the cross-key clobber.
- writeWorkflowGroupState bypasses the stale-worker guard for `queued` (new
  scheduler stamp) and `cancelled` (authoritative cancel) writes — those ARE
  the new authority for the cell. Previously the new run's stamp was being
  rejected by the same guard meant to block the OLD worker's writes.
- skipScheduler flag on UpdateRowData / BatchUpdateByIdData lets the cancel
  path and runWorkflowGroupsInternal opt out of the implicit auto-fire pass
  (cancel was waking up siblings; manual-run was racing its own scheduler).
- CELL_CONTENT pinned to h-[22px] so status badges don't grow rows.

* chore(table): remove table-row sockets, both sides

Tables don't use realtime sockets in prod — strip the dead path so we stop
paying the per-row HTTP forward + socket emit on every cell write. Polling on
running execs already covers reconciliation.

Sim side:
- service.ts: drop notifyTableRowUpdated/Deleted, notifyTableDeleted, the
  postRealtimeBridge helper, and all callsites.
- hooks/queries/tables.ts: drop the socket subscription block in useTableRows;
  poll-on-running stays. Remove useEffect / useSocket imports.
- app/.../tables/[tableId]/hooks/use-table.ts: drop the merge-on-event
  useEffect and unused imports.
- app/workspace/providers/socket-provider.tsx: drop joinTable/leaveTable,
  onTableRowUpdated/Deleted/onTableDeleted, currentTableId state, related
  events + types.

Realtime side:
- handlers/tables.ts deleted; index.ts no longer wires it.
- routes/http.ts: drop /api/table-row-updated, /api/table-row-deleted,
  /api/table-deleted endpoints.
- rooms/{memory,redis}-manager.ts: drop emitToTable, handleTableRowUpdated/
  Deleted, handleTableDeleted, related imports.
- rooms/types.ts: drop method declarations, TableRowUpdatedPayload type,
  tableRoomName helper.
- middleware/permissions.ts: drop unused verifyTableAccess.

Bonus from parallel work:
- cell-content typewriter trigger refinement.

* fix(table): clearing a workflow output cell also clears its exec record

When the user wipes a workflow output column value, the auto-fire reactor
needs to be re-armed for that group. Previously, a stale cancelled / error
exec record blocked the eligibility predicate (gate at line 79 hard-rejects
those statuses on auto-fire) and the cell stayed stuck in its old terminal
state — visible as "Cancelled" cells that wouldn't re-run no matter what.

Both updateRow and batchUpdateRows now derive an `executionsPatch[gid] = null`
for any output column the patch sets to empty. The data clear and the exec
clear ride the same SQL transaction, so the row never lands in a stale-
status-with-empty-data state.

Symmetric to how `completed` already worked via `areOutputsFilled` in the
predicate — clearing the cell wins over the prior exec status, regardless of
what that status was.

(Also revert typewriter-trigger experiment from a parallel session that was
in-progress on this branch.)

* fix(table): waiting state, optimistic UX, schema-mutation polling, exec cleanup

A bundle of small UX + correctness fixes around workflow-cell run state.

cell-render.tsx
- In-flight (queued/running/pending) now wins over the existing value, so
  re-runs surface immediately instead of looking like nothing happened until
  the worker writes the new value.
- "Waiting on X" wins over a stale `cancelled` / `error` exec when deps are
  unmet — clearing a dep now reads as actionable instead of stuck.

useRunColumn (hooks/queries/tables.ts)
- onSettled now cancels in-flight polls before invalidating. Stops a poll
  that landed mid-mutation from clobbering the optimistic state with stale
  data, which produced the queued → cancelled → queued flicker.

addWorkflowGroup / updateWorkflowGroup (autoRun toggle on)
- Awaits scheduleRunsForTable instead of fire-and-forget. The route returned
  before the queued exec stamps committed, so the post-mutation refetch saw
  no in-flight cells and polling never started — cells looked stuck even
  though the server eventually stamped them.

deleteColumn / deleteColumns
- Strip orphaned executions[gid] keys when deleting a column orphans its
  parent group. Without this, stale running/queued exec records lingered on
  every row forever and inflated the page-header "N running" counter even
  on tables with no actually-running cells.

UI
- Action-bar leading label: "Selected N workflow cell(s)".
- Context menu: Run / Refresh items mirror the action bar's Play / Refresh
  split, gated on the same selection-status flags so both surfaces show the
  actions that match the current state.

* refactor(table): consolidate exec-status helpers + fix N-running counter

Cleanup pass on the recent table changes — pulls duplicated predicates and
SQL snippets into shared helpers and fixes one drift bug along the way.

- isExecInFlight: now single export from lib/table/deps.ts. Removed the
  duplicate in components/table-grid/utils.ts. Used by isGroupEligible
  (server eligibility) and runningByRowId (client counter).
- isOptimisticInFlight: kept local to hooks/queries/tables.ts — renamed from
  isInFlight to disambiguate from the stricter isExecInFlight. The two
  predicates differ on `pending` without a jobId: optimistic patches and
  poll-trigger want the broader version, eligibility wants the strict one.
- areOutputsFilled: single export from lib/table/deps.ts, dropped duplicate
  from workflow-columns.ts.
- classifyExecStatusMix: shared row × group walker in table-grid/utils.ts.
  Replaces two copies of the same loop in table-grid.tsx (selectionStats +
  contextMenuStats). Both surfaces now have the same short-circuit
  semantics, including the seen-all-selected-rows early break that
  contextMenuStats was missing.
- stripGroupExecutions: SQL helper in service.ts. Replaces three copies of
  the `UPDATE user_table_rows SET executions = executions - $gid::text`
  pattern across deleteColumn / deleteColumns / deleteWorkflowGroup.

Drift bug:
- runningByRowId / totalRunning counted only `running` and `queued`. Every
  other in-flight check in the codebase treats post-stamp `pending` as
  in-flight too, so the page-header "N running" badge briefly dropped to 0
  between scheduler stamp and worker pickup. Now uses isExecInFlight.

* fix(table): address pr review (drop dead workflowNameById prop, reset didDragRef on dragend, align sidebar width)

* fix(table): scope post-clear schedule to targeted groups, forward mode

Multi-group manual runs (Run row, gutter Play, action-bar Play across mixed
completed + cancelled cells) re-fired completed-and-filled siblings.
runWorkflowGroupsInternal cleared only the groups it filtered, then called
scheduleRunsForRows with isManualRun: true and no group / mode filter — so
the post-clear pass walked every group on the table with default mode 'all',
and any autoRun=true completed sibling whose deps were satisfied got queued
again. Scope the post-clear call to targetGroups and forward mode.

* fix(table): meta-cell drag-leave flicker guard + plumb unique on create

* fix(table): strip sibling deps when removing workflow output via updateWorkflowGroup

deleteWorkflowGroup already stripped removed-column deps from sibling
groups, but updateWorkflowGroup (the path the UI takes when deleting one
output of a multi-output group) didn't — schema validation then rejected
the update with 'Group X depends on missing column Y'.

* improvement(table): debug logs at every cascade decision branch

* improvement(table): parallelize queued-stamp writes within concurrency-cap chunks

* Simplify stripping column names

* fix lint, ci

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 15:43:05 -04:00
Theodore Li 31cfb74dc2 feat(table): add workflow execution column type (#4338)
* Add table triggers for columns and row added

* Add async batching job for running column

* Add ui improvements, stop mechanism

* Use trigger dev for workflow runs

* Use unified column sidebar for table

* Add socket waits for tables, multi column workflow support

* change back to cell based trigger jobs

* Reuse code, add view log inline in table view

* reorganize db to treat each column separately

* adjust column naming strategy

* Column ui improvements

* fix live update on table

* Change new column behavior

* fixed errored workfows not showing as stopped

* Column sidebar improvements

* fix table column swapping behavior

* fix bugs

* add prompting, fix lint

* fix ui stuff

* flip feature flag

* Add zod contracts fix initial auto-run of columns

* ui improvements

* Use db filter to query unran rows

* Use live workflow run

* Change wording for deleting workflow column

* Update tools

* Add tool to run selected rows

* Add mothership tools

* adjust col width

* Restore ff

* fix drizzle migration

* fix test
2026-05-02 15:48:12 -04:00
Vikhyath Mondreti af859cd508 feat(workflows): lock/duplicate improvements for workflows (#4387)
* feat(workflows): lock/duplicate improvements

* fix duplicate var remap bug

* address comments

* remove dead vars

* fix tests

* address comments

* code cleanup

* address comments

* address comments

* minor change

* remove dead code
2026-05-01 20:42:10 -07:00
Vikhyath Mondreti b8959eb20d improvement(repo): zod based client-server boundary (#4355)
* improvement(repo): centralized zod contracts (#4336)

* improvement(repo): zod schema contracts

* type checks

* fix(notion): correctly register tool (#4337)

* fix func blokc

* more improvements

* fix tests

* type check

* remove v3 refs

* minor type improvements

* address comments

* update jira contract

* remove validateJsonBody

* improvement(repo): consolidation of boundary helpers + better unknown usage (#4352)

* improvement(repo): consolidation of boundary helpers + better unknown usage

* address comments

* improve file transfer error messaging

* fix docs listing schema drift

* fix inocrrect type casting

* address council comments

* remove prefix
2026-04-30 12:16:28 -07:00
Waleed ccb5f1e690 fix(db): revert statement_timeout startup options breaking pooled connections (#4284) 2026-04-23 23:31:52 -07:00
Theodore Li c22ac38ab0 Set statement timeout of 90 seconds (#4276) 2026-04-23 18:27:15 -04:00
5f0f0edd63 improvement(repo): separate realtime into separate app (#4262)
* improvement(repo): restructuring to make realtime image narrower scoped

* improvements

* chore(repo): rebase fixes and quality improvements for realtime split

Addresses merge-time issues and gaps from the realtime app split:
- Retarget stale vi.mock paths to @sim/workflow-persistence/subblocks
- Restore README branding, fix AGENTS.md script reference
- Restore TSDoc on workflow-persistence subblocks helpers
- Use toError() from @sim/utils/errors in save.ts
- Add vitest config + local mocks so @sim/audit tests run standalone
- Move socket.io-client to devDependencies in apps/realtime
- Add missing package COPY steps to docker/app.Dockerfile
- Add check:boundaries/check:realtime-prune scripts and wire into CI

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(security): consolidate crypto primitives into @sim/security

Move general-purpose crypto primitives out of apps/sim into the
@sim/security package so both apps/sim and apps/realtime can share them.

@sim/security exports (all pure, dependency-free):
  ./compare    safeCompare (constant-time HMAC-wrapped equality)
  ./encryption encrypt/decrypt (AES-256-GCM, iv:cipher:tag format)
  ./hash       sha256Hex
  ./tokens     generateSecureToken (base64url)

Migrate apps/sim call sites to use these + @sim/utils helpers:
  crypto.randomUUID()            -> generateId() from @sim/utils/id
  createHash('sha256').digest    -> sha256Hex
  timingSafeEqual on hashed hex  -> safeCompare
  new Promise(setTimeout)        -> sleep from @sim/utils/helpers

No behavior change: encryption format, digest output, and token
length are preserved exactly.

* refactor(copilot): use toError in remaining otel/finalize sites

Replace the last two `error instanceof Error ? error : new Error(String(error))`
patterns with toError from @sim/utils/errors. Completes the sweep of clean
candidates — no behavior change.

* refactor(security): consolidate HMAC-SHA256 primitives into @sim/security

Adds hmacSha256Hex and hmacSha256Base64 to @sim/security/hmac and migrates
15 webhook providers plus 5 other hot paths (deployment token signing,
outbound webhook requests, workspace notification delivery, notification
test route, Shopify OAuth callback) off bare `createHmac` calls. Secret
parameter accepts `string | Buffer` to cover base64-decoded Svix-style
secrets (Resend) and MS Teams' HMAC scheme. AWS SigV4 signing in S3 and
Textract tools intentionally retains direct `createHmac` usage — its
multi-step key derivation chain doesn't fit a generic helper.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(packages): post-audit test + packaging polish

- Add safeCompare unit tests (identity, length mismatch, hex-nibble diff).
- Add Buffer-secret cases to hmac tests to lock in Svix/MS-Teams contract.
- Declare `reactflow` as a peerDependency on @sim/workflow-types — only used for type imports.
- Add a barrel export to @sim/workflow-persistence for consumers that prefer package-level imports; subpath exports retained.
- Document the data-field invariant in load.ts for loop/parallel subflow patching.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(realtime): address PR review feedback

- Remove redundant SOCKET_PORT=3002 env from Dockerfile runner stage
  (env.PORT already defaults to 3002 via zod schema).
- Reorder PORT fallback so an explicitly-set SOCKET_PORT wins over
  the schema default for PORT; keeps SOCKET_PORT functional as an
  override instead of dead code.
- Add dedicated type-check CI step for @sim/realtime so TS errors
  surface pre-deploy (the Dockerfile runs source TS via Bun and has
  no implicit build-time type check).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(realtime): remove unused SOCKET_PORT env var

SOCKET_PORT has lived in the socket server since the June 2025 refactor
but was never actually set in any deploy config — docker-compose.prod,
helm values/templates, .env.example, and docs all use PORT or the 3002
default exclusively. No self-hoster was ever pointed at SOCKET_PORT, so
removing it is safe.

Simplifies realtime port resolution to `env.PORT` (zod-validated with a
3002 default) and drops the orphaned sim-side schema entry.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Waleed Latif <walif6@gmail.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-04-22 23:06:16 -07:00