Commit Graph
5657 Commits
Author SHA1 Message Date
Theodore Li 0dcbc56ef6 feat(tables): per-table mutation locks (schema/insert/update/delete) (#5960)
* feat(tables): per-table mutation locks (schema/insert/update/delete)

Adds four independent, admin-only locks to a table so a workspace can make it
append-only, read-only, or schema-frozen. Enforced at the lib/table service
layer rather than the routes, so the API, workflow blocks, and Mothership (which
calls the services directly) are all covered by one assert. Violations return
423; changing a lock requires workspace admin and returns 403.

* fix(tables): address review — CI snapshot, lock-aware jobs, UI gating

- Add the drizzle meta snapshot the hand-written migration was missing, which
  made CI regenerate the lock columns as a phantom 0271
- Stop a queued or retried delete/update job that starts after its lock is
  enabled; a run that has already committed pages still finishes, since pages
  are never rolled back and aborting would leave an uncancelled half-done state
- Map TableLockedError to 423 on the internal single-row DELETE (was a 500)
- Gate the lock settings UI behind NEXT_PUBLIC_TABLE_LOCKS so it can't open
  onto a Save that 403s against the server-side flag
- Respect the delete lock on column drops in the grid, matching
  assertColumnDestructive, instead of failing only on click

* fix(tables): close lock gaps in async import and undo/redo

- Assert insert (and delete for replace) at the start of the import worker.
  Replace deleted every row before the first insert assert, so an insert-locked
  table was wiped and then failed, leaving it empty.
- Run the delete/update worker start checks through assertRowDelete /
  assertRowUpdate instead of reading the lock flags directly, so the worker and
  its enqueue site apply identical rules — including the workflow-column
  exemption a bulk update was being cancelled despite.
- Make undo/redo verb mapping direction-aware for column ops: dropping or
  retyping a column needs the delete lock clear too, matching
  assertColumnDestructive, so those steps are skipped instead of stranded.

* fix(tables): stop locks over-blocking workflow runs, duplicate, and imports

- Run / Re-run / Stop no longer inherit the cell-edit gate: they write only
  workflow-output columns (which the update lock exempts) and Stop is a cancel
- Duplicate needs only the insert lock — it inserts a full copied row, so it
  stays available on an append-only table
- updateWorkflowGroup asserts the destructive rule only when a patch actually
  drops or remaps output columns; rename / autoRun / mapping edits need just
  the schema lock
- Gate the import-async enqueue on insert (and delete for replace) so a locked
  table 423s instead of claiming the write-job slot and failing in the worker
- Disable Import CSV under an insert lock, and wire the expanded cell editor
  and the previously-unused stable add-row handler to the lock-aware flags

* fix(tables): keep column menu usable and let locks always be cleared

- Pass the blocked-delete handler instead of undefined so ColumnOptionsMenu
  still mounts on a delete-locked table; withholding it hid the entire menu,
  including Insert column, which the delete lock must not affect
- Restore ColumnHeaderMenu readOnly to permission-only: it swaps the header for
  a static label, which was also disabling pin, column-select and open-config
- Assert the schema lock at the import-async enqueue when createColumns is set
- Allow a locks PATCH that only clears locks even when the feature flag is off,
  and keep the settings entry reachable on a locked table, so flipping the kill
  switch can't strand a table with locks nobody can remove

* fix(tables): scope the update-lock exemption and re-check locks per page

- Make the workflow-output carve-out opt-in via `computedWrite`, set only by
  the cell-write path. It was caller-agnostic, so an ordinary API caller could
  PATCH a workflow-output column on an update-locked table
- Re-read the lock before every page in the delete and update runners. The
  single worker-start check went stale immediately, so enabling a lock could
  not stop a job already deleting or overwriting rows. Committed pages still
  stay committed, exactly as with an explicit cancel. Import keeps its
  one-shot check: a half-imported table needs the deletes the lock now forbids
- Route Insert column to the locked-action modal under a schema lock instead
  of leaving it live (regular header) or hiding the whole menu (group header)

* fix(tables): revalidate locks in-transaction and unbreak backfill + partial unlock

- Re-assert the lock inside each page-write transaction, under the same
  advisory lock updateTableLocks takes, so check-then-write is atomic against
  a lock change. The per-page check alone still left the page-selection await
  between the assert and the write
- Pass computedWrite from the backfill runner: batchUpdateRows is the
  workflow-output backfill path, so scoping the exemption had made rebuilding
  outputs 423 on an update-locked table while live cell writes still worked
- Gate the feature flag on a lock actually going off->on rather than on an
  all-false payload. The settings UI always submits all four flags, so with
  the flag off an admin could not clear one lock while another stayed on
- Hoist the lock-kind -> flag map to lib/table/types as TABLE_LOCK_FLAGS

* fix(tables): guard import batches in-transaction and split paste by verb

- Re-assert the insert lock inside each import batch's insert transaction,
  under the same advisory lock updateTableLocks takes. Reversing the earlier
  "let the file finish" call: an admin can lift the lock to clean up a partial
  import, so honouring it beats letting the rest of the file land
- Gate paste per verb. Overwriting existing rows is an update and extending
  past the last row is a full-row insert, so a paste-append still works on an
  append-only table while an overwrite explains itself. Refuses the whole
  paste rather than applying half of it
- Space opens the same row editor as double-click, so it now follows the
  update lock instead of filling in a form that 423s on save

* fix(tables): guard the remaining import writes and explain blocked Shift+Enter

- Route the replace-mode wipe, the inferred-schema write and createColumns
  through the same guardBatch transaction as the batch inserts, so a delete or
  schema lock committed while the file is downloading or being sampled is seen
  instead of the job-start snapshot
- Shift+Enter takes the manual-add path, so it now opens the lock modal like
  the Add row button instead of returning silently

* improvement(tables): surface blocked lock actions as a toast, not a modal

- Replace TableLockedModal with a warning toast carrying a "Lock settings"
  action button for admins. Being told you can't edit shouldn't cost a dismiss
  click, and the button still routes admins straight to the panel
- Dedupe by id so repeated attempts on a locked cell replace one notice
  instead of stacking a column of them
- Move the copy into lock-copy.ts alongside the rest of the lock vocabulary
  and tighten it for toast length

* fix(tables): don't show the update-lock notice to users without write access

Double-click checked the update lock before any permission check, so a
read-only member on a locked table got "Editing rows is locked" — misleading
(the lock isn't why they can't edit) and it swallowed the expanded-cell
viewer, which is a legitimate read-only affordance.

* improvement(tables): trim the update-lock toast copy

* fix(tables): map append-import locks to 423 and stop conflating locks with permissions

- The sync append branch returns instead of rethrowing, so the outer catch's
  mapper never saw a TableLockedError and every lock violation became a 500.
  Map it in that catch, with a regression test (replace mode already rethrows)
- Stop mounting the workflow-group column menu for users without edit access:
  passing the blocked handlers unconditionally made it appear for read-only
  members and report a lock even on an unlocked table
- Enter/F2 now raises the same lock notice as double-click and Space instead
  of silently doing nothing
- guardBatch returns the freshly-read definition so addTableColumnsWithTx
  asserts live state; a schema lock cleared mid-import no longer fails the
  createColumns step. The snapshot pre-asserts now only run as the fallback
  for callers that pass no revalidator
- Append-only no longer labels a table whose schema is also locked, which
  claimed columns were mutable when they weren't
- Import dialog withholds Replace on a delete-locked table and create-column
  on a schema-locked one, instead of offering a configuration that only 423s

* fix(tables): revalidate sync imports under the advisory lock, explain every blocked key

- The sync append/replace paths asserted only the request-start snapshot, so a
  lock committed while the CSV was parsed still let the write through. Both now
  re-read under the schema advisory lock at the top of their own transaction,
  taken before acquireRowOrderLock so the order stays advisory -> rows_pos ->
  definitions
- Delete/Backspace, Cmd+D, typeahead and cut raised no notice on an
  update-locked table; they now explain the lock, and stay silent for users
  without write access
- The multipart import fetch threw a plain Error, dropping the 423 status, so
  the lock self-heal never ran and the stale detail cache survived. It throws
  ApiClientError now and both import-into-table hooks call
  handleTableLockRejection
- Force append at submit when the delete lock landed while the dialog was open
  with Replace already selected

* fix(tables): restore the workspace ownership check on copilot row deletes

Switching deleteRow/deleteRowsByIds to take a TableDefinition dropped the
workspaceId argument that previously scoped the query, and the two branches
I added loaded the table without the ownership comparison every other
operation in the tool performs — letting a caller delete rows from a table
in another workspace. Both now reject a foreign table as not found.
2026-07-25 14:21:05 -04:00
Theodore LiandClaude Opus 4.8 8329dac4c5 feat(tables): add select & multi-select column types (#5873)
* feat(tables): add select/multiselect column types (backend)

Adds two enum-style column types where the column declares a fixed set of
options (stable id + name + palette color) and every cell is constrained to
them.

- COLUMN_TYPES gains `select` / `multiselect`; SELECT_COLORS is a fixed,
  theme-aware palette mapping 1:1 to Badge color variants (no raw hex)
- ColumnDefinition.options carries the option set. Cells store option *ids*
  (a string for select, string[] for multiselect) so renaming or recoloring
  an option never rewrites row data
- Row validation enforces membership; coercion tolerantly maps an option
  *name* to its id for tool/import writes and drops unmatched entries
- validateColumnDefinition enforces non-empty, unique-id, unique-name and
  valid-color option sets, and rejects options on non-select columns
- Contract gains selectOptionSchema plus a cross-field refine requiring
  options exactly on select types; routes thread options through type
  changes and a new options-only updateColumnOptions path

No migration: column config already lives in the user_table_definitions
schema JSONB blob.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014e4MvFV1szLhNLap1CCNUd

* feat(tables): select/multiselect column UI

Surfaces the new enum column types in the tables grid.

- New `select-field/` module: SelectPill (colored option via the shared Badge
  palette), SelectValueEditor (one ChipDropdown-backed picker reused by every
  edit surface), SelectOptionsEditor (add/rename/recolor/remove)
- Option colors are picked from inline squircle swatches — no labels, no
  nested dropdown. Each swatch's fill is a Badge in that variant, so the
  palette stays single-sourced and theme-aware
- Cells render option pills; an empty select cell shows a muted "None" so it
  reads as a dropdown. The single-select menu always offers "None" to clear
- Wired into all three edit surfaces (inline cell, expanded popover, row
  modal) plus the type picker and column-type icons
- ChipDropdown gains `defaultOpen`/`onOpenChange` so the inline cell editor
  can open on mount and commit when the menu closes; open state is now
  controlled in both single and multi modes

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014e4MvFV1szLhNLap1CCNUd

* fix(tables): auto-fit select columns on option labels, not ids

Column auto-resize measured `String(val)` for every non-json/date column,
which for a select cell is the opaque option id (and for multiselect the
comma-joined id array). Selecting auto-fit on a select column therefore
sized it to ids the user never sees. Measure the resolved option names
instead, matching what the pills actually render.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014e4MvFV1szLhNLap1CCNUd

* feat(tables): default select options to grey, drop color picker for now

Removes the per-option swatch picker and defaults every option to the
neutral gray pill. The `color` field stays in the data model and the
SELECT_COLORS palette/contract are untouched, so a picker can be re-added
later as a pure UI change with no migration.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014e4MvFV1szLhNLap1CCNUd

* fix(tables): guard select type conversion, escape, and required clear

Addresses review findings on the select/multiselect column types:

- Column type conversion now checks each existing value against the target
  option set (resolve by id or name); a `select`/`multiselect` change is
  blocked when values don't fit, instead of accepting shapes that later
  coercion would strand or silently drop
- Inline select editor discards its draft on Escape (matching the text/date
  editors) rather than committing on menu close
- A required single-select no longer offers "None" — clearing to null could
  never be committed

Exports resolveSelectOptionId from validation for the conversion gate.

* fix(tables): block emptying required multiselect, skip no-op cell writes

Round 2 review fixes:

- A required multiselect can no longer be emptied — the toggle that would
  remove the last option is ignored, since an empty selection can't be
  committed (server rejects it). Mirrors the required single-select "None"
  guard.
- Inline cell save now compares old vs new value structurally, so a no-op
  edit (e.g. opening a multiselect and closing it unchanged, producing a new
  array reference) no longer writes a row update or pushes an undo entry.
  Uses JSON compare for arrays/objects, matching the existing optionsEqual
  convention; primitives keep the === fast path.

* fix(tables): expanded multiselect string value + same-type options drop

Round 3 review fixes:

- Expanded-cell popover normalizes a multiselect's initial value with
  toSelectedIds, so a single option-id string (the normal shape right after a
  select->multiselect conversion, before the row is rewritten) is no longer
  collapsed to an empty array and cleared on save.
- Column PATCH now only routes to updateColumnType on a real type change.
  updateColumnType early-returns on an unchanged type, so a payload repeating
  the current type together with new options previously applied neither; an
  unchanged type with options now routes to the options-only update. Fixed in
  both the internal and v1 column routes.

* improvement(tables): inline select edit + idiomatic options editor

- Double-clicking a select/multiselect cell now opens the inline option
  dropdown (like date/number) instead of the fixed-height text popover, which
  rendered just a small dropdown floating in a large empty container. Removes
  the now-unreachable ExpandedSelectEditor from the expanded-cell popover.
- Options editor's add/remove controls match the sibling Filter UI: a ghost,
  muted "Add option" button with a small plus, and an X remove icon.

* improvement(tables): inline select cell uses bare DropdownMenu

The inline select editor used a ChipDropdown pill, which reads as a foreign
form control inside a grid cell. It now renders the canonical DropdownMenu
anchored to the cell — an invisible full-cell trigger (no pill, no label
text), options as their colored pills with a check on the selected ones, and
"None" as a proper menu item. The cell keeps showing its pills while the menu
is open instead of going blank.

SelectValueEditor (the ChipDropdown pill) is now used only by the row modal,
where a pill form control is idiomatic. This lets the defaultOpen/onOpenChange
additions to the shared ChipDropdown be reverted — it's back to its original
behavior with no consumers depending on the change.

* improvement(tables): unify select into one type with a multiple flag

- Collapse `select` + `multiselect` into a single `select` column type with a
  `multiple` boolean. A cell holds one option id when single, an array when
  multiple. Validation/coercion, the contract, routes, and all cell/editor
  code now branch on `column.multiple` instead of a separate type.
- The column-config sidebar exposes an "Allow multiple" toggle (create and
  edit) and no longer offers "Unique" on select columns.
- Give the closed select cell a dropdown affordance: it renders emcn's
  `comboboxVariants` chrome (border + chevron, compacted to the row height),
  with the pills inside and the chevron flipping while the menu is open.
- Single↔multiple toggles reconcile lazily on the next row write (single
  tolerates an array by taking its first resolved option).

* improvement(tables): revert select cell to chip-only view

Drop the combobox chrome (border + chevron) from the select cell — the plain
option pills read better. Reverts the comboboxVariants barrel export too, since
nothing else uses it.

* fix(tables): block multiple→single select switch when cells have >1 option

Switching a select column from multiple to single would silently keep only the
first option of any multi-valued cell. updateColumnOptions now scans the rows on
that transition and errors ("N row(s) have multiple options selected") instead
of dropping data, mirroring the type-change compatibility guard.

* improvement(tables): rename select "Allow multiple" toggle to "Multiselect"

* fix(tables): drop removed select options from cells instead of stale pills

When an option is deleted, cells that referenced it previously rendered a gray
fallback pill labeled with the raw internal id. They now drop the orphaned id
and fall back to empty ("None"). Editors seed their selection from the still-
valid ids too, so editing/saving such a cell writes the cleaned value.

* improvement(tables): use dashed add-row button for select options

Match the sub-block list editors (filter/sort builders): a full-width,
dashed-border ghost "Add option" button instead of a plain text button.

* perf(tables): don't refetch rows on metadata-only column saves; fix copy log

- Add/update column are metadata-only server-side (row data never changes), so
  they now invalidate just the schema, not the rows. Saving a column (e.g.
  editing select options) no longer triggers a full rows refetch/flash; cells
  re-render from the refetched schema. Delete-column and workflow-group ops keep
  the rows invalidation since they do change row data.
- "Failed to copy rows" logged an empty {} for DOMExceptions; log the extracted
  message so the real reason (e.g. lost transient activation) is visible.

* fix(tables): migrate legacy multiselect columns; readable contract error

- Tables created before the select/multiselect merge stored `type:
  'multiselect'`, which the removed type no longer accepts — every response
  carrying such a column failed contract validation and column edits threw.
  getTableById now maps `multiselect` → `select` + `multiple: true` on read
  (covers reads, column ops via withLockedTable, and their responses); the
  migrated shape persists on the table's next schema write.
- requestJson's "Response failed contract validation" now appends a short
  "field: reason" summary of the Zod issues instead of a bare message, so the
  failing field is visible.

* revert(tables): drop legacy multiselect read-migration

Only local dev tables ever held the removed `multiselect` type (the feature
never shipped), so the read-time backward-compat mapping isn't warranted.
Keeps the readable contract-validation error from the same change.

* improvement(tables): add select options by typing into a trailing row

Replace the "Add option" button with a trailing empty row: the first keystroke
materializes the option and focus jumps into it (cursor at end) so typing flows
straight through. Enter in an option jumps back to the trailing row to add the
next one. Removing a row is unchanged.

* fix(tables): export select columns as option names, not ids

CSV/JSON export serialized the raw stored value for select cells — an option
id (or an array of ids for multi) — so downloads showed opaque `opt_…` ids.
Export now resolves ids to option names: CSV comma-joins multi names; JSON
resolves values before the id→name key translation. Ids with no matching
option (deleted) are dropped.

* fix(tables): resolve select option ids to names across all read/consume surfaces

Stored select cells hold opaque option ids; only the cell renderer and write
coercion handled id<->name. Every other read/serialize/search boundary leaked
the id. Centralize translation in lib/table/select-values and apply it at each
boundary:

- reads return names: sync export (CSV+JSON), function/snapshot mounts,
  clipboard copy/cut, workflow-tool + mothership reads, v1 API rows
- filter/search: tolerant name->id on filter operands (row-wire INTERNAL_JWT,
  mothership, v1), select-aware SQL predicates (operator whitelist, multiselect
  empty), grid Cmd-F matches resolved names, UI filter option picker sends ids
- sort: select columns order alphabetically by option name (id->name CASE)
- multiselect writes split a comma-delimited string; mothership add/update_column
  forward options/multiple

* refactor(tables): remove vestigial per-option color from select columns

Colors were never exposed (no picker; every option hardcoded to gray), only
kept in the model for a future picker. Drop the field entirely: SelectOption is
now { id, name }, pills render a fixed neutral badge, and validation/contract no
longer carry a color. Existing stored options keep a harmless color key that is
stripped on the next read/write.

* feat(tables): generate select option ids in the copilot handler from agent-supplied names

The mothership agent authors select options by name only; the copilot tool
handler now mints the stable option id (preserving any id it's re-sent, so
existing cell data survives an options edit) before calling the table service —
the model never authors the cell key. Applied to create (per select column in
the schema), add_column, and update_column.

* refactor(tables): collapse row read boundaries onto one outbound seam

Turning a stored row into its outward form needs two translations — keys
(column id to name) and values (select option id to name) — and every boundary
composed them by hand. That shape leaked twice: table-change triggers and
workflow-column/enrichment inputs both built name-keyed rows without resolving
select values, so workflows saw opt_a1b2 instead of "Open". Neither had any
test coverage; neither is live (select is unreleased on staging).

lib/table/cell-format.ts now fuses both into a single pass:
- namedRowMapper(columns) — keys + values together, so they cannot drift
- fillMissingColumns(named, columns) — widens to match a headers array
- mapInputValues(data, columns, mappings) — for id-keyed input mappings

Migrated the 12 hand-composed pairings and the inlined copy in export-runner,
then deleted rowDataIdToName, resolveRowSelectValues and selectColumnsOf so the
keys-without-values shape is unrepresentable. Also formats dates in the
read-only expanded popover, which rendered raw storage to viewer-role users.

CSV is untouched (export-format.ts has no diff), so export bytes are unchanged.

* fix(tables): stream the persisted cell value, not the raw workflow output

A workflow column writing into a select column persists correctly — updateRow
coerces the option name to its stored id — but it coerces its own merged copy,
so the live SSE snapshot still carried the raw value. The grid resolved that
name as an option id, found nothing, and rendered the cell EMPTY from the
moment the workflow completed until the next refetch: the write looked like it
had failed at exactly the moment the user was watching it land.

Coerce a copy of the event outputs through the same inbound switch the persist
path uses. A copy because the patch object is identity-compared for the
progress writer's retry bookkeeping. This aligns every type, not just select —
a date now streams canonically, and a value the DB would reject streams as null
instead of briefly showing a value that was never stored.

* fix(tables): keep option ids stable when the agent edits a select column

The agent authors select options as bare names, so normalizeSelectOptionsInput
minted a fresh id for every one. On update_column that replaced the whole
option list with new ids — orphaning every cell that referenced the old ones
and silently clearing the column's data on what looks like an additive edit
(adding "Medium" to ["Low", "High"] wiped both existing values).

Match incoming names against the column's current options (case-insensitively)
and reuse those ids; only genuinely new options get a fresh one. Found while
writing the catalog description that promised this behavior.

Also regenerates the tool catalog now that the copilot registry advertises
select columns, options and multiple.

* fix(tables): filter multi-select columns by membership, not equality

A multi-select cell stores an array of option ids, and Postgres containment
does not match a scalar against an array — `{"t":["a"]} @> {"t":"a"}` is
false. So every multi-select filter compiled to a predicate that could never be
true and silently returned nothing.

Split the select operator whitelist by cardinality: single-select keeps
eq/ne/in/nin, multi-select takes contains/ncontains, both keep isEmpty. The
membership clause wraps the operand (`data @> '{"t":["a"]}'`), which still
uses the same GIN index. The equality shorthand (`{ tags: 'opt_a' }`) compiles
to membership too — it bypasses the operator whitelist, so it had to be handled
or it would keep failing silently.

Also resolves $contains/$ncontains operands name → id, so an agent or API
caller filtering by option name reaches the right rows, and narrows the filter
UI's operator list per cardinality.

* chore(tables): drop a stale doc comment and trim two restating ones

The expanded-cell-popover TSDoc claimed workflow and boolean cells are
read-only there, but the editability rule has no workflow check and the inline
comment beside it says the opposite — a comment that contradicts the code is
worse than none. The accurate rule already sits next to what enforces it.

* fix(tables): close review findings on select column conversion and editing

Bugbot round on the select work — eight findings, all real:

- Converting away from select left opaque option ids in every cell. Migrate
  ids to option names in the same transaction (multi joins comma-separated),
  and run the compatibility check against the name, not the id.
- Multiselect to text silently dropped arrays; text targets now reject
  structured values instead of nulling them on the next write.
- Converting to single-select accepted multi-valued cells, which the next
  coerce would quietly truncate to the first id. Same guard updateColumnOptions
  already had.
- Empty strings had become compatible with every target type, so a text column
  with blank cells could convert to number. Narrow that back to select.
- Copilot update_column rejected a multiple-only payload with 'options is
  required', contradicting the catalog. Fall back to the column's options.
- Empty multiselect rendered ChipDropdown's 'All' label, reading as if every
  option were selected.
- Opening and dismissing an empty multiselect wrote a row update, since null
  and [] compared unequal.
- Converting a unique column to select stranded the constraint with no way to
  clear it.

Adds the first tests for the conversion rules.

* fix(tables): migrate select cells in both directions and refetch rows after

Follow-up round on the conversion fix — the migration was one-directional.

- Converting TO select left cells holding option names, which every reader
  resolves by id, so populated cells rendered as None and matched no filter.
- Toggling single→multi left scalar ids while multi filters compile to array
  containment, which a scalar never matches — pre-toggle rows silently dropped
  out of their own column's filters.
- useUpdateColumn settled with the schema-only invalidation, whose premise
  ('stored row data never changes') these migrations broke, so the grid kept
  serving pre-migration cells.

Both directions now share set-based helpers keyed off a jsonb map, and the
to-select map keys ids as well as names so re-running is a no-op. The client
refetches rows when the payload carries a type or multiple change.

* fix(tables): align select cell migration with the compatibility check

Two spots where the check and the migration disagreed, each leaving a cell the
check had already accounted for:

- resolveSelectOptionId matches option names case-insensitively, so a cell like
  'open' against option 'Open' passed the convert-to-select check, but the
  migration map was case-sensitive and left it as the raw name. The map now
  keys the folded name too (duplicate names are already rejected
  case-insensitively, so it is unambiguous) and lookups try the exact form
  first.
- Converting away from select, the check treats an orphaned option id as null
  while the migration passed it through, so a column with deleted-option cells
  could land an opaque opt_ id in a number/date/boolean cell. Orphans now null,
  matching selectValueForConversion; a multi cell drops them and nulls when
  nothing survives.

This reverses the pass-through choice from the earlier round — consistency with
the check is what matters, and an orphan is a value the UI never rendered.

* fix(tables): treat an emptied multiselect as empty everywhere

`[]` is the canonical empty multiselect but it is not falsy, so every
emptiness check that only looked for null/undefined/'' read a cleared selection
as filled:

- dependent workflow groups became eligible off an empty cell, and enrichments
  with a required input ran on rows they should have skipped
- marking a column required did not count `[]` rows, even though write-path
  validation rejects an empty required multiselect

One `isEmptyCellValue` predicate now backs the dep, output-filled and
enrichment checks, and the required-constraint guard counts '[]' too.

Also normalizes '' on conversion to select — compatibility admits it, so it has
to land as null (single) or [] (multi) rather than a bare string the column's
own validation would then reject.

* fix(tables): route an unchanged type with options to the options update

updateColumnType early-returns when the type is unchanged, so an agent that
restated newType: 'select' alongside a new option set had those options
silently dropped while the tool still reported success. The HTTP columns route
already guards this; the copilot path now mirrors it — only a genuine type
change goes to updateColumnType, anything else with options or multiple goes to
updateColumnOptions.

A payload that only restates the current type is a no-op and now reports the
live schema rather than an undefined one.

* fix(tables): restore select options on column-delete undo; map v1 option errors to 400

- The delete-column undo snapshot never captured options or multiple, so
  re-creating a select column was rejected outright (it is invalid with no
  option set) and the saved cell data — which is option ids — had nothing to
  attach to. Undo of a select column deletion simply could not succeed.
- The v1 add-column route mapped only 'already exists' and 'maximum column' to
  400, while the internal route also maps invalid-column and option errors. A
  bad select option set surfaced as a 500 on the public API.

* fix(tables): clear unique in the service when converting a column to select

The clearing lived only in the column-config sidebar, so a conversion through
the v1 API or the copilot tool left unique: true stranded on a select column —
where the constraint compares the stored option id, capping each option at one
row for the whole table, and the sidebar hides the toggle so it could never be
cleared again.

Moving it into updateColumnType gives all three callers the same behavior; the
sidebar's payload is now belt-and-braces rather than the only enforcement.

* fix(tables): clear removed select options, reject unique selects, resolve pasted names

- updateColumnOptions now drops ids for options that no longer exist, so a
  removed option can't leave orphans that block edits on required columns
- validateColumnDefinition and updateColumnConstraints reject unique on select
- cleanCellValue resolves pasted option names to ids client-side so the
  optimistic cache holds ids and the pill renders immediately

* fix(tables): refetch rows when a select option is removed

Removing an option rewrites cells server-side, so the schema-only
invalidation left the cache holding orphaned ids — hidden by the grid but
still visible to emptiness checks, filters, and dependent-group eligibility.

* fix(tables): let a multiselect round-trip through text

Converting multiselect to text flattens cells to `Alpha, Beta`, but
converting back read that as one unknown option and rejected the change.
Both the compatibility check and the cell migration now split a comma
string the way the write path already did, through one shared helper.

* fix(tables): scope the empty-selection guard to multiselect; order the option ref map

- cellValuesEqual treated null and [] as equal for every type, so clearing a
  stored [] on a json cell (or writing one into a null cell) never saved
- migrateCellsToSelectIds built one flat id/name map, letting an option whose
  name equals another option's id overwrite that id and repoint its cells

* fix(tables): prune stale select filters; reject unique+select before converting

An applied filter outlives the schema it was built against, so converting a
column to select (or toggling multiple) could strand an operator the server
rejects, failing every subsequent rows query until the filter was cleared by
hand. Prune those conditions in useTable, above every consumer of the rows
query key, and share one operator whitelist with the filter picker.

A payload setting both newType: "select" and unique: true ran two separate
locked transactions, committing the conversion and then failing the
constraint. Both routes now reject that pair before either runs.

* fix(tables): scope bulk ops to the pruned filter; guard unique+select in copilot

useTable pruned the filter for the rows query but kept it private, so
select-all run/stop/delete still sent the raw one — targeting a predicate the
grid wasn't displaying and that the server rejects. Return it and use it for
every server-bound scope; the filter button and editor keep showing what the
user configured so the stale rule can still be repaired.

The copilot update_column path ran the same two-transaction sequence the HTTP
routes now reject, so a convert-to-select plus unique payload committed the
conversion and then failed. Same guard applied.

Also adds the filter to selectedRunScope's deps — it was read but not listed,
so a filter change while select-all was active carried a stale scope.

* fix(tables): reject the multiple flag on non-select columns

A create/add payload could persist multiple: true on a string column, where
it sits inert until updateColumnType inherits it via `data.multiple ??
column.multiple` — silently turning an intended single-select into a
multiselect and rewriting every cell as an array. Rejected now in both the
service validator and the shared contract refine, which the internal and v1
column routes both consume.

* fix(tables): block required blank conversions; sort multiselect by option name

`required` only rejects null/undefined on a write, so a required string column
legitimately holds ''. Converting it to a required select stored null (or []
for a multi), and every later update of that row then failed its own required
check. Reject a blank source value when the target select is required.

Multiselect sort compared the raw id-array text, since `["opt_b","opt_a"]`
matches no single-id CASE branch. It now resolves elements to names and joins
them in stored order — the same text the grid renders and an export writes.

* fix(tables): validate a conversion against the required flag the request sets

A single PATCH converting an optional column to select while also setting
required: true checked blanks against the column's CURRENT required flag, so
the conversion committed and the constraint write then failed — an error
response with the type change already persisted.

updateColumnType now takes the requested `required` and validates against the
constraint the column ends up with. Empty rows are also counted when the
target is required (they were skipped outright), so the null case fails
up-front too rather than at the constraint write.

* fix(tables): one emptiness predicate, guard required option removal, order aggregates

The conversion pre-flight and the constraint write each had their own idea of
"empty", and rows missing the column key are filtered out of the conversion's
row set entirely — so a combined type + required PATCH passed the pre-flight,
committed the conversion, then failed the constraint. Both now call one
countEmptyCells helper, which is the actual fix: they can no longer disagree.

Removing the last option a required cell holds nulled it, producing exactly the
state updateColumnConstraints refuses to create. Blocked with a count of the
rows that would be stranded.

The multiselect rewrite aggregates carried no ORDER BY while sort and export
preserve stored order via WITH ORDINALITY; they now do too.

* fix(tables): restore in/nin on select filters and assert the whitelists agree

The client operator set was a hand transcription of the server's and dropped
$in/$nin, so the picker never offered them and — worse — pruneFilterForColumns
silently discarded an existing $in filter the server would have accepted.

Both server sets are now exported and a test maps the UI operators onto them
through UI_TO_WIRE_OPERATOR, so the two can't drift again by hand.

* fix(tables): stop Find crashing on a null multiselect cell

jsonb_array_elements_text throws "cannot extract elements from a scalar" on a
JSON null, which is exactly what a multiselect cell holds once it is cleared,
cut, or has its last option removed — so search failed for the whole table.
Gate the array arm on jsonb_typeof, falling back to the single mapping so a
scalar left over from a single->multi toggle stays searchable.

Swept the other array expansions: the rest are inside UPDATEs whose WHERE
already filters to arrays, or are CASE-guarded. This was the only unguarded one.

* fix(tables): clear removed options before the cardinality guard and migration

The multi->single guard counted options the same request was dropping, so
trimming a cell to one option and turning multiselect off in one save was
rejected. Reordering fixes that and a latent corruption: the migration keeps a
multi cell's FIRST element, which could be a removed id sitting ahead of a kept
one, so the surviving option would be discarded and the dead one kept.

Removal now runs first, against the pre-toggle cell shape, then the guard
counts what actually remains, then the shape migration runs.

Find also joined multiselect names unordered while sort and export preserve
stored order, so a search for the displayed label could miss.

* fix(tables): validate option removal against the pending required flag; show the applied filter

The stranded-options guard read the column's current `required`, so a removal
paired with `required: true` cleared cells and then failed the constraint write
— options change committed behind an error — while a removal paired with
`required: false` was blocked despite the column no longer being required. It
now takes the requested value, and also pre-flights already-empty rows when
required is newly imposed.

The filter chip and editor keyed off the raw filter while rows, runs, stops and
deletes used the pruned one, so the UI could claim an active filter the grid
wasn't reflecting. Both now read the applied filter.

* fix(tables): gate the unique guard on the resulting type, not on a type change

Pairing unique: true with an options or multiple update on a column that is
ALREADY select skipped the guard, so updateColumnOptions committed and the
separate constraint write then failed — the options change landing behind a
400. Gating on `updates.type ?? currentColumn.type` covers the conversion and
the options-only case with one condition, on both column routes and the
copilot update_column path.

* fix(tables): stop coercing select option ids in filter values

Option ids are caller-supplied strings, so parseScalar turned an id of "1" or
"true" into a number or boolean; JSONB containment then compared the wrong type
and matched nothing while the picker still showed a valid option.

filterRulesToFilter takes the columns and keeps a select value as text, for $in
lists too. pruneFilterForColumns forwards them as well — its round-trip through
rules would otherwise re-coerce the ids it just preserved.

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-25 13:33:37 -04:00
Theodore Li 892750b4be fix(autolayout): keep branches in stable rows across layers (#5877)
* fix(autolayout): keep branches in stable rows across layers

* fix(autolayout): container overlap on insert and error-branch alignment

Lay out contained groups before their containers so a container carries the
height it renders at before siblings are placed beside it. Targeted layout
sized containers from a hypothetical child layout it never applied, so a
container whose children stay frozen rendered taller than the space reserved
for it and overlapped its neighbour.

Also mark containers resized by a child edit so following blocks shift clear,
and derive the error handle offset from the layout's own block height instead
of the raw height field.
2026-07-25 05:05:16 -04:00
Theodore Li 19c3b6f47d feat(setup): setup wizard with browser-based Chat key handoff (#5911)
* feat(setup): setup wizard with browser-based Chat key handoff

Adds `bun run setup` and `bun run doctor` for local installs, and replaces
the wizard's paste-your-Chat-key step with a browser handoff that never puts
the key in a URL.

* improvement(setup): drop the paste-a-key fallback, simplify consent copy

The browser handoff is now the only path — the wizard waits on a spinner
instead of racing a paste prompt. Consent card leads with "Connect your
terminal" and moves the match-the-code disclaimer into the description.

* fix(setup): pin kube context, keep secrets out of argv, validate reused keys

Review findings from #5911:
- helm/kubectl now run against the validated context instead of the ambient one
- helm values are piped on stdin rather than passed as --set arguments
- ENCRYPTION_KEY/API_ENCRYPTION_KEY are checked for the 64-hex format the app
  requires, not just length, so an unusable key is replaced rather than kept
- the managed Redis container's published port is read back instead of assumed

* refactor(copilot): one module for Chat API key operations

list/generate/delete each repeated the same /api/validate-key envelope in
their route. They now share callValidateKey in lib/copilot/server/api-keys.ts,
which also keeps the display masking server-side so the full key can only ever
leave at creation.

* improvement(setup): reuse shared helpers, parallelize probes, drop dead code

- PKCE verifier/state/pairing code now use generateSecureToken, generateRandomHex
  and generateShortId instead of hand-rolled randomBytes; the pairing loop's
  modulo was unbiased only because 256 % 32 == 0
- new sha256Base64Url in @sim/security/hash so both sides of the PKCE exchange
  derive the challenge from one implementation
- isUsableSecret moved beside SECRET_KEYS so setup and doctor apply the same
  rule; doctor previously passed a key setup would replace
- isTruthy narrowed to true/1, matching the app it claims to mirror — it accepted
  yes/on, so a flag could read on in doctor and off in the app
- checkLive runs its five probes concurrently (~17s serial worst case)
- detection overlaps the banner animation instead of queueing behind it
- glyph.fail/glyph.warn at 13 sites that bypassed the constant; removed unused
  prompter exports, a dead ENV_PATHS re-export, and an unused export keyword

* fix(setup): make doctor understand the compose env layout

Compose writes a single root .env (what docker-compose reads via env_file) but
the checks required the three per-app files, so a successful compose install was
followed by doctor printing three failures and exiting 1 — and the whole
coherence catalog was skipped because it keyed off apps/sim/.env existing.

Layout is now derived from what's on disk and every check consults it: file and
schema checks iterate the layout's targets, consistency reports skip when
there's only one file to mirror, and coherence/live read the layout's primary
file. The wizard's existing-config detection counts root for the same reason —
a compose install used to read as unconfigured and re-run from scratch.

* feat(cli-auth): device-authorization poll flow, drop the loopback listener

The CLI no longer binds a local port. It generates a request id + poll secret,
opens /cli/auth, and polls /api/cli/auth/poll over TLS while the user approves
in the browser — so the flow works over SSH and inside containers, where the
browser and terminal don't share a machine.

- approve stores the approval keyed by request id (session-authed, userId from
  the session only); poll verifies the secret before an atomic claim, so an
  observer of the semi-public request id can neither mint nor cancel it
- pairing code stays as the anti-phishing compare; no key ever crosses the
  browser; done page just confirms
- removes the loopback listener, /token exchange, buildCliHandoffUrl, and
  validateCliCallbackUrl (+ its tests) — nothing hands a key to a URL anymore

* fix(setup): reuse an existing managed Postgres container instead of colliding

A running sim-postgres fell through to `docker run --name sim-postgres` and died
on the name conflict; a stopped one failed with "no DATABASE_URL to reach it"
because the generated password only lived in the env files a fresh clone lacks.

Both facts are recoverable from Docker: the ladder now reads the published port
and password back via `docker inspect` and reuses the container (starting it if
stopped). A container that won't answer prompts before recreating, and never
drops the data volume silently.

* improvement(setup): audience-first run-mode hints

Each run mode now names who it's for — compose for self-hosting/evaluating, dev
for contributing to Sim, k8s for rehearsing a production deploy — with the live
detection state (Docker/kube/VM) appended.

* fix(cli-auth): retry a failed mint, port container/port fixes to Redis + k8s

Review findings from #5911:
- poll now reserves the mint with an atomic NX lock instead of deleting the
  approval up front, so a failed mint (e.g. mothership blip) is retried by the
  next poll instead of forcing a fresh browser approval; the lock still prevents
  a double-mint and its TTL frees the slot if the caller dies
- setup reuses/recreates an unhealthy managed sim-redis instead of colliding on
  the name (Redis has no data volume, so it removes and recreates without a prompt)
- k8s failure-path hints carry --context, matching the success-path hints, so a
  changed ambient context can't send diagnostics to the wrong cluster
- compose port-free waits for a killed port to actually release before
  re-checking; SIGKILL is async, so the immediate re-check re-saw the port

* fix(setup): harden mint cleanup, Windows browser, container detection, helm cwd

Review findings from #5911:
- a post-mint completeApproval failure no longer routes into releaseMint — the
  mint lock now outlives the approval (shared TTL), so a cleanup blip can't leave
  a re-mintable window and orphan a key; cleanup is best-effort after the key ships
- compose doctor --fix writes the feature-flag twin to the layout's primary env
  (root .env on a compose install), not always apps/sim/.env
- Windows opens the browser via `cmd /c start "" <url>` — `start` is a shell
  builtin, so spawning it directly ENOENT'd and the handoff never opened
- managed-container detection filters loosely and pins the exact name in code;
  Docker's `name=^x$` anchor matches the internal `/x` form and often missed,
  skipping the reuse branch
- the shared helm/kind run helper pins cwd to the repo root, matching helm test,
  so `helm upgrade --install ./helm/sim` works from any working directory

* feat(chat-keys): standalone manage page, drop from settings nav, refresh README

- Add /account/settings/chat-keys — a linkable page to view, create, and revoke Chat API keys
- Remove Chat keys from the settings sidebar (account + unified nav) and its render branches
- README: replace Docker Compose + Manual Setup with the bun run setup wizard; drop the manual COPILOT_API_KEY step, point to the manage page

* fix(setup): per-key reason in the secret-replacement warning

Cursor: the warn hardcoded '64-character hex key', but only ENCRYPTION_KEY/API_ENCRYPTION_KEY require that — BETTER_AUTH_SECRET/INTERNAL_API_SECRET only need length >= 32. Use the existing secretRequirement(key) helper so each replaced key reports its actual requirement.

* fix(setup): compose doctor schema, cross-platform binary detection, quoted context hints

- Doctor: for the compose (root) env layout, require only the secrets compose has no interpolation default for (BETTER_AUTH_SECRET/ENCRYPTION_KEY/INTERNAL_API_SECRET). DATABASE_URL/BETTER_AUTH_URL/NEXT_PUBLIC_APP_URL come from docker-compose ${VAR:-default}, so a healthy compose install no longer fails doctor.
- Binary detection: use Bun.which instead of which (which is absent on Windows), so kubectl/helm/kind/docker resolve cross-platform.
- k8s diagnostic hints: POSIX-quote the kube-context so a context with whitespace/metacharacters can't break or inject into a copied command.

* fix(setup): quote kube-context in the helm uninstall tear-down hint too

The tear-down hint used --kube-context ${context} raw while the sibling kubectl hints already used shq(); a context with whitespace/metacharacters could break or inject into the copied command. All copyable k8s hints now go through shq(context).

* feat(setup): sim lifecycle CLI — start/stop/status/logs/down/reset

Turn the setup entry into a 'sim' command umbrella so there's one place to run everything, not scattered docker/bun commands. Adds a global bin (bun link) + a bun run sim fallback.

- Detects how you're running (compose file / managed dev containers / helm release) from disk + docker/helm state — no persisted mode. Ambiguous installs prompt.
- start/stop/restart/logs work per mode; down removes containers (volumes kept); reset archives .env + wipes managed data; both destructive verbs confirm first.
- status shows detected mode, container states, and app/realtime health.
- Wizard outro + README now point at the sim commands and the one-time bun link.

* feat(setup): 'bun run sim' is the primary entry; bare invocation prints help

- Lead usage/wizard-outro/README with 'bun run sim <cmd>' (works with zero PATH setup); global bare 'sim' via bun link is an optional upgrade, with the ~/.bun/bin PATH caveat spelled out (Homebrew's bun omits it).
- Bare 'sim' now prints help instead of launching the wizard; the wizard is 'sim setup'. The 'setup' npm script passes the keyword so 'bun run setup' is unchanged.

* fix(setup): quote the auth URL for cmd /c start on Windows

Cursor (High): cmd re-parses the command line and treats & in the query string as a command separator, so cmd /c start opened a URL truncated at the first &, breaking the key flow on win32 (the handoff URL always has request/challenge/pairing). Quote the URL and pass args verbatim so & stays literal.

* fix(setup): verify kube-context is really local; lengthen CLI handoff wait

- k8s: a context named like a local cluster (kind-*, docker-desktop) can actually point at a remote API server. Verify the server host is loopback/docker-internal before defaulting the 'use this context?' confirm to yes; otherwise warn and default to no, so generated secrets can't ship to a remote cluster on a blind Enter.
- cli-auth: bump the device-flow wait from 3 to 15 minutes so first-time users have time to sign up, wait for the email OTP, and approve before the terminal stops polling. The server-side approval record keeps its own short TTL, so a longer client wait only costs cheap rate-limited polls.

* fix(setup): only manage k8s lifecycle on a verified-local context

Greptile: sim down/reset used the ambient kube-context, so switching context after setup could uninstall a same-named sim-dev release from the wrong cluster. Gate k8sInstall on the same locality check the wizard uses (API server is loopback/docker-internal) via a shared isLocalKubeContext helper — the wizard only ever deploys locally, so a remote current-context is never treated as a Sim install.

* fix(setup): doctor skips placeholder secrets when seeding; reset names its target

- checks: the missing-file autofix copied shared keys from apps/sim/.env whenever truthy, including .env.example placeholders — doctor --fix could seed unusable secrets into realtime/db env files. Skip placeholders, matching autofixForMissing.
- lifecycle: reset now names the exact install (k8s context / compose file / dev containers) in its confirm, so a destructive reset can't silently hit the wrong same-named install after a context switch (down already names the context).

* fix(cli-auth): size the poll rate limit to the poll cadence; honor Retry-After

The poll route used the default public-IP bucket (10 burst, 5/min) but the CLI polls every 2s (30/min), so it 429'd within ~20s — worse behind a slow dev cold-compile. Give the endpoint a bucket matched to its cadence (60 burst, 60/min); it's not a brute-force surface (unknown request id returns pending, minting needs the 256-bit verifier). Also make the CLI honor Retry-After and back off on 429 so a shared-NAT per-IP limit degrades gracefully instead of hammering.

* fix(setup): check ports before starting the dev server, not just compose

Local dev auto-start spawned bun run dev:full with no port check, so it silently started a server that couldn't bind when 3000/3002 were already taken (e.g. another worktree's dev server). Extract compose's port-conflict resolver into a shared ensurePortsFree(ports) and run it before the dev start too — kill/recheck/leave, same as compose. Leaving the ports skips the auto-start with guidance instead of failing; compose still treats it as fatal.

* fix(setup): verify the kube cluster is reachable, not just local

A kubeconfig context can outlive its cluster — a kind cluster gets deleted or its Docker container stops (Docker/machine restart), but the context entry remains, pointing at a dead API-server port. The wizard checked the context looked local and handed it to helm, which failed with 'cluster unreachable'.

Add a clusterReachable() liveness probe: only offer the current context when it actually answers; if a local context is dead, fall through to the kind path. There, if kind still knows 'sim' but it's stopped, start its node containers and wait for the API; if it's gone, create fresh. Either way the user gets a working cluster instead of a cryptic helm failure.

* fix(helm): point appVersion at published image tags (v-prefixed, current)

The chart's appVersion was "0.6.73", but CI publishes GHCR tags with a v prefix (its release-commit regex captures v0.7.45). Since sim.image defaults every image tag to Chart.AppVersion, a default helm install requested ghcr.io/simstudioai/{simstudio,realtime,migrations}:0.6.73 — a tag that has never existed — so app and realtime sat in ImagePullBackOff and helm --wait failed with 'progress deadline exceeded'. Any self-hoster installing with default values hit this, not just the setup wizard.

Set appVersion to v0.7.45 (latest release on main; all three images verified present on ghcr) and bump the chart version to 1.1.1. Verified with helm lint, helm template (all images render as v0.7.45), and a live helm upgrade on a kind cluster where the new pods pull successfully while the old 0.6.73 pods remain in ImagePullBackOff.

* Revert "fix(helm): point appVersion at published image tags (v-prefixed, current)"

This reverts commit 28b6047d1d.

* chore(api-validation): rebaseline route count to 977 after staging merge

Staging moved the baseline to 975; this branch's two CLI-auth routes (approve, poll) make 977. The clean merge absorbed the earlier +2 adjustment.

* fix(settings): don't highlight a sibling nav item on nested settings pages

/account/settings/chat-keys is a real page but deliberately not a nav item, so the sidebar's parseSettingsPathSection fell through to defaultSection ('general') and highlighted General — the page read as though it lived inside General.

Resolve the sidebar's active item with a null default so an unmatched nested route highlights nothing, and widen SettingsSidebar's activeSection to string | null. The section feeding the title/description provider keeps its default (pages override title/description anyway), and /account/settings/billing/credit-usage still correctly highlights Billing.

* fix(setup,auth): manage explicitly-confirmed k8s contexts, fail loudly on reset, clear stale post-auth redirect

- lifecycle: detection is now factual — a sim-dev release either exists on the current context or it doesn't. Gating on locality stranded a release the user explicitly confirmed during setup (status/start/stop/down/reset all claimed no k8s install). Locality is recorded instead and surfaced through describeInstall, which every destructive confirm renders, so acting on a non-local cluster is named and defaulted to no rather than silently blocked or silently allowed.
- lifecycle: reset no longer discards helm uninstall's exit status. Env files are archived by that point, so claiming 'Reset complete' while the release still runs is the worst outcome — it now throws with retry/inspect commands.
- auth: signup clears POST_AUTH_REDIRECT_STORAGE_KEY when it has no callbackUrl, and the verification-disabled path consumes it, so a stale CLI/invite destination can't leak into a later flow in the same tab.
2026-07-25 04:24:36 -04:00
Waleed ca77908e20 fix(ci): skip lifecycle scripts on CI installs (#5959)
* fix(ci): set up Node 22 for every job that runs bun install

isolated-vm's install script now runs on install (#5935), and upstream only
publishes prebuilds for Node 22 (ABI 127) and Node 24 (ABI 137). Jobs without
an explicit setup-node inherit the runner default, Node 20, where
prebuild-install finds nothing and falls back to node-gyp — which crashes on
Node 20 with "webidl.util.markAsUncloneable is not a function", failing
bun install --frozen-lockfile outright.

This broke Create GitHub Release on main.

- add setup-node 22 to ci.yml create-release and deploy-trigger-dev
- add setup-node 22 to migrations.yml migrate
- bump publish-cli.yml from the EOL Node 18 to 22, matching publish-ts-sdk.yml

All eight jobs that run bun install now pin Node 22, matching the repo's
declared engines.node >= 22.19.0.

* fix(ci): skip lifecycle scripts on CI installs

No CI job needs a compiled native module. isolated-vm landed in December 2025
and .npmrc blocked all lifecycle scripts until #5935, so CI ran green for
~7 months with it never built: next.config.ts and trigger.config.ts both
external it, every test mocks it, and the only real require lives in
isolated-vm-worker.cjs, which no CI job spawns. All three Dockerfiles already
install with --ignore-scripts and rebuild it by hand.

Building it in CI therefore buys nothing and couples every job to prebuild
availability for the pinned Node. Upstream ships prebuilds for two ABIs only
(Node 22/24), so the next setup-node bump would resurface the same opaque
node-gyp failure that broke Create GitHub Release on main.

- pass --ignore-scripts to all 8 CI bun install invocations
- retarget the Setup Node comments at engines.node >= 22.19.0, which is the
  standalone reason for the pin now that scripts no longer run

* fix(ci): drop redundant setup-node from bun-only jobs

Once lifecycle scripts are skipped, nothing in create-release, migrate, or
deploy-trigger-dev invokes node: they run bun run scripts/create-single-release.ts,
bun run db:push plus bun run scripts/migrate.ts, and bunx trigger.dev deploy.
All three were green without setup-node for months — the isolated-vm install
script was the only thing that ever needed it, and --ignore-scripts covers that.

setup-node stays where it is load-bearing: publish-cli and publish-ts-sdk need
it for npm publish and its registry-url auth, and test-build/docs-embeddings
already had it.

publish-cli keeps the Node 18 -> 22 bump: that setup-node is required, and 18
has been EOL since April 2025, so it now matches publish-ts-sdk.
2026-07-24 20:43:33 -07:00
Waleed cafc9aa140 test(data-drains): cover the migrated detail and create surfaces (#5958)
Executes the two new views rather than reasoning about them, and locks the
behaviours this migration could silently drop:

- the detail carries every column the table used to show (source,
  destination, cadence, last run) and never renders credentials
- run sizes report sub-kilobyte writes instead of flooring to '0 Bytes', and
  a multi-gigabyte run stays in GB
- Run now stays disabled while a drain is disabled, as the row menu did
- delete leaves only after the request resolves, and stays put on failure
- create gates on name plus a complete destination, sends the right
  destination branch, and toasts on failure
- every labelled field in all seven destination forms resolves to a real
  control id, and every select carries an accessible name

The first run caught a defect: dropping toLowerCase from humanizeConfigKey
left keys rendering Title Case ('Force Path Style') while the TSDoc still
promised sentence case. Now sentence case with initialisms preserved —
'Access key ID', 'Service account JSON'.
2026-07-24 20:34:23 -07:00
Vikhyath Mondreti 8a2ae25d78 improvement(agent-streaming): add way to opt in for workflow executions (#5956) 2026-07-24 20:10:06 -07:00
Waleed 49bf3f2c68 fix(settings): a11y labels, URL-backed member search, and design-system cleanup (#5955)
* fix(settings): a11y labels, URL-backed member search, and design-system cleanup

* fix(secrets): use useId for autofill salt so it survives hydration
2026-07-24 19:50:46 -07:00
Waleed e39045d94e improvement(data-drains): move drains to the fullscreen list/detail pattern (#5954)
* improvement(data-drains): move drains to the fullscreen list/detail pattern

Data drains was the last settings surface still on a Table plus a create
modal. It now matches Skills and Custom Tools:

- the list is SettingsResourceRow rows in a clickable button, with the
  source, destination, cadence, and last run on the row and a Disabled tag
  when a drain is paused
- clicking a row opens a detail sub-view: the drain's actions (Run now /
  Test connection / Delete), its enabled toggle, resolved destination
  config, and recent run history — replacing the expanding table row
- creating a drain is a fullscreen view instead of a modal, reusing the
  destination form registry so the two never fork
- a drain's detail is deep-linkable via `data-drain-id`, pushed on open and
  replaced on close, matching custom-tool-id and custom-block-id

The detail is read-only apart from the enabled toggle: destination
credentials are never returned by the API, so changing one means recreating
the drain.

Alignment work the migration surfaced:

- the destination registry rendered its 36 fields with ChipModalField, whose
  label is byte-identical to a SettingsSection header — on a page that made
  section titles and field labels indistinguishable. All of them, and the
  create view's own fields, now use SettingRow like every other page-level
  detail view
- the shared source/destination/cadence label maps moved to labels.ts, now
  that three files need them
- run status uses Badge with a dot rather than hand-coloured text, and byte
  counts use formatFileSize — the old /1024 math showed a 5 GB export as
  "5242880.0 KB"
- copy follows the sibling surfaces: Create drain, Disabled, Run queued,
  Connection test passed, and "No drains found matching …"
- seed the list cache on create so landing on the new drain's detail doesn't
  depend on the invalidation refetch succeeding

* fix(data-drains): wire destination field labels to their controls

The ChipModalField -> SettingRow migration dropped the label/control
association that ChipModalField generated for free, so screen readers
announced the destination fields unnamed and clicking a label focused
nothing. All 35 id-capable controls now carry an id with a matching htmlFor
on their row; the one ChipSelect has no id prop, matching how the existing
SettingRow + combobox rows render.

Also pass includeBytes to formatFileSize — without it any run writing under
1 KB reported '0 Bytes'.

* fix(data-drains): name the select fields for assistive tech

ChipSelect takes an aria-label that lands on its trigger button, so the four
select fields no longer announce as unnamed buttons — the Datadog site plus
the create view's source, cadence, and destination type. The visible
SettingRow label stays; this only gives the trigger an accessible name,
since ChipSelect exposes no id for htmlFor to point at.
2026-07-24 19:49:16 -07:00
Waleed 2272d4c0f0 fix(skills): show the new skill after creating it (#5949)
* fix(skills): show the new skill after creating it

The create page navigated to the new skill's detail route, but the
unsaved-changes guard immediately undid it: `isDirty` was derived from
`createSkill.isSuccess`, so on success the guard's effect fired
`history.back()` to pop its sentinel entry and cancelled the navigation
that had just run. The header stayed on "New skill" with the fields
intact, and nothing confirmed the save.

The guard now exposes `release()` to retire itself for the rest of the
mount, and the create page calls it before navigating with `replace` so
the sentinel entry is consumed rather than stacked.

Also from a cleanup pass over the same surfaces:
- `useCreateSkill` resolves the created row, so the page reads `created.id`
  instead of re-deriving it from the response list
- seed the list cache with the upsert's authoritative list verbatim; the
  id-merge kept the previous ordering and appended the new skill last
- treat placeholder data as loading in the detail page, which could
  otherwise flash "Skill not found." on a workspace switch
- seed credential-detail drafts on id change rather than in a value-keyed
  effect, so a background refetch can't clobber an in-progress edit
- toast on save and on delete failure, which were both silent

* improvement(skills): align the custom tools row and harden the guard lifecycle

Follow-ups from an audit of the skills and custom tools surfaces.

Custom tools row, now matching the skills row exactly:
- the trailing arrow could be squeezed by a long description; the shared row
  owns `flex-shrink-0` for its trailing slot, so the five callers that
  hand-rolled it (and the two that forgot) all get it
- drop a redundant `cursor-pointer` — Tailwind v3's preflight already sets it
  on `button`
- the tile chrome was defined twice, in ResourceTile and again in
  SettingsResourceRow, despite ResourceTile's doc claiming to be the single
  source. Both now share RESOURCE_TILE_BASE/RESOURCE_TILE_FILL, and
  ResourceTile's redundant wrapper div is gone
- skills' row uses the text-sm/text-caption tokens instead of literal pixels;
  the added gap-[1px] offsets the 21px->20px line-height so the row stays 39px

Guard and cache correctness:
- the delete path released the guard after the await, but the delete is
  optimistic — the row leaves the cache first, so the form went clean
  mid-flight and popped the sentinel before the release landed. Release up
  front and rearm if the request fails
- hold the loading frame across the whole optimistic delete instead of
  flashing "Skill not found." on the way out
- custom tools had the same placeholder-data bug just fixed in skill detail,
  and worse: a deep-linked id could resolve against the workspace just left
- keep cached rows the create response omits, so a concurrent create isn't
  dropped until the refetch lands

* fix(skills): retire the in-app back guard on release too

release() suppressed the unload warning and the browser Back trap, but the
in-app back link still keyed only on isDirty — so the Skills chip could open
the unsaved-changes modal after a successful create, while the drafts were
still populated and the navigation was already in flight.

* fix(skills): track a sentinel consumed while the guard is released

release() intentionally leaves the seeded history entry in place, but it also
drops the popstate listener — so Back during the released window (an optimistic
delete's round-trip) consumed that entry with nothing to record it. hasSentinelRef
stayed true, and a failed delete's rearm() then skipped re-seeding, leaving the
surface with no Back confirm despite unsaved edits.

The released branch now keeps a bookkeeping-only popstate listener that clears
the ref, so rearm() seeds a fresh entry when the old one is gone.
2026-07-24 19:20:11 -07:00
Waleed 919a98d00f refactor(settings): fold verified domains into SSO, move group-detail state to nuqs, design-system cleanup (#5950)
* refactor(settings): fold verified domains into SSO and use shared primitives

Verified domains only gates SSO, so managing it on a separate page meant
discovering the requirement after filling out the whole IdP form and then
navigating away mid-setup. Move it into the SSO page as a section above the
provider config and drop the standalone page, its nav entry, and both route
branches.

Align the surfaces with the shared settings primitives rather than bespoke
chrome, matching whitelabeling/custom-blocks/access-control:
- SSO's local FormField (muted labels) is replaced by the shared SettingRow, so
  its fields read like every other settings page. SettingRow gains optional
  `optional` and `error` props to absorb what FormField did — additive, so
  existing consumers are untouched.
- The domains section is built from SettingsSection, SettingRow,
  SettingsResourceRow, and SettingsEmptyState instead of hand-rolled cards.

Also drop the redundant Upload/Change buttons in whitelabeling: the logo and
wordmark thumbnails were already clickable, so the button was a second control
for the same action. Remove still appears once an image is set.

* fix(settings): move group-detail view state to nuqs and clean up design-system drift

Access control's group detail kept its tab, three search boxes, and three status
filters in useState, so a `?group-id=` link always landed on General and a filter
was lost on reload — the parent already puts the group id in the URL. The three
tabs never render together, so search and status share one param each rather than
carrying three mutually-exclusive keys, and switching tabs resets both. Closing
the detail clears all three alongside group-id in one batched write, so nothing
lingers on the list URL.

Design-system fixes from a cleanup pass over the surfaces this branch touched:
- Restore accessible names lost when the whitelabeling Upload buttons were
  removed. The thumbnail is now the only click target, and it contained just an
  icon, so it announced as an unlabeled button; the icon-only Remove had the same
  problem. Both now carry aria-labels reflecting their state.
- Use Chip, not the legacy Button, for the domain actions — Button is ~26px
  against the 30px ChipInput beside it, so "Add domain" sat visibly short.
- SettingRow now uses the emcn Info component; a bare svg as Tooltip.Trigger was
  neither focusable nor nameable. It also stops re-specifying Label's own default
  styling.
- Use the new SettingRow error prop for the group name instead of a hand-rolled
  error paragraph, which is what the prop was added for.
- Hoist the block-category lookup out of a sort comparator, size-* over h/w, name
  the staleTime constants the rules require, and import RowActionsMenu from its
  barrel.

* fix(settings): alias the old /settings/domains path to SSO

Folding verified domains into the SSO page dropped /settings/domains, so
bookmarks and shared links 404'd instead of landing where domains now live. Both
alias maps already exist for exactly this (organization/'members',
subscription/'billing'); add domains -> sso to each.

* chore(settings): adopt ChipCopyInput, named staleTime constants, and a11y labels

* fix(settings): reset group detail params on open and drop issuer mono styling
2026-07-24 19:11:34 -07:00
Vikhyath MondretiandClaude Fable 5 f43b52c569 fix(realtime): evict revoked collaborators from live workflow rooms (#5917)
* fix(realtime): evict revoked collaborators from live workflow rooms via periodic read-access re-validation

* fix

* fix(realtime): close join/eviction race and make sweep cleanups independent

Re-authorize immediately before socket.join so an in-flight join cannot
reverse a sweep eviction, and run the sweep's best-effort cleanups
independently so a room-state failure cannot skip the presence broadcast.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): single-flight role resolution and retry failed eviction cleanup

Coalesce concurrent role resolutions per (user, workflow) so a slow stale
read can never overwrite a recorded revocation, and defer failed eviction
room-state cleanups into a per-sweep retry queue so collaborators are not
left with a stale presence entry.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): scope room removal to the target workflow and detect swallowed cleanup failures

Honor the workflowIdHint as the target room in both room managers so
removing a stale room cannot clobber the mapping of a room the socket has
since moved to, and confirm eviction cleanup via the returned workflowId
plus an unswallowed mapping read so Redis failures actually defer into the
retry queue.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): treat unconfirmed removals as failures and refresh role cache on fresh verify

Treat any null removal result as a failed cleanup (the sweep always passes
the target room, so null only means failure — including with expired
mapping keys), move the same-room rejoin guard to a synchronous check
immediately before the removal, and record verifyWorkflowAccess's fresh
decision into the role cache so a re-granted user is not blocked by a
stale cached revocation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): prefer a mid-flight recorded role decision over the in-flight query result

If a fresh authoritative read (join-time verify) records a decision while
a single-flighted resolution's query is in flight, keep the recorded
decision instead of overwriting it with the potentially stale result.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): isolate the security scan from the Redis cleanup lane

Run the revocation scan (local sockets + DB only) and the best-effort
room-state cleanup as independently-guarded lanes so a hanging Redis
command can stall only presence cleanup, never revocation enforcement.
Evictions now enqueue cleanup instead of awaiting it, and the scan no
longer reads presence for a fallback role.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): bound authorization waits in the revocation scan

Race each socket's authorization check against a per-socket timeout and
cap the whole pass with a budget below the sweep interval, so a hanging
DB query skips that socket for the pass (never evicting on uncertainty)
instead of wedging the scan lane and starving subsequent ticks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(realtime): round-robin the revocation scan so hung checks cannot starve later sockets

Resume each scan pass after the last target the previous pass processed,
so a fixed prefix of hanging authorization checks can never repeatedly
consume the pass budget and leave sockets behind it unexamined.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 16:45:40 -07:00
Theodore Li 9666533d00 fix(chat): sweep the text shimmer left-to-right again (#5926)
#5671 rewrote the shimmer-sweep keyframes to run 100% -> -100%, which is
already left-to-right, but left the `reverse` from #5650 on the animation
shorthand. The two flips cancel into a right-to-left sweep on every
ShimmerText consumer: subagent labels, tool-call rows, and the new
agent-stream thinking chrome.

Drop `reverse` so direction lives in exactly one place — the keyframes.
2026-07-24 19:20:03 -04:00
Vikhyath Mondreti fe184d3695 improvement(whatsapp): validate + improve integration skill for file inputs/outputs (#5942)
* improvement(whatsapp): validate + improve integration skill for file inputs/outputs

* fix lint

* add whatsapp subblock migration
2026-07-24 16:14:27 -07:00
Waleed d64739cf4f fix(ci): unblock @next/swc, lockfile-keyed node_modules, per-image runner sizing (#5945)
* fix(ci): key node_modules sticky disk on the lockfile hash

* improvement(ci): per-image Blacksmith runner sizing + cold-build memory preflight

* fix(deps): exclude @next/swc binaries from the release-age gate

* fix(deps): pin @next/swc binaries so frozen installs get a compiler

* docs(ci): explain ARM runner sizing rationale
2026-07-24 16:12:22 -07:00
Vikhyath Mondreti 17d77795b4 feat(providers): prompt caching capability + usage-based cache pricing (#5922)
* improvement(providers): validation pass, and stream tool loop improvements

* remove deploy options correctly

* fix

* feat(providers): prompt caching capability and usage-based cache pricing

Replace the arbitrary cached-rate heuristic with a single cache-aware pricing
function, and add prompt caching as an opt-in capability for Anthropic.

Pricing: priceModelUsage in cost-policy.ts is now the only place cache
arithmetic happens. Provider adapters normalize their wire shape into
ModelUsage (input always excludes cache buckets); the pricing function never
branches on provider. This removes five divergent behaviors, including the
!!request.context heuristic that gave Router and Evaluator an unearned 10x
input discount, and the overwrite that silently billed Anthropic cache reads
and writes at zero. Also parses OpenAI cache_write_tokens, previously ignored.

Caching: Anthropic gets a capability-gated advanced switch that places
cache_control on the last tool and last system block; system is now always a
TextBlockParam array. OpenAI gets a stable per-block prompt_cache_key with no
UI, since its caching is automatic.

* fix(providers): route OpenAI and Gemini block cost through cache-aware pricing

Cache-aware pricing only reached trace segments. The billable block cost still
called calculateCost on the cache-inclusive prompt total, so OpenAI cache hits
and Gemini implicit-cache hits were charged at the full input rate and GPT-5.6+
cache writes went unbilled.

Both providers now accumulate cache buckets and price through priceModelUsage,
matching the Anthropic token convention where input excludes cache reads and
writes. Cached counts are clamped to the prompt total so an over-reporting
payload cannot bill more input than the request contained.

* fix(streaming): redact tool payloads on selected outputs in public chat

Redaction only ran on the empty-selection branch, but a deployment almost
always selects outputs, so it was dead in the case it exists for. Selecting
toolCalls streamed the raw arguments and results to a public chat client in a
chunk frame, and providerTiming carried thinking content the same way.

Both paths now extract from the sanitized block output rather than the raw log:
the streamed selected output, which is the reachable vector, and the final
envelope. Sanitizing the source rather than per selected path means a newly
selectable field cannot reopen the hole.

* refactor(providers): drop unreachable billing fallbacks

Every provider pricing helper took a policy parameter no caller passed. Worse
than dead: passing one would have double-applied the margin the central layer
already applies. Removed, so providers can only price at list.

Also removed guards that cannot fire. The central fallback normalized cache
buckets no provider can reach it with (all three that report cache usage price
themselves) and did so at a 1x write multiplier no vendor charges.
priceModelUsage re-validated token counts the adapter had already clamped, and
applyModelCostPolicy defaulted a required total field.

Validation now happens once, in the adapter that parses the vendor payload and
is the only layer that knows cache buckets are a subset of the prompt total.
2026-07-24 15:46:11 -07:00
Waleed 290e52c3f3 fix(ci): build app on 16vcpu runner to stop cold-cache OOM kills (#5944) 2026-07-24 15:37:35 -07:00
Waleed 6ec2385928 chore(ci): run companion-pr-check on Blacksmith (#5943) 2026-07-24 15:34:13 -07:00
Waleed 90f6708270 improvement(helm): hygiene pass — CI gating, strict values schema, ESO v1 default (#5939)
* improvement(helm): hygiene pass — CI gating, strict values schema, ESO v1 default, ci values

Closes the gaps from a best-practices audit of the chart (template-level
conformance was already clean: full label set, 82 unit tests, kubeconform-
valid renders):

- new Helm Chart workflow gates every chart change: helm lint, the 82
  helm-unittest cases, kubeconform validation of default + all-components
  renders (k8s 1.29 strict, CRD catalog), and a render of all 10 example
  values files
- helm/sim/ci/ values files (chart-testing convention) so the chart lints
  and templates cleanly out of the box with dummy secrets
- values.schema.json declares all 30 top-level keys (16 were invisible) and
  sets root additionalProperties: false, so top-level typos fail fast
- externalSecrets.apiVersion defaults to v1: current ESO releases removed
  the v1beta1 compatibility path in 2026, so the old default produced
  rejected manifests on new installs; NOTES/values comments updated
- wait-for-postgres init container gets requests/limits (the only container
  in the chart without them; broke ResourceQuota'd namespaces)
- drop the telemetry Prometheus scrape config for app/realtime — neither
  exposes /metrics, so it was dead config that also rendered a realtime
  target with realtime disabled
- chart 1.2.0 with README upgrade notes

* fix(helm): version-agnostic schema-error assertions in the secret-length suite

Helm v4 phrases schema rejections as 'minLength: got N, want 32' while
v3 says 'String length must be greater than or equal to 32'. The suite
grepped for 'minLength' only, so it passed on local helm v4 and failed on
CI's v3.16.4 — the enforcement itself works on both. Patterns now assert
the key name plus either wording. Caught by the new Helm Chart workflow
on its very first run.

* improvement(helm): kind install test + chart version-bump gate in CI

Benchmarked against the flagship OSS charts (ingress-nginx, argo-cd,
kube-prometheus-stack, grafana, bitnami, cert-manager): sim already exceeds
most of them on validation rigor (strict schema — 4 of 6 ship none; 82 unit
tests vs argo-cd's zero; kubeconform manifest validation none of them run),
but every top community chart repo actually installs the chart on a kind
cluster in chart CI — the one majority practice we lacked. Adds:

- install job: kind cluster, helm install with ci/default + a new
  small-footprint ci/kind-values.yaml overlay (default app requests of 4Gi
  can't schedule on a CI node), --wait, then the chart's helm test hook,
  with pod/event/log diagnostics on failure
- version-bump job (PR-only): fails when helm/sim/** changes without a
  Chart.yaml version increment — argo/kps/grafana all enforce this

* fix(helm): declare naming overrides in the strict schema; SHA-pin CI actions

- nameOverride/fullnameOverride are consumed by sim.name/sim.fullname but
  were never in values.yaml, so root additionalProperties: false rejected
  Helm's standard naming overrides — both now declared (a helper-wide sweep
  confirmed they were the only template-read keys missing), with a smoke
  regression test that installs under the strict schema and asserts the
  override lands in resource names
- CI supply-chain hardening: checkout/setup-helm/kind-action pinned to full
  commit SHAs (tag comments retained) and the kubeconform archive verified
  against its published sha256

* fix(helm): immutable unittest runner and fail-closed version gate in CI

- helm-unittest runs via the project's official docker image pinned by
  immutable sha256 digest instead of a plugin install from a mutable git tag
- the version-bump gate fetches the full base ref (a --depth=1 fetch could
  leave no merge base), computes the merge-base and diff outside the if so
  any git failure fails the job instead of falling into the skip branch

* fix(ci): run the helm-unittest container as the runner UID

The digest-pinned image runs as a non-root user that cannot write into the
runner-owned bind mount (it creates tests/__snapshot__, absent from the
checkout since empty dirs aren't tracked). Standard bind-mount pattern:
--user "$(id -u):$(id -g)" with HOME=/tmp for helm's cache.

* fix(helm): appVersion points at a real GHCR tag (v0.7.44)

The kind install job caught this on its first full run: the default image
tag (Chart.AppVersion 0.6.73) returns 404 on GHCR for all three images —
the registry's tags are v-prefixed — so an unpinned default install could
never pull. Updated appVersion to v0.7.44 (verified 200 for simstudio,
realtime, and migrations manifests), stale values comment refreshed, and
an upgrade note added. The CI kind values deliberately stay tag-free so
the job keeps exercising the true default path.
2026-07-24 14:25:46 -07:00
mzxchandra f0f98708fc refactor(blocks): rename Sim block to 'Sim Chat' for clarity (#5933) 2026-07-24 12:15:43 -07:00
Waleed f224d02e27 improvement(library): align enterprise copy with canonical source (#5936)
Reconcile compliance and traction wording across three library guides
against the canonical data in lib/compare/data/sim.ts, and tighten a couple
of related phrasings for consistency.
2026-07-24 12:08:08 -07:00
Waleed 91a9c20acc fix(deps): build isolated-vm on install via root trustedDependencies (#5935)
.npmrc set ignore-scripts=true globally, which overrode Bun's
trustedDependencies allowlist and blocked every lifecycle script —
including the repo's own root scripts. isolated-vm therefore shipped
unbuilt and each contributor had to npm rebuild it by hand.

The .npmrc line landed in the May 2025 Bun migration, seven months before
isolated-vm was introduced. Two later commits added isolated-vm to
trustedDependencies (root, then apps/sim) and both were silently inert.

- delete .npmrc; lifecycle policy now lives solely in root trustedDependencies
- add isolated-vm to root trustedDependencies
- drop the workspace-level trustedDependencies from apps/sim (Bun only reads
  the root list, so it never had any effect)
- drop the .npmrc entry from CODEOWNERS

Naming any package in trustedDependencies replaces Bun's curated
default-trust list, so only isolated-vm and sharp may run install scripts.
Docker is unaffected — all three Dockerfiles pass --ignore-scripts
explicitly and rebuild isolated-vm against the Node ABI by hand.
2026-07-24 12:03:55 -07:00
Waleed 5a8d21904d fix(sso): surface DNS verification failures and the provider auto-append gotcha (#5931)
* fix(sso): surface DNS verification failures and the provider auto-append gotcha

Post-merge audit follow-ups for verified domains (#5909):

- The host field handed admins the FQDN `_sim-challenge.acme.com`. GoDaddy,
  Namecheap, Hover and most cPanel panels append the zone to whatever is typed,
  yielding `_sim-challenge.acme.com.acme.com` — the record looks right in their
  panel but never verifies, and our 422 tells them to wait 48 hours. Add a hint
  under the field (and a docs callout) telling those admins to enter just the
  label.
- DNS failures were logged at debug, but production log level is ERROR, so a
  blocked-egress or SERVFAIL condition was invisible to us and misreported to the
  admin as "record not found yet". Log infrastructure-class failures at warn with
  the DNS error code; keep the genuinely-absent codes at debug.
- Trim the joined TXT value before comparing: several DNS panels pad the stored
  string, which otherwise fails an exact match forever.
- Lower the resolver to 2s/1 try. c-ares multiplies timeout across servers and
  retries by ~7x, so the previous 5s/2-try config could block a verify request
  for ~35s when resolvers are unreachable.
- Cover `checkDomainTxtRecord` — the function that decides whether the gate opens
  had no tests. Adds exact match, chunk-joined value, match among unrelated
  records, padded value, near-miss, another org's token, absent record,
  infrastructure failure, and empty-response cases.

* fix(sso): log DNS infrastructure failures at error so prod actually surfaces them

Production's default minimum log level is ERROR, so the warn introduced in the
previous commit was still filtered out — the fault stayed invisible exactly as
before. A resolver failure that is not 'record absent' is a genuine
infrastructure error, so ERROR is both the visible and the honest severity.

* fix(sso): state the zone-removal rule instead of a wrong subdomain hint

The hint computed the bare label as the first segment of the challenge host, so
for a subdomain like eng.acme.com it advised entering `_sim-challenge` when the
host relative to the acme.com zone is `_sim-challenge.eng` — following it would
publish the record on the wrong name and verification would never succeed, the
exact failure the hint exists to prevent. Deriving the real zone needs the Public
Suffix List (acme.co.uk defeats naive label-stripping), so state the rule instead:
enter the host with the trailing zone removed. Docs show both the apex and the
subdomain form.
2026-07-24 11:32:27 -07:00
Waleed 6105376593 improvement(ship): pre-flight regenerate artifacts + parallel audit suite (#5927)
* improvement(ship): pre-flight regenerate artifacts + parallel audit suite

Add a two-phase pre-ship step: (A) regenerate every committed artifact
(agent-stream-docs, skills, contract syncs) in parallel so generated files
never drift into a CI failure, then (B) run lint plus the full audit suite
CI's Lint and Test job enforces, in parallel, aborting ship on any failure.
Regenerate the ship command projections.

* fix(ship): scope Phase A to in-repo generators + propagate failures

Cursor review: (1) mship:generate is an umbrella over the 9 contract
generators and biome-formats the shared generated dir, so running it
alongside its constituents races/corrupts; it also reads an external
copilot-contract source that ENOENTs in a bare worktree. (2) bare wait
swallowed generator exit codes. Narrow Phase A to the always-in-repo
generators (agent-stream-docs, skills) and collect each job's exit
status in both phases; document domain generators as on-demand only.

* fix(ship): make pre-flight status checks exit non-zero on failure

Cursor/Greptile: grep && echo ❌ || echo ✅ always exits 0, so a failed
generator/audit didn't actually gate ship (regressed the sequential
bun-run checks whose non-zero exit agents rely on). Replace both Phase A
and Phase B aggregations with an if-grep that exits 1 on any failure.

* fix(ship): gate lint exit in Phase B pre-flight

Cursor: bare bun run lint before the audit grep meant a non-zero lint
was ignored — if audits then passed, the block exited 0 and ship
continued. Gate lint with || { echo …; exit 1; } so an unfixable lint
error aborts before the audits run.
2026-07-24 11:01:20 -07:00
Waleed e83d8b84b4 content(library): add observability, procurement, and MCP security guides (#5928)
* content(library): add observability, procurement, and MCP security guides

Three AEO guides, each dated into a recent empty day in the posting
calendar (2026-07-19, 07-21, 07-22) so publication is spread across days
rather than dumped on one.

- ai-agent-observability: what observability is, why APM falls short for
  non-deterministic agents, what to instrument per lifecycle stage, and the
  signals to track
- ai-agents-in-procurement: what procurement agents do, buy-vs-build, where
  they add value, and how to start narrow
- mcp-security: tool poisoning, confused-deputy/OAuth flaws, and supply-chain
  risk, plus how to build and govern MCP servers securely

Every third-party factual claim carries an outbound citation to a verified
primary source (OpenTelemetry semconv, Fiddler, LangChain and PwC surveys,
Icertis/ProcureCon, Ironclad, GEP, the MCPoison/CurXecute CVEs on NVD,
Anthropic's MCP announcement, RFC 8707/9700, OAuth 2.1, the MCP auth spec),
and each post carries 3 internal links to related library posts. One claim
from the source copy — a "November 2025" date on the WhatsApp MCP exfil case
— could not be verified and was dropped; the case itself is cited to Docker's
writeup. Covers come from the autogeneration pipeline via the standard
/library/<slug>/cover.jpg path.

* fix(library): correct two unverifiable claims in the new guides

Accuracy pass on the three new posts turned up two claims that could not be
substantiated:

- HIPAA compliance: the procurement post claimed Sim has "SOC2 and HIPAA
  compliance," but the canonical compliance data (lib/compare/data/sim.ts)
  states SOC2 only, and explicitly that Sim offers self-hosting "beyond SOC2,
  rather than additional certifications." Removed HIPAA; kept SOC2 plus
  self-hosting for data residency. (The same claim exists in ~5 pre-existing
  library posts and should be corrected separately.)
- "more than 18,000 servers were listed on MCP Market": no source
  substantiates this figure. Replaced with "thousands of community-built
  servers," which the ecosystem supports, keeping the Anthropic and MCP Market
  links.

Every other third-party claim was verified against a primary source (both
CVEs on NVD, RFC 8707 title, RFC 9700 as January 2025, and all six survey
statistics).
2026-07-24 10:56:44 -07:00
Waleed 70313cd2ad feat(providers): add Claude Opus 5 model (#5925)
* feat(providers): add Claude Opus 5 model

* fix(providers): route Opus 5 through adaptive thinking

Opus 5 supports only adaptive thinking (no extended thinking / budget_tokens),
same as Opus 4.8/4.7. Add opus-5 to supportsAdaptiveThinking so thinking
requests use thinking.type: adaptive instead of falling through to budget_tokens
extended thinking, which Opus 4.7+ rejects with a 400.

* docs(agent): regenerate agent-stream docs for Opus 5
2026-07-24 10:33:47 -07:00
Waleed 513292f17b feat(sso): DNS domain verification gating org SSO registration (#5909)
* feat(sso): DNS domain verification gating org SSO registration

Add org-scoped domain ownership verification (DNS TXT challenge) as the
security precondition for configuring SSO. Closes the first-come domain-claim
vuln where any org could wire another company's domain to its own IdP.

- New sso_domain table + migration 0266; existing org SSO domains are
  grandfathered as verified so live tenants are unaffected
- Verified-domains settings UI (enterprise-gated) with add/verify/remove
- Register route now requires a verified domain for org-scoped registration;
  personal SSO and already-grandfathered domains are unaffected
- Self-host register script writes the verified sso_domain row directly, so
  script-driven registration stays backwards compatible

* fix(sso): harden domain verification against concurrency + fix CI lint

Addresses review findings on state invariants under concurrent/failed writes:
- Add unique index on (organization_id, domain) so concurrent claims can't
  create duplicate pending rows; POST re-reads and stays idempotent on conflict
- Verify flips the row only if it's still the exact pending challenge checked
  (guards deletion/token-rotation mid-DNS-lookup) and maps the partial unique
  index violation to 409 instead of an unhandled 500
- Wrap the self-host script's provider write + verified-domain upsert in a
  transaction so a failed ownership write can't leave a provider committed
- Format 0266 snapshot/journal with biome (fixes @sim/db lint:check)

* fix(sso): re-check domain verification before provider write (TOCTOU)

The register gate checked the verified sso_domain row only at handler entry,
then ran OIDC discovery before writing the provider. A verified row removed
during that window could still complete registration. Extract the check into a
closure and call it both as an entry fast-fail and authoritatively right before
registerSSOProvider, alongside the existing domain-conflict re-check.

* fix(sso): stop rotating verification token on idempotent re-add

Re-adding a pending domain rotated its verification token, which invalidated a
TXT record the admin may have already published and — under two concurrent
re-adds — could return a token the racing write had already superseded, so the
admin's DNS record would never verify. Return the existing row unchanged
instead; the pending token is always shown in the UI, so it is never lost.

* fix(sso): close register TOCTOU with compensating delete + harden edges

Audit-driven hardening:
- Close the residual register TOCTOU: registerSSOProvider is create-only (throws
  if the providerId exists), so a compensating delete after the write is provably
  safe — it can only remove the just-created row. If verification was revoked
  during the write, roll the provider back and 403.
- Verify is now idempotent under concurrency: a same-org row already flipped to
  verified by a racing request returns 200, not a confusing 409.
- Grandfather backfill + self-host script now match normalizeSSODomain's dominant
  transforms (lower + trim + strip leading wildcard) so a non-canonical legacy
  domain can't miss the runtime gate's lookup. Prod backfill result is unchanged.
- Cleanup: drop dead default export, align card radius to sibling convention.

* fix(sso): redact domain tokens from non-admins + fix script stale-update

Round-5 review findings:
- GET /domains redacted the pending TXT verification token (a management
  secret) to any org member. Now only owner/admins read it; members see the
  list and status without it. Non-Enterprise orgs get an empty list (entitlement
  flag only), never the domains/tokens.
- Self-host script decided update-vs-insert from a read taken OUTSIDE the
  transaction; a provider deleted mid-flight made the UPDATE match zero rows
  silently while the verified-domain upsert still committed (orphaned domain).
  The decision now happens inside the transaction from the UPDATE's row count.

* docs(sso): drop unshipped enforce-SSO / auto-join copy from verified domains

Verified domains currently only gate SSO configuration. Remove the
forward-looking references to enforcing SSO and auto-joining members (deferred
to a later release) from the docs, settings copy, nav description, and schema
comment so we don't promise unshipped features.

* fix(sso): guard rollback to new providers only + Enterprise-gate domain removal

Round-6 review findings:
- The compensating provider rollback now only fires when the provider did not
  exist before this request (providerExistedBefore). registerSSOProvider is
  create-only today so reaching the rollback already implies a fresh create, but
  this makes the safety local and future-proof: if Better Auth ever allowed
  updating an existing provider, a revoked-verification rollback must not delete
  that pre-existing row.
- DELETE /domains now requires an Enterprise plan like add/list/verify, so all
  domain mutations share one entitlement (the UI already hides removal from
  non-Enterprise orgs). Adds a delete-route test.

* fix(sso): roll back the SSO provider by row id, not logical keys

The compensating rollback deleted by (providerId, orgId). providerId is unique,
so if this request's row were deleted and recreated by a concurrent registration
in the narrow window before the rollback, the logical-key delete would remove
that other request's provider. Delete by the primary-key id registerSSOProvider
returns instead, so only the exact row this request created is ever removed.

* chore(sso): final-review polish — trim script read, unify copy, doc migration edge

Cosmetic cleanup from a final 4-track adversarial review (no bugs found in the
new logic):
- Self-host script: narrow the pre-transaction existence read to select({ id })
  instead of SELECT * (it only feeds a log line now).
- Unify invalid-domain copy ("for example acme.com") and the verified-elsewhere
  409 wording ("is already verified by another organization") across routes.
- p-3 shorthand on the domain row card.
- Document the migration's rare two-orgs-share-a-domain grandfather behavior
  (login unaffected; validated no such duplicates in prod).

* fix(sso): apply attribute mapping + make SSO edit work; drop dead guard

Two pre-existing SSO bugs the final review surfaced (prod has one SSO org, RVW,
script-registered with the default mapping, so neither change affects it):

- Attribute mapping was passed at the top level of the register payload, which
  Better Auth ignores — it reads oidcConfig.mapping / samlConfig.mapping. Nest it
  so custom mappings actually apply. (Default mapping is unchanged, so existing
  logins are unaffected.)
- Editing an SSO provider was broken: registerSSOProvider is create-only and
  threw on the existing providerId → generic 500. Route now detects a provider
  the caller already owns and updates it via Better Auth's updateSSOProvider, and
  surfaces Better Auth's own error status/message instead of a blanket 500.

Also drops the now-unnecessary providerExistedBefore guard (the rollback deletes
by the created row's primary-key id and register is create-only) and the earlier
final-review polish (script read, unified copy, migration edge note).

Smoke-test SSO login + edit on staging before merge (auth-path change).

* fix(sso): require null org on personal-mode provider lookups (gate bypass)

The personal branch of both provider-ownership lookups keyed on
(providerId, userId) without requiring organizationId IS NULL. Because org
providers store userId = their creator and providerId is globally unique, an org
admin could send a personal-mode request (no orgId) — which skips the membership
check and the domain-verification gate — yet still match, and then via the new
update path move, their org's provider to an unverified domain. Add
isNull(organizationId) to the personal branch of both clauses so it can only
match a genuinely personal provider, matching the route's own isOwnedByCaller.

Found by an adversarial review of the update path added in 394bda9f7.

* fix(sso): script updates the observed provider by id, not providerId

Inside the registration transaction the script updated WHERE providerId — the
logical key. If the observed provider was deregistered and a replacement created
with the same providerId before the transaction ran, that update would clobber
the replacement's config and ownership. Update the specific observed row by its
primary-key id instead; if it's gone we insert, which fails cleanly on the
providerId unique constraint rather than overwriting the replacement.

* fix(sso): script upserts provider via delete-then-insert (no unique constraint)

sso_provider.provider_id is a plain (non-unique) index and prod holds legitimate
duplicates, so the previous "update by id, else insert" could create a duplicate
provider when the observed row was deregistered and replaced before the
transaction — the fallback insert would succeed. Delete every row for the
providerId then insert exactly one, inside the transaction, so the providerId
ends up as exactly this config atomically. Linked accounts key on the providerId
string (not the row id), so existing logins are unaffected.

* fix(sso): guard compensating-delete row id so rollback can't silently no-op

* chore(sso): regenerate migration as 0268 after merging staging

Staging landed migrations 0266/0267, colliding with our 0266. Removed our
migration, merged staging, and regenerated cleanly with drizzle-kit as
0268_sso_domain_verification (identical sso_domain table + indexes), then
re-appended the grandfather backfill. api-validation baseline reconciled to 973
(staging 970 + our 3 domain routes). Also make the register-route test's
registerSSOProvider mock return an id so the guarded compensating delete runs.

* refactor(sso): share normalizeSSODomain via @sim/utils so script matches gate

The self-host script canonicalized SSO domains with a minimal inline transform
(lower+trim+wildcard) that diverged from the app's full normalizeSSODomain
(protocol, port, path, trailing dot, email local part) — equivalent spellings
could store a different ownership key than the runtime gate looks up. Move
normalizeSSODomain into @sim/utils/sso-domain (a pure function) so the register
route, the domain-claim route, and the script all use the identical canonicalizer.
The script now skips the verified-domain record when SSO_DOMAIN isn't a valid
registrable domain instead of storing a malformed key.
2026-07-24 00:34:16 -07:00
Theodore Li 444c415a0b improvement(data-retention): docs for overrides + PII redaction, fix wedged saves (#5905)
* fix(data-retention): clamp sub-day retention values so saves aren't wedged

A stored value under 12 hours rounded to '0' days on load and was re-sent as 0, which the contract rejects (min 24) — blocking every save on the page, including unrelated fields.

Clamp hours->days into the contract's range on read, and throw instead of emitting 0/NaN on write.

* improvement(docs): rewrite data retention for workspace overrides + PII redaction

- Document PII redaction (Logs / Workflow input / Block outputs stages, entity types, languages, custom regex patterns)
- Replace the stale 'no per-workspace overrides' section with the retention-policies list and override inheritance
- Correct log retention (also covers background job logs) and soft deletion (adds Chat conversations, KB documents)
- Add PII + override screenshots, refresh the main one
2026-07-24 02:08:12 -04:00
Vikhyath MondretiandSim Pi Agent e6eef4a4bf feat(library): What Is an MCP Server? (#5919)
Co-authored-by: Sim Pi Agent <pi@sim.ai>
2026-07-23 21:23:26 -07:00
Vikhyath Mondreti 7120cdc901 fix(providers): final regenerated stream must not re-call tools (empty chat answers) (#5915)
* fix(providers): stop the final regenerated stream from re-calling tools and clobbering the answer

After the silent tool loop settles, OpenAI Responses and Gemini re-issue a
streaming request purely to stream the answer as prose — but with tools still
attached and auto tool choice, a reasoning model can re-decide to call a tool
there. Streamed calls are never executed on this path, so the run ends with a
dead function call, an empty streamed answer, and the stream callback
clobbering the tool loop's settled text with '' (deployed chat rendered
{"content": ""}).

Force tool_choice 'none' / functionCallingConfig NONE on the regeneration and
keep the tool loop's settled answer whenever the stream ends without text.

* feat(openai): settled tool chips on the regenerated answer stream

The silent Responses tool loop has no live stream while tools run, so opted-in
consumers saw no tool chips at all for OpenAI. The loop now records each
executed call and prepends settled tool_call_start/end pairs (name + status
only) to the agent-events stream ahead of the regenerated answer. Runs without
a sink never see these events, so legacy output is unchanged.

* fix(providers): apply the regeneration guard fleet-wide

Audit of every silent tool loop for the same race fixed for OpenAI/Gemini
(final regenerated stream re-calls a tool that is never executed, ending with
an empty answer that clobbers the settled text):

- anthropic (both implementations): tool_choice {type:'none'} on the
  regeneration (tools must stay — history carries tool_use blocks) + keep the
  settled answer when the stream ends without text
- groq: was re-applying the ORIGINAL tool_choice, so forced-tool runs
  re-forced the tool on the regeneration — guaranteed dead call; now 'none'
  + fallback
- deepseek, mistral, cerebras, azure-openai (legacy chat path), openrouter,
  xai: 'auto' -> 'none' + fallback
- bedrock: fallback only — Bedrock's ToolChoice has no 'none' and toolConfig
  is required when history carries toolUse blocks

Already guarded (no change): meta, sakana, nvidia, vllm, litellm, baseten,
together, fireworks, kimi, zai, ollama.
2026-07-23 20:20:25 -07:00
Waleed bee45621ec refactor(skills): one definition of the skill fields across every surface (#5916)
* refactor(skills): one definition of the skill fields across every surface

The Name / Description / Content trio was implemented three times — once in the
canvas modal and once each in the create and detail pages — and had already
drifted: the name hint appeared on create, was missing entirely on detail, and
sat in a different slot in the modal; the content editor was 260px on the pages
and 200px in the modal; only the modal marked the fields required.

Extract SkillFields, the DetailSection form both full-page surfaces render, and
move the shared copy into skill-copy.ts. The modal keeps ChipModalField — that
is required inside a ChipModalBody, so it cannot share the pages' JSX — but now
reads the same placeholder, hint, and max-length constants, so the wording can
no longer diverge.

The detail page's local FieldLockTooltip moves into SkillFields and is now
available to both pages, and the dynamic rich-editor import drops from three
copies to two.

No behavioral change beyond the drift being resolved: the name hint now shows on
the detail page as it already did on create.

* refactor(skills): derive modal saving state, collapse the field message line

From the cleanup passes over this branch:

- SkillModal mirrored `mutation.isPending` into a local `saving` useState, which
  .claude/rules/sim-hooks.md forbids. It also carried a latent bug: the
  `finally { setSaving(false) }` ran after `onSave()` had closed the modal, so it
  set state on an unmounting component, and a throw from `onSave()` would leave
  the modal stuck in a saving state. Derived from the two mutations instead.
- The hint/error line under a field was written three times in SkillFields, in
  three slightly different shapes. One FieldMessage helper now renders all three.
- The name hint is an authoring instruction, so it no longer shows under a field
  that cannot be edited (a built-in skill, or a viewer who is not an editor).
2026-07-23 20:14:37 -07:00
6dcc65be89 feat(skills): add skill editors (#5705)
* feat(skills): permissions layer

* chore(db): drop skill_member migration 0261 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0262 on latest staging

Same DDL as the dropped 0261 (skill_member table, enums, indexes, skill.workspace_shared)
plus the hand-written write-user backfill, renumbered after staging's 0261.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0263 after staging merge

Staging claimed 0262 (strong_storm); same DDL plus the hand-written
write-user backfill, renumbered on the merged snapshot chain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(db): drop skill_member migration 0263 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0264 after staging merge

Staging claimed 0263 (workflow_fork_sync_excluded); same DDL plus the
hand-written write-user backfill, renumbered on the merged snapshot chain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* make editing skills full page

* fix disclaimer

* edit access msg

* fix lint

* chore(db): drop skill_member migration 0264 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0265 after staging merge

Staging claimed 0264 (fat_ikaris); same DDL plus the hand-written
write-user backfill, renumbered on the merged snapshot chain.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix tests

* simplify system

* fix

* fix lint

* add mship skills docs

* chore(db): drop skill_member migration 0265 for regeneration on latest staging

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(db): regenerate skill_member migration as 0266 after staging merge

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(deps): override zod to 4.3.6 to dedupe nested copies breaking type-check

better-auth 1.6.23 and fumadocs-mdx resolve ^4.3.6 to a nested zod 4.4.3,
which makes @sim/auth's inferred betterAuth types non-portable (TS2883) and
split docs onto a second zod instance. Both ranges accept the repo-wide
pinned 4.3.6, so a single hoisted copy satisfies everything.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix lint

* feat(skills,tools): fullscreen skill create + shared custom tool editor

Moves the rich-markdown and custom-tool editing surfaces out of modals and
onto full-page surfaces, and collapses the duplicated chrome behind shared
components.

Skills
- Add /skills/new, a full-page create surface mirroring the skill detail page
  (CredentialDetailLayout + DetailSection + unsaved-changes guard). "Add to
  Sim" navigates there instead of opening a modal.
- Import moves to a header action (SkillImportButton) backed by a shared
  readSkillFile helper; the GitHub-URL import and its /api/skills/import route
  are removed.
- Skill name validation is now one shared validateSkillName, replacing three
  copies of the kebab-case rule and its messages.
- The skill editor roster renders through the shared MemberRow instead of
  re-deriving its identity block, with a locked role control and a lock-reason
  tooltip explaining inherited workspace-admin access.

Custom tools
- Extract the canvas modal's schema/code editors into a shared
  custom-tool-editor module (fields, wand generation, schema helpers), cutting
  custom-tool-modal.tsx by ~900 lines.
- Settings > Custom tools gains a full-page detail sub-view (SettingsPanel +
  SettingsSection + saveDiscardActions), deep-linkable via ?custom-tool-id.
  Rows are clickable; delete now lives only in the detail view.
- Replace legacy Button/Input/Badge/Label with the chip family, move chip-field
  chrome into CodeEditor behind an error prop, and delete its dead wand button.

Rich markdown field
- maxHeight is now opt-in: omit it on a page and the editor grows with its
  content so the page owns the only scrollbar. Modals pass explicit caps.
- The field variant drops to font-weight 400 to match adjacent chip fields.

* fix(skills): address review round on create navigation, 409 copy, and editor audit

- Skill create navigated using the first element of the upsert response, but
  that endpoint returns the caller's whole skill list (built-ins prepended) —
  match the new skill by its workspace-unique name instead.
- The suggested-skill 409 toast claimed the skill existed but was not shared
  and told the user to ask a skill admin. Every workspace member can already
  see and use every skill, so a 409 only means the name is taken.
- Adding an editor emitted the skill_shared event and SKILL_MEMBER_ADDED audit
  even when onConflictDoNothing skipped the insert on a concurrent add. Gate
  both on the insert actually returning a row.

* chore: format skills-resolver test import

* fix(skills,tools): audit fixes — autocomplete boundary, resize clipping, error routing

Two real regressions introduced while simplifying the extracted editor:

- The schema-param autocomplete's trigger was rewritten to match a trailing
  identifier, but the completion still split on separators. The two disagreed,
  so typing `data.ci` opened the menu and selecting replaced `data.ci` whole —
  eating the member-access prefix. Both now share one SCHEMA_PARAM_WORD regex.
- The uncapped markdown field measured its height only on value change while
  always setting overflow-hidden, so any width change that re-wrapped lines
  clipped the tail with no scrollbar to reach it. Now re-measures via
  ResizeObserver.

Also from the audit:
- Generation writes bypass the code field's change handler, so an open
  autocomplete stayed over a disabled streaming editor; close it on busy.
- Delete failures rendered in the Schema section's error slot on both custom
  tool surfaces; route them to a toast instead.
- Skill create navigated away while still dirty, stranding the unsaved-changes
  guard's history sentinel so Back landed on an empty create form.
- The Description field on skill create never received its error border.
- Drop a double-applied opacity-50 (the editor already dims when disabled), a
  dead try/catch around a non-throwing call that also shadowed the error prop,
  and a stale reference to a /tools page that does not exist.
- Docs still described the removed GitHub-URL import and the old Add Skill
  dialog; rewrite for the create page and file/paste import.

* feat(tools): read-only tool detail, create lands on the new tool, drop dead wand prompt API

- Viewers without edit rights could not open a custom tool at all, while the
  equivalent skill and custom-block surfaces both offer a read-only view. The
  detail page now takes `readOnly`: editors inert, no Save/Discard/Delete, no
  Generate. Creating still requires edit rights.
- Creating a tool bounced back to the list while creating a skill lands on the
  new skill. Tools now do the same. The upsert returns the workspace's whole
  tool list (newest first) rather than just the new row, so the id is matched
  by title instead of by index — the same trap that produced the skill-create
  navigation bug.
- Remove `openPrompt`/`closePrompt` from useWand. `closePrompt`'s last callers
  went away with the custom-tool-modal extraction and `openPrompt` had none
  before it; nothing reads `isPromptVisible` any more either.

* fix(tools): read-only editors, design-system wrench, skills-matching tool identity

- readOnly never reached the editors: the prop gated actions and Generate but
  the schema and code fields were still typable for viewers without edit rights.
  Wire disabled through both fields into CodeEditor.
- The row icon used lucide's Wrench (strokeWidth 2) where @sim/emcn/icons ships
  one drawn for this system (1.55, tuned viewBox), and it inherited body text
  colour instead of --text-icon. Swap it.
- Give the tool detail page the same identity heading as skill detail: tile,
  name, and description at the top left, instead of only a header title.
- Extract ResourceTile so the skills and tools tiles share one definition
  (SkillTile now composes it), and add an opt-in `iconFilled` to
  SettingsResourceRow so the tools list tile matches the skills gallery. Both
  default to today's behaviour for every existing consumer.

* fix(mentions): use the product's own glyph for every @ mention kind

The `@` menu and the inserted chip mapped kinds to arbitrary lucide icons —
`Sparkles` for a skill, a generic `File` for every file — while the rest of the
product has a settled glyph per resource. Mirror CHAT_CONTEXT_KIND_REGISTRY,
which Chat's `@` menu already renders from:

- skill now uses AgentSkillsIcon, the same glyph SkillTile shows everywhere
- workflow / folder / table / knowledge use the @sim/emcn/icons set the sidebar
  and the chat registry use
- file derives its icon from the filename extension, so a .pdf and a .csv are
  distinguishable, matching the file list and Chat's context chips
- integration keeps the block's brand icon from the registry

Also drop the generic placeholder. `kind` is untrusted — the node schema
defaults it to `''` and a hand-written `sim:` link can carry anything — but an
unrecognized kind now yields no icon instead of a meaningless box, which is what
the chat registry does. The menu already guarded a missing icon; the chip now
does too, so this cannot crash on a malformed link.

* chore(db): drop skill_member migration 0266 for regeneration on latest staging

* feat(db): regenerate skill_member migration as 0267 after staging merge

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Waleed Latif <walif6@gmail.com>
2026-07-23 19:52:00 -07:00
d24bc7eccb feat(agent-stream): thinking and tool streaming (#5671)
* feat(agent-stream): add agent-events thinking/tool streaming for chat and canvas

Ship the agent-events-v1 protocol with provider tool loops, dual-gated chat thinking, DeepSeek/Groq/OpenAI reasoning wiring, and ChatGPT-like thinking chrome.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): clear stuck streaming UI and format db snapshot

Biome was failing CI on migrations/meta/0261_snapshot.json. Also settle
assistant streaming/tool flags when SSE ends without a terminal frame,
without clobbering Stop's finalized content.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): satisfy biome format and import order

Auto-format the sim package for CI lint:check, and repair the Anthropic
streaming tool-loop payload after an unsafe delete-to-undefined rewrite.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): keep drained answer on abort and update migration journal test

Treat AbortError from reader.cancel as a cancelled pump result so soft-complete
retains answerText. Point the workspace storage migration journal assertion at
0261_chat_include_thinking.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(chat): keep Stop notice when server emits cancel error

Ignore terminal SSE error frames after the user aborts so
"Client cancelled request" cannot overwrite "Response stopped by user".

Co-authored-by: Cursor <cursoragent@cursor.com>

* improvement(chat): ChatGPT-style thinking shimmer and stick-to-bottom scroll

Add left-to-right shimmer on live thinking label/body, keep scroll working by
shimmering an inner node, and follow the answer only while near the bottom.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): stop pump on client disconnect; soft-complete agents only

Abort the agent stream pump when the projected HTTP body is cancelled so
provider work does not continue after disconnect. Limit AbortError soft-success
to Agent blocks so Function/HTTP cancels still fail in logs.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): persist includeThinking across pause snapshots

Paused chat runs with Include thinking enabled were dropping the flag when
serializing the pause snapshot, so resume always rebuilt streams without
thinking/tool SSE frames.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): keep drained answer text when stream times out

Persist pump answerText onto the streaming execution before throwing on
timeout, and carry that partial content into the failed block output so
logs match what the client already saw.

Co-authored-by: Cursor <cursoragent@cursor.com>

* improvement(chat): auto-collapse tools chrome when tool streaming ends

Match thinking UX: open while tools run, collapse when finished, and keep
the panel open only if the user manually reopens it.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): settle canvas stream chrome on failure paths

Clear agentStreamActive and settle running tool chips when blocks error,
timeouts cancel runs, or execution ends without stream:done so the output
panel does not stay on live Thinking/Using tools chrome.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(lint): organize imports in terminal console store

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent-stream): mark open tools cancelled on HITL pause

Pause can interrupt a tool loop without tool end events; settling those
chips as success incorrectly showed unfinished tools as complete.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(db): drop branch-local 0261 migration ahead of staging merge

* chore(db): regenerate include_thinking migration as 0266 post staging merge

* fix(providers): resolve type errors in streaming tool loop call sites

* fix(agent-stream): gate agent events opt-in and correct provider loop behavior

- streamToolCalls and provider thinking requests now require run-level
  agentEvents opt-in (canvas on, chat dual-gated, API off) so existing
  runs keep pre-agent-events behavior exactly
- OpenAI reasoning summaries opt-in + strip-and-retry on unverified-org 400
- streaming loops run tool postProcess again (firecrawl/exa async results)
- bedrock live loop falls back to silent path for responseFormat
- deepseek: reasoning_content pass-back unconditional, 'none' sends disabled
- groq: x_groq.usage fallback, reasoning params gated, qwen none disables
- gemini: functionCall parts echoed verbatim, local ids only for events
- truncated turns (max_tokens/length) no longer execute partial tool calls
- MAX_TOOL_ITERATIONS exit flushes last turn text as final answer
- iterations reports actual model calls; shared loop plumbing extracted

* refactor(agent-stream): consolidate protocol, dedupe client/server plumbing, hygiene

- canonical ChatStreamFrame union + type guards consumed by server emitters
  and the chat client; stream_error restored to legacy log-only handling
- strip thinking/tool args from providerTiming on public final envelopes
- shared tool-chip lifecycle module for chat, canvas, and console store
- shared sink-to-execution-events forwarder replaces the copy-pasted
  adapter in the execute route and HITL manager; LIVE_ONLY event set shared
- stream:thinking payload field renamed data->text; canvas thinking batched
- abort reasons carried as AbortError DOMExceptions so raw fetch consumers
  classify correctly; thinking cap renamed to chars and scope-documented
- kimi wired for agent events like the other compat providers
- deleted dead exports/step-N comments; fixtures match real wire shapes;
  loop tests use explicit mocks instead of importOriginal

* test(agent-stream): cover the dual-gated execution path and typed abort reasons

- chat route tests assert agentEvents reaches executeWorkflow only when
  policy and protocol header agree
- execution-limits tests assert AbortError-typed reasons
- executor metadata type carries agentEvents

* fix(deploy-modal): align include-thinking spacing with the modal's 6.5px rhythm

* docs(agent-stream): autogenerate per-model thinking/tool stream support on the Agent block page

- capabilities.thinking.streamed ('full' | 'summary' | 'none') on models.ts,
  explicit for the Anthropic family where visibility varies per generation;
  getThinkingStreamVisibility exposes the derivation for docs and UI alike
- scripts/sync-agent-stream-docs.ts regenerates the support tables between
  markers in workflows/blocks/agent.mdx from the model registry and
  STREAMING_TOOL_CALL_PROVIDERS; --check fails on drift or missing metadata
- wired agent-stream-docs:check into CI next to the other sync gates

* feat(anthropic): request summarized thinking display for omitted-default Claude models

The newest Claude generations (Fable 5, Sonnet 5, Opus 4.8/4.7) default
thinking.display to omitted — empty thinking blocks, no deltas. On
agent-events runs Sim now opts back in with display: 'summarized', driven
by the registry's streamed metadata; legacy runs keep the exact
pre-agent-events request shape. Registry, generated docs, and the family
capability table updated accordingly.

* docs(skills): cover thinking.streamed and agent-stream docs sync in model skills

* chore(deps): upgrade @anthropic-ai/sdk to 0.114.0 and adopt official types

- adaptive thinking, display, and output_config are now SDK-typed; the only
  remaining custom payload field is output_format (beta-header structured
  outputs, which the SDK models as output_config.format instead)
- anthropic stream events narrow on the SDK's discriminated unions instead
  of anonymous casts; compat deltas type content/tool_calls from the OpenAI
  SDK with vendor reasoning fields as an explicit optional extension
- @sim/auth exposes an explicit VerifyAuth contract so its declarations no
  longer reference better-auth's nested zod instance (TS2883 under fresh
  install layouts); realtime consumer aligned
- docs app zod pinned to the repo's exact 4.3.6 so ai SDK types bind the
  same zod instance (docs type-check was latently broken)
- knowledge embedding tests made hermetic against local .env keys and
  hosted rotation fallback

* refactor(providers): replace legacy as-any stream casts with annotated typed casts

* refactor(providers): finish provider audit — remove dead byte-stream helper, annotate remaining legacy casts

Audit of all 26 providers for the agent-events feature confirmed every
streaming execution declares agent-events-v1 and every adapter emits
AgentStreamEvent objects. Cleanup from the audit: the unconsumed legacy
createOpenAICompatibleStream byte helper is deleted, and the remaining
streamResponse-as-any casts (xai, nvidia, kimi, meta, zai, sakana) are
annotated typed casts matching the groq/deepseek fix.

* feat(streaming): stream answer text live during tool loops via turn_end protocol

The live tool loops buffered all answer text per model turn (classification
of intermediate vs final is only known at turn end), so gated surfaces saw
thinking stream, then dead air with the thinking chrome stuck open, then the
whole answer at once.

Loops now emit text deltas live as `turn: 'pending'` plus a `turn_end`
event per turn. The pump buffers pending text and projects it to the byte
path (answerText/logs/memory/legacy clients) only on a final turn_end, so
all settled semantics are unchanged. Gated surfaces render the pending text
as it streams and reconcile with a reset when a turn resolves to tools:

- public chat: live `chunk` frames from the sink + dual-gated `chunk_reset`;
  byte-path frame emission is suppressed to avoid duplicates (kept for
  response-format transformed streams via clientStreamTransformed)
- canvas: forwarder emits live `stream:chunk` + `stream:chunk_reset`; the
  execute route and HITL resume readers stop re-emitting byte chunks; panel
  chat tracks per-block segments and replaces content on flush
- chat client: per-block text segments, chunk_reset handling, and thinking
  chrome now settles on tool start as well as first answer chunk

* fix(streaming): address validated review findings across provider gating and reset reconciliation

Three-reviewer pass over the branch, findings validated against staging:

- agent-handler forwards agentEvents to executeProviderRequest — the flag was
  computed but dropped in the field-by-field copy, so provider-side thinking
  requests (OpenAI summaries, Gemini includeThoughts, Anthropic summarized
  display) never activated on opted-in runs
- openai: restore summary:'auto' alongside explicit reasoning effort — staging
  always paired them; gating summary purely on agentEvents changed legacy
  payloads
- gemini: Gemini 2 + tools + responseFormat falls back to the silent path;
  the live loop never applied the deferred responseSchema for AUTO tools
- openai-compat loop: malformed tool-argument JSON fails the call instead of
  executing with defaulted {} args (staging parsed inside the execution try)
- openai-compat parser: a vendor id arriving after a synthesized start no
  longer renames the call (start/end ids stayed consistent)
- stream-pump: abort closes the byte projection so a drain blocked on
  backpressure cannot deadlock teardown
- chunk_reset removes the block from the client text order (deployed chat +
  panel chat) so a reset block re-registers at arrival position — fixes
  separator/order corruption when parallel blocks stream around a reset
- resume route echoes the negotiated X-Sim-Stream-Protocol response header
  (parity with the chat route); docs: [DONE] wire shape + final-vs-error
  terminal semantics corrected

* chore(deps): exempt pinned @anthropic-ai/sdk 0.114.0 from the release-age gate

CI's bun install --frozen-lockfile blocks 0.114.0 (published 2026-07-23,
younger than the 7-day supply-chain gate). The pin is exact and was vetted
for the agent-events streaming work; following the existing bunfig pattern,
the exclusion ages out on 2026-07-30 and should be dropped then.

* chore(providers): fix double-cast-allowed annotation placement for the strict boundary audit

The audit only recognizes the annotation on the line directly above the cast;
two annotations had drifted behind intervening code lines (groq stream params,
deepseek loop messages) and the OpenAI reasoning-summary widening cast was
never annotated. No behavior change.

* fix(chat): settle straggler tool chips as error when final reports failure

A failed run can still terminate with a `final` frame carrying success: false;
running chips previously settled green regardless of the outcome.

* fix(canvas): wire agent stream chrome into run-from-block

Run-from-block executions emit the same live stream:thinking/stream:tool
events as full runs but registered none of the handlers, so the terminal
never showed thinking or tool chips on that path. The per-run chrome
(batched thinking writes + tool chip lifecycle + settlement on stream done,
block error, and every terminal execution state) is extracted into a shared
createAgentStreamChrome factory consumed by both paths.

---------

Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai>
2026-07-23 19:39:03 -07:00
Justin Blumencranz 03adc8fe1d fix(chat): rotate fork chat split icon 90 degrees (#5914) 2026-07-23 19:33:15 -07:00
Waleed d48722a04e fix(helm): correct chart docs, examples, and dead config across the board (#5907)
* fix(helm): correct chart docs, examples, and dead config across the board

Audit-driven accuracy pass over the chart's entire documentation surface,
verified by rendering every example against the templates:

- migrations run as an init container on the app pod, not a Job — fix the
  README component list, troubleshooting commands, and sim-helm skill refs;
  drop the dead migrations-job NetworkPolicy ingress rule
- referenced-but-never-created resources: document the GKE ManagedCertificate
  creation (values-gcp), comment out the key-file Secret mount that stuck all
  pods in ContainerCreating (values-gcp), enable certManager for the postgres
  TLS issuerRef (values-production), add the cert-manager cluster-issuer
  annotation nginx needs (values-azure)
- values-external-db: networkPolicy.egress is a list, not a map (the map
  rendered an invalid manifest); fill schema-failing placeholder host/username
- realtime >1 replica requires REDIS_URL (Socket.IO Redis adapter) — default
  examples to 1 replica with the scaling note, and warn where autoscaling
  HPAs override replicaCount
- pod anti-affinity selectors matched nothing (simstudio vs sim name label)
- kubernetes.mdx: install commands were missing required CRON_SECRET and
  postgresql password (failed at template time), wrong deployment name in
  port-forward, stale version requirements, unsupported key-remapping claim
- remove unimplemented app.secrets.existingSecret.keys from values + schema;
  fix README PDB default, cronjob list, /metrics caveat, NOTES secret count,
  Azure-only StorageClass in generic examples, dead SOCKET_SERVER_URL and
  GOOGLE_CLOUD_* env, ESO apiVersion mismatch, and skill-reference drift
- bump chart to 1.0.1

* fix(helm): review round 1 — scoped example egress, in-tab secret note, copilot Job wording

- external-db example egress scopes to a placeholder database CIDR instead of
  to: [] (which allowed every destination on 5432, defeating the isolation
  the example teaches)
- kubernetes.mdx cloud tabs state explicitly that they reuse the variables
  generated in the Installation block
- Copilot migrations really do run as a Helm-hook Job — restore Job wording
  there (only the app migrations are an init container)

* fix(helm): template-sweep fixes — telemetry validity, ESO rollout checksums, dead passwordKey knob

- telemetry: memory_limiter gets the required check_interval (collector
  failed startup validation whenever telemetry.enabled=true); the jaeger
  exporter was removed from collector-contrib in v0.86 — export to Jaeger
  via its native OTLP endpoint instead (otlp/jaeger, default port 4317)
- app/realtime rollout checksums now hash the ExternalSecret manifest too,
  mirroring the copilot pattern — with ESO enabled the inline Secret renders
  empty, so remoteRefs changes never rolled the pods
- remove the unimplemented existingSecret.passwordKey knob (values, schema,
  README, dead helpers): nothing consumed it, and a non-default value
  silently produced a DATABASE_URL with an unexpandable placeholder; secrets
  must use the standard POSTGRES_PASSWORD / EXTERNAL_DB_PASSWORD keys
- drop the orphaned sim.migrations.labels helper (its only consumer was the
  dead NetworkPolicy rule removed earlier)
- helm test pod image resolves through sim.image so global.imageRegistry
  mirroring applies; NetworkPolicy realtime-ingress comment reflects actual
  traffic direction; smoke unittest suite loads the newly referenced
  external-secret template

* chore(helm): bump chart to 1.1.0 with upgrade notes

Removing (inert) documented values keys and changing the rollout-checksum
inputs is a values-surface change — per SemVer chart conventions that is
more than a patch. Adds an Upgrading section documenting the one-time pod
roll, the removed no-op keys, and the Jaeger-over-OTLP change.

* fix(helm): review round — external-db NP opt-in with real-CIDR-first flow, prod Jaeger OTLP endpoint

- external-db example ships networkPolicy disabled so a verbatim install
  always reaches the database; the scoped egress rule stays as the documented
  opt-in (set your CIDR first, then enable)
- values-production still pointed telemetry.jaeger at the legacy 14250
  collector port — now Jaeger's OTLP gRPC endpoint to match the otlp/jaeger
  exporter

* feat(helm): autoscaling.realtime.enabled toggle so examples can scale the app without unsafe realtime replicas

Cursor correctly flagged that comment-level warnings didn't stop a verbatim
production/external-db install from running the realtime HPA at minReplicas 2
without REDIS_URL (silent cross-pod event loss). Adds an opt-out toggle
(default true — existing deployments unchanged): the realtime HPA renders only
when autoscaling.realtime.enabled, and the realtime Deployment keeps
spec.replicas under its control when the HPA is excluded. The three
autoscaling examples set it false with the Redis rationale; README and
upgrade notes document the toggle.

* fix(helm): review round — whitelabeled realtime HPA opt-out, external-db isolation on with required CIDR in install flow

- values-whitelabeled now actually sets autoscaling.realtime.enabled: false
  (the earlier batch aborted before reaching this file — Cursor caught it)
- external-db keeps networkPolicy enabled (no isolation regression); the DB
  egress CIDR is marked REQUIRED and wired into both documented install
  commands via --set, so the copy-paste flow sets the real subnet in the same
  breath as the DB host

* fix(helm): use a syntactically valid example CIDR in the external-db install commands

An unreplaced <YOUR_DB_CIDR> literal fails Kubernetes CIDR validation and
aborts the install; every sibling placeholder in the same command (host,
username) installs fine and simply doesn't connect until replaced. The CIDR
now behaves the same way: valid example value (10.20.0.0/24), explicitly
marked as the operator's database subnet in both commands and the values
comment.

* fix(helm): remove text after continuation backslashes in external-db install commands

A trailing comment after the line-continuation backslash (and equally a
comment line spliced mid-command) breaks the shell command when copied —
the remaining --set overrides run as separate commands and required-value
validation fails. Both documented commands now reconstruct to bash-clean
multi-line invocations (verified with bash -n); the CIDR guidance lives in
the networkPolicy section comment.

* fix(docs): cloud-tab installs are alternatives via helm upgrade --install

Following Installation and then a cloud tab ran helm install twice for the
same release and failed on the second. The tabs now state they replace the
generic install and use helm upgrade --install, which is idempotent and also
converts an existing generic install to the cloud values.

* fix(docs): honest conversion caveat for cloud-tab upgrades over an existing install

The cloud values rename the bundled Postgres database to simstudio, which
Postgres only applies at first initialization — an in-place conversion of a
generic install would point DATABASE_URL at a nonexistent database. Document
the two safe paths: keep the original name via --set, or uninstall + delete
PVCs and install fresh.

* fix(helm): external-db header install command declares its secrets and includes CRON_SECRET

The primary documented command failed the chart's required-value validation
(missing app.env.CRON_SECRET with cronjobs default-on) and referenced an
undeclared DB_PASSWORD. It now shows the export lines for every variable it
uses and sets CRON_SECRET; both commands verified with bash -n.

* chore(helm): restore values.schema.json formatting — surgical deletions only

The earlier programmatic edit reformatted the whole file (~390 lines of
whitespace churn hiding the 12 real deleted lines). Re-applied the removal
of the dead keys/passwordKey properties as text-level deletions preserving
the original style.

* fix(helm+docs): explicit secret exports in external-db header; conversion must reuse original secrets

- all five export lines are written out (three were only named in a trailing
  comment, so a verbatim copy passed empty required values)
- the cloud-conversion caveat now leads with reusing the original secret
  values (helm get values) — a regenerated ENCRYPTION_KEY makes previously
  encrypted credentials undecryptable
2026-07-23 18:28:11 -07:00
Waleed ca2ea0066d improvement(library): add citations and internal links, correct stale pricing (#5913)
* improvement(library): add citations and internal links, correct stale pricing

The library had zero third-party citations across 16 posts (~51k words) and
averaged 1.2 internal links per post, with six posts at zero. Auditing every
dollar figure against the vendor's own pricing page while adding the citations
turned up several claims that had gone stale, plus one product that is being
shut down.

Corrections, each verified against the vendor's page on 2026-07-23:
- Relay.app is winding down (signups closed 2026-07-16, free accounts end
  2026-08-15, paid 2026-09-14). It was recommended as the human-in-the-loop
  pick in best-zapier-alternatives; that recommendation now points to a
  platform you can still sign up for
- Make Core is $12/mo billed monthly, not $10.59; annual saves ~15%. One post
  also described $12 as the annual rate
- Pabbly Connect lifetime starts at $349, not $249, and the Standard/Pro/
  Ultimate tier structure quoted no longer exists
- Workato publishes no pricing at all, so the "~$1,000/month" figure is
  replaced with the fact that every deal is quoted through sales
- Dify is at ~149k GitHub stars, not 131k
- n8n cloud Starter is EUR 20/mo billed annually for 2,500 executions

Verified-correct figures were left alone and given a source link: Zapier
$29.99/mo monthly for 750 tasks, Activepieces 10 free flows then $5/flow/mo,
Power Automate $15/user/mo, Lindy $49.99/mo, and the Grand View Research RPA
market figures.

Also: 59 outbound citations and 3-6 internal links per post (zero posts left
without internal links), and MDX external links now carry target=_blank plus
rel=noopener noreferrer, which the landing SEO/GEO rule requires but the MDX
anchor was not applying.

* fix(content): treat only non-Sim hosts as external links in MDX

The first pass classified any http(s) href as external, so the 33 absolute
Sim URLs already in content (sim.ai, sim.ai/slack, www.sim.ai/blog/*, and 14
docs.sim.ai pages) would have opened in a new tab with rel=noopener, which is
wrong for first-party links and contradicted the comment above the check.

Classification now compares the hostname against the apex derived from
SITE_URL, so the apex and any Sim subdomain stay same-tab. SITE_URL is used
rather than getBaseDomain() so a post renders identically in dev, preview,
and production instead of varying with NEXT_PUBLIC_APP_URL. The leading dot
in the suffix check keeps lookalikes such as evil-sim.ai external.
2026-07-23 18:10:00 -07:00
Waleed 1ffd4f0739 improvement(landing): close remaining SEO keyword gaps on home and enterprise (#5912) 2026-07-23 17:55:28 -07:00
Waleed 877dc9e719 fix(landing): lead enterprise summary with Sim so product schema names the module (#5910) 2026-07-23 17:28:37 -07:00
Waleed dbbfce88c4 fix(landing): restore single-paragraph hero copy on home and enterprise (#5906)
* fix(landing): restore single-paragraph hero copy on home and enterprise

* fix(landing): tighten enterprise hero description

* revert(landing): restore original home and enterprise hero headings

* improvement(landing): simplify enterprise hero description

* improvement(landing): restore building-and-managing homepage headline
2026-07-23 17:17:18 -07:00
Waleed 79c57bfaf6 improvement(content): surface last-verified and updated dates, codify citation rules (#5904)
* improvement(content): surface last-verified and updated dates, codify citation rules

Comparison pages already computed a latest-verified date from every fact
source's asOf and emitted it as JSON-LD dateModified, but never showed it.
Library and blog posts had the same gap: dateModified existed only as an
invisible meta tag. A freshness signal that no reader can see does nothing
for the reader deciding whether to trust the page.

- /comparisons/[provider] renders "Last verified <date>" from the existing
  getLatestVerifiedDate(), so the visible date and the JSON-LD read the
  same value and cannot drift
- /library and /blog posts render "Updated <date>" next to the publish
  date, only when the modified day actually differs; otherwise the meta
  fallback stays. dateModified is emitted exactly once either way
- collapse three duplicated toLocaleDateString blocks in
  content-post-page into one module-scope formatDate (same UTC pinning,
  identical output)
- landing-seo-geo rules: add citation/internal-linking and freshness
  sections, and extend the rule's paths to content MDX so it attaches when
  authoring posts, not just TSX

* fix(content): wrap post metadata row so the Updated date cannot overflow on narrow screens

The published/updated/authors/share row was a non-wrapping flex. With the
Updated label active and two authors it overflowed at 390px; the row also
already overflowed at 320px on staging before this PR, with no Updated
label at all. flex-wrap fixes both and leaves desktop identical.

* improvement(comparisons): fold last-verified date into the intro sentence

The paragraph already ended "Every fact below is sourced and dated" and a
separate line then stated the date, which read as redundant and added a
standalone metadata row to the header. Folding it into that sentence drops
the extra line and keeps the date as real server-rendered text in a <time>
element, so crawlers and AI answer engines still see it. A tooltip would
not: Tooltip.Content renders null during SSR and while closed.
2026-07-23 16:47:00 -07:00
Theodore Li 96c67e9123 feat(sandbox): add Daytona as a manual-flip failover for E2B (#5860)
* feat(sandbox): add Daytona as a manual-flip failover for E2B

E2B was a hard single point of failure: lib/execution/e2b.ts had no retry
and no fallback, so a failed Sandbox.create() killed Python function blocks,
JS-with-imports, shell, doc generation and the Pi cloud agent outright.

Extract a SandboxRunner boundary (lib/execution/remote-sandbox) with an E2B
runner and a Daytona runner, selected once per execution by the
sandbox-provider-daytona AppConfig flag. Everything above the provider
boundary — marker parsing, mount materialization, file export, corruption
handling — is unchanged.

Selection resolves before create() and never mid-execution, since user code
has side effects. Each sandbox kind fails closed when its snapshot id is unset.

Notes on the Daytona adapter:
- language binds at create(), not per call: Daytona applies it as a sandbox
  label and silently runs JS through Python if passed to codeRun
- Python routes via CodeInterpreter for its {name,value,traceback} error shape,
  which matches E2B's and keeps formatE2BError's line offsets correct
- timeouts convert ms to seconds
- the streaming path delivers env via the filesystem API, as
  SessionExecuteRequest has no env field and secrets must not reach a command line

Drops the dead E2BExecutionResult.images field (populated, never consumed).

* improvement(sandbox): select the provider by SANDBOX_PROVIDER env var

Replaces the boolean sandbox-provider-daytona feature flag with a
SANDBOX_PROVIDER env var naming the provider ('e2b' default, or 'daytona').
A boolean doesn't scale to a third adapter; a keyed registry does.

- PROVIDERS is a Record<SandboxProviderId, SandboxProvider>, so adding an
  adapter is one entry plus one id-union member — an unhandled provider is a
  compile error, not a runtime surprise
- resolveProvider() reads env synchronously and throws on an unknown value
  (fail fast) instead of an async feature-flag lookup
- drops the sandbox-provider-daytona flag and SANDBOX_PROVIDER_DAYTONA fallback

Verified end-to-end: a Python function block through the running app routes to
Daytona (Creating Daytona sandbox, kind: code) with SANDBOX_PROVIDER=daytona.

* fix(sandbox): gate remote execution by provider availability, not E2B

Addresses the review round on #5860.

- Availability was gated on isE2bEnabled / isE2BDocEnabled, so a Daytona-only
  deployment (E2B_ENABLED unset) had its Python/shell/JS-with-imports and doc
  paths rejected before the provider-neutral sandbox call could run. Replace both
  with provider-aware flags (isRemoteSandboxEnabled / isDocSandboxEnabled) derived
  from the selected SANDBOX_PROVIDER's own credentials + image. E2B behavior is
  unchanged (the E2B branch mirrors the old definitions exactly).
- Make the function-block gate error messages provider-neutral.
- Daytona's streaming runCommand (Pi) returned empty stdout/stderr and delivered
  output only via callbacks, so the Pi cloud flow — which parses markers from
  stdout and formats errors from stderr — saw nothing. Accumulate the streamed
  chunks and return them while still forwarding to the callbacks.

Renames the env-flag exports (and the @sim/testing mock) to match. Adds a
conformance test that the streamed Pi output lands in stdout/stderr.

* fix(sandbox): resolve SANDBOX_PROVIDER case-insensitively

Addresses the round-2 review on #5860. env-flags lowercased SANDBOX_PROVIDER
for the availability gate, but resolveProvider looked up the raw value in a
lowercase-keyed map — so 'Daytona' passed the gate then threw Unknown
SANDBOX_PROVIDER at create. resolveProvider now normalizes casing identically.

* fix(sandbox): use getErrorMessage in build/verify scripts

check:utils flagged the inline `error instanceof Error ? error.message : ...`
pattern in the two new scripts. Use getErrorMessage from @sim/utils/errors,
matching the repo convention the check enforces.

* fix(sandbox): fall back to stdout for Daytona failure text

Daytona merges both streams into stdout and returns an empty stderr, but the
shell-error, base64-export, and URL-mount error builders read only
result.stderr — so Daytona failures surfaced a generic 'Process exited with
code N' / 'base64 failed' / 'curl exited N' instead of the real command output
that the API and agents rely on. Fall back to stdout before the generic message
(provider-agnostic: E2B still populates stderr). Strengthens the shell-error
conformance test to assert the real output surfaces.

* fix(sandbox): enforce timeout on Daytona streaming; drop leftover E2B copy

- Daytona's streaming path (Pi) started the command with runAsync:true and then
  awaited getSessionCommandLogs with no bound, so a hung command never timed out
  the way E2B's commands.run({ timeoutMs }) does. Race the log stream against the
  timeout; on expiry return exit 124 with the accumulated output, and the finally's
  deleteSession terminates the still-running command.
- Two user-facing strings still named E2B after the provider-neutral rename (the
  isolated-vm sandboxPath remediation and the disabled-xlsx message). Made both
  provider-neutral.

Adds a streaming-timeout conformance test.

* fix(sandbox): handle orphaned stream promise + preserve error detail

Two regressions from the previous timeout fix:
- When the timeout won the race, the abandoned getSessionCommandLogs promise
  would reject on deleteSession with no handler (unhandledRejection). Attach a
  .catch that records the error and yields an 'error' outcome, so a late
  rejection is always handled.
- The streaming catch dropped the thrown error, so failures before any chunks
  (env write, executeSessionCommand, missing cmdId) surfaced as empty output.
  Fall back to getErrorMessage(error) when nothing streamed.

Adds conformance tests for the stream-reject and start-throw paths.
2026-07-23 19:19:00 -04:00
Waleed 3796e9db4f improvement(access-control): edit group details, filter by status, and colocate chat deploy auth (#5902)
* improvement(access-control): editable group details, status filters, block tooltips, chat auth colocation

- General tab gains editable Name and Description fields wired into the
  existing dirty buffer, so the header Save/Discard chips and the
  unsaved-changes guard cover them. Save only sends changed fields and
  surfaces the route's duplicate-name 409.
- Blocks, Model Providers, and Platform tabs gain an All/Enabled/Disabled
  filter beside their search field, evaluated against the editing buffer.
- Every block row carries an Info badge with the block's description. The
  badge sits outside the label/expand button so it never toggles the row.
- The chat deploy toggle moves out of Deploy Tabs into the Chat section
  alongside its allowed-auth-modes dropdown, mirroring the Files section.

* improvement(access-control): polish group detail filters, tooltips, and details fields

Follows the cleanup review:
- reconcile the post-save baseline from the server response instead of
  local values, matching the scope/default writes
- pin the status-filter dropdown width so the search field stops resizing
- restore flex-1 on the block-name button and move the row hover surface
  to the wrapper so Info badges align in a column
- add empty states for filter-empty lists and neutralize Select All there
- surface a 'Name is required' message next to the disabled Save
- align hint text on the field-hint tokens and drop a redundant TSDoc

* refactor(access-control): keep chat and files toggles in the platform registry

The first pass pulled hideDeployChatbot out of the declarative platformFeatures
array and hand-rolled a Chat section beside the existing bespoke Files one. That
forfeited search, status filtering, Select All, category grouping and the Info
hint, and the replacement platformSectionVisible re-implemented two of those
with different semantics — searching 'deploy' or 'deployment' hid the very
control named Deployment, and Select All silently skipped both toggles.

Both toggles are now ordinary registry entries under their own Chat and Files
categories, with an id-keyed featureExtras map supplying the nested auth-mode
dropdown. Search, filtering, Select All, hints and the empty state are correct
by construction, and the parallel filter pipeline is gone.

Also from the review:
- index the allow-lists into Sets so per-row membership checks are O(1)
- split the search and status passes so the common 'all' filter returns the
  searched list by reference and a checkbox toggle no longer re-sorts ~180 rows
- extract StatusFilterChip and AuthModeField instead of stamping out the
  dropdowns three and two times
- derive nameChanged/descriptionChanged once instead of repeating the
  comparisons in the save payload
- lock the config key-order invariant the dirty check depends on with a test

* polish(access-control): apply the second cleanup round

- indent the nested auth-mode field so it lines up under its toggle's label
  instead of reading as a sibling row, and label the dropdown for screen readers
- order Chat right after Deploy Tabs so the three deploy targets stay adjacent
- flush the trailing Select All chips on the providers and platform rows
- drop the doubled margin on the name error (SettingRow already gaps it)
- hoist PLATFORM_FEATURES and PLATFORM_CATEGORY_ORDER to module scope
- drop the useCallback on the two save/discard handlers; nothing observes their
  identity and their deps changed on every keystroke anyway
- fix a comment that still pointed at a 'Hide Chat' toggle that no longer exists

* fix(access-control): keep the block disclosure chevron inside its toggle button

Splitting the chevron out so the Info badge could sit beside the name left the
chevron with no click handler — the visible expansion affordance did nothing.
It goes back inside the button; Info stays outside it, since an Info trigger is
itself a button and cannot nest.

* chore(access-control): use structuredClone in the config key-order test

check:utils bans JSON.parse(JSON.stringify(...)). The clone only needs to hand
the schema a distinct object; structuredClone preserves key order the same way.

* fix(access-control): trim descriptions so a padded value can't wedge the form

descriptionChanged compared a trimmed draft against the raw saved description,
so a group whose stored description carried padding opened dirty and could
never be cleared — Discard restored the same padded string, and the
unsaved-changes guard then blocked navigation until a save rewrote it.

The contract now trims description on create and update, matching what name
already did, and the dirty check trims both sides so existing padded rows
behave too.

* improvement(access-control): seed the description buffer trimmed

Keeps the editing buffer and the dirty baseline normalized the same way, so a
legacy row with padding no longer shows stray whitespace in the input and the
buffer never round-trips padding a save would strip anyway.

* fix(emcn): forward aria-label/aria-labelledby from ChipDropdown to its trigger

ChipDropdown destructures only its known props, so an aria-labelledby passed by
a consumer never reached the trigger button — AuthModeField's wiring to its
visible label was silently dropped and the control had no accessible name.

Both attributes are now explicit, typed props forwarded to the trigger. Kept as
two named props rather than a rest spread so the component still owns its
chrome and consumers can't smuggle arbitrary attributes onto the button.

* fix(access-control): correct the nested field indent, platform row hover, and aria name

- aria-labelledby REPLACES the content-derived name, so naming the auth-mode
  dropdown with its caption alone dropped the selected value from the accessible
  name — worse than no attribute. It now references caption + trigger, which
  needed ChipDropdown to forward id as well.
- The nested field's pl-[30px] assumed a 14px checkbox; the default is 16px, so
  the caption sat 2px left of the label it hangs under while the dropdown (not
  flush, so mx-0.5) sat at 32. Both are pl-8 + flush now.
- Platform feature rows kept their hover surface on the label, so the highlight
  stopped 20px short of the row edge while the Blocks tab ran flush. Moved to
  the wrapper, matching the core-blocks cell.
- Split the platform search and status passes like the provider and block lists,
  so a toggle no longer recomputes three chained memos.
- Dropped an unreachable allLabel and a comment that would go stale on merge.

* fix(access-control): trim the name comparison the same way as description

nameChanged compared a trimmed buffer against an untrimmed baseline — the exact
asymmetry already fixed for description. A group stored before the name schema
gained .trim() opens permanently dirty: Save/Discard visible and the
unsaved-changes modal firing on back, until a save rewrites it.

Both buffers now seed trimmed and compare trimmed on both sides. Also dims the
Info badge along with the row it belongs to when the block is disallowed.
2026-07-23 15:44:13 -07:00
Waleed 6f90d8976d fix(mcp): stream pinned transport under Bun (providers + self-hosted-private MCP) (#5901)
* fix(mcp): stream pinned transport via undici.request + redirect interceptor

Extends the Bun undici-streaming fix to createPinnedFetchWithDispatcher (providers,
A2A, self-hosted-private MCP over SSE). It now routes through undiciRequestAsResponse
like the guarded builder, so streaming bodies deliver under Bun. Unlike the guarded
path it has no followRedirectsGuarded wrapper (it's handed straight to provider SDKs),
so redirects are followed via undici's redirect interceptor composed onto the pinned
Agent — every hop still dispatches through the pinned connect.lookup (resolvedIP), so a
redirect can't escape to another address, matching the old fetch guarantee. secureFetchWithPinnedIP (raw Node http, tools path) is untouched.

* fix(mcp): honor redirect mode + drop cross-origin credentials on pinned fetch

Replaces the always-on redirect interceptor with redirect-mode-aware handling:
- redirect:'manual' returns the 3xx without following (detectMcpAuthType inspects it)
- redirect:'error' throws on a 3xx
- default 'follow' uses followRedirectsGuarded, which drops ALL headers on a
  cross-origin hop (so a redirect can't disclose a provider api-key to another
  origin — Greptile P1) and stamps the final response.url + redirected flag.
Extracts the shared Request-lift helper used by both guarded and pinned builders.

* fix(mcp): don't block private IP-literal URLs on the pinned fetch path

Routing the pinned fetch through followRedirectsGuarded added an initial
assertGuardedRedirectTarget check the old undici.fetch path never ran, which
would block a self-hosted MCP configured with a private IP-literal URL
(e.g. http://10.0.0.5:3000/mcp) — its own transport. The pinned path's callers
already validate the target and the private carve-out intentionally pins to a
private IP, so skip the initial-target check (validateInitialTarget: false) while
still validating every redirect hop. Adds a regression test.

* fix(mcp): carry redirect mode from a Request input in liftFetchArgs

liftFetchArgs copied method/headers/body/signal from a Request but omitted
redirect, so a Request({ redirect: 'manual' }) on the pinned path defaulted to
'follow' and was transparently followed. Copy input.redirect (explicit init still
wins). Adds a Request-input redirect-mode test.

* fix(mcp): permit the pinned IP as a redirect target (initial + hops), block other private IPs

Consolidates the pinned-path redirect policy into one mechanism. followRedirectsGuarded
took validateInitialTarget to skip the initial private-IP check, but per-hop checks still
blocked a self-hosted MCP redirecting to its own pinned private IP (e.g. a trailing-slash
301 to http://10.0.0.5/mcp/). Replace it with allowRedirectToIp: the pinned fetch permits
exactly its own validated IP as a target — initial URL and any hop that stays on it — while
every OTHER private target (e.g. the 169.254.169.254 metadata IP) stays blocked. Tests cover
the same-IP hop (followed) and the metadata-IP escape (still refused).
2026-07-23 15:27:16 -07:00
Waleed a5e1030e9e fix(mcp): stream guarded transport via undici.request so SSE responses work under Bun (#5897)
* fix(mcp): stream guarded transport via undici.request so SSE responses work under Bun

undici's fetch exposes its response body as a WHATWG ReadableStream whose
bridge is broken under the Bun runtime the standalone server runs on: headers
arrive but response.body never yields data, hanging every MCP streamable-HTTP
(text/event-stream) tools/list read to its 30s timeout. undici's lower-level
request() returns a Node Readable, which Bun streams natively.

createSsrfGuardedFetchWithDispatcher (the MCP transport builder) now routes
through undiciRequestAsResponse: same guarded Agent + connect.lookup (SSRF
unchanged), request() instead of fetch(), and a hand-rolled Node->Web body
bridge (not Readable.toWeb, which throws ERR_INVALID_STATE on the redirect
body.cancel()). Buffered reads, followRedirectsGuarded, maxResponseSize, and
abort all preserved and verified on both Node and Bun.

* fix(mcp): use a single cast for header-iterable detection to satisfy boundary audit

* fix(mcp): handle URLSearchParams OAuth bodies and copy stream chunks

- toUndiciRequestBody serializes a URLSearchParams body (the MCP SDK's OAuth
  token/refresh exchange sends one); undici.request rejects it otherwise.
- Default content-type application/x-www-form-urlencoded when the caller didn't
  set one (fetch parity).
- nodeReadableToWebStream enqueues a copy (new Uint8Array(chunk)) not a view, so
  undici recycling the pooled source buffer can't corrupt queued chunks.
- Tests for both.

* fix(mcp): decode Content-Encoding on the undici.request transport (fetch parity)

undici.request returns raw bytes; fetch auto-decompresses. Restore that: pipe a
gzip/deflate/br response body through the matching zlib decoder before the WHATWG
bridge and strip content-encoding/content-length. maxResponseSize still caps the
compressed wire bytes; errors forward into the decoder and the source is torn down
when the decoded stream ends or is cancelled. Adds a gzip decode test.

* fix(mcp): guard decoder errors and null-body drain against unhandled crashes

- Attach the stream bridge's error listener before piping into the zlib decoder, so
  a synchronous zlib error (server mislabeling a non-gzip body as gzip) rejects the
  reader instead of crashing the process.
- Attach an error listener before draining a null-body response, and wrap Response
  construction in try/catch that destroys the source (no socket leak) on an
  out-of-range status. Adds an invalid-gzip regression test.
2026-07-23 14:17:52 -07:00
Vikhyath Mondreti e737901b0b chore(blocks): rename webhook block (#5900) 2026-07-23 14:14:16 -07:00
Waleed 5c427795a2 fix(confluence): preserve panel/callout macro semantics through sync (#5896)
* fix(confluence): preserve panel/callout macro semantics through sync

Confluence's rendered view HTML wraps Info/Note/Warning/Tip and custom Panel
macros in divs whose class/color convey meaning that the shared
htmlToPlainText tag-stripper discards along with the tags — a red "do not
use" warning panel becomes indistinguishable from a plain paragraph once
flattened, so RAG has no signal that a bullet under it is an exclusion rule
rather than a normal one.

Adds preserveConfluenceCallouts, a Confluence-specific pre-pass that rewrites
each detected panel into a single bracketed label (e.g. "[WARNING] Do NOT use
this form for: GitLab") before the generic plain-text conversion runs, so the
callout semantic survives both the tag strip and htmlToPlainText's trailing
whitespace collapse. Bumps the connector's content-representation marker so
already-synced pages get one automatic re-hydration under the new extraction,
rather than silently keeping their stale flattened content until their next
edit.

* fix(confluence): preserve word boundaries when extracting callout body text

Greptile P1: cheerio's .text() concatenates every descendant text node with
no separator, so pulling a macro body's text in one call fused adjacent
blocks together (e.g. a paragraph ending in "for:" immediately followed by a
list item "GitLab" became "for:GitLab"), corrupting the exact word boundaries
RAG chunking and keyword matching depend on.

extractBlockJoinedText now extracts each paragraph/list-item/heading/cell/quote
individually and joins them with a single space, keeping every block's text
intact and properly separated, matching how htmlToPlainText already treats
the rest of the page.

* fix(confluence): fix nested-block duplication in callout text extraction

Greptile P1: filtering the found blocks to only top-level ones still wasn't
enough — a nested block (an outer <li> containing its own nested <ul><li>, a
<td> containing a <blockquote>) matched the selector once, but its .text()
call recurses into and flattens its own matched descendants with no
separator, reproducing the exact word-fusion bug one level deeper (and any
duplicate-selection would have double-counted the same text).

Replaces the block-selector approach with a recursive text-node walk:
extractBlockJoinedText now visits every text node individually and joins them
all with a single space, so word boundaries are preserved at every nesting
depth with no double-counting, matching the pattern html-parser.ts already
uses elsewhere in this codebase for the same class of problem.

* fix(confluence): apply the same word-boundary-safe extraction to panel headers

Greptile: panelHeader text extraction was left on the plain .text() call
while panel/macro body extraction was already fixed to use
extractBlockJoinedText, so a rich multi-node header (e.g. <b>Warning:</b>
followed by a sibling <span>) could still fuse into "Warning:Do not use"
with no space. Panel headers now go through the same recursive text-node
walk as bodies, for consistency across every text extraction in this file.

* fix(confluence): distinguish inline formatting from block boundaries in extraction

Greptile: the recursive text-node walk unconditionally inserted a space
between every text node, which fixed block-boundary fusion but broke
genuinely inline-formatted text — "un<b>believe</b>able" became
"un believe able" and "Hello<b>!</b>" became "Hello !", corrupting valid
callout content on its way into the index.

Adds an INLINE_FORMATTING_TAGS allowlist (b, strong, i, em, span, a, etc.):
text flowing through those tags accumulates with no artificial separator,
preserving exact source adjacency, while every other tag boundary (p, li,
td, headings, br, ...) still flushes to a new segment — a block always
implies a break even with no literal whitespace in the source, but an inline
tag never does. Fixed one test that had encoded the old, incorrect
expectation for two genuinely adjacent inline tags with no source whitespace
between them, and added regression tests for mid-word inline formatting,
punctuation attached to an inline tag, and a header with real source spacing.

* fix(confluence): process nested panels/macros innermost-first

Cursor: processing matches in document order (outermost first) read a
nested, not-yet-converted panel/macro as plain body text before it ever got
its own bracketed label, silently dropping the inner callout's type. Worse,
an untitled outer panel's `.find('.panelHeader')` could reach past its own
missing header into a nested panel's header and adopt it as its own title.

Replaces the two independent .each() passes with a loop that converts only
"leaf" macros (no remaining nested macro/panel inside them) and repeats
until none are left. This processes innermost-first, so a nested macro is
already its own bracketed <p> by the time its parent's body/header text is
read, and an untitled outer panel's .find() can no longer reach a header
that isn't its own, since a leaf by definition has no nested panel left to
reach into.

* test(confluence): add explicit regression for the exact reported <br> repro

Formalizes an explicit test for the exact <br>-separated string Greptile's
review cited as broken (verified manually not to reproduce, but wasn't
directly asserted in the suite before this).
2026-07-23 14:08:09 -07:00
Theodore LiandClaude Opus 4.8 84e7aad633 improvement(slack): request the approved mention/assistant/DM scopes, advertised on v2 only (#5898)
Slack app review has since approved `app_mentions:read`, `assistant:write`, and
`im:history` — they are live in the prod app manifest's `oauth_config.scopes.bot`
— but the repo still had them commented out behind a stale "re-add once approved"
TODO. Sim was therefore requesting a narrower grant than the app is entitled to,
so newly connected accounts got tokens missing exactly the scopes backing three
`simSubscribed` events on the native Sim app trigger: `app_mention`
(app_mentions:read), assistant threads (assistant:write), and DMs (im:history).

The requested scope set now matches the manifest's 20 approved bot scopes exactly.

Scopes are per-credential, not per-block (one Slack app -> one `slack` provider ->
one shared token, requested server-side via getCanonicalScopesForProvider), so the
grant itself cannot be scoped to v2. What is per-block is what each picker
advertises and treats as missing, so:

- slack_v2 + the slack_oauth trigger advertise the full set (they host the native
  Sim app trigger that needs it).
- The legacy v1 block stays pinned to the pre-expansion 17 scopes. It has no
  feature needing the new three, and advertising them there would flag every
  existing Slack credential as missing scopes and prompt a needless reconnect.


Claude-Session: https://claude.ai/code/session_018asmKsWQ5Vi7T7wD9uHofz

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 16:44:37 -04:00
Waleed bc237a0dcc improvement(shopify): pin Admin API to supported 2025-10 via shared constant (#5895)
* improvement(shopify): pin Admin API to supported 2025-10 via shared constant

Every Shopify tool and the credential validator hardcoded the retired 2024-10
Admin API version. Retired versions still work (Shopify forward-falls to the
oldest supported version) but the served version drifts silently and emits
deprecation signals.

- Add apps/sim/tools/shopify/constants.ts exporting SHOPIFY_API_VERSION as the
  single source of truth; bump to the supported 2025-10.
- Wire all 21 Shopify tools, the token-service-account validator, and the OAuth
  store route to the shared constant (no more per-file inline version).
- Derive the validator test's expected URL from the constant so it can't drift.

Our operations already use modern 2024-10+ shapes (ProductCreateInput /
ProductUpdateInput new product model, orderCancel new signature, CustomerInput,
InventoryAdjustQuantitiesInput, fulfillmentCreate), none of the removed inline
ProductInput.variants pattern — so the bump is a small, explicit step forward
from the version Shopify already forward-serves.

* chore(shopify): remove dead ShopifySetInventoryParams type

Surfaced by /validate-integration: unexported, unused (no set_inventory tool).
2026-07-23 12:12:44 -07:00
Waleed d39b193201 feat(sidebar): search the workspace switcher when a user has many workspaces (#5893)
* feat(sidebar): search the workspace switcher when a user has many workspaces

Shows a search input in the workspace dropdown once the list exceeds
WORKSPACE_SEARCH_THRESHOLD (3). ArrowUp/Down move through results, Enter
switches, and the query resets on close. All the search machinery is gated on
showSearch so users with few workspaces get no extra re-renders. The input
reuses the emcn ChipInput chrome (icon prop) rather than hand-rolling the field.

The highlight tracks the highlighted workspace by identity (a stored id, not a
numeric position). An effect keeps that id pinned to a workspace that is actually
in the current results — seeding the first result on open and re-seeding when a
query filters the current one out — so even the default highlight is identity-
stable. A live list change while the menu is open (shrink, grow, or reorder from
a membership change or background refetch) therefore carries the highlight and
Enter along with the same workspace instead of stranding them on whatever now
sits at that row. `activeIndex` derives from the id and is the single source of
truth for Enter, the visual highlight, and the scroll target.

Search and highlight reset via an effect keyed on the menu-open state, so closing
by any path (selecting a workspace, Escape, click-away) clears them — not only
the Radix-driven close that routes through onOpenChange. Hover-highlight is wired
to mousemove rather than mouseenter so a keyboard-driven scrollIntoView can't fire
a synthetic enter that hijacks the keyboard selection, and composition keys are
ignored while an IME is active so confirming a candidate can't switch workspace.

* fix(settings): pad settings-sidebar scroll list so the last item's hover isn't clipped

The scrollable settings-nav container had `pt-1.5` but no bottom padding, so
the last item's rounded `--surface-active` hover pill was clipped against the
container's overflow edge (visible on "Custom blocks", the final Enterprise
item). Give it symmetric vertical padding (`py-1.5`) so the last row gets
equal breathing room. Padding sits on the scroll container, not between
items, so item spacing is unchanged.
2026-07-23 12:02:31 -07:00
Waleed c8fea6bd3b fix(knowledge): replace deletion-safety heuristics with a two-phase tombstone (#5884)
* fix(knowledge): replace deletion-safety heuristics with a two-phase tombstone

PR #5883 merged an interim fix (sourceConfirmedEmpty bypassing two
static safety heuristics) before this follow-up redesign was ready.
This supersedes that approach with a properly general fix, matching
how production sync systems (Entra Connect, SCIM/Entra ID
deprovisioning, Cassandra/Couchbase tombstones) handle this exact
problem: never let a single observation trigger an irreversible mass
deletion, no matter how confident the signal looks.

A document missing from a normal sync's listing is now soft-deleted
(marked pending-removal) rather than hard-deleted immediately. It's
only actually purged once a *later* sync confirms it's still absent.
If it reappears in between, it's resurrected automatically — this
self-heals a transient outage or a bad API response without needing
to distinguish 'real' emptiness from 'ambiguous' emptiness at all,
which is what the removed heuristics were trying (and failing) to do
from a single observation.

This removes shouldSkipEmptyListing, exceedsDeletionSafetyThreshold,
and sourceConfirmedEmpty entirely — Google Sheets no longer needs a
connector-specific bypass flag, since a genuinely trashed spreadsheet
now reconciles through the exact same general path as every other
connector, with no special-casing and no new misuse surface for
future connectors. A forced fullSync still purges everything absent
in one pass, preserving the existing 'trigger a full sync to force
cleanup' escape hatch.

Uses the existing (previously unused for individual documents)
document.deletedAt column as the tombstone marker — no schema
migration required. shouldReconcileDeletions (the isIncremental /
listingCapped / listingTruncated gate) is unchanged; it still governs
whether reconciliation may run at all. Resurrection runs
unconditionally even when that gate is closed, since presence is
trustworthy evidence regardless of whether the listing was complete.

* fix(knowledge): resurrect atomically on content update, sweep tombstoned docs on connector teardown

Two real gaps found by review, both stemming from the same root
cause: other code paths assumed deletedAt IS NULL means 'the only
real rows' and were never updated for the new tombstone semantics.

- updateDocument's guard required deletedAt IS NULL, so a tombstoned
  document reappearing with CHANGED content failed its content
  update (rejected by the guard) while the separate resurrect step
  still cleared deletedAt regardless — the document became active
  again but kept serving stale pre-tombstone content. Fixed by
  clearing deletedAt as part of the same update statement and
  dropping the guard, so content and resurrection land atomically.
- Both connector-teardown cleanup paths (the ConnectorDeletedException
  handler in sync-engine.ts, and the connector DELETE API route) only
  swept documents with deletedAt IS NULL, so pending-removal documents
  escaped cleanup entirely and were orphaned once their connector was
  gone. Fixed by including tombstoned docs in both sweeps — there's no
  future sync left to confirm or resurrect them once the connector
  itself is deleted.

* fix(knowledge): don't resurrect a tombstoned document whose content refresh failed

Critical race found by adversarial audit: resurrectIds was derived
purely from 'externalId was seen in the listing', independent of
whether the paired content update actually succeeded. If updateDocument
threw (hydration failure, storage upload failure, or any other
transient error) or the deferred-hydration fetch itself failed, the
document's row was never touched — yet the separate unconditional
resurrect step still cleared deletedAt for it anyway, reproducing the
exact bug this PR fixes (visible again, serving stale pre-tombstone
content), just triggered by a failed write instead of a gated one.

Track externalIds whose refresh attempt failed (hydration rejection or
write rejection) and exclude them from partitionSyncReconciliation's
resurrectIds. A failed refresh leaves the document tombstoned as-is —
not soft-deleted, not hard-deleted, not resurrected — so a later sync
gets a clean retry instead of the row landing in an inconsistent state
either way.

* fix(knowledge): exclude fulfilled-but-unverified hydration outcomes from resurrection too

Both bots independently found a third instance of the same bug class:
when deferred hydration for an update fulfills but has no usable
content (skipped as oversized, or an empty re-fetch), the code falls
back to keeping the stored content as last-known-good and counts it
unchanged — correct for an already-visible document, but for a
tombstoned one it means content was never actually verified as
current. That fallback wasn't added to failedExternalIds, so
reconciliation still resurrected it with stale pre-tombstone content
despite hydration never actually confirming anything.

Fixed by adding both fallback branches to failedExternalIds, same as
the rejected-promise cases. Verified across every connector's
skippedReason call site that this only ever happens inside
getDocument (the deferred hydration path already covered here) —
no connector sets skippedReason directly in listDocuments, so there's
no equivalent listing-time gap to fix.

* fix(knowledge): force a full listing when pending-removal documents exist

A subtler instance of the same bug class, this time architectural
rather than a code-path gap: an incremental listing only includes
documents whose content changed since the last sync. A tombstoned
document that's still genuinely present at the source but unchanged
would never appear in an incremental delta at all, so it could never
be resurrected — and on a connector that runs incrementally from here
on (its normal syncMode), it would stay tombstoned indefinitely with
no self-correcting path, only a manual full resync.

Added shouldRunIncrementalSync (extracted as a testable pure function,
matching shouldReconcileDeletions' existing pattern) and a cheap
existence check for any pending-removal document on this connector.
Whenever one exists, this sync forces a full listing instead of an
incremental one, guaranteeing every tombstoned document gets a real
resurrect-or-confirm decision. This only affects which listing mode
runs — it doesn't touch options.fullSync, so the deletion-safety grace
period for other, unrelated documents in the same sync is unaffected.

* fix(knowledge): bound tombstone-forced full syncs, resurrect kept docs on connector delete

Two more real findings from this review round:

- A document whose refresh keeps failing every sync (e.g. permanently
  oversized) never resurrects and never hard-deletes (it's present in
  the listing, just unreadable) — correct on its own, but it also never
  stops being counted by hasTombstonedDocs, so it would force a full
  listing for this connector forever, permanently disabling incremental
  sync on account of one stuck document. Bounded the check to the same
  RETRY_WINDOW_DAYS already used for the stuck-document retry sweep
  below: past the window, this connector stops forcing full syncs on
  the stuck document's account. The document itself is unaffected —
  it stays tombstoned either way, matching the existing 'last-known-good
  forever' tolerance this codebase already accepts for any document
  whose hydration keeps failing.

- The connector-DELETE route's deleteDocuments=false path (kept docs)
  counted tombstoned documents but never resolved them one way or the
  other. With the connector gone, there's no future sync left to ever
  confirm or resurrect them, so they'd become permanent invisible
  orphans holding storage forever. Since 'kept' documents become normal
  standalone KB entries once detached from their connector, resurrect
  any pending-removal ones as part of that transition — consistent
  with what happens to their non-tombstoned sibling documents.

* fix(knowledge): serialize reconciliation writes against a concurrent connector delete

An independent adversarial audit (not just Greptile/Cursor) found the
one genuinely critical gap 6 rounds of bot review missed: resurrect/
soft-delete/hard-delete writes applied raw document IDs snapshotted at
the top of the sync, with no re-check immediately before the write. A
connector-DELETE request choosing to keep documents detaches them
(connectorId set to NULL) via the exact same FOR UPDATE lock on the
connector row that this fix now also takes before applying any
reconciliation write — serializing the two: whichever transaction
commits first wins, and the loser's re-check sees the up-to-date
connectorId and skips any document the other request already claimed.
Without this, a sync racing a 'delete connector, keep documents'
request could silently resurrect-then-strand or soft/hard-delete a
document the user explicitly chose to keep, with a secondary effect of
misclassifying it for storage-billing decrement (which keys off
whether connectorId is still set).

Also tightened the excludedDocs query: it previously required
deletedAt IS NULL, so a document that was both userExcluded and
tombstoned (reachable via excludeConnectorDocuments, which has no
deletedAt filter) fell out of the exclusion set and could be silently
un-excluded and re-indexed on reappearing. Dropped that requirement so
userExcluded is honored regardless of tombstone state, consistent with
how existingDocs/tombstonedDocs are already merged for classification.

Documented (not code-changed) the remaining lower-severity finding: a
document that outlives the 7-day hasTombstonedDocs bound on a
persistently-incremental connector can stay unresolved indefinitely.
Deliberately not hard-deleting it after the window expires — that
would delete a document with no positive evidence it's actually gone,
reintroducing the exact risk this whole design exists to avoid. It's
already fully excluded from search/billing/listings either way, so
this is an accepted, bounded, orphaned-row trade-off, not a
correctness or security issue.

* fix(knowledge): close remaining hard-delete race window, guard listing-time skip/drop resurrection

Two more real findings, both closing gaps in the previous round's fixes:

- The FOR UPDATE lock protected resurrect/soft-delete (applied inside
  the same transaction) but hardDeleteDocuments still ran after that
  transaction committed, using IDs snapshotted under the lock. A
  concurrent 'delete connector, keep documents' request could still
  detach those same documents in the gap between commit and the
  hardDeleteDocuments call. Added an optional expectedConnectorId
  parameter to hardDeleteDocuments/hardDeleteDocumentBatch — when
  provided, it re-verifies connectorId at the moment of the actual
  delete query, not just the caller's earlier snapshot. Every other
  caller is unaffected (parameter is optional, defaults to no filter).

- Two more listing-time paths could resurrect a tombstoned document
  without ever verifying its content: a listing-time skippedReason
  short-circuits classification straight to 'unchanged' before the
  hash comparison ever runs, and empty non-deferred content classifies
  as 'drop' unconditionally regardless of hash. Both are now added to
  failedExternalIds when reappearing on an existing (possibly
  tombstoned) document, same treatment as the deferred-hydration
  equivalents from the prior round.

* refactor(knowledge): dedupe retry-cutoff computation, extract ownership-filter helper

/simplify pass: hoist the shared RETRY_WINDOW_DAYS cutoff into one computation
reused by both the tombstone-retry bound and the stuck-document retry query,
and pull the FOR-UPDATE-lock reconciliation's ID filtering into a pure,
directly-unit-tested filterStillOwnedReconciliationIds function matching this
file's existing convention for its other decision logic.

* fix(knowledge): re-verify connectorId at the actual hard-delete, fix docsDeleted count

Cursor findings: expectedConnectorId was only checked on the pre-transaction
SELECT in hardDeleteDocumentBatch, not on the DELETE itself — the billing
lookups and KB locking in between are async and can span a concurrent
"delete connector, keep documents" request, so the delete (and its embedding
cleanup) now re-verifies against a fresh in-transaction snapshot instead of
the stale existingIds. Also fixed docsDeleted to use hardDeleteDocuments'
actual returned count instead of the pre-filter candidate count, so a sync
log no longer overreports deletions that expectedConnectorId skipped.

/cleanup: dropped two comments that only restated the line below them.

* fix(knowledge): re-verify connectorId in updateDocument's atomic write

Greptile P1: updateDocument's content-update/resurrect write only checked
document.id and archivedAt, never connectorId — despite connectorId being
a parameter — so a document a concurrent "delete connector, keep documents"
request already detached could still be matched, resurrected, and
overwritten with connector-sourced content after the connector was deleted.
Adds the same connectorId ownership check already used by the reconciliation
transaction and hardDeleteDocuments in this PR.

* fix(knowledge): close connectorId race in the stuck-document retry path

Final independent audit: the stuck-document retry block selected candidate
IDs filtered by connectorId, then reset their processing state and deleted
their embeddings using that stale ID set with no re-check — the same
SELECT-then-write race already patched in updateDocument and
hardDeleteDocumentBatch. A concurrent "delete connector, keep documents"
request could null out connectorId in between, so this now re-verifies
ownership immediately before the embedding delete/document update/re-enqueue
and only acts on documents still owned by this connector.

* fix(knowledge): lock the connector row for stuck-doc retry ownership too

Cursor + Greptile (same root cause, two reports): the previous round's fix
re-checked connectorId via a separate SELECT before the embedding delete and
processing-state reset, but a bare re-SELECT only narrows a TOCTOU window, it
never closes it — a concurrent "delete connector, keep documents" request
could still commit its detach in between. Wraps the ownership re-check and
both writes in a transaction that takes the same knowledge_connector FOR
UPDATE lock the DELETE route takes before nulling connectorId, so the two
requests serialize instead of racing, matching the pattern already used by
the reconciliation transaction elsewhere in this PR.
2026-07-23 11:59:24 -07:00