Commit Graph
4679 Commits
Author SHA1 Message Date
Waleed 5c193442f7 improvement(kb-connectors): align connector UI surfaces (#4728) 2026-05-22 15:39:04 -07:00
Waleed 500f35ac87 fix(landing): remove cursor lerp causing laggy tracking in collaboration section (#4727) 2026-05-22 12:58:29 -07:00
Vikhyath Mondreti 209ca5f121 fix(large-refs): cleanup based on table read (#4716)
* fix(large-refs): cleanup based on table read

* address comments

* address comments

* bubble up storage ref errors

* cleanup code

* do not attempt blob deletion for infra outage

* cleanup dup helper
2026-05-22 12:08:22 -07:00
Theodore Li 786c6f0607 fix(db): disable statement_timeout for migrations (#4714)
* fix(db): disable statement_timeout for migrations

* fix(ci): route migration workflow through guarded migrate.ts
2026-05-22 14:23:06 -04:00
Waleed afcbcf2b54 fix(tools): pin resolved IP in DB connectors to prevent DNS-rebinding SSRF (#4725)
* fix(tools): pin resolved IP in DB connectors to prevent DNS-rebinding SSRF

`validateDatabaseHost` resolved an IP that was then discarded — drivers re-resolved
the hostname at connect time, enabling DNS-rebinding TOCTOU.

- mongodb: pass resolved IP via MongoClient `lookup` option
- mysql: pin TCP socket via `stream` factory; keep hostname for TLS servername
- postgresql: connect to resolved IP; pass `ssl` object with `servername` for SNI
- redis: parse URL explicitly and pass options-only (URL+options breaks override
  due to ioredis's lodash.defaults); pin host and set `tls.servername` for rediss
- neo4j: pin IP for plain `bolt://`; leave `bolt+s`/`neo4j+s` unchanged to keep
  Aura cert validation working (driver hardcodes servername with no override)

* chore(tools): remove explainer comments from DB connector SSRF fix

* fix(tools): add explicit TCP timeout to mysql stream factory

* fix(tools): unify postgres ssl handling to send SNI in preferred mode

* fix(tools): preserve postgres 'preferred' fallback behavior for backward compat

* fix(tools): reject non-numeric Redis URL db segment instead of silently using db 0
2026-05-22 10:33:36 -07:00
Waleed b2ad5e9127 improvement(mcp): post-merge hardening — protocol negotiation + distributed OAuth lock + typed errors (#4722)
* improvement(mcp): post-merge hardening — protocol negotiation + distributed OAuth lock + typed error dispatch

- Inbound MCP server now negotiates protocolVersion per MCP 2025-06-18: echoes
  the client's requested version when supported, falls back to our latest.
  Previously hardcoded to the oldest spec version (2024-11-05).
- withMcpOauthRefreshLock now takes a Postgres advisory transaction lock in
  addition to the in-process Promise chain, so concurrent processes (multi-task
  ECS) serialize on a per-OAuth-row basis. Previously a refresh race across
  processes could rotate a token under another process's feet and force re-auth.
- categorizeError dispatches on McpOauthAuthorizationRequiredError /
  UnauthorizedError / McpConnectionError first, only falling back to substring
  matching for SDK / third-party errors. Adds 502 for connection failures and
  503 for cooldown. Tests cover all four typed cases.
- discoverTools no longer pretends to handle deferred-side-effect rejections
  via a dead allSettled().catch() — each side-effect already self-logs; we just
  swallow per-promise to silence unhandled-rejection warnings.

* fix(mcp): bugbot review on hardening PR

- Replace `as never` cast in negotiateProtocolVersion with `as readonly string[]`
  on the array — preserves TypeScript narrowing on the comparison while
  satisfying the readonly-tuple `.includes()` constraint properly.
- Document the pg_advisory_xact_lock tradeoff: session-level locks
  (`pg_advisory_lock`) would release the connection earlier, but PgBouncer
  transaction-pooling mode breaks them. xact_lock is the correct choice for
  Sim's deployment; if pool pressure becomes real, the comment notes the
  Redlock escape hatch.

* chore(mcp): trim verbose comments + reuse SDK Tool type in McpTool

- McpTool now extends `Pick<Tool, 'name' | 'description'>` from
  @modelcontextprotocol/sdk so name/description fields stay in sync with
  the SDK; serverId/serverName remain Sim-specific additions.
- Drop file-header restatements ("MCP Types - for connecting to external
  MCP servers"), one-line wrapper docstrings ("Get connection status"),
  and narrative comment blocks that just restate visible code.
- Keep only comments that document non-obvious "why" — OAuth refresh-lock
  tradeoff, in-flight dedup key composition, SDK Tool.inputSchema typing,
  preregistered-client semantics, postMessage handshake contract.

* improvement(mcp): swap PG advisory lock for Redis mutex on OAuth refresh

withMcpOauthRefreshLock now uses `coalesceLocally` + Redis acquireLock/
releaseLock with polling — the same primitives backing regular OAuth
refresh (`app/api/auth/oauth/utils.ts`). No more pinning a Postgres
connection for the duration of the SDK's OAuth HTTP refresh.

- In-process dedup: shared promise via `coalesceLocally`.
- Cross-process: Redis SET NX EX mutex; followers poll until the leader
  releases (30s max wait, 100ms poll), then acquire and run fn().
- Each MCP caller still constructs its own client (semantics preserved).
- Falls open when Redis is unavailable — same behavior as the regular
  OAuth refresh code path.

* improvement(mcp): use SDK protocol versions + pool pinned undici agents + cover OAuth lock

- McpClient.SUPPORTED_VERSIONS removed; getVersionInfo() and the inbound
  serve route both import LATEST_PROTOCOL_VERSION / SUPPORTED_PROTOCOL_VERSIONS
  directly from @modelcontextprotocol/sdk/types.js. New protocol revisions
  ship automatically with SDK upgrades.
- pinned-fetch now caches undici Agents in a module-level LRU keyed by
  resolvedIP (max 64). Back-to-back MCP calls to the same server reuse the
  keep-alive connection pool instead of opening fresh TCP + TLS each time.
- New integration tests for withMcpOauthRefreshLock covering: in-process
  dedup via coalesceLocally, cross-process serialization via Redis mutex,
  fall-open on Redis unavailable, lock release on throw, release-failure
  swallow, per-row key isolation.

* fix(mcp): serialize OAuth refresh callers; do not share McpClient

withMcpOauthRefreshLock previously wrapped fn() in coalesceLocally, which
returns the SAME promise (and the same resolved value) to all in-process
callers. fn() returns a stateful McpClient — sharing it meant whichever
caller finished first would disconnect the client while another was still
mid-call, leaving in-flight RPC on a closed connection.

Swap coalesceLocally for a per-row Promise chain: each caller waits for
the previous to settle, then runs its OWN fn() (gets its own client).
Cross-process Redis mutex semantics unchanged.

The "shareable scalar" assumption that makes coalesceLocally correct for
regular OAuth refresh (returns an access token string) does not hold for
MCP, where each caller needs an independent connection.

* fix(mcp): bugbot — TTL watchdog on OAuth lock + don't close evicted pinned agents

- Redis refresh lock now uses a 15s TTL with a watchdog that extends every
  5s while fn() runs. Long-running OAuth refreshes no longer lose the lock
  mid-flight and let another process race the same refresh.
- Pinned-agent LRU eviction no longer calls `agent.close()`. Existing
  `createMcpPinnedFetch` closures hold the dispatcher reference and were
  using a closed Agent after eviction. We drop from the cache and let GC
  release the dispatcher once the last closure dies; undici closes idle
  keep-alive connections via its own internal timeout.
- New tests: watchdog extends while fn() runs and stops once it settles;
  evicted agents are not closed and captured closures still work.

* fix(mcp): throw instead of falling open when refresh lock wait exceeds deadline

When the Redis refresh lock can't be acquired within REFRESH_MAX_WAIT_MS
the previous code ran fn() uncoordinated — but another process can still
be holding the lock (watchdog-extended) and refreshing the same OAuth
row, recreating the exact race the lock prevents.

Throw on deadline. The caller can retry; the Redis-down branch remains
the only path that runs fn() uncoordinated (no coordination is possible
there).

* docs(mcp): restore TSDoc that documented intent on exported types/methods

Earlier comment-trim pass went too far on a few exports — restored
the TSDoc that explained non-obvious "why" decisions:

- SimMcpOauthProviderInit.preregistered: when set, the SDK skips DCR.
- McpServerConfig.userId: required for OAuth; selects which user's
  stored tokens to use.
- McpOauthAuthorizationRequiredError: benign pending state vs failure.
- McpToolsChangedCallback / McpClientOptions: notification semantics,
  DNS-rebinding pinning rationale, OAuth provider contract.
- StoredMcpToolReference / StoredMcpTool: minimal vs extended use.
- McpClient.connect: documents listChanged handler registration.
- McpService.executeTool: documents session-error retry behavior.

Pure-restatement comments ("Disconnect from MCP server") stay trimmed.
2026-05-22 10:19:18 -07:00
WaleedandClaude Opus 4.7 19b5099d19 fix(hubspot): selector fetchOptions default + credentialId validation (#4723)
* fix(hubspot): fall back to objectType default in selector fetchOptions

useSubBlockStore.getValue returns null for default-valued dropdowns
until the user interacts with them. The properties, pipelines, stages,
and ownerId selectors were treating that as "no selection" and
short-circuiting, so the dropdowns appeared empty even though the
trigger uses 'contact' as the visible default.

Adds resolveSelectedObjectType to mirror the rendered default, so the
selectors fire on first paint with a valid objectType.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(hubspot): validate credentialId in selector routes

Mirrors the Gmail/Webflow/Jira selector route security pattern by
rejecting non-alphanumeric credentialId values before authorization
or token refresh.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(hubspot): use resolveSelectedObjectType in pipelineId/stageId fetchOptions

Both selectors used inline `?? 'contact'` fallbacks while properties and
targetPropertyName already routed through the resolver. Switch to the
shared helper so custom-object handling stays consistent across every
cascading selector.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-22 09:07:48 -07:00
Waleed b6d08fb0e6 fix(combobox): show selected values in multi-select trigger label (#4721)
* fix(combobox): show selected values in multi-select trigger label

The collapsed trigger was reading only `selectedOption` (the single-value
path) and falling back to the placeholder when nothing matched, so a
multi-select dropdown with 1+ checked items still rendered "Select one or
more channels" instead of the actual selections.

Added `multiSelectLabel` derived from `multiSelectValues`:
- 1 value  → that label
- 2 values → "A, B"
- 3+       → "A, B +N"

Trigger now prefers `multiSelectLabel` when present and falls back to the
single-select label / placeholder otherwise. Muted-text color also flips
off when multi has any selection.

* chore(kb-connectors): strip redundant field-level descriptions

Removed 41 inline `description:` lines from configFields across 16 connectors
(Slack, MS Teams, GCal, Gmail, Notion, Linear-adjacent, Discord, Dropbox,
Evernote, Fireflies, Google Sheets, Intercom, Obsidian, Outlook, Reddit,
ServiceNow, WordPress, Zendesk). They mostly restated the field title (e.g.
"Channels to sync messages from" under a "Channels" label) and cluttered
the add/edit modal. Field titles + placeholders already communicate intent.

Connector-level `description` (used in the connector picker grid) is
unchanged.

* test(leader-lock): use fake timers to deterministically test follower polling

The "follower does a final read after timeout" test (and the
"follower returns null after timeout" test) relied on real-clock
`setTimeout` and `Date.now()` with very tight bounds (pollIntervalMs=5,
maxWaitMs=9). Any CI scheduler jitter of >4ms would cause the second
in-loop poll to be skipped, the polls counter to end at 2 instead of 3,
and the assertion `expect(result).toBe('late-leader')` to fail.

Switched both tests to `vi.useFakeTimers()` so the schedule is driven by
mocked time advanced via `vi.advanceTimersByTimeAsync`. The intent is
unchanged — verify that the in-loop deadline triggers exactly one
post-deadline last-chance call to `onFollower` — but the assertions no
longer depend on wall-clock timing.

Verified across 5 sequential runs with zero flakes.

* improvement(kb-connectors): restore field descriptions as info-icon tooltips

Restores the 41 field-level `description` lines stripped in fc644210d, but
instead of rendering them as inline muted-text paragraphs they're shown via a
small Info icon next to each field title. Hovering or focusing the icon
reveals the description in the existing emcn Tooltip. Keeps the modal layout
tight while preserving the per-field guidance.

Used <button type="button"> as the tooltip trigger (Radix asChild) rather
than <span tabIndex={0}> to satisfy a11y/noNoninteractiveTabindex.
2026-05-22 08:52:18 -07:00
Waleed 1631b36b38 perf(copilot): narrow getAccessibleCopilotChat projection (#4720) 2026-05-22 08:35:39 -07:00
Waleed 1afa8815c2 improvement(branding): dark og image matching landing surface (#4719) 2026-05-22 06:40:07 -07:00
Waleed 3d9a1c4410 fix(oauth): follower last-chance read after poll deadline (#4718)
* fix(oauth): follower last-chance read after poll deadline

* test(oauth): exercise last-chance read in follower timeout test
2026-05-21 23:07:17 -07:00
WaleedandClaude Opus 4.7 0c96964f8f improvement(mcp): per-server tool queries + negative cache (#4715)
* improvement(mcp): per-server tool queries + negative cache so one slow server can't block the workspace

Move MCP tool discovery off the workspace-aggregated `Promise.all` fan-out and
onto per-server React Query keys, matching how Cursor and Claude Code render
remote MCP. `useMcpToolsQuery` is now a `useQueries` combiner: each server has
its own cache entry, its own loading state, and a slow neighbor never gates the
others. Public shape stays compatible with existing consumers.

Add a short-TTL negative cache: when `listTools` fails (timeout, connection
error, etc.) we mark the server unhealthy for 30s so subsequent discovery calls
short-circuit instead of re-paying the timeout. OAuth-required errors are
exempt so re-auth retries immediately. Drop `LIST_TOOLS_TIMEOUT_MS` from 30s
to 10s to bound the worst-case first failure.

Invalidations are per-server where the action is per-server (OAuth popup,
per-server SSE event, refresh, update, delete). Bulk operations stay
workspace-broad. Adds tests for the negative-cache behavior and the OAuth
exemption.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* improvement(mcp): in-flight dedup + 2min negative TTL + no refetch on window focus

Three small follow-ups on top of per-server tool queries:

- Coalesce concurrent `discoverServerTools(userId, serverId, workspaceId)`
  calls into a single upstream `tools/list`. Races between OAuth-callback
  cache priming, post-OAuth UI refetch, and multi-tab loads no longer
  double-fetch the same server.
- Bump negative-cache TTL from 30s to 2 minutes. Cleared on listChanged,
  OAuth completion, manual refresh, and the next successful discovery, so
  this floor only matters for genuinely dead servers — drops their floor
  traffic by 4x.
- Disable `refetchOnWindowFocus` on per-server tool queries. listChanged
  SSE + mutation invalidations already cover real schema changes; alt-tab
  no longer triggers N parallel `tools/list` calls.

* fix(mcp): address bugbot/greptile review on per-server tool discovery

- Workspace-scoped `mcpKeys.serverToolsWorkspace(workspaceId)` prefix for bulk
  invalidations (create-server, refresh-all, SSE workspace fallback, OAuth
  fallback). The previous `mcpKeys.serverTools()` prefix was global and
  invalidated every workspace's tools cache.
- `useMcpToolsQuery` folds `useMcpServers().isLoading` into the aggregate
  `isLoading` so mounting no longer flashes an "empty tools" state during the
  servers-list fetch. Aggregate `error` is suppressed when any per-server
  query already returned data so one slow server can't blank out the others.
- `useForceRefreshMcpTools` invalidates the per-server query keys of servers
  whose force refresh failed, so stale tools don't linger.
- `DiscoveryOutcome` error variant carries the original error, restoring the
  OAuth-exemption check that `getErrorMessage(...)` previously erased.
- `discoverServerTools(userId, serverId, workspaceId, forceRefresh = false)`
  now consults the positive + negative cache by default. Per-server React
  Query refetches hit the cache instead of re-paying the listTools timeout;
  callers that explicitly bypass cache (refresh route, OAuth callback, bulk
  POST refresh) pass `forceRefresh: true`. Negative-cache hits throw a typed
  `McpConnectionError` so the route layer can surface a fast 503.

* update icons

* chore(mcp): remove dead query keys and trim verbose comments

- Drop unused `mcpKeys.tools()` / `mcpKeys.toolsList()` — replaced by
  per-server keys, no remaining callers.
- Trim narrative comments to keep only the non-obvious "why" notes.

* fix(mcp): map negative-cache cooldown error to HTTP 503

`McpConnectionError` thrown when a server is in cooldown previously
fell through `categorizeError` to a generic 500. Cooldown is a
transient-unavailability condition, so route it to 503.

* test(mcp): cover cooldown error → 503 categorization

* fix(mcp): address second-round bugbot review

- useMcpToolsQuery serverIds: filter on enabled + workspaceId match.
  Disabled rows no longer trigger discover calls that get negative-cached,
  and keepPreviousData on useMcpServers no longer races a workspace switch
  into cross-workspace discover requests.
- Aggregate skips per-server data when that server's latest refetch errored,
  so a broken server's last-known tools no longer linger in the workspace
  view while its card shows an error.
- discoverServerTools failure path drops the positive cache alongside writing
  the negative-cache marker. A cache-respecting follow-up now fails fast via
  cooldown instead of returning stale tools from a now-broken server.
- useMcpTools.refreshTools drops the dead forceRefresh param — the per-server
  queryFn always sends refresh=false, so the flag was never effective. Callers
  wanting cache-bypass should use useForceRefreshMcpTools.

* chore(mcp): trim verbose comments

* fix(mcp): third-round bugbot review

- discoverTools failure path now drops the per-server positive cache alongside
  writing the negative-cache marker, matching discoverServerTools' behavior so
  a workspace-aggregate failure doesn't leave stale tools cached.
- useForceRefreshMcpTools filters disabled and out-of-workspace rows before
  fan-out so disabled servers don't 404 → negative-cache themselves.
- Remove unused useMcpServerTools export — the aggregate goes through
  useQueries directly, no external consumer exists.

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-21 22:40:11 -07:00
Waleed e0551b3a1e improvement(kb-connectors): multi-select fields + Slack bot/app message extraction (#4711)
* improvement(kb-connectors): multi-select fields + Slack bot/app message extraction

Adds multi-value support to KB connector configuration fields and applies it
across 8 connectors: Jira (projects), Confluence (spaces), Slack (channels),
Microsoft Teams (channels), Google Calendar (calendars), Gmail (labels),
Notion (databases), and Linear (teams + projects). Each connector emits
byte-identical externalId for legacy single-value configs so existing rows
reconcile in place via the sync engine's externalId-keyed matching.

Framework changes:
- ConnectorConfigField gains `multi?: boolean`
- New `parseMultiValue` helper in @/connectors/utils
- useConnectorConfigFields state model upgraded to string|string[]
- ConnectorSelectorField renders Combobox in multi-select mode when `field.multi`
- Add/edit connector modals handle array values end-to-end

Per-connector specifics:
- Jira: JQL `project in (...)` for 2+ keys, `project = X` for one
- Confluence: routes through CQL `space in (...)` when multi; v2 fast path stays
  for single+no-label; also fixes selector returning space.id instead of space.key
- Slack: loops per channel emitting one document each; extracts text from
  attachments and Block Kit blocks (incl. nested attachment.blocks where GitHub
  embeds PR bodies); contentHash bumped to slack-v2: to force one-time re-embed
- Microsoft Teams: loops per channel within a single team
- Google Calendar: compound cursor across calendars; single-calendar keeps
  legacy externalId/contentHash for zero-churn
- Gmail: (label:A OR label:B) with quoted-form for labels with spaces
- Notion: sequential walk via JSON compound cursor; single-DB keeps bare cursor
- Linear: GraphQL IdComparator.in for multi, eq for single

* fix(kb-connectors): valuesEqual treats legacy scalar as equal to multi-array

Existing connectors created before multi-select store sourceConfig values as
scalars (e.g. projectKey: "ENG"). With the field now declared multi: true,
resolveSourceConfig returns an array (["ENG"]), and the original valuesEqual
fell through to a strict reference comparison — falsely flagging unsaved
changes on open and triggering an unnecessary string→array shape rewrite on
save.

valuesEqual now normalizes both sides to string[] via CSV-split when either
is an array, so persisted scalar and in-memory array of the same content
compare equal. Single-value (non-multi) fields keep strict string equality.

* fix(kb-connectors): GCal externalId on config downgrade, Slack silent skip, valuesEqual order

- google-calendar getDocument: derive isMultiCalendar from the externalId's
  `:` separator instead of the current config count. Prevents duplicates when
  a user downgrades from multi to single calendar — previously the returned
  doc lost its `calendarId:` prefix and was treated as a new row by the sync
  engine, orphaning the original.
- slack listDocuments: throw on unresolvable channel instead of silently
  skipping. Matches MS Teams behaviour. Silent skip would let the sync
  engine orphan-delete the previously indexed channel content if a bot was
  removed or a channel was archived/renamed mid-life.
- edit-connector-modal valuesEqual: order-insensitive comparison for multi-
  select arrays via Set membership. Multi-select UI doesn't guarantee
  insertion order matches the server-returned order, so `["A","B"]` vs
  `["B","A"]` would otherwise flag false unsaved changes.

* chore(kb-connectors): use emptyValue() fallback in isFieldPopulated for consistency

Behavior unchanged — isValuePopulated('') and isValuePopulated([]) both return
false — but reading the field-typed fallback inline matches the convention
used elsewhere in the hook (coerceForField, handleFieldChange, resolveSourceConfig).

* fix(kb-connectors): Linear projects selector loads across all selected teams

When the team selector is in multi-select mode, the basic-mode projects
dropdown was passing only the first team ID into the linear.projects
selector context (via readFirst in resolveDepValue), so projects from other
selected teams were invisible.

resolveDepValue now joins multi-value parents into a CSV string so dependent
selectors receive every selected parent ID. The /api/tools/linear/projects
route splits the CSV teamId, fetches projects from each team in parallel,
and dedupes by project ID. Single-team configs pass through unchanged
(`split(",")` on a bare ID yields a one-element array).

The AND-of-filters semantics in buildIssuesQuery is intentional and matches
standard GraphQL filter behavior — a user filtering on teams [A,B] and
projects [X,Y] gets issues in (A or B) AND (X or Y). With this fix the
project dropdown now shows every project under any selected team so the
user can compose the right project set.

* fix(gmail-connector): always wrap OR-containing custom query, not just unwrapped ones

The previous check `!/^\(.*\)$/.test(trimmedCustom)` was supposed to avoid
double-wrapping an already-parenthesized expression, but it false-positives
on inputs like `(from:alice) OR (from:bob)` where the parens don't bracket
the whole string. Those would skip wrapping and the top-level OR would bind
across the preceding label / category / date filters instead of the custom
clause.

Always wrap when an OR is present — double-parens are a no-op in Gmail
search syntax, so `((from:a OR from:b))` parses the same as `(from:a OR
from:b)`. Simpler than walking parens depth and provably safe.
2026-05-21 19:41:45 -07:00
Waleed 952eb1216f feat(mailer): add AWS SES and SMTP providers with auto-detect fallback (#4710)
* feat(mailer): add AWS SES and SMTP providers with auto-detect fallback

* fix(mailer): cast SES options to bridge duplicate @aws-sdk type identities

* fix(mailer): dedupe aws-sdk-sesv2, address review feedback

- Force a single @aws-sdk/client-sesv2 install via root package.json overrides; @types/nodemailer pulled in a nested copy whose nominal class brand made the two SDK type identities incompatible, breaking the CI build. With one install the cast disappears.
- Batch result message now reports successCount instead of sendable.length when entries are skipped, so "5 emails sent" no longer overstates delivery on partial failures.
- SMTP provider now warns when SMTP_HOST is set without SMTP_PORT, and when only one of SMTP_USER/SMTP_PASS is set — both previously silent misconfigurations.
- SMTP_SECURE schema is z.boolean() to match every other boolean in env.ts; runtime parsing is still handled by envBoolean.
- Strip the verbose TSDoc comments I had added.

* fix(mailer): exact sent counts in batch results, restore SES type cast

- mergeBatchResults: data.count and the message now report only emails that were actually delivered, not skipped-unsubscribed ones (they returned success: true and inflated the count). Empty-sendable branch distinguishes "all unsubscribed" from "mixed skip/failure" so the message stops lying when some entries fail validation.
- ses.ts: revert the package.json override approach (bun honors it locally but CI still installs a nested @types/nodemailer copy). Reinstate the `as unknown as` cast with a single-line WHY comment.

* fix(mailer): annotate double-cast in ses provider for strict api-validation

* fix(mailer): batch degrades isUnsubscribed errors to per-entry failures

A transient DB error in isUnsubscribed used to abort the whole batch
because the call sat outside the per-email try/catch in prepareBatch.
Wrap the unsubscribe check inside the same catch so a rejection becomes
a per-recipient failure, matching sendEmail's behavior. Lock it in
with a regression test.
2026-05-21 19:27:33 -07:00
Vikhyath Mondreti 1af65381f5 improvement(search-replace): pass down to subblocks (#4712)
* improvement(search-replace): pass down to subblocks:

* fix local lowercasing bug
2026-05-21 19:16:18 -07:00
Waleed 543d2faa70 fix(files): RFC 5987 encode Content-Disposition filenames (#4713) 2026-05-21 18:58:56 -07:00
Waleed 21c956cf97 improvement(hubspot): OAuth-native polling trigger replacing webhook flow (#4705)
* improvement(hubspot): OAuth-native polling trigger replacing webhook flow

* feat(hubspot): property autocomplete, multi-filter, property-changed, list-membership, pipeline/owner dropdowns

* fix(hubspot): freeze cursor on failure + request full OAuth scopes

* chore(api): bump API route baseline from 749 to 753 for HubSpot selector routes

* fix(hubspot): make eventType required conditional on visibility

* improvement(hubspot): align trigger name and longDescription with poll-trigger conventions

* fix(hubspot): encodeURIComponent on search path segment for defense-in-depth

* fix(hubspot): cursor-based seed for list_membership polling

* fix(hubspot): Map-backed property snapshot + drop redundant filter parse
2026-05-21 17:25:19 -07:00
Waleed 92764631ed improvement(oauth): coalesce token refresh + cache terminal failures (#4706) 2026-05-21 17:21:54 -07:00
Theodore LiandClaude Opus 4.7 4002242d84 feat(tables): virtualize data grid with bounded copy and chunked delete (#4693)
* feat(tables): virtualize data grid with bounded copy and chunked delete

* fix(tables): keyboard scroll past sticky header, copy toast row count, clipboard error handling

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(tables): keep copy/cut progress toast until the operation completes

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(tables): stop bulk cut chunks on first failure, reconcile grid on error

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 19:17:20 -04:00
Vikhyath Mondreti 47bd7fa345 fix(logs-cleanup): listing active workspaces into mem + download time streaming lims (#4692)
* fix(logs-cleanup): listing active workspaces into mem + download time streaming lims

* fix

* fix ci

* address comments

* add skill

* fix client-server sep

* fix parse bytes enforcement

* address comments, antipatterns

* address slides ssrf comment

* more fixes

* fix tests

* fix
2026-05-21 16:16:03 -07:00
Waleed 3f7698c66b perf(db): reduce read/write fanout across hot paths (#4704)
* perf(db): reduce read/write fanout across hot paths

* fix(templates): include name/workflowId/status/tags in DELETE projection

* fix(db): address review feedback (joinedAt backfill, mcp delete error, chatDeploy projection, sql newline)

* fix(webhooks): include NULL failedCount in markWebhookSuccess guard
2026-05-21 15:25:31 -07:00
mini d7ed3c24c0 fix(sidebar): pass showDelete to hide delete menu for non-admin members (#4697)
The ContextMenu component already has a showDelete prop with conditional
rendering, but workflow-item and folder-item never pass it, leaving it
at the default value of true. This causes write members to see an active
Delete option that always fails with 403, since the DELETE API requires
admin permission.

Pass showDelete={userPermissions.canAdmin} from both workflow-item and
folder-item so that non-admin users no longer see the Delete menu.
Simplify disableDelete to only check canDeleteSelection and
effectiveLocked, since permission gating is now handled by showDelete.
2026-05-21 14:27:42 -07:00
Waleed 89695de81d fix(mcp): cache result of discoverServerTools to prevent post-OAuth refetch storm (#4701)
* fix(mcp): write per-server cache from discoverServerTools to prevent post-OAuth refetch storm

* improvement(mcp): update server status from discoverServerTools + cap list-tools timeout at 30s
2026-05-21 11:10:09 -07:00
Waleed d892165337 fix(copilot): default SIM_AGENT_API_URL to www.copilot.sim.ai to avoid redirect path drop (#4700) 2026-05-21 10:50:35 -07:00
Theodore Li 896eee3758 fix(table): derive typewriter slice from elapsed time (no full-text flash) (#4694)
The reveal used a lagging nullable `revealed` state with a `revealed ?? kind.text`
fallback in the caller. Under React 18 concurrent rendering a committed render
could observe `revealed === null` while `text` was the full value, so the
fallback painted the entire string for one frame before the type-on — an
intermittent flash, reproducible on a large Run-all (verified in-browser: 60+
cells flashing).

Derive the revealed slice from `text` + elapsed time during render instead of
holding it in state. For a non-null value the result is never `null` and never
the full string on the frame `text` changes (elapsed ≈ 0 → 0 chars), so the
fallback can't fire. `prevText` is tracked in state (not a ref) so a discarded
render rolls it back and the change is re-detected on the committed render.
Verified via DOM MutationObserver: 0 flashes across 213 animated cells.
2026-05-21 05:52:34 -04:00
Waleed d730015022 perf(mcp): per-server tool cache + surface OAuth start errors (#4691)
* fix(mcp): surface authorization-server error instead of generic toast

* perf(mcp): cache tool discovery per-server with test coverage

* refactor(mcp): use SDK's typed OAuthError subclasses for error surfacing

* test(mcp): unit-test surfaceOauthError typed and fallback paths

* chore(mcp): remove unused __where mock export in service.test.ts

* chore(mcp): drop redundant clearCache TSDoc and stray trailing comment

* fix(mcp): don't leak non-OAuth errors; clearCache covers disabled servers

* chore(mcp): tighten clearCache comment to one block
2026-05-20 23:31:49 -07:00
Waleed 11ad8918a9 improvement(cleanup): batchTrigger fan-out, chunked queries, batched S3, faster outlier drain (#4688)
* improvement(cleanup): batchTrigger fan-out, chunked queries, batched S3, faster outlier drain

- Fan cleanup-tasks/logs/soft-deletes out via tasks.batchTrigger (500 ws/chunk); bump to large-1x with concurrencyLimit: 5
- Chunk bulk DELETEs (1000 IDs/stmt) and collectChatFiles JSONB SELECT (500 chats/stmt) to bound worker memory and lock duration
- Replace per-key position() table scans with one LATERAL unnest scan per 200-key chunk
- Route storage deletes through StorageService.deleteFiles (S3 DeleteObjects: 1000 keys/HTTP)
- Raise per-run row cap to 100K so long-tail tenants (one prod workspace has 723K doomed rows) drain in days, not weeks

* improvement(cleanup): chunk-index labels, clarify upper-bound failure counter

Addresses Greptile review feedback:
- Disambiguate downstream logs when a plan splits into multiple workspace chunks (e.g. 'free/1', 'free/2')
- Document that deleteRowsById's failed counter is an upper bound (chunk rolls back to 0 deletes on error)
2026-05-20 19:35:40 -07:00
Waleed e27afaad0f fix(mcp): probe-based OAuth detection in test-connection (#4689)
* fix(mcp): probe-based OAuth detection in test-connection

* fix(mcp): guard undefined url in probe and use canonical McpAuthType
2026-05-20 19:35:33 -07:00
Theodore Li 57b9a2fb38 fix(table): typewriter flash, Run-row completed-skip, dispatch-scope running count (#4687)
* fix(table): no typewriter flash; Run-row skips completed workflows

- Typewriter: reset the revealed text synchronously during render when the
  value changes (not in an effect), so a cell going from running→value no
  longer flashes the full text for one frame before animating.
- Run row / manual incomplete runs now treat a `completed` group as done even
  if an output column is blank — only "Run all" re-runs completed cells. The
  auto cascade keeps re-filling blank outputs (completedAndFilled). Client
  optimistic stamp mirrors: incomplete skips `completed` cells.

* fix(table): incomplete bulk-clear is per-group, not per-row

bulkClearWorkflowGroupCells in incomplete mode wiped EVERY targeted group's
output data + exec on any row that wasn't fully filled across all targeted
groups. So Run-row on a row with one completed group and one cancelled group
wiped the completed group's outputs + exec too, and the dispatcher re-ran it.

Now incomplete-mode clears per-group: only error/cancelled groups get their
output columns + exec cleared; completed and in-flight groups are left intact
(never-run groups have nothing to clear and run via eligibility). Combined
with the classifyEligibility guard, a completed workflow is never re-run by
Run-row — only Run-all re-runs it.

* fix(table): X-running count from dispatch scope so reload matches live

The "X running" badge read countRunningCells (sidecar in-flight), but the
dispatcher only stamps one ~20-cell window at a time. During a 1000-row
Run-all the client optimistically showed ~1000 while a reload showed ~20 —
the sidecar never holds more than a window.

Derive the count from the active dispatches instead: rows in scope ahead of
the cursor × |groupIds| (exact for Run-all, upper bound for incomplete/new).
Both scope and cursor are persisted, so a reload computes the same number.

- countActiveRunCells (dispatcher.ts): dispatch-scope total, sidecar fallback
  when no dispatch is active. byRowId stays sidecar-based (the client overlay
  renders queued rows ahead of the cursor).
- Live: applyDispatch re-syncs the badge from the server on every dispatch
  event (one per window, after its cells finish + cursor advances), so the
  badge steps down per window and matches reload. applyCell no longer touches
  runningCellCount (still keeps runningByRowId live for the gutter).
- Optimistic on click: useRunColumn seeds the full run scope (totalCount ×
  groups) so the badge is right before the first window lands.

* fix(table): address review — parallel dispatch counts, unfiltered rowCount

- countActiveRunCells: run the per-dispatch COUNT queries + the sidecar count
  in parallel instead of serially (one round-trip per dispatch).
- Optimistic Run-all estimate now reads the table definition's maintained,
  unfiltered rowCount (detail cache) instead of the rows query's filter-scoped
  totalCount — the dispatcher runs every row regardless of the active filter.
2026-05-20 18:55:27 -07:00
Waleed 48cf200ccd fix(helm): allow host[:port][/path] form in global.imageRegistry schema (#4686)
The values.schema.json constrained global.imageRegistry to JSON Schema
hostname format (RFC 1123), which forbids '/'. That rejected the
host+path form required by Artifactory virtual repos, Harbor projects,
GCR (gcr.io/project-id), and ECR-with-namespace — all of which the
chart's image-rendering helper already supports (it prints '%s/%s:%s').

Drop the format constraint and document the supported shapes. Matches
the bitnami common-chart convention of validating image registry as a
plain string and deferring to Docker for the actual reference parse.
2026-05-20 17:52:52 -07:00
Theodore Li 4445e31768 fix(table): bump run counter on edit/auto-run so Stop shows for queued cells (#4682)
The "X running" badge + per-row gutter Stop only updated on manual Run
(useRunColumn bumped the run-state counter). Edit-triggered auto-runs
(useUpdateTableRow, useBatchUpdateTableRows, useCreateTableRow) stamped cells
pending in the rows cache but never bumped runningCellCount/runningByRowId, so
Stop stayed hidden even though cells were queued (the counter is already
queued-inclusive). Extracted countNewlyInFlight + bumpRunState helpers and
wired them into all the optimistic auto-fire paths with onError rollback;
reused them in useRunColumn.
2026-05-20 17:22:26 -04:00
WaleedandClaude Opus 4.7 9347da5e5b improvement(branding): white-background sim wordmark for og image (#4683)
Swap the default Open Graph and Twitter card image from the purple "sim"
lettermark on purple to the brand wordmark (green icon + dark "sim" text)
centered on a white background. Pulled directly from the sidebar's
wordmark-dark.svg so the asset stays in sync with the in-app brand.

- New asset at /logo/426-240/reverse/small.png (2130x1200, matches
  declared OG dimensions)
- Default branded metadata + landing-page-specific overrides both updated;
  all sub-pages that inherit the default pick up the new image
  automatically

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-20 14:22:11 -07:00
Waleed 4ca7651ea6 improvement(knowledge): eliminate N+1 on tag definitions in bulk upload (#4681)
* improvement(knowledge): eliminate N+1 on tag definitions in bulk upload

createDocumentRecords previously called processDocumentTags per-doc, each
running a SELECT against knowledge_base_tag_definitions — N queries that
all returned the same kbId-scoped rows. Worse, those reads used the
global db pool while the tx held a FOR UPDATE lock on the KB row, risking
pool contention on large bulk uploads.

Split the helper into loadTagDefinitions (single query, accepts the tx as
executor) and resolveDocumentTags (pure, takes the pre-loaded Map). The
bulk path loads once inside the transaction; createSingleDocument loads
once outside its tx. Same throw-on-validation-error semantics preserved.

* improvement(knowledge): fold processDocumentAsync prefetches into one JOIN

processDocumentAsync was issuing three separate SELECTs per processed
document: knowledge_base (config), workspace (billing settings), and
document (tag values). For a typical Trigger.dev fleet processing
thousands of docs, that's thousands of redundant pool checkouts.

Collapsed into a single JOIN at the top of processDocumentAsync that
fetches kb config + billed account user + document tag values in one
roundtrip. The post-embedding tag SELECT (which previously held tags
through the full embedding-generation wait) is gone; tags from the
initial prefetch are reused.

Behavior:
- Missing/archived/deleted document or KB → same 'failed' status outcome
  as before, single consolidated error message.
- Missing billed account → preserves existing error.
- All 208 KB tests pass (test mock extended for innerJoin/leftJoin).

* improvement(knowledge): skip tag-definitions load when no doc carries tags

Trim verbose comments in the same pass.

* lint
2026-05-20 14:18:30 -07:00
Waleed b6679a9880 feat(google-slides): complete API surface for branded slide generation (#4678)
* feat(google-slides): complete API surface for branded slide generation

* fix(google-slides): address PR review — explicit videoId mapping, fast base64 export, remove dead utility

* fix(google-slides): declare z-order operation output in block outputs map
2026-05-20 13:46:40 -07:00
Waleed 7e678552d6 improvement(elevenlabs): wire stability and similarity_boost end-to-end (#4679)
* improvement(elevenlabs): wire stability and similarity_boost end-to-end

* fix(elevenlabs): guard NaN in voice settings and always send both knobs together
2026-05-20 13:46:29 -07:00
Waleed a1b2130256 improvement(knowledge): batch trigger dispatch, prune redundant DB roundtrips (#4680)
* improvement(knowledge): batch trigger dispatch, prune redundant DB roundtrips

Connector sync was dispatching Trigger.dev document-processing jobs one
HTTP roundtrip at a time. processDocumentsWithQueue now uses
tasks.batchTrigger when Trigger.dev is available, collapsing N roundtrips
to ceil(N/1000). Idempotency keys protect against duplicate runs on retry.

Also trims DB roundtrips inside the sync loop:
- Per-batch isConnectorDeleted + isKnowledgeBaseDeleted collapsed into a
  single checkSyncLiveness JOIN (one SELECT instead of two per batch).
- Dropped redundant pre-upload isKnowledgeBaseDeleted checks from
  addDocument/updateDocument: the batch-boundary liveness check already
  catches pre-batch deletions and the in-tx FOR UPDATE is authoritative
  for races during the batch.
- Removed dead processDocumentsWithTrigger helper (never called).

* refactor(knowledge): split dispatch helpers, drop dead trigger branch

- Use the canonical DocumentProcessingPayload from the task module instead
  of the duplicate DocumentJobData interface in service.ts
- Pass typeof processDocumentTask as a generic to tasks.batchTrigger so the
  payload shape is type-checked against the task definition
- Inline TRIGGER_BATCH_SIZE provenance (Trigger.dev SDK 4.3.1+ doc'd cap,
  we're on 4.4.3)
- Split direct vs trigger dispatch into dispatchInProcess and
  dispatchViaBatchTrigger; collapse the all-failed throw into a single
  check on the combined dispatched counter
- Remove dispatchDocumentProcessingJob — its trigger branch is no longer
  reachable now that batchTrigger handles the trigger path, and the direct
  branch is inlined

* improvement(knowledge): log Trigger.dev batchIds for audit trail

tasks.batchTrigger returns a batchId per call. Collecting and logging
them after dispatch makes it possible to look up or cancel batches in
the Trigger.dev dashboard when investigating stuck or missing documents.

* improvement(knowledge): thread requestId through direct dispatch logs

Symmetry polish: dispatchInProcess now includes [requestId] in its error
log so direct-mode failures are correlatable the same way trigger-mode
failures already are.

* improvement(knowledge): trim verbose comments

Tightens TSDoc on processDocumentsWithQueue, TRIGGER_BATCH_SIZE,
checkSyncLiveness, and the idempotency-key inline comment.
2026-05-20 13:45:42 -07:00
Waleed c3815509b2 fix(landing-nav): scroll to top on route change in shared shells (#4676)
* fix(landing-nav): scroll to top on route change in shared shells

* fix(landing-nav): hoist popstate flag to module scope and skip on hash anchors

* fix(landing-nav): use popstate timestamp window so flag self-expires

* fix(landing-nav): use -Infinity sentinel and consume popstate timestamp on use

* fix(landing-nav): skip scroll on initial mount so reload restoration wins

* fix(landing-nav): gate initial-mount skip on document.readyState

* refactor(landing-nav): simplify scroll-to-top to conventional pattern
2026-05-20 12:48:48 -07:00
Theodore Li af8025c411 fix(table): dispatcher cold-start, live run counter, smooth typewriter (#4675)
* fix(table): cut dispatcher cold-start by lazy-loading heavy import chains

The trigger.dev table-run-dispatcher spent ~6s in module init before its
first batchTriggerAndWait — it imports lib/table/service for getTableById,
which eagerly imported lib/table/trigger → @/lib/webhooks/processor →
webhook-execution + executor, dragging the entire workflow-execution stack
into the dispatcher container even though it never fires a trigger.

- trigger.ts lazy-imports the webhook processor + polling utils inside
  fireTableTrigger (the only consumer), so importing service no longer pulls
  the executor.
- buildEnqueueItems only imports the cell job (for the inline `runner`) on the
  database backend; the trigger.dev backend triggers by task id and ignores
  runner.

* fix(table): run counter + gutter Stop update instantly on Run

The "X running" badge, per-row gutter Stop button, and runningByRowId map
stayed at zero after clicking Run until a manual refetch. useRunColumn
optimistically stamped cells pending in the rows cache but never bumped the
activeDispatches counter — so when the dispatcher's real pending SSE arrived,
applyCell saw the cell was already in-flight (wasInFlight === isInFlight) and
skipped the counter delta. The optimistic stamp ate the transition.

- onMutate now bumps runningCellCount / runningByRowId by the cells it stamps,
  snapshotting prior run-state for rollback on error.
- onSuccess seeds the dispatch into the overlay list from the response instead
  of invalidating activeDispatches (a refetch would reset the optimistic
  counter to the server's still-zero count before the dispatcher stamps).

* fix(table): drive cell typewriter with rAF so concurrent reveals stay smooth

The character-by-character reveal used a per-cell setInterval. When many cells
reveal at once (a Run-all completing in waves), the independent interval
callbacks fire at uncoordinated times and each forces its own render +
layout/paint — O(cells) reflows over an un-virtualized grid, so it degrades as
more cells fill. Switch to requestAnimationFrame: all cells' callbacks run
before one paint, so React batches them into a single render + paint per frame
regardless of cell count. Reveal length is derived from elapsed time, so a
dropped frame catches up instead of slowing the animation.

* fix(table): roll back optimistic run counter when no dispatch is created

useRunColumn.onSuccess returned early on a null dispatchId (no matching
groups / eligible rows) without undoing the onMutate counter bump — and no
SSE would arrive to correct it, leaving the counter permanently inflated.
Restore the pre-mutation run-state on that path, mirroring onError.

* chore(table): tighten inline comments on dispatcher cold-start fixes
2026-05-20 15:15:45 -04:00
WaleedandClaude Opus 4.7 46db40620f feat(mcp): OAuth 2.1 + PKCE for outbound MCP servers (#4441)
* feat(mcp): OAuth 2.1 support for outbound MCP servers

* fix(mcp): tighten OAuth refresh race and session-error detection

Re-load the OAuth row inside withMcpOauthRefreshLock so concurrent
callers observe predecessor-written tokens instead of a stale snapshot
loaded before lock acquisition. Without this, the second caller's
provider held a rotated-out refresh token and the SDK tripped
invalid_grant, forcing reauthorization.

Switch isSessionError to match the SDK's typed StreamableHTTPError
(code 404/400) instead of substring-checking arbitrary error messages,
removing false positives on URLs that happen to contain those digits.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(mcp): tighten OAuth callback contract and registration metadata

- Validate callback query params via mcpOauthCallbackContract instead of
  raw searchParams.get, matching the rest of the MCP route surface.
- Drop non-RFC-7591 application_type field from dynamic client registration
  to avoid rejection by strict authorization servers.
- Collapse the pre-lock OAuth row load in createClient — the row is now
  loaded exclusively inside withMcpOauthRefreshLock, removing a redundant
  query and a stale-snapshot path.

* fix(mcp): narrow workspaceId before async closure in OAuth createClient

* fix(mcp): return authType from create-server endpoint

The POST /api/mcp/servers handler omitted authType from the success
response, so useCreateMcpServer always saw data.data.authType as
undefined and never triggered the OAuth popup after creating an
OAuth-protected server. Thread authType through performCreateMcpServer
into the response so the client can decide whether to auto-start OAuth.

* fix(mcp): mirror server null normalization in optimistic oauthClientId update

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(mcp): revert optimistic oauthClientId to undefined to match McpServer type

The response contract preprocesses null → undefined, so McpServer.oauthClientId
is string | undefined. Using null broke type checking.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(mcp): tighten OAuth probe signal and clear stale popup interval

- probe: only classify as OAuth on resource_metadata or scope params.
  Bare `Bearer error="invalid_token"` is generic and used by API-key servers,
  so it must not auto-flip the auth type to OAuth.
- popup hook: clear any existing close-watcher interval before overwriting
  when startOauthForServer is invoked twice for the same serverId.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(mcp): normalize empty-string oauthClientId at route boundary

Orchestration already converts falsy → null via `|| null` (server-lifecycle.ts),
so the DB was never receiving an empty string. Tightening the route layer to
match the same convention keeps the boundary contract consistent and avoids
relying on downstream normalization.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(canvas): expand MCP tool params into per-row labels on block tile

The MCP Tool block on the workflow canvas previously crammed every selected-
tool parameter into a stringified blob under the `Tool` row. Now, when a tool
is selected, the tile reads the cached `_toolSchema` and emits one labeled
SubBlockRow per parameter (matching the Exa block's per-param layout). Labels
reuse `formatParameterLabel` for parity with the editor panel; values pass
through the existing `getDisplayValue` so booleans/numbers/arrays render
identically to other blocks. Deterministic tile height counts expanded rows
so the tile sizes correctly.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* feat(logs): show MCP icon and strip prefix in trace tool spans

Tool spans for MCP calls were rendering the raw id (e.g.
`mcp-f908f259-planetscale_list_organizations`) with the default blank-
square icon. Now they read just the tool name and render the MCP block's
icon and bgColor, matching how workflow-execute tools render.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(logs): lift near-black trace icon backgrounds for dark-mode contrast

Block bgColors below a small luminance threshold (e.g. the MCP block's
#181C1E) rendered nearly invisible against the dark-mode surface
(--bg: #1b1b1b). Adds a tiny adjustBgForContrast helper that floors each
RGB channel at 0x33 only when luminance is below 30,000, leaving every
branded color above that band untouched. Applied to both the trace tree
row and the detail pane.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(logs): fall back to neutral gray for near-black trace icon bgs

#333333 was still too close to the dark-mode surface to read. For bgs
below the luminance threshold (e.g. the MCP block's #181C1E) we now fall
back to DEFAULT_BLOCK_COLOR (#6b7280) — the same neutral the renderer
uses for blocks with no distinct identity. Clearly visible in both
themes; brighter brand colors still pass through.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(db): drop 0209_mcp_oauth migration ahead of staging merge

Staging shipped 0209_smiling_fixer; the MCP OAuth migration will be
regenerated on top of staging as 0210.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(db): regenerate MCP OAuth migration as 0210

Re-runs drizzle-kit generate on top of staging's 0209_smiling_fixer.
Same schema (mcp_server_oauth table + mcp_servers.auth_type / oauth_*
columns) as the dropped 0209_mcp_oauth.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(audit): bump route baseline 748 → 749 after staging merge

The post-merge route count is 749 (this branch's OAuth start/callback
plus staging's new route). I had set the baseline to 748 in the merge
conflict resolution — bumping to match reality so the strict audit
passes.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore: remove source-command skill files committed by accident

These were untracked-then-accidentally-staged in 05c4bc19e via a wide
`git add -A`. They aren't part of this PR's scope.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-20 12:10:04 -07:00
WaleedandClaude Opus 4.7 d9dd7a3e55 fix(cors): re-enable credentials on chat/form embed CORS policy (#4673)
* fix(cors): re-enable credentials on embed CORS policy

Chat and form embeds authenticate via the chat_auth_<id> / form auth
cookie set by setDeploymentAuthCookie. The previous PR set
Access-Control-Allow-Credentials: false on these routes, which made the
browser drop the auth cookie and produce 401s on subsequent embed calls
after login. Restore credentials: true (matching pre-consolidation
behavior) while keeping reflected origin and Vary: Origin.

The wildcard fallback when Origin is absent now also drops credentials
to stay CORS-spec-compliant.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(cors): trim verbose comments in proxy

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(cors): restore concise TSDoc on proxy CORS helpers

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-20 10:55:26 -07:00
Theodore LiandClaude Opus 4.7 f0311a6f5e feat(table): chunked dispatcher + workflow cascade (#4672)
* feat(table): chunked dispatcher for workflow-column runs

Replaces the all-rows-at-once runWorkflowColumn with a row-window dispatcher
backed by a new table_run_dispatches row. Each user click inserts a dispatch
row and triggers a trigger.dev task that crawls the table 20 rows at a time,
re-enqueueing itself between windows. The HTTP/Mothership entrypoints return
{ dispatchId } immediately instead of holding the request open for minutes
on multi-thousand-row dispatches.

- Per-row cancel stamps cancelledAt; the dispatcher skips cells whose
  cancelledAt > dispatch.requestedAt so a mid-cascade cancel sticks even
  under isManualRun.
- Table-wide cancel marks active dispatches cancelled atomically so the
  dispatcher bails on its next iteration.
- New 'dispatch' SSE event variant plumbed; client ignores for v1.

* fix(table): eager bulk clear on column run so cells flip immediately

Run-column with run-mode 'all' wasn't visually flipping rows that already
had data — the cell renderer's "value wins" branch kept showing the prior
output behind the queued/running state. The dispatcher only cleared one
window of rows at a time, so most of the column stayed stale until the
cursor walked to it.

Now:
- Dispatcher's `pending → dispatching` transition runs a single SQL UPDATE
  that wipes targeted `data` output columns and `executions[gid]` across
  every targeted row (mode-aware: 'incomplete' skips fully-filled rows).
- Per-window clear in `dispatcherStep` is gone — rows are pre-cleared,
  the loop only filters cancel tombstones / unmet deps and enqueues.
- Optimistic patch in `useRunColumn` mirrors the bulk clear by nulling
  output values in the cached row, so the UI flips queued/running
  instantly without waiting for the SSE catch-up.

* fix(table): bulk clear honors in-flight execs under mode: 'incomplete'

The eager bulk clear for mode: 'incomplete' only skipped rows that were
already fully filled, so two overlapping dispatches could race — dispatch B
would nuke executions[gid] on a row dispatch A had just stamped 'queued',
flickering the cell and potentially confusing the worker.

Skip any row whose targeted group is currently queued/running/pending — an
'incomplete' run shouldn't touch what another dispatch is actively working
on. The per-walk 'in-flight' eligibility skip already handles rows that
flip in-flight between the clear and the cursor reaching them.

* refactor(table): dispatcher uses batchTriggerAndWait + tag-based cancel

Switch the per-window cell fan-out from fire-and-forget tasks.trigger to
tasks.batchTriggerAndWait. The dispatcher is now a single long-lived
trigger.dev task that loops dispatcherStep until the table is exhausted;
trigger.dev CRIU-checkpoints the parent during each wait so we don't pay
compute while cells execute. Queue depth is bounded at WINDOW_SIZE per
dispatch — no more flooding trigger.dev with a million queued runs.

- dispatcher.ts builds payloads via the new shared buildPendingRuns helper
  and calls tasks.batchTriggerAndWait directly. Pre-stamps each cell to
  `queued` (jobId=null) so the UI flips instantly.
- table-run-dispatcher.ts is now a plain while-true loop. No
  RUN_BUDGET_MS, no self-re-enqueue, no cold-start tax per window.

Cancel:
- New cancelCellRunsByTags(tags) paginates runs.list + runs.cancel(id).
- cancelWorkflowGroupRuns fires the tag-sweep alongside the per-jobId
  queue.cancelJob path (preserved for auto-fire cells that have real
  jobIds from single tasks.trigger calls).
- Trigger.dev acks the cancel → batchTriggerAndWait resumes → dispatcher
  observes the dispatch-row cancel flag → exits.

Side fixes:
- getAsyncBackendType returns 'trigger-dev' whenever taskContext.isInsideTask
  is true, regardless of TRIGGER_DEV_ENABLED env. The preview/dev-sim
  worker silently routing cell jobs to DatabaseJobQueue (no poller) is
  fixed without any env config change.
- runWorkflowColumn skips the dispatcher entirely when trigger.dev is
  disabled, running cells inline via DatabaseJobQueue.runInline. HTTP
  response returns dispatchId: null in that mode.
- runColumnContract response schema updated to dispatchId.nullable().

* fix(table): show Stop button on optimistic-pending row cells

isExecInFlight required a jobId for `pending` status, gating it as "real
backend pending" vs "optimistic flag only." The row-gutter Stop button
keyed on this — so a freshly clicked Play sat as `pending` (no jobId) and
the user couldn't cancel it until the server-side `queued` stamp arrived
via SSE. With the dispatcher pre-batch stamping cells as `queued` (not
`pending`) and no per-cell jobIds under batchTriggerAndWait, the gap was
worse.

Drop the jobId requirement. `pending` now counts as in-flight everywhere.
Cancel writes `cancelled` to the cell exec authoritatively whether or not
a real trigger.dev run exists yet — cancelling an optimistic cell means
"don't run this," which is correct.

Also collapse isOptimisticInFlight into isExecInFlight since the two
helpers are now identical.

* refactor(table): loop-in-cell cascade + dispatcher-everywhere routing

Two coupled changes:

1. Cell-task runs the row's full cascade in-process. executeWorkflowGroupCellJob
   acquires a Redis lock per (tableId, rowId) with heartbeat (10s/30s TTL),
   then loops through eligible workflow groups for the row. One cell-task =
   one row's full cascade, not N. Resume worker holds the same lock and
   continues the cascade after a HITL resume. Shared withCascadeLock helper
   in lib/table/cascade-lock.ts.

2. Every cell-enqueue goes through the dispatcher. The implicit
   scheduleRunsForRows reactor in service.ts is removed — 8 callsites
   (insertRow, batchInsertRows, upsertRow, updateRowsByFilter,
   batchUpdateRows, addWorkflowGroup, updateWorkflowGroup) now fire
   runWorkflowColumn with mode: 'incomplete', isManualRun: false. HTTP
   routes that call updateRow directly also fire runWorkflowColumn
   afterwards. scheduleRunsForTable / scheduleRunsForRowIds deleted;
   scheduleRunsForRows demoted to private (only the TRIGGER_DEV_ENABLED=false
   fallback uses it). skipScheduler flag dropped from UpdateRowData /
   BatchUpdateByIdData — no longer meaningful since there's nothing implicit
   to suppress.

Plumbed isManualRun through the dispatch row (new is_manual_run column,
default true) so auto-fire callers honor autoRun: false and don't re-run
completed cells.

Stamp 'pending' (not 'queued', executionId: null) before
batchTriggerAndWait — cell-task writes its own 'queued' on lock acquire.

Small UI polish: row gutter Play button spacing, "Delete workflow" →
"Delete column" label, optimistic-pending cells now show Stop button
(isExecInFlight no longer requires jobId).

* fix(table): SQL cancellation guard allows worker to claim a null-execId cell

The dispatcher's pre-batch `pending` stamp leaves executionId unset so any
cell-task that wins the cascade lock can claim the cell. The cancellation-
guard SQL clause was rejecting these claims because it tested
`executions->gid IS NULL` (whole exec missing) but the pre-stamp leaves
the exec present with executionId=null.

Add a third carve-out: `executions->gid->>'executionId' IS NULL`. Now the
guard reads "write allowed if no exec exists, OR no executionId is set
yet, OR the executionId matches ours."

Symptom: every cell-task's first markWorkflowGroupPickedUp call would log
"SQL guard saw cancelled" and skip, leaving cells stuck at the dispatcher's
pending stamp.

* fix(table): dispatcher cursor starts at -1 so position 0 is included

The dispatcher's row-window SELECT is `position > cursor` for exclusive
lower-bound semantics. With cursor initialized to 0, position-0 rows were
never picked up — every dispatch silently skipped the table's first row.

Start cursor at -1 instead. First window's filter `position > -1` matches
position 0; subsequent iterations advance to `lastPosition` which then
correctly excludes already-processed rows.

* refactor(table): align optimistic UI with new dispatcher; sticky cancel via 'new' mode

Fix 0: new `DispatchMode = 'new'` for auto-fire callsites. Eligibility skips
rows with any prior `executions[gid]` entry — cancelled / errored / completed
cells stay sticky until a manual run. Dispatcher's windowed SELECT pushes
`NOT jsonb_exists_any(...)` to SQL so CSV imports into mostly-attempted
tables don't pay a per-window load+JS-filter. `batchInsertRows` drops its
`rowIds` payload (keeps dispatch scope tiny on big imports).

Fix A/B/D: client optimistic patches now mirror the backend's actual
invariants. `useCreateTableRow.onSuccess` stamps eligible groups via
`optimisticallyScheduleNewlyEligibleGroups` so newly-inserted rows show
`Queued` instantly. `useCancelTableRuns.onMutate` distinguishes optimistic-
only pending (`executionId == null` — strip silently) from real worker
claims (stamp cancelled; SSE will reconcile). Drop `onSettled` invalidation
on `useUpdateTableRow` / `useBatchUpdateTableRows` to kill the
delete-cell flicker.

Fix C: active-dispatches overlay. New `listActiveDispatches` helper,
contract, and `GET /api/table/[tableId]/dispatches` route. `kind:'dispatch'`
SSE events carry scope+cursor+mode on every transition. New
`useActiveDispatches` hook + `resolveCellExec` synthesize a virtual
`pending` exec for cells in an active dispatch's scope ahead of cursor —
queued indicators now survive page refresh during long Run-all dispatches.
`cancelWorkflowGroupRuns` emits `kind:'dispatch',status:'cancelled'`
events so the overlay clears without a refetch.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(table): unify trigger.dev and inline dispatcher paths

`runWorkflowColumn` now always inserts a `table_run_dispatches` row and
drives the dispatcher state machine. The trigger.dev / in-process branch
narrows to a single line: trigger.dev fires `tableRunDispatcherTask` (which
calls the new `runDispatcherToCompletion`), the inline path calls the same
helper fire-and-forget. Deletes `scheduleRunsForRows` and
`stampQueuedOrCancel` — the inline-fallback no longer duplicates window
walking, SSE emission, or cancel.

The dispatcher's window-execute call goes through `JobQueueBackend`:
- New `batchEnqueueAndWait` interface method.
- Trigger.dev impl wraps `tasks.batchTriggerAndWait` behind a
  `taskContext.isInsideTask` guard (clear error if called from outside a
  task).
- Database impl skips `async_jobs` entirely — `Promise.all` over
  `options.runner(payload, signal)` per item, with per-cell AbortControllers
  tracked by `cancelKey` for cancel.

`cancelInlineRun` moves to the interface as `cancelByKey` so
`cancelWorkflowGroupRuns` no longer reaches into the database backend.

Fix `mode: 'new'` SQL filter:
- `${array}::text[]` interpolated as a tuple-cast which Postgres rejected
  ("cannot cast type record to text[]") and every inline dispatch silently
  failed. Switched to `ARRAY[${sql.join(...)}]::text[]`.
- Predicate was `jsonb_exists_any` ("any one targeted group present"),
  which excluded rows that needed at least one group re-run after a
  downstream output was deleted. Switched to `jsonb_exists_all` — per-group
  JS eligibility handles the rest.

Cascade-loop workflowId bug: `runRowCascadeLoop` was not threading the new
group's `workflowId` when advancing across groups. The cell-task ran the
previous group's workflow against the next group's cell, terminating
`completed` with empty `accumulatedData`. Fixed by tracking
`currentWorkflowId` alongside `currentGroupId` / `currentExecutionId`.

Client optimistic-patch tightening:
- `useRunColumn.onMutate` mirrors server eligibility — skip cells with
  unmet deps so unmet rows don't flash Queued and get stuck (no SSE will
  arrive for cells the server skipped).
- `resolveCellExec` overlay synthesizes a virtual `pending` only when
  `areGroupDepsSatisfied` is true. Rows with unmet deps render Waiting,
  matching the dispatcher's actual behavior.

Cleanup from /simplify pass:
- Use `generateShortId(20)` instead of
  `generateId().replace(/-/g, '').slice(0, 20)`.
- Inline `batchEnqueueAndWait` no longer allocates synthetic ids
  (returned `string[]` is unused).
- Flattened the per-cell `tracked` array — only push entries that
  registered controllers, drop the null placeholders.
- Extracted `runDispatcherToCompletion` to share the loop between the
  trigger.dev wrapper and the in-process path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(table): backend running counter, dep-aware retrigger, sidebar polish

Counter (Fix 1): top-right "X running" + per-row badge are now
backend-bootstrapped via a count on `user_table_rows.executions ->> 'status'
= 'running'` returned alongside active dispatches. SSE `kind: 'cell'` events
compute a delta from `prev → next` status to keep the cache live; cell
events for rows outside the loaded page slice trigger a run-state refetch.
On `pruned` we invalidate the cache. Counts only worker-claimed `running`
cells — optimistic queued/pending no longer inflate the badge, and rows
outside the loaded page slice are counted too.

Sidebar (Fix 2 + 3a): `Run after` no longer ticks every column by default
for new groups (empty list). Save is disabled with an inline error when
auto-run is on with zero deps. `edit-group` mode anchors the left-of-current
filter to the group's leftmost column, so a workflow can only depend on
columns to its left.

Reorder scrub (Fix 3b): `updateTableMetadata` walks the schema's workflow
groups when `columnOrder` is in the patch and drops any dep whose new
position lands at or after the group's leftmost column (uses the existing
`stripGroupDeps` helper). Metadata + schema updates land atomically.

Server returns ordered columns (Fix 3b cont'd): `getTableById` /
`listTables` now sort `schema.columns` by `metadata.columnOrder` before
returning, via a new `applyColumnOrderToSchema` helper. Every consumer
(grid, sidebar, copilot, mothership) gets one ordered list — the sidebar's
leftmost-group-column anchor now points at the right index.

Dep-aware retrigger (Fix 4): editing a value that a downstream workflow
depends on now re-runs that workflow.
- `deriveExecClearsForDataPatch` returns
  `{ executionsPatch, inFlightDownstreamGroups }`. Walks
  `schema.workflowGroups[].dependencies.columns` for every column in the
  patch, clears terminal-state downstream entries, and reports in-flight
  entries.
- `updateRow` calls `cancelWorkflowGroupRuns` + `runWorkflowColumn`
  (`mode: 'incomplete' + isManualRun: true`) for in-flight downstream
  groups, then always fires `runWorkflowColumn({ mode: 'new' })` for the
  cleared groups. Skips both when `executionsPatch` is provided by the
  caller — those are cell-task / cancel writes that would otherwise spawn
  a recursive flood of dispatches per partial-write.
- `cancelWorkflowGroupRuns(tableId, rowId, { groupIds? })` accepts a
  per-group filter so the cancel only touches the affected groups, not
  every in-flight cell on the row.
- `pickNextEligibleGroupForRow` now treats a dispatcher pre-stamp
  (`pending` + `executionId: null`) as claimable — the cascade-loop is the
  real owner. Without this, the dispatcher's pre-stamp of downstream
  groups made the cascade-loop see them as "in-flight" and skip them,
  stranding `pending` cells forever.
- `optimisticallyScheduleNewlyEligibleGroups` extends the cache patch to
  flip dep-touched groups to `pending` regardless of their current status,
  matching the server's cancel-then-rerun behavior.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): paused workflow cells route through executeResumeJob; render Pending + viewable

Three connected issues with workflows that pause mid-cell (e.g. wait blocks):

1. `/api/resume/poll` (the time-pause auto-resumer) called
   `PauseResumeManager.startResumeExecution` directly, bypassing
   `executeResumeJob` from `background/resume-execution.ts`. The wrapper is
   where the cell-context restoration + cascade-loop continuation lives —
   without it, the resumed workflow ran to completion but never wrote the
   terminal state back to the table cell. Cell stays `pending` forever
   even though the underlying execution finished.

   Fix: dynamically import `executeResumeJob` and use it for the
   `'starting'` branch. Same primitive the trigger.dev `resumeExecutionTask`
   wraps — calling it directly handles both trigger.dev-disabled local dev
   and trigger.dev-enabled prod identically.

2. The cell renderer mapped `status: 'pending'` to `kind: 'queued'` (gray
   "Queued" badge) regardless of whether the run had started. A HITL-paused
   run has `status: 'pending'` + `jobId` prefixed `paused-` + a real
   `executionId` — semantically very different from "queued, hasn't run."
   Now renders as `pending-upstream` (the existing Pending pill) for
   paused-jobId rows.

3. Right-click "View execution" was disabled for `pending` cells (gated to
   `completed | error | running`), so users couldn't open the trace for a
   paused execution. Paused runs have a viewable trace (the executionId is
   real and the log row exists). Both the per-row context menu and the
   action-bar derivation now recognize `pending` + `paused-` jobId as a
   started run.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(table): typewriter reveal for SSE-driven workflow cell values

Workflow-output cells now reveal their text character-by-character when an
SSE update lands, while page reloads and virtualization remounts still paint
the value instantly. A first-render guard inside the new useTypewriter hook
distinguishes hydration from live updates with no plumbing through the cell
tree.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(table): address bugbot/greptile review feedback

Two P1 issues + one cleanup from the bot reviewers:

1. **Double-dispatch + completed-output wipe.** Both PATCH row routes
   (`app/api/table/[tableId]/rows/[rowId]` and
   `app/api/v1/tables/[tableId]/rows/[rowId]`) were firing a second
   `runWorkflowColumn({ mode: 'incomplete' })` after `updateRow` returns.
   `updateRow` already fires `mode: 'new'` internally for user edits, so
   the second call created a concurrent dispatch. Worse, the
   `mode: 'incomplete'` path's `bulkClearWorkflowGroupCells` wipes ALL
   targeted output columns on any row where any one column is empty —
   meaning sibling-group completed outputs could be erased. Removed both
   route-level calls; auto-dispatch lives entirely in `updateRow`.

2. **`runWorkflowColumn` log-spamming on plain tables.**
   `if (targetGroups.length === 0) throw new Error(...)` fired on every
   row insert/update for tables without any workflow groups (the
   majority). Every caller wraps with `.catch(logger.error)`, so each
   PATCH produced an error-level log. Return `{ dispatchId: null }`
   silently — manual `runWorkflowColumn` callers pass `groupIds`
   explicitly so they can't reach this branch.

3. **`isManualRun` plumbed through dispatch SSE events.** Late-arriving
   `kind: 'dispatch'` events for dispatches not in the initial fetch
   were hardcoding `isManualRun: false`. Added the field to the event
   shape, emit it from `dispatcherStep` (pending → complete, dispatching
   transitions) and `markActiveDispatchesCancelled`, and consume it in
   the SSE handler with a sensible fallback for legacy emits.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(table): row executions sidecar + left-to-right dep retrigger + cancel counter refresh

Split per-row workflow-group execution state out of the user_table_rows.executions
JSONB column into a new table_row_executions sidecar keyed by (row_id, group_id).
Dispatcher filters, "X running" counter, bulk clears, and the cancellation guard
all hit indexed columns instead of walking JSONB. Wire shape unchanged — server
merges sidecar rows back into row.executions on the way out.

Also:
- deriveExecClearsForDataPatch now walks workflowGroups left-to-right with a
  propagating dirtied-column set so transitive dep chains (edit col A → group 1
  re-runs → group 2 depends on group 1's output → group 2 re-runs) collapse to
  a single forward pass.
- useCancelTableRuns.onSettled invalidates the activeDispatches query so the
  top-right counter and row gutter Stop button refetch from the server after
  any Stop (per-cell, row, or table-wide). countRunningCells is the source of
  truth; client no longer needs duplicate state.

Three migrations on this branch (0209 + 0210 + new sidecar) collapsed into one
since the feature is unreleased.

* fix(table): address remaining cursor/greptile review feedback

- Mothership update_row no longer double-dispatches. updateRow already fires
  the auto-cascade internally; the second `mode: 'incomplete'` call here
  raced with it and could bulk-clear sibling-group outputs.
- SSE dispatch events no longer dropped when the activeDispatches cache is
  cold. Seed an empty TableRunState if the initial fetch hasn't landed yet
  so the queued overlay doesn't lose the first dispatch event.
- batchUpdateRows now runs cancel+rerun for per-row in-flight downstream
  groups, mirroring updateRow. Without this, dep edits in a batch left
  running workflows reading stale upstream values.

* fix(table): cancel prior runs, scope batch insert dispatch, recover orphan pre-stamps

Addresses cursor + greptile review feedback on table dispatcher edge cases:

- Manual table-wide Run-all / Run-column now cancels prior active dispatches
  AND in-flight cell workers before bulk-clearing. Without this, mode:'all'
  deleted running sidecar rows out from under their workers (which kept
  writing into the wiped state) and a second Run-all could enqueue overlapping
  cells racing on the same rows. Row-scoped manual calls (dep-edit cascade)
  are excluded — those already cancel their own scope.
- batchInsertRowsWithTx now scopes its auto-dispatch to the newly-inserted
  row ids. Without this, after the sidecar migration the NOT EXISTS filter
  matches every existing row (zero sidecar entries), so a CSV import would
  walk the entire table dispatching workflow runs on every pre-existing row.
- classifyEligibility carve-out: pending + executionId=null is an orphan
  pre-stamp (cascade-lock contention, batchEnqueueAndWait failure, etc.),
  treated as claimable so future dispatchers can re-stamp instead of skipping
  it as 'in-flight' forever. Matches pickNextEligibleGroupForRow's logic.
- On batchEnqueueAndWait failure, dispatcherStep now sweeps the orphan
  pre-stamps it wrote for the failed batch so the cells don't render Queued
  forever; the next user action picks them up cleanly.

* fix(table): row-scoped Refresh cancels in-flight; counter includes queued/pending

- runWorkflowColumn now cancels prior in-flight cells for row-scoped manual
  runs too (context-menu Refresh on a row subset, action-bar Refresh on
  selected rows). Previously only the table-wide path cancelled, so a
  row-scoped Refresh would bulk-clear running sidecar rows without aborting
  workers. Per-row cancel skips markActiveDispatchesCancelled so unrelated
  dispatches keep running.
- countRunningCells now counts all in-flight statuses (queued / running /
  pending) instead of just running. The row gutter Run/Stop button reads
  this map — with the old behavior, clicking Play during the queued window
  would re-enqueue an already-queued cell. SSE applyCell handler updated
  to use isExecInFlight so client deltas track the same semantics.

* fix(table): per-row Stop tombstones ahead-of-cursor rows during Run-all

Per-row Stop only cancelled sidecar rows already in flight. A row the
dispatcher hadn't reached yet had no exec record, so Stop was a no-op there
— the dispatcher would later walk to it, classify the group eligible, and
re-fire workflows the user thought they stopped.

cancelWorkflowGroupRuns now, for a per-row cancel, checks active dispatches
whose scope covers the row and writes `cancelled` tombstones (cancelledAt =
now) for the at-risk groups that don't already have a sidecar entry. The
dispatcher's existing `cancelledAt > dispatch.requestedAt` filter then skips
them when the cursor arrives. onConflictDoNothing guards against clobbering
a concurrently-written entry; the active-dispatch check avoids stamping
spurious cancels on idle rows.

* fix(table): seed dispatch overlay on Run; surface batch-enqueue failures as error

- useRunColumn.onSuccess invalidates the activeDispatches query so the
  resolveCellExec queued overlay populates immediately for ahead-of-cursor
  rows (scrolled-in / refetched), instead of waiting for the first dispatch
  SSE. Targeted at activeDispatches only — the rows cache stays owned by
  useTableEventStream.
- On batchEnqueueAndWait failure, dispatcherStep now flips the orphan
  pre-stamps to a terminal `error` state and emits a cell SSE event, rather
  than deleting them. The cursor still advances past the window, but the
  dropped cells are now visible (Error pill) instead of silently empty, stay
  out of the in-flight set, and re-run on the next manual run.

* fix(table): seed dispatch overlay on Run; surface batch-enqueue failures as error

- useRunColumn.onSuccess invalidates activeDispatches so the resolveCellExec
  queued overlay populates immediately for ahead-of-cursor rows instead of
  waiting for the first dispatch SSE. Rows cache stays owned by SSE.
- On batchEnqueueAndWait failure, dispatcherStep flips orphan pre-stamps to a
  terminal error state (+ cell SSE) instead of deleting them, so the dropped
  window is visible (Error pill) rather than silently empty and re-runs on the
  next manual run.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 05:44:09 -04:00
Waleed d55f40ce14 fix(blocks): preserve agent block color (#4671) 2026-05-19 18:05:38 -07:00
Vikhyath Mondreti 974a18dd61 improvement(media-blocks): new versions of image and video gen with latest models + fixes (#4667)
* improvement(media-blocks): new versions of image and video gen with latest models + fixes

* respect versioning for icons

* fix integration routes

* address comments

* address api mismatches

* more ltx 2.3 durations

* typing tightness
2026-05-19 16:42:04 -07:00
WaleedandClaude Opus 4.7 e40c91588f fix(workflow-search): unclip block-name highlight shadow on the left (#4670)
The editor title h2 uses truncate (overflow:hidden) so the search highlight's -3px box-shadow was getting clipped on the left, exposing the mark's sharp edge. Replaces overflow:hidden with overflow:clip plus a 3px overflow-clip-margin so the shadow can bleed past the clip boundary without shifting the title text or breaking truncation.

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-19 16:32:05 -07:00
Waleed 972ec5f216 fix(branding): align auth and deploy UI colors (#4669) 2026-05-19 16:18:07 -07:00
WaleedandClaude Opus 4.7 d7d58cec36 improvement(workflow-search): include block names in in-workflow search (#4668)
* improvement(workflow-search): include block names in in-workflow search

Adds block names to the workflow search index alongside subblock content. Selecting a block-name match navigates to the block and highlights its title in the editor header with the same orange treatment used for subblock labels.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* refactor(panel-store): derive ActiveSearchTargetKind from WorkflowSearchTarget

Replaces the hand-maintained literal union with a derived type so adding a new target kind in search-replace/types.ts automatically propagates here.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-19 15:44:34 -07:00
Waleed 6c755cbb59 feat(integrations): add Gong incident.io Railway and New Relic (#4663)
* feat(integrations): add Gong incident.io Railway and New Relic

* fix(railway): preserve explicit empty variable values

* fix(incidentio): fail on invalid workflow JSON

* fix(new-relic): validate custom attributes JSON

* fix(integrations): address incident workflow review fixes

* chore(docs): apply lint formatting

* chore: refresh integration docs and validation fixes

* more

* fix(integrations): address PR review comments

* fix(gong): align list calls block validation
2026-05-19 14:12:28 -07:00
df1e2ddc2e feat(azure-devops): block and trigger (#4664)
* azure devops logo on white

* generated ADO tool docs

* generated ADO tool docs

* added ADO to registries

* ADO workflow triggers

* ADO workflow triggers

* tool layer for ADO, checks passed and manual verified

* ADO workflow triggers

* block layer for ADO

* ADO icon svg

* generated docs for ADO triggers

* committing the tests for azure devops tools and blocks

* Update apps/sim/triggers/azure_devops/utils.ts

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* Update apps/sim/tools/azure_devops/update_work_item.ts

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* Update apps/sim/triggers/azure_devops/utils.ts

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>

* comma syntax error patched

* azure devops: validate-integration fixes + manual description

- bgColor switched from white to Azure DevOps brand color #0078D4 (block + mdx)
- WIQL query_work_items: hydrate ALL matched IDs by chunking through batches
  of 200 instead of silently truncating; check response.ok on the follow-up
  fetch and surface a clear error on 4xx/5xx; trim org/project; expose
  totalMatched in metadata so users can see pre-hydration count
- Add MANUAL-CONTENT-START:intro section to the azure_devops.mdx docs page
- Update unit tests for new chunking behavior and update-work-item validation

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* azure_devops: second-pass audit fixes + formatter cleanup

- Add types barrel export to tools/azure_devops/index.ts
- Normalize comment endpoint path casing (/workItems/ -> /workitems/)
- Update test assertions to match normalized path
- Biome formatter reflow across tools, triggers, registry, and docs icon

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* azure_devops: address PR review comments

- Fix bgColor #FFFFFF -> #0078D4 in integrations.json and triggers/azure_devops.mdx
- Bump File tool operationCount from 4 to 5 (Read, Fetch, Get, Write, Append)
- Apply .trim() to org/project across all 15 remaining tools (consistency with query_work_items)
- Fix Found ${data.count} -> Found ${data.count ?? items.length} fallback in list_builds, list_pipelines, list_pipeline_runs content strings

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* idemtpotency

* azure_devops: address bugbot review comments

- triggers/utils: match build.complete result case-insensitively, accept stopped/cancelled in addition to failed/canceled/partiallySucceeded so PascalCase and legacy Azure DevOps payloads aren't dropped
- get_work_items_batch: chunk comma-separated IDs into 200-batch loops with proper status checks (was failing or returning incomplete data on >200 IDs)
- Add tests for both behaviors

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* azure_devops: address additional bugbot comments

- Block update_work_item now forwards areaPath; the Area Path subblock condition expanded to include update operation
- get_build_timeline.failedRecords now also flags partiallySucceeded and succeededWithIssues, normalized case-insensitively. Output description and added a focused test

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* azure_devops: address more bugbot comments

- Webhook provider extractIdempotencyId returns null when subscriptionId or notificationId is missing/empty, preventing the literal "azure_devops:undefined:undefined" key from collapsing unrelated deliveries into duplicates
- Get Work Items Batch validates that at least one non-empty ID is supplied before issuing the API request, throwing a clear error instead of hitting an empty ids= query
- Tests cover both behaviors

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* azure_devops: pin add_comment to documented api-version 7.0-preview.3

Microsoft's Add Comments docs only publish 7.0-preview.3 (the 7.2 view falls back to the 7.0 page). Get Comments stays on the documented 7.2-preview.4. Matches what's strictly in the Azure DevOps REST API reference rather than relying on undocumented version behavior.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Marcus Chandra <mzxchandra@gmail.com>
Co-authored-by: mzxchandra <129460234+mzxchandra@users.noreply.github.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-19 14:10:42 -07:00
WaleedandClaude Opus 4.7 49f70bc635 fix(docker): restore NEXT_PUBLIC_APP_URL build arg with dummy fallback (#4665)
* fix(docker): restore NEXT_PUBLIC_APP_URL build arg with dummy fallback

getBaseUrl() in lib/core/utils/urls is evaluated at module load during
next build's page-data collection and throws if NEXT_PUBLIC_APP_URL is
unset. PR #4658 removed the build arg, breaking the Docker build at the
"/_not-found" page-data collection step.

Restore the dummy localhost fallback (mirroring DATABASE_URL). The CORS
fix from #4658 is preserved: next.config.ts no longer reads
NEXT_PUBLIC_APP_URL at build time, and no module-level expression
captures getBaseUrl() — every caller invokes it at request time, where
getEnv() reads the deployed container env. The dummy localhost value
cannot leak into runtime CORS response headers.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* chore(docker): trim verbose comment on build-time env args

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-19 13:34:49 -07:00
Vikhyath Mondreti 6414b9381d improvement(cleanup): cleanup refs along in logs cleanup job (#4661)
* improvement(cleanup): cleanup refs along in logs cleanup job

* address comments

* cleanup code
2026-05-19 12:19:54 -07:00