Commit Graph
4596 Commits
Author SHA1 Message Date
Siddharth Ganesan bad21cb23d improvement(agent, file-block): files in agent block, file block v4 (#4610)
* File block v4

* Support files in agent block

* Clean up attachments

* Fix start.files

* Fix

* Fix
2026-05-15 10:48:21 -07:00
Waleed 9bbbe0adcb chore(deps): bump mermaid to 11.15.0 for GHSA-ghcm-xqfw-q4vr (#4615)
* chore(deps): bump mermaid to 11.15.0 for GHSA-ghcm-xqfw-q4vr

* chore(deps): override transitive mermaid to 11.15.0
2026-05-15 10:42:42 -07:00
Waleed 62636c7eb3 improvement(gmail): replace custom html-to-text regex with library (#4613)
* improvement(gmail): replace custom html-to-text regex with html-to-text library

Resolves 4 CodeQL alerts on htmlToPlainText (incomplete tag/entity handling,
unsafe regex backtracking). Delegates to the html-to-text npm package already
used by the outlook polling trigger and the mail/send route.

* improvement(gmail): match outlook selectors config, add nbsp/anchor tests

Aligns html-to-text options with apps/sim/lib/webhooks/polling/outlook.ts:
suppress anchor hrefs when identical to text, drop bare # anchors, skip
img/script/style content. Adds tests for nbsp preservation and anchor
behavior.
2026-05-14 19:49:30 -07:00
Waleed e23c20e1eb fix(gmail): send emails as multipart/alternative so they render full-width (#4611)
* fix(gmail): send emails as multipart/alternative so they render full-width

* fix(gmail): decode & last in htmlToPlainText to avoid double-decoding compound entities

* fix(gmail): encode body parts as base64 and decode numeric HTML entities
2026-05-14 19:26:14 -07:00
Waleed 922de38d93 fix(date-picker): eliminate infinite re-render crash on re-open with existing selection (#4609)
* docs(uploads): clarify QUOTA_EXEMPT_STORAGE_CONTEXTS logs entry in JSDoc

* fix(date-picker): eliminate infinite re-render on re-open with existing selection

The useEffect that syncs picker state on open had initialStart and
initialEnd — Date objects computed on every render — in its dependency
array. Because Object.is returns false for any two distinct Date
instances, the effect fired on every render when open=true, calling
setRangeStart/setRangeEnd and triggering another render, producing an
infinite loop that crashed the page.

Fix: compute start and end as local variables inside the effect and
use the stable string props (props.startDate, props.endDate) as deps
instead.

Also removes the redundant typeof fileSize === 'number' guard in the
multipart quota check — fileSize is z.number() (required) in the
contract so it can never be undefined at that point.

* refactor(date-picker): comprehensive cleanup and reliable crash fix

The previous fix still had derived Date objects in useEffect deps.
Object.is(new Date(), new Date()) === false, so any Date in deps causes
the effect to run every render, reproducing the infinite loop on
re-open with existing time selection.

Key changes:
- useEffect deps now use only stable primitives (startDate, endDate strings)
  and compute Date values inside the effect — eliminating the loop
- Replace `rest as any` with a FlatDatePickerProps merged type for safe,
  typed destructuring across the discriminated union
- Remove initialStart/initialEnd render-scope variables; compute inline or
  inside effects to keep derivation local to each use site
- Callbacks use destructured props (onChange, onRangeChange, etc.) instead
  of props.x references
- Remove verbose TSDoc on internal callbacks — names are self-documenting
- Preserve all existing JSX structure and CalendarMonth logic unchanged
2026-05-14 18:38:15 -07:00
Waleed d0519c1503 fix(security): supabase rpc path validation, ssh stream byte cap, storage quota coverage (#4605)
* fix(security): supabase rpc path validation, ssh stream byte cap, storage quota coverage

* fix(security): scope execution log writes to owning workflow; add env-var workspace membership guard

Closes two cross-tenant vulnerabilities:

1. Workflow log cross-tenant write (route.ts + logging-session.ts):
   - Route: SELECT before creating LoggingSession to verify executionId belongs
     to the claimed workflowId; reject with 404 if owned by a different workflow.
   - LoggingSession: add workflow_id to all UPDATE/SELECT WHERE clauses
     (raw SQL marker queries, flushAccumulatedCost, loadExistingCost) so
     writes are a no-op if executionId was somehow injected.

2. Env-var workspace membership guard (environment/utils.ts):
   - getPersonalAndWorkspaceEnv now calls checkWorkspaceAccess when workspaceId
     is provided; throws if the userId is not a member, preventing any future
     caller from reading another workspace's decrypted secrets without
     explicit membership verification at the call site.

* fix(security): remove fileSize > 0 quota bypass gate; exempt logs context from quota

* chore: remove extraneous inline comments

* fix(security): scope markExecutionAsFailed UPDATE by workflowId; thread workflowId through HITL callers

* fix(security): add personal credential ownership check in sharepoint site route; scope markExecutionAsFailed by workflowId

* fix: remove logs from user-accessible upload contexts; restore distinct biome .next glob

* fix(sharepoint): migrate site route to authorizeCredentialUse

The previous fix only checked userId equality for personal credentials and
workspace membership (via getUserEntityPermissions) for workspace credentials.
authorizeCredentialUse additionally enforces credentialMember access for
workspace-scoped credentials, matching the standard pattern used by all
other tool selector routes.

* fix(logging): make workflowId required in markExecutionAsFailed

Making workflowId optional left a footgun — future callers could silently
omit it and the WHERE clause would degrade to executionId-only, losing the
cross-tenant scoping guarantee. All callers already supply workflowId, so
making it required (with string | undefined for the middle params to keep
call sites unchanged) closes the gap without touching any caller.

* test(security): add tests for cross-tenant log guard, quota bypass fix, and workflowId scoping

- log/route.test.ts: verifies cross-tenant executionId guard returns 404
  when the execution belongs to a different workflow, and passes for same
  workflow or fresh executions
- multipart/route.test.ts: verifies fileSize:0 no longer bypasses quota
  check and that the logs context is rejected at the endpoint level
- logging-session.test.ts: verifies markExecutionAsFailed scopes by both
  executionId and workflowId, and that the instance method forwards workflowId

* fix(lint): move IconComponent outside ToolInput to fix noNestedComponentDefinitions

* fix(logging): scope completeWithCancellation and completeWithPause reads by workflowId

Both SELECT queries that check execution status before writing a
terminal result were only filtering on executionId. Adds workflowId
to the WHERE clause so all seven reads and writes in LoggingSession
consistently scope by (workflowId, executionId).
2026-05-14 17:08:09 -07:00
Waleed 11fa96cac0 chore(deps): bump next to 16.2.5 for CVE-2026-44578 SSRF fix (#4606)
* chore(deps): bump next to 16.2.5 for CVE-2026-44578 SSRF fix

* chore(deps): bump next to 16.2.6 for full May 2026 security release coverage
2026-05-14 14:34:05 -07:00
Waleed 80c9a01275 fix(security): harden file access controls, webhook auth, and input bounds (#4601)
* fix(security): harden file access controls, webhook auth, and input bounds

* fix(security): extend file access checks to remaining tool routes

* fix(logs): address PR review comments on time filter

* fix(logs): set end-time milliseconds to 999 for datetime filter strings

* fix(files): return 404 instead of 500 on file access denial in utility paths

* remove tooltip from resource tabs
2026-05-14 13:32:40 -07:00
Theodore Li 4a9e248eac feat(cloudwatch): add mute and unmute alarm operations (#4602) 2026-05-14 14:54:35 -04:00
Waleed 044e034719 fix(integrations): gdrive trashed search, slack blocks-with-file, slack get_message ts (#4600)
* fix(integrations): gdrive trashed search, slack blocks-with-file, slack get_message ts

- Google Drive search/list: skip default `trashed = false` when user query
  already specifies a `trashed = ...` predicate, so trashed-file searches work.
- Slack send-message with files: forward `blocks` through to
  `files.completeUploadExternal` so Block Kit renders when files are attached.
- Slack get_message: switch from `conversations.history` (oldest lower-bound
  returned the next message after) to `conversations.replies` with `ts=`
  for exact-match lookup, plus a defensive ts-equality guard and clearer error.

* fix(google_drive): revert list.ts trashed guard — query is plain text, not gdrive syntax

* fix(slack): omit initial_comment when blocks present so Block Kit actually renders on file uploads
2026-05-14 11:33:16 -07:00
Theodore Li 642231f875 improvement(scheduler): drain due schedules in chunks (#4578)
* improvement(scheduler): drain in chunks instead of a single capped claim

Replaces the fixed MAX_CRON_CLAIMS (200) with a chunked drain loop:
claim WORKFLOW_CHUNK_SIZE + JOB_CHUNK_SIZE per iteration, process via
Promise.allSettled, repeat until both claim queries return empty or
MAX_TICK_DURATION_MS elapses. Throughput is no longer bounded by a
static per-tick ceiling; it scales until DB or trigger.dev is the
limit. Per-iteration chunk size still bounds row-lock set and fan-out
concurrency.

Extracts processScheduleItem and processJobItem so the loop body stays
readable. Existing claim semantics (FOR UPDATE SKIP LOCKED, lastQueuedAt
as the claim signal, staleness reclaim) are unchanged.

* improvement(scheduler): skip claim once a queue is exhausted and drop workflowUtils non-null assertion

Addresses Greptile review on PR #4578:
- track per-queue exhaustion when a claim returns fewer than CHUNK_SIZE
  rows; subsequent iterations skip the claim query for that queue. Saves
  one DB round-trip per iteration once one queue drains while the other
  is still working.
- narrow workflowUtils to a local const inside the loop body so the
  schedule processing branch only runs when the import has completed.
  Removes the misleading non-null assertion.
2026-05-14 14:21:41 -04:00
Theodore Li b1a87d531c Revert "improvement(db): add session statement/lock timeouts; simplify KB doc tx (#4593)" (#4599)
This reverts commit 4295a5c855.
2026-05-14 13:42:15 -04:00
Waleed 8831defd2e fix(seo): use canonical SITE_URL for robots and sitemap (#4598)
* fix(seo): use canonical SITE_URL for robots and sitemap

* fix(seo): drop /templates from sitemap and guard robots/sitemap in seo test
2026-05-14 10:25:17 -07:00
4295a5c855 improvement(db): add session statement/lock timeouts; simplify KB doc tx (#4593)
* v0.6.29: login improvements, posthog telemetry (#4026)

* feat(posthog): Add tracking on mothership abort (#4023)

Co-authored-by: Theodore Li <theo@sim.ai>

* fix(login): fix captcha headers for manual login  (#4025)

* fix(signup): fix turnstile key loading

* fix(login): fix captcha header passing

* Catch user already exists, remove login form captcha

* improvement(db): add session statement/lock timeouts; simplify KB doc tx

* fix(knowledge): close soft-delete TOCTOU on KB document insert

Fix the race the bots flagged: KB delete is soft (`deletedAt = now`) so
the FK can't catch a concurrent KB delete between the existence check
and the document insert.

- Add `insertDocumentsIfKbAlive` helper that gates the insert on
  `EXISTS(SELECT 1 FROM knowledge_base WHERE id=$kb AND deleted_at IS NULL)`
  in the same statement via INSERT...SELECT...WHERE EXISTS. Atomic at the
  MVCC snapshot — no transaction, no row lock.
- Use jsonb_to_recordset to declare column types once, avoiding per-param
  casts for nullable columns.
- Wire into both `createDocumentRecords` (bulk) and `createSingleDocument`.
- Keep the upfront KB existence check as a fast-path early-out for the
  common case; the atomic insert is the race guard.

---------

Co-authored-by: Waleed <walif6@gmail.com>
Co-authored-by: Siddharth Ganesan <33737564+Sg312@users.noreply.github.com>
Co-authored-by: Vikhyath Mondreti <vikhyathvikku@gmail.com>
2026-05-14 12:54:05 -04:00
Waleed c3ac54e0a9 fix(vfs): make copilot message ordering deterministic via WITH ORDINALITY (#4597) 2026-05-14 00:03:30 -07:00
Vikhyath Mondreti b1a9443178 improvement(billing): move overage calculations out of txes (#4595)
* improvement(billing): move calc subscription overage out of tx

* fix double billing risk

* address comments

* address comments

* share timeout const
2026-05-13 23:52:32 -07:00
Waleed b5dba82ac9 improvement(db): reduce connection saturation and egress hotspots (#4594)
* improvement(db): reduce connection saturation and egress hotspots

* fix(vfs): preserve native content type in copilot SQL projection

* fix(vfs): guard jsonb_array_elements against non-array contentBlocks
2026-05-13 23:39:59 -07:00
Waleed 104949bdc2 fix(tables): eliminate checkbox flicker on rapid cell toggle (#4592)
* fix(tables): eliminate checkbox flicker on rapid cell toggle

* fix(tables): symmetric guarded onSettled across row write mutations

* fix(tables): merge only mutated keys in onSuccess to preserve concurrent optimistic patches
2026-05-13 19:52:43 -07:00
Waleed 568a552d67 fix(rate-limit): close rate-limit bypass and tighten public route limits (#4591)
* fix(rate-limit): close rate-limit bypass and tighten public route limits

* fix(rate-limit): address PR review — drop success field from 429 body, fall back to per-IP when JWT auth lacks userId
2026-05-13 19:16:57 -07:00
Waleed 1c111ff2d7 fix(mothership): persist @-mentioned resources across send (#4587)
* fix(mothership): persist @-mentioned resources across send and merge on hydration

* fix(mship-resources): handle ADD/DELETE race and reorder during pending flush

- Track in-flight ADD promises so DELETE chains off finally(), preventing orphaned server rows when a user removes a resource before its POST resolves
- Defer reorder PATCH until pending flush completes; emit with full local order
- Clear new refs in reset paths

* fix(mship-resources): defer reorder when ADDs are in-flight on existing chat

reorderResources previously only checked pendingPersistResourceKeysRef. When a
chatId exists, addResource fires the POST immediately and only tracks the
promise in inFlightResourceAddsRef — so a reorder before those ADDs settle
shipped a PATCH the server rejected, and the silent catch lost the reorder.

Now treat in-flight ADDs like pending ones: defer the PATCH and replay it
after Promise.allSettled on the in-flight map.
2026-05-13 18:47:21 -07:00
Siddharth Ganesan ff3c8f765e fix(file-block): fix get op (#4590)
* File block get

* Lint

* Fix

* Fix auth
2026-05-13 18:39:29 -07:00
Siddharth Ganesan 214355b26e improvement(file-block): add get operation (#4588)
* File block get

* Lint

* Fix
2026-05-13 18:34:18 -07:00
Theodore Li 4de955d0b6 fix(otel): address staging pr comments for trigger otel (#4586) 2026-05-13 20:01:09 -04:00
Theodore Li ed39edb906 feat(observability): export Trigger.dev telemetry to Grafana Cloud OTLP (#4583)
* feat(observability): export Trigger.dev telemetry to Grafana Cloud OTLP

Wire OTLP HTTP exporters for traces, logs, and metrics from the
Trigger.dev runtime to Grafana Cloud. Auth uses Basic with instance ID
and API token. Gated behind GRAFANA_OTLP_ENDPOINT, GRAFANA_INSTANCE_ID,
and GRAFANA_API_TOKEN — all three must be set together or all unset;
partial config throws at startup.

* improvement(observability): use OTLP HTTP/JSON for metrics for consistency with traces and logs

* feat(observability): tag Trigger.dev telemetry with deployment.environment.name

* improvement(observability): switch Grafana telemetry vars to OTLP-shaped trio
2026-05-13 19:38:20 -04:00
Siddharth Ganesan 2441d5ad6e feat(mothership): add files to mship block (#4584)
* Add files to mship block

* Fixes

* Fix

* fix
2026-05-13 16:35:56 -07:00
Waleed 1ed3a4ec6d feat(mothership): pin tasks to keep them at the top of the sidebar (#4582)
* feat(mothership): pin tasks to keep them at the top of the sidebar

* fix(sidebar): address PR review feedback for pin tasks

* fix(posthog): register task_pinned and task_unpinned events

* fix(tasks): insert new optimistic tasks below pinned partition
2026-05-13 14:46:04 -07:00
Siddharth Ganesan cdc7513d23 improvement(mothership): allow mship to send function execute timeout (#4581)
* Improve mship fexecute

* Fix
2026-05-13 14:37:36 -07:00
Vikhyath Mondreti bdf9ffc798 fix(event-buffer): re-compact the event with preserveUserFileBase64: false (#4579)
* preserveUserFileBase64 is on and the event still exceeds the threshold, re-compact the event with preserveUserFileBase64: false

* address comments
2026-05-12 21:16:22 -07:00
Siddharth Ganesan 689e1f76d1 feat(mothership): Add conversationId to mship block (#4577) 2026-05-12 20:14:35 -07:00
WaleedandClaude Opus 4.7 c21bb915f7 improvement(grafana): align tools and block with Grafana API spec (#4574)
* improvement(grafana): align tools and block with official Grafana API spec

Validates and corrects the Grafana integration against the official API
docs: fixes wire-format field naming for provisioned alert rules
(missing_series_evals_to_resolve, keepFiringFor, orgID), adds
X-Disable-Provenance support, expands alert-rule params (isPaused,
notificationSettings, record, annotations, labels), corrects defaults
(execErrState=Error, dashboard overwrite=false), and centralizes alert-rule
output mapping in a shared utils module.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(grafana): correct wire-format casing for provisioned alert rule fields

Grafana's ProvisionedAlertRule schema (verified against upstream Go source
and swagger spec) uses keep_firing_for (snake_case) and
missingSeriesEvalsToResolve (camelCase) — the opposite of what prior audit
rounds assumed. POST/PUT bodies now send the correct field names; mapAlertRule
reads the correct primary names with the old casings kept as fallbacks.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(grafana): address PR review feedback

- Drop hardcoded orgID: 1 fallback; only send orgID when organizationId is
  provided, so token-scoped org context drives rule placement.
- Surface invalid JSON for notificationSettings/record on alert rule
  create/update instead of silently dropping the input.
- Fix execErrState description in update_alert_rule to include Error.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(grafana): surface invalid JSON for annotations/labels/data on alert rules

Match the behavior of other JSON params (data, notificationSettings, record):
return a descriptive error instead of silently falling back to {} (create)
or keeping the existing value (update).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* docs(grafana): expose alert-rule output fields in generated docs

Move ALERT_RULE_OUTPUT_FIELDS from utils.ts to types.ts and rename to
SCREAMING_SNAKE_CASE so scripts/generate-docs.ts (which only resolves const
references from types.ts matching [A-Z][A-Z_0-9]+) can inline the per-field
rows into the generated alert-rule output tables.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-12 19:42:20 -07:00
WaleedandVikhyath Mondreti 773cd84e3f fix(mothership): reconcile stuck conversation_id against Redis lock to clear stuck-yellow task tiles (#4556)
* fix(mothership): reconcile stuck conversation_id against Redis lock to clear stuck-yellow task tiles

copilot_chats.conversation_id has no TTL/heartbeat, so when a stream
process dies before the clear path runs (pod OOM, SIGKILL, uncaught
throw, deploy mid-stream) the column is orphaned and the task tile
renders yellow forever. The Redis lock at copilot:chat-stream-lock:<chatId>
is the canonical liveness signal and self-heals via 60s TTL + 20s
heartbeat, but the mothership APIs weren't consulting it.

Adds read-time reconciliation: a batched MGET helper checks whether
each persisted conversation_id still has a live Redis lock, and both
GET /api/mothership/chats and GET /api/mothership/chats/[chatId]
rewrite the marker to null when the lock has expired. No DB writes;
stuck rows self-heal on next fetch.

* test(mothership): clarify test name to reflect that getActiveChatStreamIds is called with empty candidateIds

* address comments

* fix state machine issue

* cleanup code and fix types

---------

Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai>
2026-05-12 18:54:49 -07:00
Waleed d5c2ead5d4 fix(console): match child-workflow inner blocks by instanceId when reconciling dropped SSE events (#4575)
* fix(console): match child-workflow inner blocks by instanceId when reconciling dropped SSE events

* fix(console): drop noisy warn when reconcile finds no matching entry
2026-05-12 17:41:50 -07:00
Waleed 6503671102 fix(security): harden findings — path traversal, SSRF, IDOR, file auth, credential access (#4571)
* fix(security): harden HIGH deepsec findings across multiple attack surfaces

- Supabase tools (get_row, delete, update): validate table name with strict
  identifier regex and encodeURIComponent to prevent LLM-controlled path
  traversal to admin endpoints; add missing empty-filter guard to update
  matching the delete.ts pattern

- SFTP/SMTP/SharePoint upload routes: add verifyFileAccess ownership check
  before downloadFileFromStorage, matching the WordPress reference pattern;
  rejects files the requesting user does not own with 404

- Gmail labels, OneDrive folders, Wealthbox items (×2): replace bare
  resolveOAuthAccountId + workspace-only membership check with
  authorizeCredentialUse which enforces credentialMember table; use
  credentialOwnerUserId for token refresh instead of bare accountRow.userId

- A2A utils: thread pre-resolved IP from validateUrlWithDNS into A2A SDK
  via pinnedFetch (secureFetchWithPinnedIP) for JsonRpcTransportFactory,
  RestTransportFactory, and DefaultAgentCardResolver, closing the TOCTOU
  DNS rebinding window

- SSH utils: cap stdout/stderr accumulation at 16 MB with truncation marker
  to prevent OOM from unbounded command output

- Form DELETE route: replace db.delete() with db.update({archivedAt}) for
  true soft delete matching the schema's archivedAt column

- Workflow admin import: fix Array.isArray() guard that silently dropped
  all variables (export format is Record, not Array)

- Multipart upload: apply checkStorageQuota and MAX_WORKSPACE_FILE_SIZE to
  mothership context, closing the quota bypass for workspace-scoped storage

* fix(security): eliminate workspace env lost-update race with atomic JSONB ops

PUT: use `variables || excluded.variables` in onConflictDoUpdate so
concurrent writes merge atomically in the DB instead of last-writer-wins
at the application layer.

DELETE: replace the read-modify-write upsert with a single UPDATE that
removes keys via the JSONB `-` operator, preventing concurrent deletes
from resurrecting previously-removed secrets.

* fix(security): address audit findings from security fix review

- SMTP send: restructure attachment loop from Promise.all to sequential
  for...of so verifyFileAccess denial returns 404 instead of propagating
  as a generic 500 via the SMTP error classifier

- Supabase tools: extend table-name validation and encodeURIComponent to
  the five previously missed tools — insert, upsert, count, query,
  text_search — completing coverage across all nine Supabase tools

- Credential routes: remove unnecessary `request as any` casts in Gmail,
  OneDrive, and Wealthbox routes; authorizeCredentialUse already accepts
  NextRequest directly

- Form soft delete: also set isActive=false alongside archivedAt so that
  any future code paths querying by isActive see a consistent state

- SSH utils: fix exit code fallback from 0 to -1 so an abnormally closed
  connection that supplies no exit code is not reported as success

- Workspace env: capitalize EXCLUDED.variables in the onConflictDoUpdate
  set clause to make the pseudo-table reference unambiguous

* fix(security): address PR review comments and harden deepsec fixes

- fix(env): replace jsonb operators with transaction+FOR UPDATE read-modify-write
  - PUT: uses db.transaction + SELECT FOR UPDATE + JS merge to avoid lost-update race
  - DELETE: same pattern; fixes variable scope bug where current was referenced outside tx
  - removes broken || and - jsonb operators that fail on json-typed column

- fix(ssh): trim truncated output consistently with non-truncated path

- fix(gmail): remove redundant resolveOAuthAccountId call
  - adds credentialType field to CredentialAccessResult
  - authorizeCredentialUse now returns credentialType in all success paths
  - gmail/labels route uses authz.credentialType and authz.resolvedCredentialId directly

- fix(supabase): centralize table identifier validation
  - adds validateDatabaseIdentifier() to input-validation.ts
  - all 8 supabase tools use the shared util instead of inline regex

* fix(workflows): fix VariableType assignment in admin workflow import route

The intermediate Record cast used 'string' for the type field which TypeScript
correctly rejected — WorkflowVariable.type is 'VariableType', not string.
Changed the cast to use VariableType so both branches typecheck correctly.

* fix(a2a): handle Request objects in pinnedFetch URL extraction

* fix(security): extract shared file-access guard; merge workspace/mothership branch

* fix(security): advisory lock for env first-insert race; handle all BodyInit types in pinnedFetch

* chore: remove inline comment from advisory lock

* fix(security): remove stray comment; narrow credentialType to literal union

* fix(security): add credentialId validation to wealthbox oauth route; fix null body override in pinnedFetch

* fix(security): stream A2A response body to unblock SSE; keep text/json/arrayBuffer for non-streaming callers

* fix(security): resolve credentialId guard on OneDrive, use assertToolFileAccess in WordPress, memoize body buffer to prevent silent empty reads, fix ArrayBuffer type cast

* fix(security): handle string[][] HeadersInit format in pinnedFetch

* fix(security): keep abort listener alive during body streaming; clean up in stream end/error/cancel

* chore: remove extraneous inline comment

* fix(security): cleanup abort listener when maxResponseBytes limit is exceeded
2026-05-12 17:40:25 -07:00
Waleed ec936be5ca improvement(workflow-block): support manual workflow ID via advanced mode (#4573)
* improvement(workflow-block): support manual workflow ID via advanced mode

* fix(input-mapping): resolve workflowId via canonical hook for advanced mode

* fix(input-mapping): fall back to manualWorkflowId in preview context

* refactor(input-mapping): resolve workflowId via useDependsOnGate canonical pattern
2026-05-12 16:39:28 -07:00
Vikhyath Mondreti 43f53bb7d3 feat(execution): payload size bottlenecks with lazy execution value hydration, safer materialization, and batched parallel execution (#4560)
* improvement(resolver): lazy resolution for underlying fields greater than 10MB

* progress

* feat(parallel): batching

* codegen to allow inline substitution

* address comments

* ui inconsistencies

* cleanup redundant code

* address more comments

* address comments

* replace helper

* fix tests
2026-05-12 14:57:32 -07:00
Waleed f6b246ba44 fix(docs): restore media centering and full-width intro image (#4570)
* fix(docs): restore media centering and full-width intro image

* fix(docs): drop overflow-hidden from intro media wrappers so focus ring is not clipped

* fix(docs): use inset focus ring on lightbox media so parent overflow-hidden cannot clip it

* fix(docs): drop focus ring on lightbox media to match original UI
2026-05-12 14:56:19 -07:00
Waleed d1eb79ecd3 fix(helm): preserve STS serviceName + networkPolicy.egress back-compat (#4569)
* fix(helm): preserve STS serviceName + networkPolicy.egress back-compat

Greptile flagged two real upgrade-breaking changes vs the prior chart:

1. statefulset-postgresql spec.serviceName flipped from <name>-postgresql
   to <name>-postgresql-headless. spec.serviceName is immutable, so any
   existing install would hit 'Forbidden: updates to statefulset spec ...'
   on helm upgrade. Revert to the original name (the headless Service in
   services.yaml is added alongside, not as a swap).

2. networkPolicy.egress changed from a list to a map ({extraRules, exceptCidrs}),
   silently dropping any custom egress list set by existing users. Restore
   the original list semantics for networkPolicy.egress and move cloud-metadata
   blocking to a sibling top-level field networkPolicy.egressExceptCidrs.

Adds NOTES.txt upgrade-notes entry covering both + the ESO v1→v1beta1 default
flip (functionally a no-op, but worth surfacing).

* docs(helm): update README egress reference to new key name

* fix(helm): revert copilot-postgresql STS serviceName too (same immutability issue)

Audit caught that the main fix in d5c2e8ef5 missed statefulset-copilot-postgres.yaml,
which had the identical immutable-field rename from -copilot-postgresql to
-copilot-postgresql-headless. Same upgrade-break vector for anyone running
copilot.enabled=true on a prior chart version. Mirrors the fix and comment
from the main postgresql STS.

* improvement(helm): postgres startupProbe + otel-collector NetworkPolicy

- add startupProbe defaults for both postgresql + copilot-postgresql STSs
  to shield liveness from slow first-boot (pgvector init, WAL replay)
- render a dedicated NetworkPolicy for the otel-collector when
  telemetry.enabled=true (OTLP ingress from app/realtime/copilot, DNS +
  HTTPS egress for forwarding to external observability backends)
- document why copilot + copilot-postgresql intentionally do NOT ship
  dedicated NetworkPolicies (Redis URL is unknowable at render time)
- regression test pins the otel-collector NP at documentIndex 3

* test(helm): assert custom egress applied to realtime NP too

The prior test claimed coverage of both app and realtime NPs but only
asserted documentIndex 0. Split into two tests so a regression that drops
custom egress from realtime would fail loudly.

* docs(helm-skill): trim narrative bloat in values-model

Cut the historical 'Layer 2 was added in chart 1.0.0' note and the
generic 'single source of truth' framing. Kept the two actionable
points: ESO requires mapping Layer 1 keys; app.env overrides
envDefaults.
2026-05-12 14:23:40 -07:00
Theodore Li 05892f74f2 improvement(scheduler): raise per-tick claim budget to drain backlog (#4567)
* improvement(scheduler): raise per-tick claim budget to drain backlog

MAX_CRON_CLAIMS 20 -> 100; reserved workflow/job slots 10/10 -> 50/50.
Throughput was capped at 20 schedules/tick which created a 20+ hour
backlog when due work exceeded ~1 item per cron-second.

* improvement(scheduler): raise per-tick claim budget to 200

Bumps MAX_CRON_CLAIMS 100 -> 200 (workflow/job split 100/100). Pairs
with the fire-and-forget cron Lambda change so per-tick processing
time is no longer bounded by the Lambda's 50s HTTP timeout.
2026-05-12 16:05:57 -04:00
Waleed 9d2dd8f550 improvement(helm): helm chart updates with security, ESO, and docs overhaul (#4565)
* improvement(helm): production-ready chart with security, ESO, and docs overhaul

Comprehensive Helm chart improvements bringing the chart up to industry
standards for security, secret management, and documentation.

Security
- Pod Security Standards "restricted" defaults on every pod and container
  (runAsNonRoot, allowPrivilegeEscalation=false, capabilities.drop=[ALL],
  seccompProfile=RuntimeDefault)
- automountServiceAccountToken=false on ServiceAccount and every pod
- NetworkPolicy egress blocks cloud metadata endpoints by default
- Sensitive app/realtime env keys auto-partitioned into chart-managed Secret
  via envFrom; no more plaintext secrets on container specs

Secret management
- Three modes: inline, existingSecret, ExternalSecrets Operator (ESO)
- ESO sync supports arbitrary sensitive keys
- Fail-fast template rendering when ESO enabled but sensitive key unmapped
- AWS/Azure/GCP example files document all three modes

Reliability
- Headless Services for both Postgres StatefulSets
- HPA-aware replicas (omits spec.replicas when autoscaling.enabled)
- PodDisruptionBudget auto-activates when replicaCount > 1
- Startup / liveness / readiness probes with distinct timings
- CronJob ttlSecondsAfterFinished for automatic cleanup

Chart hygiene
- Image tags default to Chart.AppVersion; pullPolicy IfNotPresent
- Optional image.digest pin for content-addressed deploys
- kubeVersion >=1.25.0-0 enforced
- Ollama pinned to 0.23.2; mount moved to /data

Documentation
- README rewritten in cert-manager / Bitnami style
- NOTES.txt with post-install guidance
- Example values files annotated with usage and secret-strategy guidance

* fix(helm): correct resource names in README (sim-sim-* → sim-*)

The sim.fullname helper collapses to the release name when the release
name contains the chart name. With the documented release name 'sim',
actual resources are 'sim-app', 'sim-postgresql', etc. — not the
'sim-sim-*' form previously documented. Fixes copy-paste commands in the
pre-1.0.0 upgrade walkthrough and several troubleshooting snippets.

Also expands the cronjobs component description to reflect the full set
of 13 scheduled jobs (was understated as just Gmail/Outlook polling).

* improvement(helm): split app/realtime env into Secret-bound + inline defaults

- Add app.envDefaults / realtime.envDefaults for chart-shipped operational
  tunables (rate limits, timeouts, IVM, feature-flag defaults, localhost URL
  fallbacks). Rendered inline on the container, not into the Secret
- Remove operational defaults from app.env / realtime.env so the chart-managed
  Secret stays minimal and External Secrets Operator users only map keys they
  actually set, not every chart default
- Skip an envDefaults key when the user explicitly sets it in env (K8s `env`
  overrides `envFrom`, so an inline default would otherwise mask a Secret
  value at runtime)
- Relax values.schema.json to allow empty strings on NEXT_PUBLIC_APP_URL,
  BETTER_AUTH_URL, NEXT_PUBLIC_SUPPORT_EMAIL (defaults supplied via envDefaults)

* fix(helm): address PR review — cronjob validation, ESO apiVersion, secret merge order, image guard

- CronJobs reference CRON_SECRET via secretKeyRef; fail-fast at template
  time when cronjobs.enabled=true and app.env.CRON_SECRET is empty so users
  get a clear error instead of a CreateContainerConfigError loop
- Default externalSecrets.apiVersion to "v1beta1" (supported by every ESO
  release since v0.7). The previous "v1" default targets only ESO v0.17+
- Swap merge order in secrets-app.yaml so app.env wins over realtime.env
  for shared keys (BETTER_AUTH_SECRET, BETTER_AUTH_URL, …) — both pods
  consume the same Secret via envFrom, so the app value must be canonical
- Add `required` guard on sim.image so an empty tag + empty digest +
  empty Chart.AppVersion surfaces as a clear template-time error instead
  of rendering an invalid `repo:` reference

* fix(helm): require critical secrets to be mapped when ESO is enabled

Previously, enabling externalSecrets without mapping BETTER_AUTH_SECRET /
ENCRYPTION_KEY / INTERNAL_API_SECRET (and CRON_SECRET when cronjobs are
on) rendered cleanly but produced CrashLoopBackOff at runtime with
cryptic missing-env errors. Fail at template time instead.

* fix(helm): auto-enable PDB when HPA minReplicas > 1

Previously the auto-enable predicate only checked the static
app.replicaCount, which defaults to 1 even when autoscaling is on
(HPA owns spec.replicas). PDB now also activates when
autoscaling.enabled=true and minReplicas > 1.

* fix(helm): prevent realtime envDefaults from masking app.env Secret values; add StatefulSet upgrade NOTES

- Realtime override-skip now considers keys set in either app.env or
  realtime.env. The shared app Secret is mounted via envFrom on the
  realtime pod, so a key set in app.env (e.g. NEXT_PUBLIC_APP_URL) would
  previously be masked by the realtime envDefault (inline env overrides
  envFrom in K8s).
- NOTES.txt now prints a StatefulSet orphan-delete reminder on upgrade,
  surfacing the immutable serviceName issue documented in the README.

* feat(helm): add Claude Skill for chart deployment

Adds a skill at helm/sim/.claude/skills/sim-helm/ that teaches agents how
to deploy and troubleshoot the Sim Helm chart: install path selection
(inline / existingSecret / ESO), secret generation, the values.yaml
four-layer mental model, common-failure troubleshooting, and the
pre-1.0.0 StatefulSet orphan-delete upgrade procedure.

Skill is loadable by Claude Code, Codex, and OpenCode via the standard
skills convention (directory name matches frontmatter name).

* docs(helm): add CRON_SECRET to TL;DR, dry-run, and example install headers

The validateSecrets guard requires CRON_SECRET when cronjobs.enabled=true
(the default), but the quickstart and example file install commands
omitted it — users following the docs hit a hard template-render failure.
Adds CRON_SECRET to README TL;DR, validate-the-install dry-run snippet,
and the install command headers in all example values files.

* fix(helm): require INTERNAL_API_SECRET in inline secret mode

The ESO coverage validator already required INTERNAL_API_SECRET, but the
inline validateSecrets path only checked BETTER_AUTH_SECRET, ENCRYPTION_KEY,
and CRON_SECRET — letting inline installs render successfully and then
crash at runtime when the realtime↔app shared auth secret was missing.
Adds the same fail-fast check to the inline path.

* docs(helm): surface INTERNAL_API_SECRET upgrade requirement in NOTES.txt

The new validateSecrets check makes app.env.INTERNAL_API_SECRET mandatory
on upgrade. Existing installs that never set it would hit a template
render failure with no in-context guidance. Adds an upgrade-only note
with the generation snippet and storage guidance alongside the existing
StatefulSet orphan-delete instructions.

* fix(helm): NetworkPolicy egress to OTEL collector + external-db example format

- Add app/realtime NetworkPolicy egress rules for the OpenTelemetry
  collector pod on ports 4317 (OTLP gRPC) and 4318 (OTLP HTTP) when
  telemetry.enabled=true. Without these, traces and metrics were silently
  dropped with connection-refused errors when both telemetry and
  networkPolicy were enabled.
- Migrate values-external-db.yaml from the legacy list-shaped egress
  format to the new {exceptCidrs, extraRules} object. The list form would
  replace the default object on merge and crash template rendering when
  the chart tried to access .exceptCidrs on a list.

* fix(helm): NOTES.txt no longer prints false secret warning for ESO users

The secrets-empty warning only checked app.secrets.existingSecret.enabled
before scanning app.env. ESO users intentionally leave app.env empty —
secrets come from the ESO-synced Secret — so every ESO install/upgrade
printed a misleading 'pods will fail to start' warning.

Reorders the branches so externalSecrets.enabled takes precedence: ESO
users now see a confirmation message with kubectl commands to verify the
ExternalSecret has synced. The empty-app.env warning only fires when
both ESO and existingSecret are disabled.

* fix(helm): existingSecret mode no longer drops app.env / realtime.env values

In existingSecret mode the chart-managed Secret is not rendered, so non-empty
values in app.env / realtime.env had nowhere to land — yet the envDefaults
skip logic still suppressed the matching defaults. Result: keys like
NEXT_PUBLIC_APP_URL, BETTER_AUTH_URL, and NODE_ENV silently went missing
on both pods (the example values-existing-secret.yaml hit this directly).

Both app and realtime deployments now inline non-empty values from app.env
(plus realtime.env on the realtime container) when existingSecret is enabled
and ESO is not. Inline / ESO modes are unchanged: inline still flows through
the chart-managed Secret, ESO still owns the synced Secret.

* fix(helm): correct realtime env overlay + filter chart-computed keys in existingSecret mode

Realtime: Sprig merge gives the first source precedence and treats "" as a
real value, so realtime.env empty defaults for shared keys shadowed
non-empty app.env values. Replace with deepCopy($appEnv) base + manual
non-empty overlay of $rtEnv.

Both deployments: exclude DATABASE_URL/SOCKET_SERVER_URL/OLLAMA_URL from
the existingSecret inline path so user-supplied values can't override
chart-computed ones via last-wins env semantics.

* fix(helm): skip envDefaults in existingSecret mode + document egress rename

In existingSecret mode the user's pre-existing Secret is the source of
truth (loaded via envFrom). Inlining localhost envDefaults for URL keys
(BETTER_AUTH_URL, NEXT_PUBLIC_APP_URL, ALLOWED_ORIGINS) silently shadowed
the Secret-bound values because K8s env always wins over envFrom. Skip
envDefaults entirely on both deployments when existingSecret is enabled.

Also call out the networkPolicy.egress shape change (list -> map with
exceptCidrs + extraRules) in the NOTES.txt upgrade block so operators
migrate their custom rules rather than silently losing them.

* fix(helm): copy-pasteable install commands in copilot + ESO examples

values-copilot.yaml: the install header was missing every required
copilot.server.env.* secret (AGENT_API_DB_ENCRYPTION_KEY, INTERNAL_API_SECRET,
LICENSE_KEY, SIM_BASE_URL, SIM_AGENT_API_KEY, REDIS_URL, one model key) plus
copilot.postgresql.auth.password. Pasting it as-is failed at template render.

values-external-secrets.yaml: NEXT_PUBLIC_APP_URL, BETTER_AUTH_URL, etc. were
declared under app.env / realtime.env. In ESO mode the chart-managed Secret
isn't rendered, so the validator (rightly) rejects keys in app.env that
aren't mapped under externalSecrets.remoteRefs. Moved non-secret URL/config
to envDefaults, which is inlined and not subject to the ESO mapping rule.

* polish(helm): configurable NetworkPolicy ingress peers + clearer API_ENCRYPTION_KEY comment

- networkPolicy.ingressFrom lets operators scope the ingress-controller
  rule to a specific namespace/podSelector. Defaults to a single empty
  peer (`- {}`), which is the explicit form of "any source" — same
  effective behavior as the old `from: []` but unambiguous across CNIs.
  To restrict, override with e.g.:
    networkPolicy:
      ingressFrom:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: ingress-nginx

- API_ENCRYPTION_KEY comment: drop the "must be exactly 64 hex
  characters" phrasing that sat awkwardly next to `openssl rand -hex 32`.
  The generation command already produces the required length.

* test(helm): add helm-unittest suites + CI workflow + ci values matrix

- 7 helm-unittest suites covering smoke, validators, secret modes,
  envDefaults secret-mode-aware inlining (round-9 regression net),
  chart-computed env keys (round-8 regression net), NetworkPolicy
  shape, and PDB/HPA conditional rendering (38 tests, ~265ms).
- ci/*.yaml render fixtures for default, production, existingSecret,
  ESO, and external-db install modes.
- GitHub Actions workflow runs helm lint --strict, helm unittest,
  helm template across the ci matrix, and kubeconform validation
  against Kubernetes 1.30 schemas.
- CONTRIBUTING.md documents how to run the same gates locally.

* test(helm): add helm test hook + kind apiserver dry-run in CI

- New templates/tests/test-connection.yaml renders a Pod with
  helm.sh/hook=test that wgets the app Service (and realtime when
  enabled). Lets users run `helm test <release>` after install for
  a real in-cluster connectivity check. Restricted PSS context.
- tests.* values block (image, timeoutSeconds, resources) is the
  knob to disable or tune the probe; documented in values.schema.json.
- 3 helm-unittest tests cover the hook annotations, PSS context,
  and tests.enabled=false skip path (41 tests total).
- New CI job spins up a kind v1.30 cluster and runs
  `kubectl apply --dry-run=server` against the rendered manifests
  for the CRD-free ci fixtures (default / existing-secret /
  external-db). Catches admission and validation issues the static
  kubeconform schema check can't see.

* chore(helm): remove pre-1.0.0 upgrade fluff + tighten .helmignore

This is the 1.0.0 release of the chart — there is no pre-1.0.0
predecessor for users to upgrade from, so all of the dedicated upgrade
narration was hypothetical.

- Drop the 'Upgrading from a pre-1.0.0 build' README section and the
  matching troubleshooting entry.
- Drop the .Release.IsUpgrade block from NOTES.txt: items 5 (StatefulSet
  orphan-delete), 6 (INTERNAL_API_SECRET 'new in 1.0.0'), 7
  (networkPolicy.egress shape change). Each described a migration off a
  chart version that never shipped.
- Delete references/upgrade-pre-1.0.0.md and remove the corresponding
  pointers from SKILL.md.
- Anchor .helmignore patterns to chart root so /tests/ (unit suites)
  and /examples/ are dropped from the packaged tarball without also
  dropping templates/tests/ (the helm test hook).

* chore(helm): drop CI workflow + ci/ fixtures + CONTRIBUTING.md

The helm-unittest suites in helm/sim/tests/ and the helm test hook
in helm/sim/templates/tests/ stay — those are chart-internal quality
scaffolding, not CI. Removed:

- .github/workflows/helm-chart.yml
- helm/sim/ci/*.yaml (5 render fixtures used only by the workflow)
- helm/sim/CONTRIBUTING.md (mostly documented those gates)
- dead /ci/ and /CONTRIBUTING.md entries in .helmignore

* feat(helm): pod rollout on Secret change + topologySpreadConstraints

- Add checksum/secret pod annotations on app, realtime, and copilot
  Deployments (plus checksum/config on app when branding ConfigMap is
  enabled). Closes the long-standing footgun where 'helm upgrade' with
  a changed Secret would silently leave pods running the old values
  until a manual rollout restart.
- New top-level topologySpreadConstraints value (and sim.topologySpreadConstraints
  helper) applied to app and realtime Deployments. Mirrors how affinity
  and tolerations are plumbed; users supply their own labelSelector
  to mirror Bitnami convention.
- 5 helm-unittest cases cover the checksum annotations and topology
  spread rendering (46 tests total).

* fix(helm): drop empty-string shadowing in app/realtime env merge

Sprig 'merge' treats "" as a real value, so a default-empty
app.env.BETTER_AUTH_URL would shadow a non-empty realtime.env override
and the URL would never reach the rendered Secret. Replace 'merge'
with an explicit two-pass overlay that filters empties before writing,
mirroring the same pattern already used in deployment-realtime.yaml's
existingSecret block.

Adds two regression tests: realtime.env-only value reaches the Secret
when app.env is empty, and app.env still wins on collision when both
are non-empty (48 tests total).

* fix(helm): make topologySpreadConstraints per-component to match docstring

Greptile flagged that sim.topologySpreadConstraints helper docstring promised
per-component config (.Values.app, .Values.realtime, ...) but call sites
passed .Values, so any app.topologySpreadConstraints / realtime.topologySpreadConstraints
set by the user was silently dropped. The single global key also prevented
distinct app-vs-realtime spread rules.

Pass .Values.app / .Values.realtime to the helper at each call site; move
the top-level topologySpreadConstraints key into both component sections in
values.yaml. Adds a regression test that app constraints don't leak onto
the realtime pod.

* fix(helm): allow cron pods through app NetworkPolicy

Cursor flagged that when networkPolicy.enabled=true and cronjobs.enabled=true
(the recommended production config), the app NetworkPolicy only allowed
ingress from realtime and the ingress controller — silently blocking every
cron pod's HTTP call to /api/schedules/execute, webhook polls, etc. All 13
default cronjobs would fail.

Tag cron pods with a stable simstudio.ai/component-group: cronjob label so
the app NetworkPolicy can allow them with a single rule (no per-job
enumeration). Rule is conditional on cronjobs.enabled. Adds positive and
negative regression tests.
2026-05-12 10:44:33 -07:00
Waleed 1b94424b4b improvement(mothership): align markdown blockquote, img, em, del with design tokens (#4566)
* improvement(mothership): align markdown blockquote, img, em, del with design tokens

* fix(mothership): correctly scope blockquote paragraph margin reset to first/last child

* improvement(mothership): restore italic on blockquotes

* fix(mothership): widen img component prop type to satisfy Streamdown Components
2026-05-12 10:24:50 -07:00
Siddharth Ganesan 9b333eaae9 fix(mothership): fix tool hidden in ui (#4564) 2026-05-11 19:39:41 -07:00
Waleed 22b5a1e577 fix(tables): cmd+a always selects all cells on the table page (#4562)
* fix(tables): cmd+a always selects all cells on the table page

* fix(tables): move preventDefault inside non-empty guard
2026-05-11 18:24:59 -07:00
Waleed b7c34d4323 fix(file-viewer): prevent scroll jump to top during Mothership streaming (#4559)
* fix(file-viewer): prevent scroll jump to top during Mothership streaming

- Fix root cause: MarkdownCheckboxCtx.Provider was conditionally rendered,
  causing the scroll container to unmount/remount when isStreaming flipped,
  resetting scrollTop to 0 on every stream start
- Add useScrollAnchor hook with spacer element to preserve scroll position
  when streamed content temporarily shrinks the scroll container
- Linger active session ID on complete to prevent streamingContent→undefined
  flicker between consecutive tool calls on the same file
- Gate upsert activation on incoming session having renderable content
- Fix shouldShowStreamingFilePanel to keep panel mounted during linger
- Fix use-chat post-write navigation to work with lingered completed session
- Fix useAutoScroll to check proximity before pinning to bottom on stream start

* fix(file-viewer): preserve session linger in hydrate/upsert and clear spacer after stream

* chore(file-viewer): trim verbose comments to match codebase style

* fix(file-viewer): re-engage auto-scroll when user scrolls back to bottom

* chore(file-viewer): update scroll-anchor tsdoc, remove test separators, add hydrate linger tests

* fix(file-viewer): prevent false re-engage when spacer restoration triggers onScroll

* refactor(file-viewer): extract shouldReengage as pure tested function

Pulls the spacer-guard re-engage condition out of onScroll into an
exported pure function so the false-re-engage invariant (spacer active
→ no re-engage) is covered by automated tests rather than relying on
manual QA. Adds 8 unit tests for shouldReengage alongside the existing
15 for computeSpacerShortage.

* chore(file-viewer): trim verbose inline comments in use-scroll-anchor

* chore(file-viewer): final comment cleanup before merge
2026-05-11 17:49:07 -07:00
Waleed 39f74aa353 feat(data-drains): add GCS, Azure Blob, BigQuery, Snowflake, and Datadog destinations (#4552)
* feat(data-drains): add GCS, Azure Blob, BigQuery, Snowflake, and Datadog destinations

* fix(data-drains): address PR review comments

* fix(data-drains): extract sleepUntilAborted, honor abort across all destinations

* fix(data-drains): widen BigQuery projectId max and dedupe parseServiceAccount

* fix(data-drains): tighten GCS bucket contract and expose Azure endpointSuffix

* improvement(data-drains): extract normalizePrefix and buildObjectKey to shared utils

* fix(data-drains): retry BigQuery network errors; tighten Azure accountKey contract

- BigQuery insertAll now wraps the fetch in try/catch inside the retry loop so DNS failures, socket resets, and timeouts are retried with backoff instead of propagating immediately.
- Align azureBlobCredentialsBodySchema with the runtime schema (min 64 / max 120 / base64 regex) so obviously invalid keys are rejected at the API boundary rather than at drain-run time.

* improvement(data-drains): consolidate parseRetryAfter; add Datadog NDJSON line context

- Extract a single parseRetryAfter helper (capped at 30s, returns number | null) into lib/data-drains/destinations/utils.ts and remove the five local copies in bigquery, datadog, gcs, snowflake, and webhook.
- Datadog parseNdjson now wraps JSON.parse in try/catch and surfaces the failing line index, matching BigQuery's parser.

* fix(data-drains): correct Datadog size guard and Snowflake VARIANT limit

- Datadog payload guard now checks the uncompressed size against the 5 MB limit and the wire size against the 6 MB compressed limit, so gzip cannot smuggle an oversized body past the client-side check.
- Snowflake VARIANT limit is 16 MiB (16,777,216 bytes), not 16,000,000 bytes — small payloads between 16 MB and 16 MiB were being rejected unnecessarily.
- Drop the unused apiKey field on Datadog PostInput; the key is already embedded in the prepared request headers.

* improvement(data-drains): consolidate backoffWithJitter into shared utils

Datadog, GCS, and webhook each had byte-identical backoff helpers (BASE 500ms, MAX 30s, jitter ±20%, Retry-After floor). Lift the helper into lib/data-drains/destinations/utils.ts alongside parseRetryAfter and sleepUntilAborted, and drop the per-file copies and their BASE_BACKOFF_MS/MAX_BACKOFF_MS constants.

* fix(data-drains): align destinations with live provider specs

Audited every destination against live AWS/GCS/Azure/BigQuery/Snowflake/
Datadog/webhook docs and applied spec-correctness fixes:

- S3: reserved bucket prefix amzn-s3-demo-, suffixes --x-s3/--table-s3;
  metadata byte formula excludes x-amz-meta- prefix per AWS spec
- GCS: reject -./.- adjacency; UTF-8 prefix cap; forbid .well-known/
  acme-challenge/ prefix; ASCII-only x-goog-meta-* enforcement
- BigQuery: insertId is 128 chars (not bytes); split DATASET_RE (ASCII)
  and TABLE_RE (Unicode L/M/N + connectors); UTF-8 byte cap on tableId
- Snowflake: disambiguate org-account vs legacy locator account formats;
  requestId+retry=true for idempotent retries; server-side timeout=600;
  default column DATA uppercase to match unquoted canonical form
- Azure: endpoint suffix allowlist (4 sovereign clouds); accountKey
  length(88) base64
- Webhook: url max(2048); CRLF/NUL rejection on bearer/secret/sig header

* fix(data-drains): address PR review on snowflake poll + shared NDJSON parsing

- snowflake pollStatement: per-attempt timeout via AbortSignal.any, retry on 429/5xx with Retry-After + jitter
- bigquery parseNdjson error messages now 1-indexed
- consolidate parseNdjson variants into shared parseNdjsonLines/parseNdjsonObjects in utils

* fix(data-drains): per-attempt fetch timeouts in gcs/bigquery, snowflake poll double-sleep

- gcs.fetchWithRetry + bigquery.postInsertAll now use AbortSignal.any with a per-attempt timeout so a hung TCP connection cannot stall the drain
- snowflake.pollStatement skips the next interval sleep when it just slept for retry backoff

* fix(data-drains): bigquery probe timeout + jittered retries, align Snowflake column default UI/docs

- bigquery test() probe now uses AbortSignal.any + per-attempt timeout
- bigquery insertAll retry switches to backoffWithJitter for thundering-herd avoidance
- Snowflake column placeholder + docs say DATA (uppercase) to match the code default

* fix(data-drains): mirror webhook signingSecret min length in form gate

isComplete now requires signingSecret >= 32 to match the contract/runtime
schema so the Save button can't enable on a value that will fail server-side.

* fix(data-drains): validate JSON client-side for Snowflake before binding

Switch Snowflake to parseNdjsonObjects so malformed rows are caught locally
with 1-indexed line numbers instead of failing the whole INSERT server-side.
Re-stringify each parsed object before binding to PARSE_JSON(?).
Drop the now-unused parseNdjsonLines helper.

* fix(data-drains): cross-cutting audit pass against live provider docs

- Azure: bound retryOptions on BlobServiceClient (SDK default tryTimeoutInMs is per-try unbounded; cap at 30s x 5 tries)
- Webhook contract: mirror runtime — signingSecret.max(512), bearerToken.max(4096) + CRLF/NUL refine, signatureHeader charset + CRLF/NUL refine
- S3 (lib + contract): reject bucket names with dash adjacent to dot; require https:// endpoint at the schema layer
- Snowflake: bind original NDJSON line bytes (re-stringifying a JSON.parse'd value loses bigint precision beyond 2^53-1); check pollStatement 200 body for the SQL error envelope (sqlState/code)
- Datadog: entry builder writes defaults first then user attrs then forced ddtags/message so user rows can't clobber routing fields; validate config.tags as comma-separated key:value pairs
- registry.tsx: tighten isComplete predicates to mirror contract minimums (GCS bucket >= 3, Azure containerName >= 3 / accountKey === 88, BigQuery projectId >= 6, Snowflake account >= 3)

* fix(data-drains): force ddsource/service overrides on Datadog entries

Previous fix placed ddsource/service before ...attrs, leaving them clobberable
by a user row field. Per Datadog docs, service + ddsource pick the processing
pipeline, so a drain's routing config must not be overridable per-row. Spread
attrs first, then force all four reserved fields (ddsource, service, ddtags,
message).

* fix(data-drains): preserve row-distinguishing index when BigQuery insertId overflows

Truncating from the left dropped the index suffix, so any overflow would
collapse all rows in a chunk to the same insertId and BigQuery would silently
dedupe them. Path is unreachable today (UUIDs keep raw ~85 chars), but the
overflow branch is now correct: hash the prefix, keep the index intact.

* fix(data-drains): refresh GCS token per retry, tighten Azure key regex

- gcs: rebuild Authorization header per attempt via buildHeaders so token
  refresh from google-auth-library kicks in if a 5xx retry crosses the
  hour-long token lifetime
- azure_blob: pin account-key regex to {0,2} trailing '=' (base64 of 64
  bytes = exactly 88 chars with up to two '=' pad chars)

* fix(data-drains): address bugbot review of 6336948f6

- gcs: allow 1-char dot-separated bucket components (e.g. "a.bucket")
  to match GCS naming rules — overall name is 3-63 (or up to 222 with
  dots), but per-component minimum is 1 per Google's spec
- bigquery: drain the 401 response body before re-issuing the request
  with a refreshed token so undici can return the socket to the
  keep-alive pool
- snowflake: hoist getJwt() above the perAttempt timer in
  executeStatement so JWT signing doesn't eat the network budget
  (matches the order already used in pollStatement)

* fix(data-drains): allow org-account Snowflake identifier with region suffix

The account validation rejected `<orgname>-<acctname>.<region>.<cloud>`
because `ACCOUNT_LOCATOR_RE`'s first segment forbade hyphens, while
`ACCOUNT_ORG_RE` forbade dots. `normalizeAccountForJwt` already handles
this composite form. Widen the first segment of `ACCOUNT_LOCATOR_RE` to
allow hyphens so the boundary contract and the runtime schema accept
what the JWT layer was already designed to process.

* fix(data-drains): drain retryable response bodies in datadog/gcs loops

Mirrors the bigquery 401 fix. Without consuming the body before
sleeping, undici can't return the socket to the keep-alive pool, so
each retry leaks a TCP connection instead of reusing it.

* fix(data-drains): drain snowflake poll bodies on 202 and retryable status

Mirrors the bigquery/datadog/gcs drains. Long async statements can poll
many times against the same connection; without consuming the body
undici can't return the socket to the keep-alive pool, so each iteration
leaks a connection until GC.

* fix(data-drains): consume success bodies; check Snowflake sqlState on 200

- gcs: drain the body on success paths so undici can return the socket
  to the keep-alive pool
- snowflake: drain the body on synchronous 200 OK and run the same
  sqlState envelope check pollStatement already does — otherwise a
  statement-level failure that completes synchronously would silently
  return success

* fix(data-drains): drain datadog and bigquery probe success bodies

Same undici keep-alive issue as the prior fixes: postWithRetries
returned the Response on success without draining (callers only read
headers); the BigQuery `test()` probe returned without consuming the
body. Both now drain before returning.

* chore(data-drains): regenerate enum migration as 0206 after staging rebase

* fix(data-drains): cap snowflake poll retries; tighten datadog tags min length
2026-05-11 17:37:10 -07:00
Siddharth GanesanandTheodore Li 0b2cfaf7f7 feat(mothership): add superuser env selection (#4558)
* feat(table): live cell updates via SSE + per-table event buffer

Replaces the polling-based row refetch with a push-based SSE stream that
patches the React Query cache directly as cell-state events arrive.

Architecture:
- New per-table event buffer in apps/sim/lib/table/events.ts. Redis sorted-set
  with monotonic eventId, 1h TTL, 5000-event cap, in-memory fallback. Modeled
  after apps/sim/lib/execution/event-buffer.ts but stripped of complexity
  tables don't need (no per-execution lifecycle, no id-batching, no write
  queue serialization). ~150 lines instead of 700.
- writeWorkflowGroupState appends a fat event after each successful 'wrote'.
  Status transitions carry executionId + jobId; terminal/partial transitions
  also include the new output values inline so the client can patch row data
  without a follow-up refetch.
- New SSE route at /api/table/[tableId]/events/stream?from=<lastEventId>.
  Replays from buffer on connect, polls at 500ms (mirrors workflow execution
  stream), heartbeat every 15s, signals 'pruned' if the caller fell off the
  back of the buffer.
- Client hook useTableEventStream subscribes via EventSource. Reconnect-resume
  with last-seen eventId. On 'pruned', invalidates the rows query and resumes
  from the new earliest. Cache patches walk every cached query under
  rowsRoot(tableId) so filter/sort variants all stay live.
- Removes refetchInterval from useTableRows and the per-page polling effect
  from useInfiniteTableRows. React Query's refetchOnWindowFocus +
  refetchOnReconnect cover the durability gap if any push is dropped.

Out of scope:
- Bulk-cancel events (cancellation path is being redesigned separately).
- Generalizing the workflow event-buffer module to a shared primitive (defer
  until a third use case appears; for now the table buffer is the simpler
  cousin of the workflow one).

* fix(table): drop run-mutation refetch so SSE patches aren't overwritten

useRunColumn.onSettled was canceling in-flight queries and invalidating the
rows query — leftover behavior from the polling era. With the SSE stream
now keeping the cache live via incremental patches, this refetch races the
stream and snaps the cache back to whatever DB shows at the refetch moment,
which can lag the just-arrived queued/running events. Cells appeared stuck
on the optimistic 'pending' even though the SSE was delivering the real
transitions.

* chore(table): simplify SSE plumbing — reuse helpers, drop dead polling code

- Reuse snapshotAndMutateRows for SSE cache patches instead of reimplementing
  the page-walk + cache-shape detection. Adds a {cancelInFlight: false} opt
  for the SSE caller (mutations still cancel as before).
- Drop client-side type duplication in use-table-event-stream — import
  TableEvent and TableEventEntry from lib/table/events directly.
- Drop the now-dead mergePagePreservingIdentity + rowEqual from tables.ts;
  their only caller was the polling effect that was removed earlier.
- Drop the defensive try/catch around appendTableEvent in cell-write — the
  function is documented as never-throwing (returns null on failure).
- Combine INCR + ZADD into one Lua eval in events.ts. Halves Redis RTT per
  cell-write. Lua returns the new eventId; the script splices it into the
  pre-built entry JSON.
- Trim refs to plain let bindings inside the effect; trim stale
  comments referencing the old polling implementation.

* fix(table): address PR review on SSE buffer

- TTL-expiry silent miss: when all keys expire, hgetall(meta) returns empty
  so earliestEventId is undefined and the prune branch was skipped. Reconnect
  with non-zero afterEventId now checks the seq counter — its absence (TTL
  expired) signals pruned so the client refetches. Memory fallback mirrors.
- Unbounded ZRANGEBYSCORE: cap reads at TABLE_EVENT_READ_CHUNK = 500 events
  per call. The route's 500ms poll loop drains chunks across ticks instead of
  flushing 5000 entries (multi-MB) in one tick after a long disconnect.
- Pruned handler closes EventSource client-side: server-side close was firing
  onerror and routing through the 500ms backoff path. Now we close
  proactively, reset the reconnect attempt counter, and reconnect immediately
  from the new earliest.

* Cross env copilot

* Force deploy

* Run migration

* Updates

* Fix migration

* Redeploy

* Make dev db push

* restore old migs

* Cross env copilot

* Add custom tools, skills, mcps to mothership

* Update migration

* Fix migs

* UPdate

* Fix types

* Fix

---------

Co-authored-by: Theodore Li <theo@sim.ai>
2026-05-11 16:58:48 -07:00
Waleed 8774f5c341 fix(deps): patch next-mdx-remote and opentelemetry CVEs (#4557)
- bump next-mdx-remote 5.0.0 → 6.0.0 (GHSA-g4xw-jxrg-5f6m / CVE-2026-0969, arbitrary code execution in MDX serialize)
- bump @opentelemetry/sdk-node and exporter-trace-otlp-http 0.200.0 → 0.217.0 (GHSA-q7rr-3cgh-j5r3 / CVE-2026-44902, Prometheus exporter DoS)
- align @opentelemetry/sdk-trace-base, sdk-trace-node, resources to ^2.7.0 to keep all @opentelemetry/* packages on a single core@2.7.1 instance
2026-05-11 15:06:12 -07:00
Waleed d895e0efee fix(security): authorize MCP subagent IDs, oauth workspace, credential admin demotion (#4551)
* fix(security): authorize MCP subagent IDs, oauth workspace, credential admin demotion

- handleSubagentToolCall and handleDirectToolCall now authorize user-supplied
  workflowId/workspaceId via authorizeWorkflowByWorkspacePermission /
  ensureWorkspaceAccess before forwarding downstream; resolvedWorkspaceId is
  derived from the authorized workflow record instead of trusted from the body
- executeOAuthGetAuthLink verifies caller membership (write level) on the
  target workspaceId before generating the OAuth link or writing
  pendingCredentialDraft, closing the cross-workspace credential injection path
- POST /api/credentials/[id]/members wraps role updates in a transaction that
  counts active admins and rejects demotion of the last admin (mirrors the
  existing DELETE guard in the same file)
- GET /api/credentials/[id]/members returns uniform 404 for both missing and
  inaccessible credentials to remove the existence oracle

* fix(security): address PR review — active-status guard, FOR UPDATE locks, workspaceId propagation

- credentials/members POST: add `current.status === 'active'` check to the
  last-admin demotion guard so re-inviting a revoked admin as a non-admin role
  no longer incorrectly hits the "Cannot demote the last admin" path
- credentials/members POST+DELETE: add `.for('update')` to the active-admin
  count SELECT inside both transactions to serialize concurrent demotions and
  eliminate the admin-count TOCTOU race under Postgres READ COMMITTED
- credentials/members POST: also lock the member row itself with `.for('update')`
  so the role+status read and the subsequent UPDATE are atomic
- mcp/copilot handleDirectToolCall: thread the DB-verified workspaceId from the
  authorization result into prepareExecutionContext instead of relying on
  user-supplied args
- oauth handler: fix error message to mention both workspaceId and userId when
  either is missing from the execution context
2026-05-11 10:08:31 -07:00
Waleed 3adbde404b fix(oauth): persist rotated Microsoft refresh tokens (#4554)
* fix(oauth): persist rotated Microsoft refresh tokens

Microsoft Entra rotates refresh tokens on every refresh and expects clients to replace the stored token with the new one. The Microsoft provider config was missing supportsRefreshTokenRotation, so the rotated refresh_token returned by Azure AD was silently discarded and the original token from initial OAuth connect was reused indefinitely — causing periodic 'Failed to refresh access token' errors for Excel, Teams, Outlook, OneDrive, SharePoint, Planner, AD, and Dataverse integrations.

* test(oauth): cover hyphenated Microsoft service IDs in rotation test
2026-05-11 10:01:39 -07:00
Waleed c47777fa33 fix(security): close cross-tenant IDOR gaps in OAuth credential and execution auth (#4549)
* fix(security): close IDOR gaps in OAuth credential and execution authorization

Routes that called resolveOAuthAccountId followed by a conditional
workspace permission check (only run when workspaceId was set) silently
skipped all ownership validation on the legacy account-ID fallback path.
Any authenticated user could supply a raw account.id to access another
tenant's OAuth credentials.

- Replace resolveOAuthAccountId + conditional perm check with
  authorizeCredentialUse in: auth/oauth/wealthbox/item, tools/gmail/label,
  tools/onedrive/files, tools/onedrive/folder, tools/outlook/folders,
  tools/wealthbox/item (routes 1, 3-7)
- Add authorizeCredentialUse ownership gate before resolveVertexCredential
  in providers/route.ts (route 2)
- Add verifyFileAccess check on the user-supplied file key before
  downloadFileFromStorage in tools/wordpress/upload (route 8)
- Add workflowId param to PauseResumeManager methods
  (enqueueOrStartResume, beginPausedCancellation, completePausedCancellation,
  blockQueuedResumesForCancellation, clearPausedCancellationIntent,
  getPausedCancellationStatus, processQueuedResumes) and filter all
  pausedExecutions lookups by workflowId so callers cannot act on another
  tenant's paused execution by supplying a foreign executionId (route 9)
- Update all call sites (cancel, resume, poll routes) to pass workflowId

* fix(security): close verifyFileAccess bypass, thread workflowId to processQueuedResumes, fix log level

- Fail closed in WordPress upload when userFile.key is present but authResult.userId is absent, preventing silent bypass of ownership check via JWT fallback path
- Thread workflowId into processQueuedResumes in the async resume error-recovery path and in pause-persistence.ts to close residual cross-tenant gap
- Change logger.error to logger.warn for credential access denial in OneDrive folder route to match all other routes in this PR

* fix(security): thread workflowId through all processQueuedResumes call sites

Closes residual cross-tenant IDOR gap where processQueuedResumes was called
without a workflowId scope in persistPauseResult, startResumeExecution (success
and error paths), and clearPausedCancellationIntent. workflowId was already in
scope at each site — this wires it through to the existing optional parameter.

* fix(security): remove any types, drop extraneous comments, normalize caught errors

- catch (error: any) → catch (error) + toError(error).message in resume and cancel routes
- Remove what-not-why inline comments from wordpress upload and onedrive/files routes
- Remove redundant debug-only item breakdown log and the file-IDs log in onedrive/files
- Trim extraneous DAG-edge comments from updateSnapshotAfterResume in HITL manager

* fix: use logger.warn for credential access denial in outlook folders route

* fix(security): make workflowId required in all HITL pause/resume methods

All 7 method signatures (processQueuedResumes, enqueueOrStartResume,
beginPausedCancellation, completePausedCancellation, blockQueuedResumesForCancellation,
clearPausedCancellationIntent, getPausedCancellationStatus) previously accepted
workflowId as optional. Every call site already supplies it — making it required
closes the vulnerability at the type level so future callers cannot accidentally
omit tenant scoping and silently fall back to an unscoped DB query.

* fix(security): thread workflowId through internal HITL cancellation calls and remove dead branches in credential-access

* fix(security): harden workflowId scoping and file key guard

Replace falsy workflowId checks in PauseResumeManager (all methods now
unconditionally apply the workflowId WHERE clause, preventing empty-string
bypass). Flip WordPress upload file guard from truthy key check to explicit
non-empty validation so key:"" fails closed with a 404 instead of silently
skipping access control.
2026-05-11 09:50:00 -07:00
Waleed badfeae432 fix(security): SSRF fixes (#4548)
* fix: address SSRF and token-leakage security vulnerabilities

- Azure TTS SSRF: validate region against /^[a-z][a-z0-9-]{1,30}[a-z0-9]$/
  in both the contract (tts.ts) and runtime guard in synthesizeWithAzure,
  preventing user-supplied region from redirecting requests to arbitrary hosts
- HubSpot token in logs: remove fullResponse from logger.info call;
  log only non-sensitive metadata (hub_id, hub_domain, user_id) instead
  of the full introspection response which included the access token
- Wealthbox account takeover: replace hardcoded email with per-user identity
  by fetching /v1/users/me; fall back to token-derived stable identifier
  so distinct Wealthbox users no longer share the same email address
- Shopify SSRF: apply shopifyShopDomainSchema (.myshopify.com allowlist)
  to shopDomain from cookie before using it to build the fetch URL

* fix(wealthbox): correct getUserInfo endpoint, auth header, and stable identity

- Bug 1: Change API endpoint from /v1/users/me to /v1/me (correct Wealthbox API path)
- Bug 2: Replace ACCESS_TOKEN header with Authorization: Bearer <token> (standard OAuth 2.0)
- Bug 3: Remove generateId() from returned id (was non-deterministic, caused duplicate accounts);
  use refresh token (stable, long-lived) instead of access token (rotates every ~2 hours)
  as the hash source for the fallback identity; return null if no token is available

* fix(security): hash wealthbox fallback token identity, guard undefined userId

- Replace base64 encoding with SHA-256 hash for fallback token-derived identity
  so raw token bytes are never stored in the DB
- Return null early when Wealthbox API response lacks an id field to prevent
  all such users colliding on the wealthbox-undefined account

* fix(auth): replace stale wealthbox userInfoUrl placeholder with actual endpoint

The dummy URL comment was rendered obsolete when getUserInfo was updated
to fetch from api.crmworkspace.com/v1/me. Align userInfoUrl with the real
endpoint used in the getUserInfo implementation.

* fix(auth): append generateId() suffix to Wealthbox account IDs to match codebase pattern

All other providers use `${stableId}-${generateId()}` so the account.create.after
hook can strip the UUID suffix, find stale sibling rows, and migrate credential FKs.
Without the suffix the migration logic is skipped and reconnections would hit
duplicate key conflicts instead of gracefully updating credentials.
2026-05-10 19:52:38 -07:00