mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
bad21cb23d64f4e030d000ec4e7e36d72a862c7f
4596
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bad21cb23d |
improvement(agent, file-block): files in agent block, file block v4 (#4610)
* File block v4 * Support files in agent block * Clean up attachments * Fix start.files * Fix * Fix |
||
|
|
9bbbe0adcb |
chore(deps): bump mermaid to 11.15.0 for GHSA-ghcm-xqfw-q4vr (#4615)
* chore(deps): bump mermaid to 11.15.0 for GHSA-ghcm-xqfw-q4vr * chore(deps): override transitive mermaid to 11.15.0 |
||
|
|
62636c7eb3 |
improvement(gmail): replace custom html-to-text regex with library (#4613)
* improvement(gmail): replace custom html-to-text regex with html-to-text library Resolves 4 CodeQL alerts on htmlToPlainText (incomplete tag/entity handling, unsafe regex backtracking). Delegates to the html-to-text npm package already used by the outlook polling trigger and the mail/send route. * improvement(gmail): match outlook selectors config, add nbsp/anchor tests Aligns html-to-text options with apps/sim/lib/webhooks/polling/outlook.ts: suppress anchor hrefs when identical to text, drop bare # anchors, skip img/script/style content. Adds tests for nbsp preservation and anchor behavior. |
||
|
|
e23c20e1eb |
fix(gmail): send emails as multipart/alternative so they render full-width (#4611)
* fix(gmail): send emails as multipart/alternative so they render full-width * fix(gmail): decode & last in htmlToPlainText to avoid double-decoding compound entities * fix(gmail): encode body parts as base64 and decode numeric HTML entities |
||
|
|
922de38d93 |
fix(date-picker): eliminate infinite re-render crash on re-open with existing selection (#4609)
* docs(uploads): clarify QUOTA_EXEMPT_STORAGE_CONTEXTS logs entry in JSDoc * fix(date-picker): eliminate infinite re-render on re-open with existing selection The useEffect that syncs picker state on open had initialStart and initialEnd — Date objects computed on every render — in its dependency array. Because Object.is returns false for any two distinct Date instances, the effect fired on every render when open=true, calling setRangeStart/setRangeEnd and triggering another render, producing an infinite loop that crashed the page. Fix: compute start and end as local variables inside the effect and use the stable string props (props.startDate, props.endDate) as deps instead. Also removes the redundant typeof fileSize === 'number' guard in the multipart quota check — fileSize is z.number() (required) in the contract so it can never be undefined at that point. * refactor(date-picker): comprehensive cleanup and reliable crash fix The previous fix still had derived Date objects in useEffect deps. Object.is(new Date(), new Date()) === false, so any Date in deps causes the effect to run every render, reproducing the infinite loop on re-open with existing time selection. Key changes: - useEffect deps now use only stable primitives (startDate, endDate strings) and compute Date values inside the effect — eliminating the loop - Replace `rest as any` with a FlatDatePickerProps merged type for safe, typed destructuring across the discriminated union - Remove initialStart/initialEnd render-scope variables; compute inline or inside effects to keep derivation local to each use site - Callbacks use destructured props (onChange, onRangeChange, etc.) instead of props.x references - Remove verbose TSDoc on internal callbacks — names are self-documenting - Preserve all existing JSX structure and CalendarMonth logic unchanged |
||
|
|
d0519c1503 |
fix(security): supabase rpc path validation, ssh stream byte cap, storage quota coverage (#4605)
* fix(security): supabase rpc path validation, ssh stream byte cap, storage quota coverage
* fix(security): scope execution log writes to owning workflow; add env-var workspace membership guard
Closes two cross-tenant vulnerabilities:
1. Workflow log cross-tenant write (route.ts + logging-session.ts):
- Route: SELECT before creating LoggingSession to verify executionId belongs
to the claimed workflowId; reject with 404 if owned by a different workflow.
- LoggingSession: add workflow_id to all UPDATE/SELECT WHERE clauses
(raw SQL marker queries, flushAccumulatedCost, loadExistingCost) so
writes are a no-op if executionId was somehow injected.
2. Env-var workspace membership guard (environment/utils.ts):
- getPersonalAndWorkspaceEnv now calls checkWorkspaceAccess when workspaceId
is provided; throws if the userId is not a member, preventing any future
caller from reading another workspace's decrypted secrets without
explicit membership verification at the call site.
* fix(security): remove fileSize > 0 quota bypass gate; exempt logs context from quota
* chore: remove extraneous inline comments
* fix(security): scope markExecutionAsFailed UPDATE by workflowId; thread workflowId through HITL callers
* fix(security): add personal credential ownership check in sharepoint site route; scope markExecutionAsFailed by workflowId
* fix: remove logs from user-accessible upload contexts; restore distinct biome .next glob
* fix(sharepoint): migrate site route to authorizeCredentialUse
The previous fix only checked userId equality for personal credentials and
workspace membership (via getUserEntityPermissions) for workspace credentials.
authorizeCredentialUse additionally enforces credentialMember access for
workspace-scoped credentials, matching the standard pattern used by all
other tool selector routes.
* fix(logging): make workflowId required in markExecutionAsFailed
Making workflowId optional left a footgun — future callers could silently
omit it and the WHERE clause would degrade to executionId-only, losing the
cross-tenant scoping guarantee. All callers already supply workflowId, so
making it required (with string | undefined for the middle params to keep
call sites unchanged) closes the gap without touching any caller.
* test(security): add tests for cross-tenant log guard, quota bypass fix, and workflowId scoping
- log/route.test.ts: verifies cross-tenant executionId guard returns 404
when the execution belongs to a different workflow, and passes for same
workflow or fresh executions
- multipart/route.test.ts: verifies fileSize:0 no longer bypasses quota
check and that the logs context is rejected at the endpoint level
- logging-session.test.ts: verifies markExecutionAsFailed scopes by both
executionId and workflowId, and that the instance method forwards workflowId
* fix(lint): move IconComponent outside ToolInput to fix noNestedComponentDefinitions
* fix(logging): scope completeWithCancellation and completeWithPause reads by workflowId
Both SELECT queries that check execution status before writing a
terminal result were only filtering on executionId. Adds workflowId
to the WHERE clause so all seven reads and writes in LoggingSession
consistently scope by (workflowId, executionId).
|
||
|
|
11fa96cac0 |
chore(deps): bump next to 16.2.5 for CVE-2026-44578 SSRF fix (#4606)
* chore(deps): bump next to 16.2.5 for CVE-2026-44578 SSRF fix * chore(deps): bump next to 16.2.6 for full May 2026 security release coverage |
||
|
|
80c9a01275 |
fix(security): harden file access controls, webhook auth, and input bounds (#4601)
* fix(security): harden file access controls, webhook auth, and input bounds * fix(security): extend file access checks to remaining tool routes * fix(logs): address PR review comments on time filter * fix(logs): set end-time milliseconds to 999 for datetime filter strings * fix(files): return 404 instead of 500 on file access denial in utility paths * remove tooltip from resource tabs |
||
|
|
4a9e248eac | feat(cloudwatch): add mute and unmute alarm operations (#4602) | ||
|
|
044e034719 |
fix(integrations): gdrive trashed search, slack blocks-with-file, slack get_message ts (#4600)
* fix(integrations): gdrive trashed search, slack blocks-with-file, slack get_message ts - Google Drive search/list: skip default `trashed = false` when user query already specifies a `trashed = ...` predicate, so trashed-file searches work. - Slack send-message with files: forward `blocks` through to `files.completeUploadExternal` so Block Kit renders when files are attached. - Slack get_message: switch from `conversations.history` (oldest lower-bound returned the next message after) to `conversations.replies` with `ts=` for exact-match lookup, plus a defensive ts-equality guard and clearer error. * fix(google_drive): revert list.ts trashed guard — query is plain text, not gdrive syntax * fix(slack): omit initial_comment when blocks present so Block Kit actually renders on file uploads |
||
|
|
642231f875 |
improvement(scheduler): drain due schedules in chunks (#4578)
* improvement(scheduler): drain in chunks instead of a single capped claim Replaces the fixed MAX_CRON_CLAIMS (200) with a chunked drain loop: claim WORKFLOW_CHUNK_SIZE + JOB_CHUNK_SIZE per iteration, process via Promise.allSettled, repeat until both claim queries return empty or MAX_TICK_DURATION_MS elapses. Throughput is no longer bounded by a static per-tick ceiling; it scales until DB or trigger.dev is the limit. Per-iteration chunk size still bounds row-lock set and fan-out concurrency. Extracts processScheduleItem and processJobItem so the loop body stays readable. Existing claim semantics (FOR UPDATE SKIP LOCKED, lastQueuedAt as the claim signal, staleness reclaim) are unchanged. * improvement(scheduler): skip claim once a queue is exhausted and drop workflowUtils non-null assertion Addresses Greptile review on PR #4578: - track per-queue exhaustion when a claim returns fewer than CHUNK_SIZE rows; subsequent iterations skip the claim query for that queue. Saves one DB round-trip per iteration once one queue drains while the other is still working. - narrow workflowUtils to a local const inside the loop body so the schedule processing branch only runs when the import has completed. Removes the misleading non-null assertion. |
||
|
|
b1a87d531c |
Revert "improvement(db): add session statement/lock timeouts; simplify KB doc tx (#4593)" (#4599)
This reverts commit
|
||
|
|
8831defd2e |
fix(seo): use canonical SITE_URL for robots and sitemap (#4598)
* fix(seo): use canonical SITE_URL for robots and sitemap * fix(seo): drop /templates from sitemap and guard robots/sitemap in seo test |
||
|
|
4295a5c855 |
improvement(db): add session statement/lock timeouts; simplify KB doc tx (#4593)
* v0.6.29: login improvements, posthog telemetry (#4026) * feat(posthog): Add tracking on mothership abort (#4023) Co-authored-by: Theodore Li <theo@sim.ai> * fix(login): fix captcha headers for manual login (#4025) * fix(signup): fix turnstile key loading * fix(login): fix captcha header passing * Catch user already exists, remove login form captcha * improvement(db): add session statement/lock timeouts; simplify KB doc tx * fix(knowledge): close soft-delete TOCTOU on KB document insert Fix the race the bots flagged: KB delete is soft (`deletedAt = now`) so the FK can't catch a concurrent KB delete between the existence check and the document insert. - Add `insertDocumentsIfKbAlive` helper that gates the insert on `EXISTS(SELECT 1 FROM knowledge_base WHERE id=$kb AND deleted_at IS NULL)` in the same statement via INSERT...SELECT...WHERE EXISTS. Atomic at the MVCC snapshot — no transaction, no row lock. - Use jsonb_to_recordset to declare column types once, avoiding per-param casts for nullable columns. - Wire into both `createDocumentRecords` (bulk) and `createSingleDocument`. - Keep the upfront KB existence check as a fast-path early-out for the common case; the atomic insert is the race guard. --------- Co-authored-by: Waleed <walif6@gmail.com> Co-authored-by: Siddharth Ganesan <33737564+Sg312@users.noreply.github.com> Co-authored-by: Vikhyath Mondreti <vikhyathvikku@gmail.com> |
||
|
|
c3ac54e0a9 | fix(vfs): make copilot message ordering deterministic via WITH ORDINALITY (#4597) | ||
|
|
b1a9443178 |
improvement(billing): move overage calculations out of txes (#4595)
* improvement(billing): move calc subscription overage out of tx * fix double billing risk * address comments * address comments * share timeout const |
||
|
|
b5dba82ac9 |
improvement(db): reduce connection saturation and egress hotspots (#4594)
* improvement(db): reduce connection saturation and egress hotspots * fix(vfs): preserve native content type in copilot SQL projection * fix(vfs): guard jsonb_array_elements against non-array contentBlocks |
||
|
|
104949bdc2 |
fix(tables): eliminate checkbox flicker on rapid cell toggle (#4592)
* fix(tables): eliminate checkbox flicker on rapid cell toggle * fix(tables): symmetric guarded onSettled across row write mutations * fix(tables): merge only mutated keys in onSuccess to preserve concurrent optimistic patches |
||
|
|
568a552d67 |
fix(rate-limit): close rate-limit bypass and tighten public route limits (#4591)
* fix(rate-limit): close rate-limit bypass and tighten public route limits * fix(rate-limit): address PR review — drop success field from 429 body, fall back to per-IP when JWT auth lacks userId |
||
|
|
1c111ff2d7 |
fix(mothership): persist @-mentioned resources across send (#4587)
* fix(mothership): persist @-mentioned resources across send and merge on hydration * fix(mship-resources): handle ADD/DELETE race and reorder during pending flush - Track in-flight ADD promises so DELETE chains off finally(), preventing orphaned server rows when a user removes a resource before its POST resolves - Defer reorder PATCH until pending flush completes; emit with full local order - Clear new refs in reset paths * fix(mship-resources): defer reorder when ADDs are in-flight on existing chat reorderResources previously only checked pendingPersistResourceKeysRef. When a chatId exists, addResource fires the POST immediately and only tracks the promise in inFlightResourceAddsRef — so a reorder before those ADDs settle shipped a PATCH the server rejected, and the silent catch lost the reorder. Now treat in-flight ADDs like pending ones: defer the PATCH and replay it after Promise.allSettled on the in-flight map. |
||
|
|
ff3c8f765e |
fix(file-block): fix get op (#4590)
* File block get * Lint * Fix * Fix auth |
||
|
|
214355b26e |
improvement(file-block): add get operation (#4588)
* File block get * Lint * Fix |
||
|
|
4de955d0b6 | fix(otel): address staging pr comments for trigger otel (#4586) | ||
|
|
ed39edb906 |
feat(observability): export Trigger.dev telemetry to Grafana Cloud OTLP (#4583)
* feat(observability): export Trigger.dev telemetry to Grafana Cloud OTLP Wire OTLP HTTP exporters for traces, logs, and metrics from the Trigger.dev runtime to Grafana Cloud. Auth uses Basic with instance ID and API token. Gated behind GRAFANA_OTLP_ENDPOINT, GRAFANA_INSTANCE_ID, and GRAFANA_API_TOKEN — all three must be set together or all unset; partial config throws at startup. * improvement(observability): use OTLP HTTP/JSON for metrics for consistency with traces and logs * feat(observability): tag Trigger.dev telemetry with deployment.environment.name * improvement(observability): switch Grafana telemetry vars to OTLP-shaped trio |
||
|
|
2441d5ad6e |
feat(mothership): add files to mship block (#4584)
* Add files to mship block * Fixes * Fix * fix |
||
|
|
1ed3a4ec6d |
feat(mothership): pin tasks to keep them at the top of the sidebar (#4582)
* feat(mothership): pin tasks to keep them at the top of the sidebar * fix(sidebar): address PR review feedback for pin tasks * fix(posthog): register task_pinned and task_unpinned events * fix(tasks): insert new optimistic tasks below pinned partition |
||
|
|
cdc7513d23 |
improvement(mothership): allow mship to send function execute timeout (#4581)
* Improve mship fexecute * Fix |
||
|
|
bdf9ffc798 |
fix(event-buffer): re-compact the event with preserveUserFileBase64: false (#4579)
* preserveUserFileBase64 is on and the event still exceeds the threshold, re-compact the event with preserveUserFileBase64: false * address comments |
||
|
|
689e1f76d1 | feat(mothership): Add conversationId to mship block (#4577) | ||
|
|
c21bb915f7 |
improvement(grafana): align tools and block with Grafana API spec (#4574)
* improvement(grafana): align tools and block with official Grafana API spec Validates and corrects the Grafana integration against the official API docs: fixes wire-format field naming for provisioned alert rules (missing_series_evals_to_resolve, keepFiringFor, orgID), adds X-Disable-Provenance support, expands alert-rule params (isPaused, notificationSettings, record, annotations, labels), corrects defaults (execErrState=Error, dashboard overwrite=false), and centralizes alert-rule output mapping in a shared utils module. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(grafana): correct wire-format casing for provisioned alert rule fields Grafana's ProvisionedAlertRule schema (verified against upstream Go source and swagger spec) uses keep_firing_for (snake_case) and missingSeriesEvalsToResolve (camelCase) — the opposite of what prior audit rounds assumed. POST/PUT bodies now send the correct field names; mapAlertRule reads the correct primary names with the old casings kept as fallbacks. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(grafana): address PR review feedback - Drop hardcoded orgID: 1 fallback; only send orgID when organizationId is provided, so token-scoped org context drives rule placement. - Surface invalid JSON for notificationSettings/record on alert rule create/update instead of silently dropping the input. - Fix execErrState description in update_alert_rule to include Error. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(grafana): surface invalid JSON for annotations/labels/data on alert rules Match the behavior of other JSON params (data, notificationSettings, record): return a descriptive error instead of silently falling back to {} (create) or keeping the existing value (update). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * docs(grafana): expose alert-rule output fields in generated docs Move ALERT_RULE_OUTPUT_FIELDS from utils.ts to types.ts and rename to SCREAMING_SNAKE_CASE so scripts/generate-docs.ts (which only resolves const references from types.ts matching [A-Z][A-Z_0-9]+) can inline the per-field rows into the generated alert-rule output tables. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
773cd84e3f |
fix(mothership): reconcile stuck conversation_id against Redis lock to clear stuck-yellow task tiles (#4556)
* fix(mothership): reconcile stuck conversation_id against Redis lock to clear stuck-yellow task tiles copilot_chats.conversation_id has no TTL/heartbeat, so when a stream process dies before the clear path runs (pod OOM, SIGKILL, uncaught throw, deploy mid-stream) the column is orphaned and the task tile renders yellow forever. The Redis lock at copilot:chat-stream-lock:<chatId> is the canonical liveness signal and self-heals via 60s TTL + 20s heartbeat, but the mothership APIs weren't consulting it. Adds read-time reconciliation: a batched MGET helper checks whether each persisted conversation_id still has a live Redis lock, and both GET /api/mothership/chats and GET /api/mothership/chats/[chatId] rewrite the marker to null when the lock has expired. No DB writes; stuck rows self-heal on next fetch. * test(mothership): clarify test name to reflect that getActiveChatStreamIds is called with empty candidateIds * address comments * fix state machine issue * cleanup code and fix types --------- Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai> |
||
|
|
d5c2ead5d4 |
fix(console): match child-workflow inner blocks by instanceId when reconciling dropped SSE events (#4575)
* fix(console): match child-workflow inner blocks by instanceId when reconciling dropped SSE events * fix(console): drop noisy warn when reconcile finds no matching entry |
||
|
|
6503671102 |
fix(security): harden findings — path traversal, SSRF, IDOR, file auth, credential access (#4571)
* fix(security): harden HIGH deepsec findings across multiple attack surfaces
- Supabase tools (get_row, delete, update): validate table name with strict
identifier regex and encodeURIComponent to prevent LLM-controlled path
traversal to admin endpoints; add missing empty-filter guard to update
matching the delete.ts pattern
- SFTP/SMTP/SharePoint upload routes: add verifyFileAccess ownership check
before downloadFileFromStorage, matching the WordPress reference pattern;
rejects files the requesting user does not own with 404
- Gmail labels, OneDrive folders, Wealthbox items (×2): replace bare
resolveOAuthAccountId + workspace-only membership check with
authorizeCredentialUse which enforces credentialMember table; use
credentialOwnerUserId for token refresh instead of bare accountRow.userId
- A2A utils: thread pre-resolved IP from validateUrlWithDNS into A2A SDK
via pinnedFetch (secureFetchWithPinnedIP) for JsonRpcTransportFactory,
RestTransportFactory, and DefaultAgentCardResolver, closing the TOCTOU
DNS rebinding window
- SSH utils: cap stdout/stderr accumulation at 16 MB with truncation marker
to prevent OOM from unbounded command output
- Form DELETE route: replace db.delete() with db.update({archivedAt}) for
true soft delete matching the schema's archivedAt column
- Workflow admin import: fix Array.isArray() guard that silently dropped
all variables (export format is Record, not Array)
- Multipart upload: apply checkStorageQuota and MAX_WORKSPACE_FILE_SIZE to
mothership context, closing the quota bypass for workspace-scoped storage
* fix(security): eliminate workspace env lost-update race with atomic JSONB ops
PUT: use `variables || excluded.variables` in onConflictDoUpdate so
concurrent writes merge atomically in the DB instead of last-writer-wins
at the application layer.
DELETE: replace the read-modify-write upsert with a single UPDATE that
removes keys via the JSONB `-` operator, preventing concurrent deletes
from resurrecting previously-removed secrets.
* fix(security): address audit findings from security fix review
- SMTP send: restructure attachment loop from Promise.all to sequential
for...of so verifyFileAccess denial returns 404 instead of propagating
as a generic 500 via the SMTP error classifier
- Supabase tools: extend table-name validation and encodeURIComponent to
the five previously missed tools — insert, upsert, count, query,
text_search — completing coverage across all nine Supabase tools
- Credential routes: remove unnecessary `request as any` casts in Gmail,
OneDrive, and Wealthbox routes; authorizeCredentialUse already accepts
NextRequest directly
- Form soft delete: also set isActive=false alongside archivedAt so that
any future code paths querying by isActive see a consistent state
- SSH utils: fix exit code fallback from 0 to -1 so an abnormally closed
connection that supplies no exit code is not reported as success
- Workspace env: capitalize EXCLUDED.variables in the onConflictDoUpdate
set clause to make the pseudo-table reference unambiguous
* fix(security): address PR review comments and harden deepsec fixes
- fix(env): replace jsonb operators with transaction+FOR UPDATE read-modify-write
- PUT: uses db.transaction + SELECT FOR UPDATE + JS merge to avoid lost-update race
- DELETE: same pattern; fixes variable scope bug where current was referenced outside tx
- removes broken || and - jsonb operators that fail on json-typed column
- fix(ssh): trim truncated output consistently with non-truncated path
- fix(gmail): remove redundant resolveOAuthAccountId call
- adds credentialType field to CredentialAccessResult
- authorizeCredentialUse now returns credentialType in all success paths
- gmail/labels route uses authz.credentialType and authz.resolvedCredentialId directly
- fix(supabase): centralize table identifier validation
- adds validateDatabaseIdentifier() to input-validation.ts
- all 8 supabase tools use the shared util instead of inline regex
* fix(workflows): fix VariableType assignment in admin workflow import route
The intermediate Record cast used 'string' for the type field which TypeScript
correctly rejected — WorkflowVariable.type is 'VariableType', not string.
Changed the cast to use VariableType so both branches typecheck correctly.
* fix(a2a): handle Request objects in pinnedFetch URL extraction
* fix(security): extract shared file-access guard; merge workspace/mothership branch
* fix(security): advisory lock for env first-insert race; handle all BodyInit types in pinnedFetch
* chore: remove inline comment from advisory lock
* fix(security): remove stray comment; narrow credentialType to literal union
* fix(security): add credentialId validation to wealthbox oauth route; fix null body override in pinnedFetch
* fix(security): stream A2A response body to unblock SSE; keep text/json/arrayBuffer for non-streaming callers
* fix(security): resolve credentialId guard on OneDrive, use assertToolFileAccess in WordPress, memoize body buffer to prevent silent empty reads, fix ArrayBuffer type cast
* fix(security): handle string[][] HeadersInit format in pinnedFetch
* fix(security): keep abort listener alive during body streaming; clean up in stream end/error/cancel
* chore: remove extraneous inline comment
* fix(security): cleanup abort listener when maxResponseBytes limit is exceeded
|
||
|
|
ec936be5ca |
improvement(workflow-block): support manual workflow ID via advanced mode (#4573)
* improvement(workflow-block): support manual workflow ID via advanced mode * fix(input-mapping): resolve workflowId via canonical hook for advanced mode * fix(input-mapping): fall back to manualWorkflowId in preview context * refactor(input-mapping): resolve workflowId via useDependsOnGate canonical pattern |
||
|
|
43f53bb7d3 |
feat(execution): payload size bottlenecks with lazy execution value hydration, safer materialization, and batched parallel execution (#4560)
* improvement(resolver): lazy resolution for underlying fields greater than 10MB * progress * feat(parallel): batching * codegen to allow inline substitution * address comments * ui inconsistencies * cleanup redundant code * address more comments * address comments * replace helper * fix tests |
||
|
|
f6b246ba44 |
fix(docs): restore media centering and full-width intro image (#4570)
* fix(docs): restore media centering and full-width intro image * fix(docs): drop overflow-hidden from intro media wrappers so focus ring is not clipped * fix(docs): use inset focus ring on lightbox media so parent overflow-hidden cannot clip it * fix(docs): drop focus ring on lightbox media to match original UI |
||
|
|
d1eb79ecd3 |
fix(helm): preserve STS serviceName + networkPolicy.egress back-compat (#4569)
* fix(helm): preserve STS serviceName + networkPolicy.egress back-compat
Greptile flagged two real upgrade-breaking changes vs the prior chart:
1. statefulset-postgresql spec.serviceName flipped from <name>-postgresql
to <name>-postgresql-headless. spec.serviceName is immutable, so any
existing install would hit 'Forbidden: updates to statefulset spec ...'
on helm upgrade. Revert to the original name (the headless Service in
services.yaml is added alongside, not as a swap).
2. networkPolicy.egress changed from a list to a map ({extraRules, exceptCidrs}),
silently dropping any custom egress list set by existing users. Restore
the original list semantics for networkPolicy.egress and move cloud-metadata
blocking to a sibling top-level field networkPolicy.egressExceptCidrs.
Adds NOTES.txt upgrade-notes entry covering both + the ESO v1→v1beta1 default
flip (functionally a no-op, but worth surfacing).
* docs(helm): update README egress reference to new key name
* fix(helm): revert copilot-postgresql STS serviceName too (same immutability issue)
Audit caught that the main fix in
|
||
|
|
05892f74f2 |
improvement(scheduler): raise per-tick claim budget to drain backlog (#4567)
* improvement(scheduler): raise per-tick claim budget to drain backlog MAX_CRON_CLAIMS 20 -> 100; reserved workflow/job slots 10/10 -> 50/50. Throughput was capped at 20 schedules/tick which created a 20+ hour backlog when due work exceeded ~1 item per cron-second. * improvement(scheduler): raise per-tick claim budget to 200 Bumps MAX_CRON_CLAIMS 100 -> 200 (workflow/job split 100/100). Pairs with the fire-and-forget cron Lambda change so per-tick processing time is no longer bounded by the Lambda's 50s HTTP timeout. |
||
|
|
9d2dd8f550 |
improvement(helm): helm chart updates with security, ESO, and docs overhaul (#4565)
* improvement(helm): production-ready chart with security, ESO, and docs overhaul
Comprehensive Helm chart improvements bringing the chart up to industry
standards for security, secret management, and documentation.
Security
- Pod Security Standards "restricted" defaults on every pod and container
(runAsNonRoot, allowPrivilegeEscalation=false, capabilities.drop=[ALL],
seccompProfile=RuntimeDefault)
- automountServiceAccountToken=false on ServiceAccount and every pod
- NetworkPolicy egress blocks cloud metadata endpoints by default
- Sensitive app/realtime env keys auto-partitioned into chart-managed Secret
via envFrom; no more plaintext secrets on container specs
Secret management
- Three modes: inline, existingSecret, ExternalSecrets Operator (ESO)
- ESO sync supports arbitrary sensitive keys
- Fail-fast template rendering when ESO enabled but sensitive key unmapped
- AWS/Azure/GCP example files document all three modes
Reliability
- Headless Services for both Postgres StatefulSets
- HPA-aware replicas (omits spec.replicas when autoscaling.enabled)
- PodDisruptionBudget auto-activates when replicaCount > 1
- Startup / liveness / readiness probes with distinct timings
- CronJob ttlSecondsAfterFinished for automatic cleanup
Chart hygiene
- Image tags default to Chart.AppVersion; pullPolicy IfNotPresent
- Optional image.digest pin for content-addressed deploys
- kubeVersion >=1.25.0-0 enforced
- Ollama pinned to 0.23.2; mount moved to /data
Documentation
- README rewritten in cert-manager / Bitnami style
- NOTES.txt with post-install guidance
- Example values files annotated with usage and secret-strategy guidance
* fix(helm): correct resource names in README (sim-sim-* → sim-*)
The sim.fullname helper collapses to the release name when the release
name contains the chart name. With the documented release name 'sim',
actual resources are 'sim-app', 'sim-postgresql', etc. — not the
'sim-sim-*' form previously documented. Fixes copy-paste commands in the
pre-1.0.0 upgrade walkthrough and several troubleshooting snippets.
Also expands the cronjobs component description to reflect the full set
of 13 scheduled jobs (was understated as just Gmail/Outlook polling).
* improvement(helm): split app/realtime env into Secret-bound + inline defaults
- Add app.envDefaults / realtime.envDefaults for chart-shipped operational
tunables (rate limits, timeouts, IVM, feature-flag defaults, localhost URL
fallbacks). Rendered inline on the container, not into the Secret
- Remove operational defaults from app.env / realtime.env so the chart-managed
Secret stays minimal and External Secrets Operator users only map keys they
actually set, not every chart default
- Skip an envDefaults key when the user explicitly sets it in env (K8s `env`
overrides `envFrom`, so an inline default would otherwise mask a Secret
value at runtime)
- Relax values.schema.json to allow empty strings on NEXT_PUBLIC_APP_URL,
BETTER_AUTH_URL, NEXT_PUBLIC_SUPPORT_EMAIL (defaults supplied via envDefaults)
* fix(helm): address PR review — cronjob validation, ESO apiVersion, secret merge order, image guard
- CronJobs reference CRON_SECRET via secretKeyRef; fail-fast at template
time when cronjobs.enabled=true and app.env.CRON_SECRET is empty so users
get a clear error instead of a CreateContainerConfigError loop
- Default externalSecrets.apiVersion to "v1beta1" (supported by every ESO
release since v0.7). The previous "v1" default targets only ESO v0.17+
- Swap merge order in secrets-app.yaml so app.env wins over realtime.env
for shared keys (BETTER_AUTH_SECRET, BETTER_AUTH_URL, …) — both pods
consume the same Secret via envFrom, so the app value must be canonical
- Add `required` guard on sim.image so an empty tag + empty digest +
empty Chart.AppVersion surfaces as a clear template-time error instead
of rendering an invalid `repo:` reference
* fix(helm): require critical secrets to be mapped when ESO is enabled
Previously, enabling externalSecrets without mapping BETTER_AUTH_SECRET /
ENCRYPTION_KEY / INTERNAL_API_SECRET (and CRON_SECRET when cronjobs are
on) rendered cleanly but produced CrashLoopBackOff at runtime with
cryptic missing-env errors. Fail at template time instead.
* fix(helm): auto-enable PDB when HPA minReplicas > 1
Previously the auto-enable predicate only checked the static
app.replicaCount, which defaults to 1 even when autoscaling is on
(HPA owns spec.replicas). PDB now also activates when
autoscaling.enabled=true and minReplicas > 1.
* fix(helm): prevent realtime envDefaults from masking app.env Secret values; add StatefulSet upgrade NOTES
- Realtime override-skip now considers keys set in either app.env or
realtime.env. The shared app Secret is mounted via envFrom on the
realtime pod, so a key set in app.env (e.g. NEXT_PUBLIC_APP_URL) would
previously be masked by the realtime envDefault (inline env overrides
envFrom in K8s).
- NOTES.txt now prints a StatefulSet orphan-delete reminder on upgrade,
surfacing the immutable serviceName issue documented in the README.
* feat(helm): add Claude Skill for chart deployment
Adds a skill at helm/sim/.claude/skills/sim-helm/ that teaches agents how
to deploy and troubleshoot the Sim Helm chart: install path selection
(inline / existingSecret / ESO), secret generation, the values.yaml
four-layer mental model, common-failure troubleshooting, and the
pre-1.0.0 StatefulSet orphan-delete upgrade procedure.
Skill is loadable by Claude Code, Codex, and OpenCode via the standard
skills convention (directory name matches frontmatter name).
* docs(helm): add CRON_SECRET to TL;DR, dry-run, and example install headers
The validateSecrets guard requires CRON_SECRET when cronjobs.enabled=true
(the default), but the quickstart and example file install commands
omitted it — users following the docs hit a hard template-render failure.
Adds CRON_SECRET to README TL;DR, validate-the-install dry-run snippet,
and the install command headers in all example values files.
* fix(helm): require INTERNAL_API_SECRET in inline secret mode
The ESO coverage validator already required INTERNAL_API_SECRET, but the
inline validateSecrets path only checked BETTER_AUTH_SECRET, ENCRYPTION_KEY,
and CRON_SECRET — letting inline installs render successfully and then
crash at runtime when the realtime↔app shared auth secret was missing.
Adds the same fail-fast check to the inline path.
* docs(helm): surface INTERNAL_API_SECRET upgrade requirement in NOTES.txt
The new validateSecrets check makes app.env.INTERNAL_API_SECRET mandatory
on upgrade. Existing installs that never set it would hit a template
render failure with no in-context guidance. Adds an upgrade-only note
with the generation snippet and storage guidance alongside the existing
StatefulSet orphan-delete instructions.
* fix(helm): NetworkPolicy egress to OTEL collector + external-db example format
- Add app/realtime NetworkPolicy egress rules for the OpenTelemetry
collector pod on ports 4317 (OTLP gRPC) and 4318 (OTLP HTTP) when
telemetry.enabled=true. Without these, traces and metrics were silently
dropped with connection-refused errors when both telemetry and
networkPolicy were enabled.
- Migrate values-external-db.yaml from the legacy list-shaped egress
format to the new {exceptCidrs, extraRules} object. The list form would
replace the default object on merge and crash template rendering when
the chart tried to access .exceptCidrs on a list.
* fix(helm): NOTES.txt no longer prints false secret warning for ESO users
The secrets-empty warning only checked app.secrets.existingSecret.enabled
before scanning app.env. ESO users intentionally leave app.env empty —
secrets come from the ESO-synced Secret — so every ESO install/upgrade
printed a misleading 'pods will fail to start' warning.
Reorders the branches so externalSecrets.enabled takes precedence: ESO
users now see a confirmation message with kubectl commands to verify the
ExternalSecret has synced. The empty-app.env warning only fires when
both ESO and existingSecret are disabled.
* fix(helm): existingSecret mode no longer drops app.env / realtime.env values
In existingSecret mode the chart-managed Secret is not rendered, so non-empty
values in app.env / realtime.env had nowhere to land — yet the envDefaults
skip logic still suppressed the matching defaults. Result: keys like
NEXT_PUBLIC_APP_URL, BETTER_AUTH_URL, and NODE_ENV silently went missing
on both pods (the example values-existing-secret.yaml hit this directly).
Both app and realtime deployments now inline non-empty values from app.env
(plus realtime.env on the realtime container) when existingSecret is enabled
and ESO is not. Inline / ESO modes are unchanged: inline still flows through
the chart-managed Secret, ESO still owns the synced Secret.
* fix(helm): correct realtime env overlay + filter chart-computed keys in existingSecret mode
Realtime: Sprig merge gives the first source precedence and treats "" as a
real value, so realtime.env empty defaults for shared keys shadowed
non-empty app.env values. Replace with deepCopy($appEnv) base + manual
non-empty overlay of $rtEnv.
Both deployments: exclude DATABASE_URL/SOCKET_SERVER_URL/OLLAMA_URL from
the existingSecret inline path so user-supplied values can't override
chart-computed ones via last-wins env semantics.
* fix(helm): skip envDefaults in existingSecret mode + document egress rename
In existingSecret mode the user's pre-existing Secret is the source of
truth (loaded via envFrom). Inlining localhost envDefaults for URL keys
(BETTER_AUTH_URL, NEXT_PUBLIC_APP_URL, ALLOWED_ORIGINS) silently shadowed
the Secret-bound values because K8s env always wins over envFrom. Skip
envDefaults entirely on both deployments when existingSecret is enabled.
Also call out the networkPolicy.egress shape change (list -> map with
exceptCidrs + extraRules) in the NOTES.txt upgrade block so operators
migrate their custom rules rather than silently losing them.
* fix(helm): copy-pasteable install commands in copilot + ESO examples
values-copilot.yaml: the install header was missing every required
copilot.server.env.* secret (AGENT_API_DB_ENCRYPTION_KEY, INTERNAL_API_SECRET,
LICENSE_KEY, SIM_BASE_URL, SIM_AGENT_API_KEY, REDIS_URL, one model key) plus
copilot.postgresql.auth.password. Pasting it as-is failed at template render.
values-external-secrets.yaml: NEXT_PUBLIC_APP_URL, BETTER_AUTH_URL, etc. were
declared under app.env / realtime.env. In ESO mode the chart-managed Secret
isn't rendered, so the validator (rightly) rejects keys in app.env that
aren't mapped under externalSecrets.remoteRefs. Moved non-secret URL/config
to envDefaults, which is inlined and not subject to the ESO mapping rule.
* polish(helm): configurable NetworkPolicy ingress peers + clearer API_ENCRYPTION_KEY comment
- networkPolicy.ingressFrom lets operators scope the ingress-controller
rule to a specific namespace/podSelector. Defaults to a single empty
peer (`- {}`), which is the explicit form of "any source" — same
effective behavior as the old `from: []` but unambiguous across CNIs.
To restrict, override with e.g.:
networkPolicy:
ingressFrom:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: ingress-nginx
- API_ENCRYPTION_KEY comment: drop the "must be exactly 64 hex
characters" phrasing that sat awkwardly next to `openssl rand -hex 32`.
The generation command already produces the required length.
* test(helm): add helm-unittest suites + CI workflow + ci values matrix
- 7 helm-unittest suites covering smoke, validators, secret modes,
envDefaults secret-mode-aware inlining (round-9 regression net),
chart-computed env keys (round-8 regression net), NetworkPolicy
shape, and PDB/HPA conditional rendering (38 tests, ~265ms).
- ci/*.yaml render fixtures for default, production, existingSecret,
ESO, and external-db install modes.
- GitHub Actions workflow runs helm lint --strict, helm unittest,
helm template across the ci matrix, and kubeconform validation
against Kubernetes 1.30 schemas.
- CONTRIBUTING.md documents how to run the same gates locally.
* test(helm): add helm test hook + kind apiserver dry-run in CI
- New templates/tests/test-connection.yaml renders a Pod with
helm.sh/hook=test that wgets the app Service (and realtime when
enabled). Lets users run `helm test <release>` after install for
a real in-cluster connectivity check. Restricted PSS context.
- tests.* values block (image, timeoutSeconds, resources) is the
knob to disable or tune the probe; documented in values.schema.json.
- 3 helm-unittest tests cover the hook annotations, PSS context,
and tests.enabled=false skip path (41 tests total).
- New CI job spins up a kind v1.30 cluster and runs
`kubectl apply --dry-run=server` against the rendered manifests
for the CRD-free ci fixtures (default / existing-secret /
external-db). Catches admission and validation issues the static
kubeconform schema check can't see.
* chore(helm): remove pre-1.0.0 upgrade fluff + tighten .helmignore
This is the 1.0.0 release of the chart — there is no pre-1.0.0
predecessor for users to upgrade from, so all of the dedicated upgrade
narration was hypothetical.
- Drop the 'Upgrading from a pre-1.0.0 build' README section and the
matching troubleshooting entry.
- Drop the .Release.IsUpgrade block from NOTES.txt: items 5 (StatefulSet
orphan-delete), 6 (INTERNAL_API_SECRET 'new in 1.0.0'), 7
(networkPolicy.egress shape change). Each described a migration off a
chart version that never shipped.
- Delete references/upgrade-pre-1.0.0.md and remove the corresponding
pointers from SKILL.md.
- Anchor .helmignore patterns to chart root so /tests/ (unit suites)
and /examples/ are dropped from the packaged tarball without also
dropping templates/tests/ (the helm test hook).
* chore(helm): drop CI workflow + ci/ fixtures + CONTRIBUTING.md
The helm-unittest suites in helm/sim/tests/ and the helm test hook
in helm/sim/templates/tests/ stay — those are chart-internal quality
scaffolding, not CI. Removed:
- .github/workflows/helm-chart.yml
- helm/sim/ci/*.yaml (5 render fixtures used only by the workflow)
- helm/sim/CONTRIBUTING.md (mostly documented those gates)
- dead /ci/ and /CONTRIBUTING.md entries in .helmignore
* feat(helm): pod rollout on Secret change + topologySpreadConstraints
- Add checksum/secret pod annotations on app, realtime, and copilot
Deployments (plus checksum/config on app when branding ConfigMap is
enabled). Closes the long-standing footgun where 'helm upgrade' with
a changed Secret would silently leave pods running the old values
until a manual rollout restart.
- New top-level topologySpreadConstraints value (and sim.topologySpreadConstraints
helper) applied to app and realtime Deployments. Mirrors how affinity
and tolerations are plumbed; users supply their own labelSelector
to mirror Bitnami convention.
- 5 helm-unittest cases cover the checksum annotations and topology
spread rendering (46 tests total).
* fix(helm): drop empty-string shadowing in app/realtime env merge
Sprig 'merge' treats "" as a real value, so a default-empty
app.env.BETTER_AUTH_URL would shadow a non-empty realtime.env override
and the URL would never reach the rendered Secret. Replace 'merge'
with an explicit two-pass overlay that filters empties before writing,
mirroring the same pattern already used in deployment-realtime.yaml's
existingSecret block.
Adds two regression tests: realtime.env-only value reaches the Secret
when app.env is empty, and app.env still wins on collision when both
are non-empty (48 tests total).
* fix(helm): make topologySpreadConstraints per-component to match docstring
Greptile flagged that sim.topologySpreadConstraints helper docstring promised
per-component config (.Values.app, .Values.realtime, ...) but call sites
passed .Values, so any app.topologySpreadConstraints / realtime.topologySpreadConstraints
set by the user was silently dropped. The single global key also prevented
distinct app-vs-realtime spread rules.
Pass .Values.app / .Values.realtime to the helper at each call site; move
the top-level topologySpreadConstraints key into both component sections in
values.yaml. Adds a regression test that app constraints don't leak onto
the realtime pod.
* fix(helm): allow cron pods through app NetworkPolicy
Cursor flagged that when networkPolicy.enabled=true and cronjobs.enabled=true
(the recommended production config), the app NetworkPolicy only allowed
ingress from realtime and the ingress controller — silently blocking every
cron pod's HTTP call to /api/schedules/execute, webhook polls, etc. All 13
default cronjobs would fail.
Tag cron pods with a stable simstudio.ai/component-group: cronjob label so
the app NetworkPolicy can allow them with a single rule (no per-job
enumeration). Rule is conditional on cronjobs.enabled. Adds positive and
negative regression tests.
|
||
|
|
1b94424b4b |
improvement(mothership): align markdown blockquote, img, em, del with design tokens (#4566)
* improvement(mothership): align markdown blockquote, img, em, del with design tokens * fix(mothership): correctly scope blockquote paragraph margin reset to first/last child * improvement(mothership): restore italic on blockquotes * fix(mothership): widen img component prop type to satisfy Streamdown Components |
||
|
|
9b333eaae9 | fix(mothership): fix tool hidden in ui (#4564) | ||
|
|
22b5a1e577 |
fix(tables): cmd+a always selects all cells on the table page (#4562)
* fix(tables): cmd+a always selects all cells on the table page * fix(tables): move preventDefault inside non-empty guard |
||
|
|
b7c34d4323 |
fix(file-viewer): prevent scroll jump to top during Mothership streaming (#4559)
* fix(file-viewer): prevent scroll jump to top during Mothership streaming - Fix root cause: MarkdownCheckboxCtx.Provider was conditionally rendered, causing the scroll container to unmount/remount when isStreaming flipped, resetting scrollTop to 0 on every stream start - Add useScrollAnchor hook with spacer element to preserve scroll position when streamed content temporarily shrinks the scroll container - Linger active session ID on complete to prevent streamingContent→undefined flicker between consecutive tool calls on the same file - Gate upsert activation on incoming session having renderable content - Fix shouldShowStreamingFilePanel to keep panel mounted during linger - Fix use-chat post-write navigation to work with lingered completed session - Fix useAutoScroll to check proximity before pinning to bottom on stream start * fix(file-viewer): preserve session linger in hydrate/upsert and clear spacer after stream * chore(file-viewer): trim verbose comments to match codebase style * fix(file-viewer): re-engage auto-scroll when user scrolls back to bottom * chore(file-viewer): update scroll-anchor tsdoc, remove test separators, add hydrate linger tests * fix(file-viewer): prevent false re-engage when spacer restoration triggers onScroll * refactor(file-viewer): extract shouldReengage as pure tested function Pulls the spacer-guard re-engage condition out of onScroll into an exported pure function so the false-re-engage invariant (spacer active → no re-engage) is covered by automated tests rather than relying on manual QA. Adds 8 unit tests for shouldReengage alongside the existing 15 for computeSpacerShortage. * chore(file-viewer): trim verbose inline comments in use-scroll-anchor * chore(file-viewer): final comment cleanup before merge |
||
|
|
39f74aa353 |
feat(data-drains): add GCS, Azure Blob, BigQuery, Snowflake, and Datadog destinations (#4552)
* feat(data-drains): add GCS, Azure Blob, BigQuery, Snowflake, and Datadog destinations
* fix(data-drains): address PR review comments
* fix(data-drains): extract sleepUntilAborted, honor abort across all destinations
* fix(data-drains): widen BigQuery projectId max and dedupe parseServiceAccount
* fix(data-drains): tighten GCS bucket contract and expose Azure endpointSuffix
* improvement(data-drains): extract normalizePrefix and buildObjectKey to shared utils
* fix(data-drains): retry BigQuery network errors; tighten Azure accountKey contract
- BigQuery insertAll now wraps the fetch in try/catch inside the retry loop so DNS failures, socket resets, and timeouts are retried with backoff instead of propagating immediately.
- Align azureBlobCredentialsBodySchema with the runtime schema (min 64 / max 120 / base64 regex) so obviously invalid keys are rejected at the API boundary rather than at drain-run time.
* improvement(data-drains): consolidate parseRetryAfter; add Datadog NDJSON line context
- Extract a single parseRetryAfter helper (capped at 30s, returns number | null) into lib/data-drains/destinations/utils.ts and remove the five local copies in bigquery, datadog, gcs, snowflake, and webhook.
- Datadog parseNdjson now wraps JSON.parse in try/catch and surfaces the failing line index, matching BigQuery's parser.
* fix(data-drains): correct Datadog size guard and Snowflake VARIANT limit
- Datadog payload guard now checks the uncompressed size against the 5 MB limit and the wire size against the 6 MB compressed limit, so gzip cannot smuggle an oversized body past the client-side check.
- Snowflake VARIANT limit is 16 MiB (16,777,216 bytes), not 16,000,000 bytes — small payloads between 16 MB and 16 MiB were being rejected unnecessarily.
- Drop the unused apiKey field on Datadog PostInput; the key is already embedded in the prepared request headers.
* improvement(data-drains): consolidate backoffWithJitter into shared utils
Datadog, GCS, and webhook each had byte-identical backoff helpers (BASE 500ms, MAX 30s, jitter ±20%, Retry-After floor). Lift the helper into lib/data-drains/destinations/utils.ts alongside parseRetryAfter and sleepUntilAborted, and drop the per-file copies and their BASE_BACKOFF_MS/MAX_BACKOFF_MS constants.
* fix(data-drains): align destinations with live provider specs
Audited every destination against live AWS/GCS/Azure/BigQuery/Snowflake/
Datadog/webhook docs and applied spec-correctness fixes:
- S3: reserved bucket prefix amzn-s3-demo-, suffixes --x-s3/--table-s3;
metadata byte formula excludes x-amz-meta- prefix per AWS spec
- GCS: reject -./.- adjacency; UTF-8 prefix cap; forbid .well-known/
acme-challenge/ prefix; ASCII-only x-goog-meta-* enforcement
- BigQuery: insertId is 128 chars (not bytes); split DATASET_RE (ASCII)
and TABLE_RE (Unicode L/M/N + connectors); UTF-8 byte cap on tableId
- Snowflake: disambiguate org-account vs legacy locator account formats;
requestId+retry=true for idempotent retries; server-side timeout=600;
default column DATA uppercase to match unquoted canonical form
- Azure: endpoint suffix allowlist (4 sovereign clouds); accountKey
length(88) base64
- Webhook: url max(2048); CRLF/NUL rejection on bearer/secret/sig header
* fix(data-drains): address PR review on snowflake poll + shared NDJSON parsing
- snowflake pollStatement: per-attempt timeout via AbortSignal.any, retry on 429/5xx with Retry-After + jitter
- bigquery parseNdjson error messages now 1-indexed
- consolidate parseNdjson variants into shared parseNdjsonLines/parseNdjsonObjects in utils
* fix(data-drains): per-attempt fetch timeouts in gcs/bigquery, snowflake poll double-sleep
- gcs.fetchWithRetry + bigquery.postInsertAll now use AbortSignal.any with a per-attempt timeout so a hung TCP connection cannot stall the drain
- snowflake.pollStatement skips the next interval sleep when it just slept for retry backoff
* fix(data-drains): bigquery probe timeout + jittered retries, align Snowflake column default UI/docs
- bigquery test() probe now uses AbortSignal.any + per-attempt timeout
- bigquery insertAll retry switches to backoffWithJitter for thundering-herd avoidance
- Snowflake column placeholder + docs say DATA (uppercase) to match the code default
* fix(data-drains): mirror webhook signingSecret min length in form gate
isComplete now requires signingSecret >= 32 to match the contract/runtime
schema so the Save button can't enable on a value that will fail server-side.
* fix(data-drains): validate JSON client-side for Snowflake before binding
Switch Snowflake to parseNdjsonObjects so malformed rows are caught locally
with 1-indexed line numbers instead of failing the whole INSERT server-side.
Re-stringify each parsed object before binding to PARSE_JSON(?).
Drop the now-unused parseNdjsonLines helper.
* fix(data-drains): cross-cutting audit pass against live provider docs
- Azure: bound retryOptions on BlobServiceClient (SDK default tryTimeoutInMs is per-try unbounded; cap at 30s x 5 tries)
- Webhook contract: mirror runtime — signingSecret.max(512), bearerToken.max(4096) + CRLF/NUL refine, signatureHeader charset + CRLF/NUL refine
- S3 (lib + contract): reject bucket names with dash adjacent to dot; require https:// endpoint at the schema layer
- Snowflake: bind original NDJSON line bytes (re-stringifying a JSON.parse'd value loses bigint precision beyond 2^53-1); check pollStatement 200 body for the SQL error envelope (sqlState/code)
- Datadog: entry builder writes defaults first then user attrs then forced ddtags/message so user rows can't clobber routing fields; validate config.tags as comma-separated key:value pairs
- registry.tsx: tighten isComplete predicates to mirror contract minimums (GCS bucket >= 3, Azure containerName >= 3 / accountKey === 88, BigQuery projectId >= 6, Snowflake account >= 3)
* fix(data-drains): force ddsource/service overrides on Datadog entries
Previous fix placed ddsource/service before ...attrs, leaving them clobberable
by a user row field. Per Datadog docs, service + ddsource pick the processing
pipeline, so a drain's routing config must not be overridable per-row. Spread
attrs first, then force all four reserved fields (ddsource, service, ddtags,
message).
* fix(data-drains): preserve row-distinguishing index when BigQuery insertId overflows
Truncating from the left dropped the index suffix, so any overflow would
collapse all rows in a chunk to the same insertId and BigQuery would silently
dedupe them. Path is unreachable today (UUIDs keep raw ~85 chars), but the
overflow branch is now correct: hash the prefix, keep the index intact.
* fix(data-drains): refresh GCS token per retry, tighten Azure key regex
- gcs: rebuild Authorization header per attempt via buildHeaders so token
refresh from google-auth-library kicks in if a 5xx retry crosses the
hour-long token lifetime
- azure_blob: pin account-key regex to {0,2} trailing '=' (base64 of 64
bytes = exactly 88 chars with up to two '=' pad chars)
* fix(data-drains): address bugbot review of 6336948f6
- gcs: allow 1-char dot-separated bucket components (e.g. "a.bucket")
to match GCS naming rules — overall name is 3-63 (or up to 222 with
dots), but per-component minimum is 1 per Google's spec
- bigquery: drain the 401 response body before re-issuing the request
with a refreshed token so undici can return the socket to the
keep-alive pool
- snowflake: hoist getJwt() above the perAttempt timer in
executeStatement so JWT signing doesn't eat the network budget
(matches the order already used in pollStatement)
* fix(data-drains): allow org-account Snowflake identifier with region suffix
The account validation rejected `<orgname>-<acctname>.<region>.<cloud>`
because `ACCOUNT_LOCATOR_RE`'s first segment forbade hyphens, while
`ACCOUNT_ORG_RE` forbade dots. `normalizeAccountForJwt` already handles
this composite form. Widen the first segment of `ACCOUNT_LOCATOR_RE` to
allow hyphens so the boundary contract and the runtime schema accept
what the JWT layer was already designed to process.
* fix(data-drains): drain retryable response bodies in datadog/gcs loops
Mirrors the bigquery 401 fix. Without consuming the body before
sleeping, undici can't return the socket to the keep-alive pool, so
each retry leaks a TCP connection instead of reusing it.
* fix(data-drains): drain snowflake poll bodies on 202 and retryable status
Mirrors the bigquery/datadog/gcs drains. Long async statements can poll
many times against the same connection; without consuming the body
undici can't return the socket to the keep-alive pool, so each iteration
leaks a connection until GC.
* fix(data-drains): consume success bodies; check Snowflake sqlState on 200
- gcs: drain the body on success paths so undici can return the socket
to the keep-alive pool
- snowflake: drain the body on synchronous 200 OK and run the same
sqlState envelope check pollStatement already does — otherwise a
statement-level failure that completes synchronously would silently
return success
* fix(data-drains): drain datadog and bigquery probe success bodies
Same undici keep-alive issue as the prior fixes: postWithRetries
returned the Response on success without draining (callers only read
headers); the BigQuery `test()` probe returned without consuming the
body. Both now drain before returning.
* chore(data-drains): regenerate enum migration as 0206 after staging rebase
* fix(data-drains): cap snowflake poll retries; tighten datadog tags min length
|
||
|
|
0b2cfaf7f7 |
feat(mothership): add superuser env selection (#4558)
* feat(table): live cell updates via SSE + per-table event buffer
Replaces the polling-based row refetch with a push-based SSE stream that
patches the React Query cache directly as cell-state events arrive.
Architecture:
- New per-table event buffer in apps/sim/lib/table/events.ts. Redis sorted-set
with monotonic eventId, 1h TTL, 5000-event cap, in-memory fallback. Modeled
after apps/sim/lib/execution/event-buffer.ts but stripped of complexity
tables don't need (no per-execution lifecycle, no id-batching, no write
queue serialization). ~150 lines instead of 700.
- writeWorkflowGroupState appends a fat event after each successful 'wrote'.
Status transitions carry executionId + jobId; terminal/partial transitions
also include the new output values inline so the client can patch row data
without a follow-up refetch.
- New SSE route at /api/table/[tableId]/events/stream?from=<lastEventId>.
Replays from buffer on connect, polls at 500ms (mirrors workflow execution
stream), heartbeat every 15s, signals 'pruned' if the caller fell off the
back of the buffer.
- Client hook useTableEventStream subscribes via EventSource. Reconnect-resume
with last-seen eventId. On 'pruned', invalidates the rows query and resumes
from the new earliest. Cache patches walk every cached query under
rowsRoot(tableId) so filter/sort variants all stay live.
- Removes refetchInterval from useTableRows and the per-page polling effect
from useInfiniteTableRows. React Query's refetchOnWindowFocus +
refetchOnReconnect cover the durability gap if any push is dropped.
Out of scope:
- Bulk-cancel events (cancellation path is being redesigned separately).
- Generalizing the workflow event-buffer module to a shared primitive (defer
until a third use case appears; for now the table buffer is the simpler
cousin of the workflow one).
* fix(table): drop run-mutation refetch so SSE patches aren't overwritten
useRunColumn.onSettled was canceling in-flight queries and invalidating the
rows query — leftover behavior from the polling era. With the SSE stream
now keeping the cache live via incremental patches, this refetch races the
stream and snaps the cache back to whatever DB shows at the refetch moment,
which can lag the just-arrived queued/running events. Cells appeared stuck
on the optimistic 'pending' even though the SSE was delivering the real
transitions.
* chore(table): simplify SSE plumbing — reuse helpers, drop dead polling code
- Reuse snapshotAndMutateRows for SSE cache patches instead of reimplementing
the page-walk + cache-shape detection. Adds a {cancelInFlight: false} opt
for the SSE caller (mutations still cancel as before).
- Drop client-side type duplication in use-table-event-stream — import
TableEvent and TableEventEntry from lib/table/events directly.
- Drop the now-dead mergePagePreservingIdentity + rowEqual from tables.ts;
their only caller was the polling effect that was removed earlier.
- Drop the defensive try/catch around appendTableEvent in cell-write — the
function is documented as never-throwing (returns null on failure).
- Combine INCR + ZADD into one Lua eval in events.ts. Halves Redis RTT per
cell-write. Lua returns the new eventId; the script splices it into the
pre-built entry JSON.
- Trim refs to plain let bindings inside the effect; trim stale
comments referencing the old polling implementation.
* fix(table): address PR review on SSE buffer
- TTL-expiry silent miss: when all keys expire, hgetall(meta) returns empty
so earliestEventId is undefined and the prune branch was skipped. Reconnect
with non-zero afterEventId now checks the seq counter — its absence (TTL
expired) signals pruned so the client refetches. Memory fallback mirrors.
- Unbounded ZRANGEBYSCORE: cap reads at TABLE_EVENT_READ_CHUNK = 500 events
per call. The route's 500ms poll loop drains chunks across ticks instead of
flushing 5000 entries (multi-MB) in one tick after a long disconnect.
- Pruned handler closes EventSource client-side: server-side close was firing
onerror and routing through the 500ms backoff path. Now we close
proactively, reset the reconnect attempt counter, and reconnect immediately
from the new earliest.
* Cross env copilot
* Force deploy
* Run migration
* Updates
* Fix migration
* Redeploy
* Make dev db push
* restore old migs
* Cross env copilot
* Add custom tools, skills, mcps to mothership
* Update migration
* Fix migs
* UPdate
* Fix types
* Fix
---------
Co-authored-by: Theodore Li <theo@sim.ai>
|
||
|
|
8774f5c341 |
fix(deps): patch next-mdx-remote and opentelemetry CVEs (#4557)
- bump next-mdx-remote 5.0.0 → 6.0.0 (GHSA-g4xw-jxrg-5f6m / CVE-2026-0969, arbitrary code execution in MDX serialize) - bump @opentelemetry/sdk-node and exporter-trace-otlp-http 0.200.0 → 0.217.0 (GHSA-q7rr-3cgh-j5r3 / CVE-2026-44902, Prometheus exporter DoS) - align @opentelemetry/sdk-trace-base, sdk-trace-node, resources to ^2.7.0 to keep all @opentelemetry/* packages on a single core@2.7.1 instance |
||
|
|
d895e0efee |
fix(security): authorize MCP subagent IDs, oauth workspace, credential admin demotion (#4551)
* fix(security): authorize MCP subagent IDs, oauth workspace, credential admin demotion
- handleSubagentToolCall and handleDirectToolCall now authorize user-supplied
workflowId/workspaceId via authorizeWorkflowByWorkspacePermission /
ensureWorkspaceAccess before forwarding downstream; resolvedWorkspaceId is
derived from the authorized workflow record instead of trusted from the body
- executeOAuthGetAuthLink verifies caller membership (write level) on the
target workspaceId before generating the OAuth link or writing
pendingCredentialDraft, closing the cross-workspace credential injection path
- POST /api/credentials/[id]/members wraps role updates in a transaction that
counts active admins and rejects demotion of the last admin (mirrors the
existing DELETE guard in the same file)
- GET /api/credentials/[id]/members returns uniform 404 for both missing and
inaccessible credentials to remove the existence oracle
* fix(security): address PR review — active-status guard, FOR UPDATE locks, workspaceId propagation
- credentials/members POST: add `current.status === 'active'` check to the
last-admin demotion guard so re-inviting a revoked admin as a non-admin role
no longer incorrectly hits the "Cannot demote the last admin" path
- credentials/members POST+DELETE: add `.for('update')` to the active-admin
count SELECT inside both transactions to serialize concurrent demotions and
eliminate the admin-count TOCTOU race under Postgres READ COMMITTED
- credentials/members POST: also lock the member row itself with `.for('update')`
so the role+status read and the subsequent UPDATE are atomic
- mcp/copilot handleDirectToolCall: thread the DB-verified workspaceId from the
authorization result into prepareExecutionContext instead of relying on
user-supplied args
- oauth handler: fix error message to mention both workspaceId and userId when
either is missing from the execution context
|
||
|
|
3adbde404b |
fix(oauth): persist rotated Microsoft refresh tokens (#4554)
* fix(oauth): persist rotated Microsoft refresh tokens Microsoft Entra rotates refresh tokens on every refresh and expects clients to replace the stored token with the new one. The Microsoft provider config was missing supportsRefreshTokenRotation, so the rotated refresh_token returned by Azure AD was silently discarded and the original token from initial OAuth connect was reused indefinitely — causing periodic 'Failed to refresh access token' errors for Excel, Teams, Outlook, OneDrive, SharePoint, Planner, AD, and Dataverse integrations. * test(oauth): cover hyphenated Microsoft service IDs in rotation test |
||
|
|
c47777fa33 |
fix(security): close cross-tenant IDOR gaps in OAuth credential and execution auth (#4549)
* fix(security): close IDOR gaps in OAuth credential and execution authorization Routes that called resolveOAuthAccountId followed by a conditional workspace permission check (only run when workspaceId was set) silently skipped all ownership validation on the legacy account-ID fallback path. Any authenticated user could supply a raw account.id to access another tenant's OAuth credentials. - Replace resolveOAuthAccountId + conditional perm check with authorizeCredentialUse in: auth/oauth/wealthbox/item, tools/gmail/label, tools/onedrive/files, tools/onedrive/folder, tools/outlook/folders, tools/wealthbox/item (routes 1, 3-7) - Add authorizeCredentialUse ownership gate before resolveVertexCredential in providers/route.ts (route 2) - Add verifyFileAccess check on the user-supplied file key before downloadFileFromStorage in tools/wordpress/upload (route 8) - Add workflowId param to PauseResumeManager methods (enqueueOrStartResume, beginPausedCancellation, completePausedCancellation, blockQueuedResumesForCancellation, clearPausedCancellationIntent, getPausedCancellationStatus, processQueuedResumes) and filter all pausedExecutions lookups by workflowId so callers cannot act on another tenant's paused execution by supplying a foreign executionId (route 9) - Update all call sites (cancel, resume, poll routes) to pass workflowId * fix(security): close verifyFileAccess bypass, thread workflowId to processQueuedResumes, fix log level - Fail closed in WordPress upload when userFile.key is present but authResult.userId is absent, preventing silent bypass of ownership check via JWT fallback path - Thread workflowId into processQueuedResumes in the async resume error-recovery path and in pause-persistence.ts to close residual cross-tenant gap - Change logger.error to logger.warn for credential access denial in OneDrive folder route to match all other routes in this PR * fix(security): thread workflowId through all processQueuedResumes call sites Closes residual cross-tenant IDOR gap where processQueuedResumes was called without a workflowId scope in persistPauseResult, startResumeExecution (success and error paths), and clearPausedCancellationIntent. workflowId was already in scope at each site — this wires it through to the existing optional parameter. * fix(security): remove any types, drop extraneous comments, normalize caught errors - catch (error: any) → catch (error) + toError(error).message in resume and cancel routes - Remove what-not-why inline comments from wordpress upload and onedrive/files routes - Remove redundant debug-only item breakdown log and the file-IDs log in onedrive/files - Trim extraneous DAG-edge comments from updateSnapshotAfterResume in HITL manager * fix: use logger.warn for credential access denial in outlook folders route * fix(security): make workflowId required in all HITL pause/resume methods All 7 method signatures (processQueuedResumes, enqueueOrStartResume, beginPausedCancellation, completePausedCancellation, blockQueuedResumesForCancellation, clearPausedCancellationIntent, getPausedCancellationStatus) previously accepted workflowId as optional. Every call site already supplies it — making it required closes the vulnerability at the type level so future callers cannot accidentally omit tenant scoping and silently fall back to an unscoped DB query. * fix(security): thread workflowId through internal HITL cancellation calls and remove dead branches in credential-access * fix(security): harden workflowId scoping and file key guard Replace falsy workflowId checks in PauseResumeManager (all methods now unconditionally apply the workflowId WHERE clause, preventing empty-string bypass). Flip WordPress upload file guard from truthy key check to explicit non-empty validation so key:"" fails closed with a 404 instead of silently skipping access control. |
||
|
|
badfeae432 |
fix(security): SSRF fixes (#4548)
* fix: address SSRF and token-leakage security vulnerabilities
- Azure TTS SSRF: validate region against /^[a-z][a-z0-9-]{1,30}[a-z0-9]$/
in both the contract (tts.ts) and runtime guard in synthesizeWithAzure,
preventing user-supplied region from redirecting requests to arbitrary hosts
- HubSpot token in logs: remove fullResponse from logger.info call;
log only non-sensitive metadata (hub_id, hub_domain, user_id) instead
of the full introspection response which included the access token
- Wealthbox account takeover: replace hardcoded email with per-user identity
by fetching /v1/users/me; fall back to token-derived stable identifier
so distinct Wealthbox users no longer share the same email address
- Shopify SSRF: apply shopifyShopDomainSchema (.myshopify.com allowlist)
to shopDomain from cookie before using it to build the fetch URL
* fix(wealthbox): correct getUserInfo endpoint, auth header, and stable identity
- Bug 1: Change API endpoint from /v1/users/me to /v1/me (correct Wealthbox API path)
- Bug 2: Replace ACCESS_TOKEN header with Authorization: Bearer <token> (standard OAuth 2.0)
- Bug 3: Remove generateId() from returned id (was non-deterministic, caused duplicate accounts);
use refresh token (stable, long-lived) instead of access token (rotates every ~2 hours)
as the hash source for the fallback identity; return null if no token is available
* fix(security): hash wealthbox fallback token identity, guard undefined userId
- Replace base64 encoding with SHA-256 hash for fallback token-derived identity
so raw token bytes are never stored in the DB
- Return null early when Wealthbox API response lacks an id field to prevent
all such users colliding on the wealthbox-undefined account
* fix(auth): replace stale wealthbox userInfoUrl placeholder with actual endpoint
The dummy URL comment was rendered obsolete when getUserInfo was updated
to fetch from api.crmworkspace.com/v1/me. Align userInfoUrl with the real
endpoint used in the getUserInfo implementation.
* fix(auth): append generateId() suffix to Wealthbox account IDs to match codebase pattern
All other providers use `${stableId}-${generateId()}` so the account.create.after
hook can strip the UUID suffix, find stale sibling rows, and migrate credential FKs.
Without the suffix the migration logic is skipped and reconnections would hit
duplicate key conflicts instead of gracefully updating credentials.
|