* fix(wait): poll partially_resumed rows so chained waits resume
The chained-pause flow leaves a row in 'partially_resumed' status (wait1 done, wait2 still waiting). The poll's WHERE filter only matched 'paused', so wait2 was never picked up. Include 'partially_resumed' in the filter.
* feat(wait): make in-process threshold env-overridable for local testing
Adds WAIT_INPROCESS_MAX_MS env var (default 300000ms = 5 min). Lower it locally (e.g. 5000) to exercise the suspend/cron-resume path with short waits.
* feat(wait): add Suspend Workflow toggle, restore 5-min in-process default
* fix(wait): address bot review — setNextResumeAt + suspend unit default
- setNextResumeAt now matches paused OR partially_resumed; otherwise the
cron poller can't null nextResumeAt after dispatching a chained-wait
row, so it keeps reappearing in every poll batch until execution ends
(flagged by both greptile and bugbot)
- Suspend mode now defaults missing timeUnitLong to 'minutes' instead of
falling back to 'seconds' and immediately erroring (flagged by bugbot)
* improvement(wait): hint the wait-amount cap on the input
Restores the pre-#4331 description on the Wait Amount field so the limit is visible before submit instead of only at runtime. Mentions both the 5 min default and the 30 day cap with Suspend Workflow.
* refactor(wait): rename Suspend Workflow toggle to Async
* fix(wait): reword cap errors to say 'async mode'
* feat(workflows): add GET /workflows/[id]/executions/[executionId] status endpoint
Normalized status (pending|running|paused|completed|failed|cancelled)
across workflowExecutionLogs and pausedExecutions in a single response.
Surfaces paused-state details (resumeAt, pauseKind, blockedOnBlockId)
when a row exists in pausedExecutions, and the error string for failed
runs. finalOutput is opt-in via ?includeOutput=true.
* feat(workflows): support ?selectedOutputs= on execution status endpoint
Returns per-block outputs filtered by selectedOutputs paths (same
shape as the execute endpoint). Reads from executionData.traceSpans,
walks children recursively, and resolves dot-paths into each block's
output. Bare blockId returns the full output.
* fix(wait): drop dead tooltip prop on Async switch
Switch sub-blocks return null from renderLabel (sub-block.tsx:238), so
the tooltip never reached the user. The trade-off explanation already
lives in longDescription and bestPractices. Flagged by bugbot.
* docs(api): document GET /workflows/[id]/executions/[executionId]
Adds the WorkflowExecutionStatus schema and the getWorkflowExecution
operation to the OpenAPI spec, including completed/paused/failed
response examples and the includeOutput + selectedOutputs query params.
Registers the page in the Workflows section of the API reference.
* docs(api): tighten getWorkflowExecution descriptions
* improvement(redis): strip idempotency body and cap mothership stream zsets
* chore(redis): trim verbose comments on idempotency body-strip
* test(buffer): pin exact ZREMRANGEBYRANK stop arg
Pinning -5_001 (= -(DEFAULT_EVENT_LIMIT) - 1) so the off-by-one
boundary is directly validated; expect.any(Number) would have passed
a wrong formula like -eventLimit.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* chore(utils): migrate to shared random/ID utilities and add enforcement linting
- Replace all Math.random(), crypto.randomUUID(), crypto.randomBytes(), nanoid, and uuid usages with shared @sim/utils/random and @sim/utils/id helpers across 72 files
- Add new @sim/utils exports: deepClone, omit, filterUndefined (object), truncate (string), backoffWithJitter, parseRetryAfter (retry), getErrorMessage (errors)
- Sweep all getErrorMessage, sleep, deepClone callsites across 500+ files to use shared utilities
- Add Biome noRestrictedImports rule to catch nanoid, uuid, and crypto named imports at lint time
- Add scripts/check-utils-enforcement.ts to catch Math.random and crypto.* global property access
- Add check:utils script to package.json
* chore(utils): replace deepClone wrapper with structuredClone built-in
deepClone() was a one-line wrapper around structuredClone(), which is
universally available in Node 17+ and all modern browsers. Removing the
abstraction reduces indirection and means contributors don't need to
learn a project-specific name for a well-known built-in.
- Remove deepClone from packages/utils/src/object.ts and index.ts
- Replace all 17 call sites with structuredClone() directly
- Update check:utils script suggestion text
- Update CLAUDE.md and global.md docs
* fix(utils): add missing biome noRestrictedImports rule and correct truncate docs
- Add noRestrictedImports to biome.json under style — bans nanoid and uuid
package imports at lint time (crypto.randomUUID/randomBytes are caught by
the check:utils grep script which handles global property access)
- Correct truncate() TSDoc and parameter name: sliceLength makes it clear
that total output length is sliceLength + suffix.length, matching the
behavior all callers were already written to expect
* fix(utils): add missing getErrorMessage imports at 4 call sites
The sweep agents added getErrorMessage calls without the corresponding
import in 4 files, causing test failures. Added the missing imports.
* fix(utils): fix build errors from getErrorMessage sweep and retry.ts Turbopack issue
- Fix retry.ts cross-file import: Turbopack cannot resolve './random.js' for
internal package imports; inline the jitter crypto call directly
- Add missing getErrorMessage imports to 32 files where the sweep added calls
without the corresponding import (caught by type-check and test runs)
- Remove accidental getErrorMessage import from crowdstrike/query/route.ts
which has its own domain-specific getErrorMessage for parsing CrowdStrike's
JSON error format
- Fix use-sub-block-value.ts type error from structuredClone narrowing:
add 'as T' cast at emitValue callsite (safe — valueCopy is always a
structural copy of newValue)
* fix(tools): use toError in crowdstrike catch block instead of local getErrorMessage
The catch block was calling the local getErrorMessage function which
parses CrowdStrike API JSON responses, not JavaScript Error objects.
Use toError(error).message to correctly extract the message from a
caught value in this context.
* improvement(providers): align attachment dispatch to vendor SDK types
Post-merge audit of #4610 surfaced three follow-ups:
1. xAI Grok vision was blocked. Grok runs through the OpenAI-compatible
chat-completions endpoint, so removing xAI from UNSUPPORTED_FILE_PROVIDERS
and routing it through the image-only branch restores image attachments
on vision models.
2. Azure OpenAI chat-completions deployments blocked any file attachment.
Added a per-message image_url parts path; documents still require the
Responses API endpoint and throw a clear, actionable error.
3. Wire shapes were loosely typed (Record<string, unknown> arrays).
Replaced with `satisfies` clauses against each vendor SDK union at every
push site: OpenAI Responses/Chat, Anthropic ContentBlockParam, Gemini
Part, Bedrock ContentBlock members. AnthropicImageMediaType now derives
from Base64ImageSource['media_type'] so it tracks SDK updates.
Also collapsed the validation cascade into an exhaustive switch with
`never` enforcement, and dropped the redundant per-provider
formatMessagesForProvider call from xai/index.ts (providers/index.ts
already runs the dispatcher centrally).
* fix(providers): restore getProviderAttachmentMaxBytes export and xAI message dispatch
- Restore `getProviderAttachmentMaxBytes` — still consumed by agent-handler.ts
for per-provider attachment size limits in file hydration
- Restore `formatMessagesForProvider(allMessages, 'xai')` — providers/index.ts
does NOT dispatch centrally on this branch; each OpenAI-compat provider
formats its own messages. Without it, xAI Grok vision drops image attachments
* fix(providers): tighten ResponsesInputItem content type to SDK ResponseInputContent
Build fix: buildOpenAIMessageContent returns ResponseInputContent[] which
isn't assignable to Record<string, unknown>[] (ResponseInputText lacks an
index signature). Align the type to the SDK shape.
- Guard completeWithError against overwriting a cancelled execution status — cancel route writes cancelled to DB optimistically, but a block error racing the 500ms Redis check could finalize with failed before the engine detects cancellation
- Add tests covering the guard: cancelled DB status skips the write, non-cancelled proceeds normally, DB failure falls through to cost-only fallback, and subsequent attempts are deduped after guard marks session complete
- Move ImpersonationBanner from workspace root into components/ folder
* v0.6.29: login improvements, posthog telemetry (#4026)
* feat(posthog): Add tracking on mothership abort (#4023)
Co-authored-by: Theodore Li <theo@sim.ai>
* fix(login): fix captcha headers for manual login (#4025)
* fix(signup): fix turnstile key loading
* fix(login): fix captcha header passing
* Catch user already exists, remove login form captcha
* feat(files): folders + vfs update
* address comments
* address comments
* cleanup unnused code
* address comments
* perf improvements
* address next set
* cycle detect
* error handling
* path improvements
* cleanup, best practices
* react query best practices: targeted invalidation, optimistic updates, key factory hierarchy
- Add workspaceLists(workspaceId) intermediate key level to both workspaceFilesKeys
and workspaceFileFolderKeys so invalidation targets only the affected workspace
instead of all workspaces
- Replace all lists() invalidation calls with workspaceLists(workspaceId) across
every mutation (upload, rename, delete, restore, update content, folder mutations)
- Add optimistic updates to useRenameWorkspaceFile and useUpdateWorkspaceFileFolder
with onMutate snapshot, onError rollback, onSettled reconciliation
- Move storage key into the content() factory as optional param so query keys are
always built through the factory (useWorkspaceFileContent, useWorkspaceFileBinary)
- Fix AnimatePresence wrapping in FilesActionBar so exit animation fires on deselect
- Fix ResourceColGroup to use percentage weights instead of pixel widths to prevent
horizontal scroll on narrow viewports
* add shift-click range selection and selection-aware context menu for files
- Extend SelectableConfig.onSelectRow with optional shiftKey param; DataRow captures shiftKey before onCheckedChange fires via a ref so the Radix Checkbox interaction chain stays intact
- Implement shift-click range selection in files.tsx using lastSelectedIndexRef; tracks last-selected index in visibleRowIds to compute the range
- Reset lastSelectedIndexRef on deselect and select-all
- Add selectedCount prop to FileRowContextMenu; hide Open and Rename when multiple items are selected, show "Delete N items" / "Download N items" labels in multi-select mode
* add Move submenu to file context menu and fix shift-click anchor update
- Add nested Move submenu to FileRowContextMenu using DropdownMenuSub/SubTrigger/SubContent; shows available folders filtered by selection, converts '__root__' -> null for moving to the root level
- Add handleContextMenuMove in files.tsx that calls moveItems.mutateAsync directly (no modal) and clears selection on success
- Fix shift-click range selection: update lastSelectedIndexRef after range select so chained shift-clicks extend from the new anchor point correctly
* fix move submenu: use folder names with tree-ordered indentation instead of stale paths
- Compute folder depth from parentId chain client-side (avoids stale server-computed path field)
- Tree-order folders so parents appear before their children, sorted by sortOrder then name
- Show folder.name instead of folder.path so optimistic renames are reflected immediately
- Indent each folder by depth * 12px in the submenu so po/shit renders as 'shit' indented under 'po'
- MoveOption gains optional depth field; contextMenuMoveOptions is a separate memo from moveFolderOptions (modal keeps its existing path-label behavior)
* fix shift-click anchor drift and remove dead stopPropagation constant
- Remove dead stopPropagation const in resource.tsx (replaced by handleSelectRowClick)
- Reset lastSelectedIndexRef when visibleRowIds changes so search/filter/folder navigation doesn't leave a stale anchor that produces wrong ranges on the next shift-click
- Update lastSelectedIndexRef in handleRowContextMenu when right-clicking resets selection to a single item, so the anchor matches the newly-selected row
- Add visibleRowIds to handleRowContextMenu deps (now reads it to compute anchor index)
- Remove moveItems.mutateAsync from handleContextMenuMove deps per project convention (.mutateAsync is stable in TanStack v5)
* complete workspace files feature: audit logs, posthog events, folder restore, empty state, keyboard shortcuts, storage indicator, breadcrumb rename
- Audit + PostHog: wire file_renamed, file_deleted, file_moved, file_bulk_deleted, folder_created, folder_renamed, folder_deleted, folder_moved events to all file/folder API routes
- Add AuditAction.FOLDER_UPDATED, FILE_MOVED, FOLDER_MOVED to audit types
- Folder restore: server function, contract, API route (POST /files/folders/[folderId]/restore), hook (useRestoreWorkspaceFileFolder), Recently Deleted integration with new File Folders tab
- Empty state: contextual emptyMessage passed to <Resource> based on search/filters/folder context
- Keyboard shortcuts: Delete/Backspace deletes selection, Escape deselects, Cmd+A selects all (list view only, input-aware guard)
- Storage indicator: useStorageInfo drives compact "used / limit" display in file list header via leadingActions
- Breadcrumb rename: current folder breadcrumb gains Rename dropdown + inline editing via breadcrumbRename (useInlineRename)
- Resource: thread leadingActions prop from ResourceProps to ResourceHeader
* cleanup: accessibility, emcn design tokens, react best practices across workspace UI
- Add sr-only ModalDescription to dialogs/modals for accessibility
- Replace hardcoded colors and z-indices with design token CSS variables
- Apply emcn design review fixes across tables, knowledge, logs, settings, workflows
* fix audit and posthog: FOLDER_RESTORED action on restore, fire folder_moved event separately from file_moved
* sidebar: add Files section with nested folder tree; polish move UX and cleanup
- Files section in sidebar shows folder/file tree with expand/collapse,
matching Workflows section structure; collapsed sidebar shows flyout menu
- Move action bar now uses nested DropdownMenuSub tree instead of flat modal
- Context menu and action bar share renderMoveOption from move-options.tsx
- FolderInput added to emcn icons barrel; all FolderInput imports migrated
- Drag ghost uses CSS vars (--border, --shadow-medium, --z-toast)
- Selection pruning converted from useEffect to render-time comparison
- Keyboard listener stabilized with handleBulkDeleteRef pattern
- toError() used consistently in restore and move route handlers
* remove Files section from sidebar
* restore Files nav item in sidebar workspace section
* fix infinite re-render on files page - revert selection pruning to useEffect
* add filefolder resource type for ingesting workspace file folders
* export filefolder tree types; add toast feedback for file/folder mutations
* regenerate migration as 0208 after rebase onto staging
* add workspaceFileFolder to schema mock
* add FILE_MOVED, FOLDER_MOVED, FOLDER_UPDATED to audit mock
* add filefolder ChatContext kind and wire through schema and resolver
* add filefolder to AgentContextType
* add filefolder to chat context kind registry; fix resolver to use workspaceFiles table
* add .deepsec to gitignore
* cleanup: effect, emcn tokens, mutation error handling
- Replace selection-pruning useEffect with inline state adjustment during render
- Fix drag overlay using invalid --accent HSL token → --brand-secondary; z-50 → z-[var(--z-dropdown)]
- Move static inline styles on context menu trigger div to className
- Add missing onError toast to useUpdateWorkspaceFileFolder, useRestoreWorkspaceFileFolder, useRestoreWorkspaceFile
* lint
* fix: remove duplicate handleCopilotStopGeneration from rebase
* feat(copilot): folder-aware file context in WORKSPACE.md
* feat(copilot): add move operation to file manage API
* fix(files): make targetFolder optional in move file contract
* perf(files): parallelize buffer fetches, fix N+1 folder queries, stabilize drag useMemo
- download route: fan out all fetchWorkspaceFileBuffer calls with Promise.all
before zip assembly so 100 files resolve in one round-trip instead of sequentially
- getWorkspaceFileFolder: replace per-ancestor SELECTs with a single workspace-wide
folder load + buildWorkspaceFileFolderPathMap, making depth irrelevant to query count
- ensureWorkspaceFileFolderPath: pre-load all workspace folders in one SELECT before
the segment loop; resolve existing segments from an in-memory map; only hit the DB
to CREATE missing segments; conflict retry path preserved and also updates the map
- files.tsx rowDragDropConfig: move activeDropTargetId into a ref so the useMemo
does not recompute on every drag-over event
* fix(files): remove files/ path stripping, fix stale path in optimistic update
- splitWorkspaceFilePath: remove the unconditional .replace(/^files\//, '')
that clobbered paths for files inside a folder literally named "files"
- useUpdateWorkspaceFileFolder: when a name update is in flight, recompute
the path field for the renamed folder (replace last segment) and propagate
the new prefix to all descendant folders so breadcrumbs stay correct
during the optimistic window
* fix(files): revert broken ref opt, clean 409 on restore, null parentId on orphaned restore
- files.tsx: revert the activeDropTargetId ref optimization — the ref doesn't
trigger re-renders so the drop-target highlight never updated during drag;
activeDropTargetId is back in state and in the rowDragDropConfig deps
- restore/route.ts: catch Postgres 23505 unique-constraint violation and
return a clean 409 instead of leaking the raw error as 400
- restoreWorkspaceFileFolder: check if the parent folder is still archived
before restoring; if it is, restore to root (parentId: null) so the folder
is never orphaned under an archived parent
* feat(search): show folder path for files in cmd-k modal, strip extraneous comments
- FileItem interface with folderPath?: string[] added to search modal utils
- MemoizedFileItem component renders folder breadcrumb identically to
MemoizedWorkflowItem — truncated path segments on the right with / separators
- FilesGroup rewritten as a dedicated memo component (was createIconGroup factory)
so it accepts FileItem[] and includes folderPath segments in the search value
- searchModalFiles in sidebar splits f.folderPath string into string[] segments
- search-modal.tsx typed to FileItem and includes folderPath in filterAndSort
- Remove self-explanatory "Phase 1" section label from download route
- Remove redundant TSDoc on the unique index in db schema
* fix(workspace-files): audit fixes — transaction, status codes, contract refinements, guards
* fix(vfs): pass folderPath separately so buildWorkspaceMd groups files correctly
* fix(types): narrow unknown fileInput with Record cast after object guard
* fix(routes): replace instanceof Error with toError() across new workspace file routes
* improvement(files): cleanup pass — remove unnecessary useCallbacks, consolidate emcn icon imports
- Remove useCallback from 5 drag-event handlers in DataRow (passed to native <tr> elements, no observer)
- Remove stable useCallback fns from 3 useMemo deps arrays in files.tsx (editingId/editValue remain)
- Merge all @/components/emcn/icons subpath imports into barrel (files.tsx, action-bar, file-row-context-menu)
* fix(files): apply activeSort to folders, reject drop onto current parent folder
- visibleFolders now respects activeSort column (name/updated/created) and direction
so folder ordering stays consistent with file ordering
- isInvalidDropTarget now returns true when all dragged items are already direct children
of the target folder, preventing a no-op move mutation
* fix breadcrumb
* add new tools to rename, create, delete folders
* move more ui actions into orchestration dir
* address comments
* fix params
* fix tests
* address comments
* improve error codes
* address comments
* address more nits
* fix mcp server error code
---------
Co-authored-by: Theodore Li <theodoreqili@gmail.com>
Co-authored-by: waleed <walif6@gmail.com>
* improvement(gmail): replace custom html-to-text regex with html-to-text library
Resolves 4 CodeQL alerts on htmlToPlainText (incomplete tag/entity handling,
unsafe regex backtracking). Delegates to the html-to-text npm package already
used by the outlook polling trigger and the mail/send route.
* improvement(gmail): match outlook selectors config, add nbsp/anchor tests
Aligns html-to-text options with apps/sim/lib/webhooks/polling/outlook.ts:
suppress anchor hrefs when identical to text, drop bare # anchors, skip
img/script/style content. Adds tests for nbsp preservation and anchor
behavior.
* fix(gmail): send emails as multipart/alternative so they render full-width
* fix(gmail): decode & last in htmlToPlainText to avoid double-decoding compound entities
* fix(gmail): encode body parts as base64 and decode numeric HTML entities
* docs(uploads): clarify QUOTA_EXEMPT_STORAGE_CONTEXTS logs entry in JSDoc
* fix(date-picker): eliminate infinite re-render on re-open with existing selection
The useEffect that syncs picker state on open had initialStart and
initialEnd — Date objects computed on every render — in its dependency
array. Because Object.is returns false for any two distinct Date
instances, the effect fired on every render when open=true, calling
setRangeStart/setRangeEnd and triggering another render, producing an
infinite loop that crashed the page.
Fix: compute start and end as local variables inside the effect and
use the stable string props (props.startDate, props.endDate) as deps
instead.
Also removes the redundant typeof fileSize === 'number' guard in the
multipart quota check — fileSize is z.number() (required) in the
contract so it can never be undefined at that point.
* refactor(date-picker): comprehensive cleanup and reliable crash fix
The previous fix still had derived Date objects in useEffect deps.
Object.is(new Date(), new Date()) === false, so any Date in deps causes
the effect to run every render, reproducing the infinite loop on
re-open with existing time selection.
Key changes:
- useEffect deps now use only stable primitives (startDate, endDate strings)
and compute Date values inside the effect — eliminating the loop
- Replace `rest as any` with a FlatDatePickerProps merged type for safe,
typed destructuring across the discriminated union
- Remove initialStart/initialEnd render-scope variables; compute inline or
inside effects to keep derivation local to each use site
- Callbacks use destructured props (onChange, onRangeChange, etc.) instead
of props.x references
- Remove verbose TSDoc on internal callbacks — names are self-documenting
- Preserve all existing JSX structure and CalendarMonth logic unchanged
* fix(security): supabase rpc path validation, ssh stream byte cap, storage quota coverage
* fix(security): scope execution log writes to owning workflow; add env-var workspace membership guard
Closes two cross-tenant vulnerabilities:
1. Workflow log cross-tenant write (route.ts + logging-session.ts):
- Route: SELECT before creating LoggingSession to verify executionId belongs
to the claimed workflowId; reject with 404 if owned by a different workflow.
- LoggingSession: add workflow_id to all UPDATE/SELECT WHERE clauses
(raw SQL marker queries, flushAccumulatedCost, loadExistingCost) so
writes are a no-op if executionId was somehow injected.
2. Env-var workspace membership guard (environment/utils.ts):
- getPersonalAndWorkspaceEnv now calls checkWorkspaceAccess when workspaceId
is provided; throws if the userId is not a member, preventing any future
caller from reading another workspace's decrypted secrets without
explicit membership verification at the call site.
* fix(security): remove fileSize > 0 quota bypass gate; exempt logs context from quota
* chore: remove extraneous inline comments
* fix(security): scope markExecutionAsFailed UPDATE by workflowId; thread workflowId through HITL callers
* fix(security): add personal credential ownership check in sharepoint site route; scope markExecutionAsFailed by workflowId
* fix: remove logs from user-accessible upload contexts; restore distinct biome .next glob
* fix(sharepoint): migrate site route to authorizeCredentialUse
The previous fix only checked userId equality for personal credentials and
workspace membership (via getUserEntityPermissions) for workspace credentials.
authorizeCredentialUse additionally enforces credentialMember access for
workspace-scoped credentials, matching the standard pattern used by all
other tool selector routes.
* fix(logging): make workflowId required in markExecutionAsFailed
Making workflowId optional left a footgun — future callers could silently
omit it and the WHERE clause would degrade to executionId-only, losing the
cross-tenant scoping guarantee. All callers already supply workflowId, so
making it required (with string | undefined for the middle params to keep
call sites unchanged) closes the gap without touching any caller.
* test(security): add tests for cross-tenant log guard, quota bypass fix, and workflowId scoping
- log/route.test.ts: verifies cross-tenant executionId guard returns 404
when the execution belongs to a different workflow, and passes for same
workflow or fresh executions
- multipart/route.test.ts: verifies fileSize:0 no longer bypasses quota
check and that the logs context is rejected at the endpoint level
- logging-session.test.ts: verifies markExecutionAsFailed scopes by both
executionId and workflowId, and that the instance method forwards workflowId
* fix(lint): move IconComponent outside ToolInput to fix noNestedComponentDefinitions
* fix(logging): scope completeWithCancellation and completeWithPause reads by workflowId
Both SELECT queries that check execution status before writing a
terminal result were only filtering on executionId. Adds workflowId
to the WHERE clause so all seven reads and writes in LoggingSession
consistently scope by (workflowId, executionId).
* fix(integrations): gdrive trashed search, slack blocks-with-file, slack get_message ts
- Google Drive search/list: skip default `trashed = false` when user query
already specifies a `trashed = ...` predicate, so trashed-file searches work.
- Slack send-message with files: forward `blocks` through to
`files.completeUploadExternal` so Block Kit renders when files are attached.
- Slack get_message: switch from `conversations.history` (oldest lower-bound
returned the next message after) to `conversations.replies` with `ts=`
for exact-match lookup, plus a defensive ts-equality guard and clearer error.
* fix(google_drive): revert list.ts trashed guard — query is plain text, not gdrive syntax
* fix(slack): omit initial_comment when blocks present so Block Kit actually renders on file uploads
* improvement(scheduler): drain in chunks instead of a single capped claim
Replaces the fixed MAX_CRON_CLAIMS (200) with a chunked drain loop:
claim WORKFLOW_CHUNK_SIZE + JOB_CHUNK_SIZE per iteration, process via
Promise.allSettled, repeat until both claim queries return empty or
MAX_TICK_DURATION_MS elapses. Throughput is no longer bounded by a
static per-tick ceiling; it scales until DB or trigger.dev is the
limit. Per-iteration chunk size still bounds row-lock set and fan-out
concurrency.
Extracts processScheduleItem and processJobItem so the loop body stays
readable. Existing claim semantics (FOR UPDATE SKIP LOCKED, lastQueuedAt
as the claim signal, staleness reclaim) are unchanged.
* improvement(scheduler): skip claim once a queue is exhausted and drop workflowUtils non-null assertion
Addresses Greptile review on PR #4578:
- track per-queue exhaustion when a claim returns fewer than CHUNK_SIZE
rows; subsequent iterations skip the claim query for that queue. Saves
one DB round-trip per iteration once one queue drains while the other
is still working.
- narrow workflowUtils to a local const inside the loop body so the
schedule processing branch only runs when the import has completed.
Removes the misleading non-null assertion.
* v0.6.29: login improvements, posthog telemetry (#4026)
* feat(posthog): Add tracking on mothership abort (#4023)
Co-authored-by: Theodore Li <theo@sim.ai>
* fix(login): fix captcha headers for manual login (#4025)
* fix(signup): fix turnstile key loading
* fix(login): fix captcha header passing
* Catch user already exists, remove login form captcha
* improvement(db): add session statement/lock timeouts; simplify KB doc tx
* fix(knowledge): close soft-delete TOCTOU on KB document insert
Fix the race the bots flagged: KB delete is soft (`deletedAt = now`) so
the FK can't catch a concurrent KB delete between the existence check
and the document insert.
- Add `insertDocumentsIfKbAlive` helper that gates the insert on
`EXISTS(SELECT 1 FROM knowledge_base WHERE id=$kb AND deleted_at IS NULL)`
in the same statement via INSERT...SELECT...WHERE EXISTS. Atomic at the
MVCC snapshot — no transaction, no row lock.
- Use jsonb_to_recordset to declare column types once, avoiding per-param
casts for nullable columns.
- Wire into both `createDocumentRecords` (bulk) and `createSingleDocument`.
- Keep the upfront KB existence check as a fast-path early-out for the
common case; the atomic insert is the race guard.
---------
Co-authored-by: Waleed <walif6@gmail.com>
Co-authored-by: Siddharth Ganesan <33737564+Sg312@users.noreply.github.com>
Co-authored-by: Vikhyath Mondreti <vikhyathvikku@gmail.com>
* fix(rate-limit): close rate-limit bypass and tighten public route limits
* fix(rate-limit): address PR review — drop success field from 429 body, fall back to per-IP when JWT auth lacks userId
* fix(mothership): persist @-mentioned resources across send and merge on hydration
* fix(mship-resources): handle ADD/DELETE race and reorder during pending flush
- Track in-flight ADD promises so DELETE chains off finally(), preventing orphaned server rows when a user removes a resource before its POST resolves
- Defer reorder PATCH until pending flush completes; emit with full local order
- Clear new refs in reset paths
* fix(mship-resources): defer reorder when ADDs are in-flight on existing chat
reorderResources previously only checked pendingPersistResourceKeysRef. When a
chatId exists, addResource fires the POST immediately and only tracks the
promise in inFlightResourceAddsRef — so a reorder before those ADDs settle
shipped a PATCH the server rejected, and the silent catch lost the reorder.
Now treat in-flight ADDs like pending ones: defer the PATCH and replay it
after Promise.allSettled on the in-flight map.
* feat(observability): export Trigger.dev telemetry to Grafana Cloud OTLP
Wire OTLP HTTP exporters for traces, logs, and metrics from the
Trigger.dev runtime to Grafana Cloud. Auth uses Basic with instance ID
and API token. Gated behind GRAFANA_OTLP_ENDPOINT, GRAFANA_INSTANCE_ID,
and GRAFANA_API_TOKEN — all three must be set together or all unset;
partial config throws at startup.
* improvement(observability): use OTLP HTTP/JSON for metrics for consistency with traces and logs
* feat(observability): tag Trigger.dev telemetry with deployment.environment.name
* improvement(observability): switch Grafana telemetry vars to OTLP-shaped trio
* feat(mothership): pin tasks to keep them at the top of the sidebar
* fix(sidebar): address PR review feedback for pin tasks
* fix(posthog): register task_pinned and task_unpinned events
* fix(tasks): insert new optimistic tasks below pinned partition
* improvement(grafana): align tools and block with official Grafana API spec
Validates and corrects the Grafana integration against the official API
docs: fixes wire-format field naming for provisioned alert rules
(missing_series_evals_to_resolve, keepFiringFor, orgID), adds
X-Disable-Provenance support, expands alert-rule params (isPaused,
notificationSettings, record, annotations, labels), corrects defaults
(execErrState=Error, dashboard overwrite=false), and centralizes alert-rule
output mapping in a shared utils module.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(grafana): correct wire-format casing for provisioned alert rule fields
Grafana's ProvisionedAlertRule schema (verified against upstream Go source
and swagger spec) uses keep_firing_for (snake_case) and
missingSeriesEvalsToResolve (camelCase) — the opposite of what prior audit
rounds assumed. POST/PUT bodies now send the correct field names; mapAlertRule
reads the correct primary names with the old casings kept as fallbacks.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(grafana): address PR review feedback
- Drop hardcoded orgID: 1 fallback; only send orgID when organizationId is
provided, so token-scoped org context drives rule placement.
- Surface invalid JSON for notificationSettings/record on alert rule
create/update instead of silently dropping the input.
- Fix execErrState description in update_alert_rule to include Error.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(grafana): surface invalid JSON for annotations/labels/data on alert rules
Match the behavior of other JSON params (data, notificationSettings, record):
return a descriptive error instead of silently falling back to {} (create)
or keeping the existing value (update).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* docs(grafana): expose alert-rule output fields in generated docs
Move ALERT_RULE_OUTPUT_FIELDS from utils.ts to types.ts and rename to
SCREAMING_SNAKE_CASE so scripts/generate-docs.ts (which only resolves const
references from types.ts matching [A-Z][A-Z_0-9]+) can inline the per-field
rows into the generated alert-rule output tables.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
* fix(mothership): reconcile stuck conversation_id against Redis lock to clear stuck-yellow task tiles
copilot_chats.conversation_id has no TTL/heartbeat, so when a stream
process dies before the clear path runs (pod OOM, SIGKILL, uncaught
throw, deploy mid-stream) the column is orphaned and the task tile
renders yellow forever. The Redis lock at copilot:chat-stream-lock:<chatId>
is the canonical liveness signal and self-heals via 60s TTL + 20s
heartbeat, but the mothership APIs weren't consulting it.
Adds read-time reconciliation: a batched MGET helper checks whether
each persisted conversation_id still has a live Redis lock, and both
GET /api/mothership/chats and GET /api/mothership/chats/[chatId]
rewrite the marker to null when the lock has expired. No DB writes;
stuck rows self-heal on next fetch.
* test(mothership): clarify test name to reflect that getActiveChatStreamIds is called with empty candidateIds
* address comments
* fix state machine issue
* cleanup code and fix types
---------
Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai>
* fix(console): match child-workflow inner blocks by instanceId when reconciling dropped SSE events
* fix(console): drop noisy warn when reconcile finds no matching entry
* fix(security): harden HIGH deepsec findings across multiple attack surfaces
- Supabase tools (get_row, delete, update): validate table name with strict
identifier regex and encodeURIComponent to prevent LLM-controlled path
traversal to admin endpoints; add missing empty-filter guard to update
matching the delete.ts pattern
- SFTP/SMTP/SharePoint upload routes: add verifyFileAccess ownership check
before downloadFileFromStorage, matching the WordPress reference pattern;
rejects files the requesting user does not own with 404
- Gmail labels, OneDrive folders, Wealthbox items (×2): replace bare
resolveOAuthAccountId + workspace-only membership check with
authorizeCredentialUse which enforces credentialMember table; use
credentialOwnerUserId for token refresh instead of bare accountRow.userId
- A2A utils: thread pre-resolved IP from validateUrlWithDNS into A2A SDK
via pinnedFetch (secureFetchWithPinnedIP) for JsonRpcTransportFactory,
RestTransportFactory, and DefaultAgentCardResolver, closing the TOCTOU
DNS rebinding window
- SSH utils: cap stdout/stderr accumulation at 16 MB with truncation marker
to prevent OOM from unbounded command output
- Form DELETE route: replace db.delete() with db.update({archivedAt}) for
true soft delete matching the schema's archivedAt column
- Workflow admin import: fix Array.isArray() guard that silently dropped
all variables (export format is Record, not Array)
- Multipart upload: apply checkStorageQuota and MAX_WORKSPACE_FILE_SIZE to
mothership context, closing the quota bypass for workspace-scoped storage
* fix(security): eliminate workspace env lost-update race with atomic JSONB ops
PUT: use `variables || excluded.variables` in onConflictDoUpdate so
concurrent writes merge atomically in the DB instead of last-writer-wins
at the application layer.
DELETE: replace the read-modify-write upsert with a single UPDATE that
removes keys via the JSONB `-` operator, preventing concurrent deletes
from resurrecting previously-removed secrets.
* fix(security): address audit findings from security fix review
- SMTP send: restructure attachment loop from Promise.all to sequential
for...of so verifyFileAccess denial returns 404 instead of propagating
as a generic 500 via the SMTP error classifier
- Supabase tools: extend table-name validation and encodeURIComponent to
the five previously missed tools — insert, upsert, count, query,
text_search — completing coverage across all nine Supabase tools
- Credential routes: remove unnecessary `request as any` casts in Gmail,
OneDrive, and Wealthbox routes; authorizeCredentialUse already accepts
NextRequest directly
- Form soft delete: also set isActive=false alongside archivedAt so that
any future code paths querying by isActive see a consistent state
- SSH utils: fix exit code fallback from 0 to -1 so an abnormally closed
connection that supplies no exit code is not reported as success
- Workspace env: capitalize EXCLUDED.variables in the onConflictDoUpdate
set clause to make the pseudo-table reference unambiguous
* fix(security): address PR review comments and harden deepsec fixes
- fix(env): replace jsonb operators with transaction+FOR UPDATE read-modify-write
- PUT: uses db.transaction + SELECT FOR UPDATE + JS merge to avoid lost-update race
- DELETE: same pattern; fixes variable scope bug where current was referenced outside tx
- removes broken || and - jsonb operators that fail on json-typed column
- fix(ssh): trim truncated output consistently with non-truncated path
- fix(gmail): remove redundant resolveOAuthAccountId call
- adds credentialType field to CredentialAccessResult
- authorizeCredentialUse now returns credentialType in all success paths
- gmail/labels route uses authz.credentialType and authz.resolvedCredentialId directly
- fix(supabase): centralize table identifier validation
- adds validateDatabaseIdentifier() to input-validation.ts
- all 8 supabase tools use the shared util instead of inline regex
* fix(workflows): fix VariableType assignment in admin workflow import route
The intermediate Record cast used 'string' for the type field which TypeScript
correctly rejected — WorkflowVariable.type is 'VariableType', not string.
Changed the cast to use VariableType so both branches typecheck correctly.
* fix(a2a): handle Request objects in pinnedFetch URL extraction
* fix(security): extract shared file-access guard; merge workspace/mothership branch
* fix(security): advisory lock for env first-insert race; handle all BodyInit types in pinnedFetch
* chore: remove inline comment from advisory lock
* fix(security): remove stray comment; narrow credentialType to literal union
* fix(security): add credentialId validation to wealthbox oauth route; fix null body override in pinnedFetch
* fix(security): stream A2A response body to unblock SSE; keep text/json/arrayBuffer for non-streaming callers
* fix(security): resolve credentialId guard on OneDrive, use assertToolFileAccess in WordPress, memoize body buffer to prevent silent empty reads, fix ArrayBuffer type cast
* fix(security): handle string[][] HeadersInit format in pinnedFetch
* fix(security): keep abort listener alive during body streaming; clean up in stream end/error/cancel
* chore: remove extraneous inline comment
* fix(security): cleanup abort listener when maxResponseBytes limit is exceeded
* improvement(workflow-block): support manual workflow ID via advanced mode
* fix(input-mapping): resolve workflowId via canonical hook for advanced mode
* fix(input-mapping): fall back to manualWorkflowId in preview context
* refactor(input-mapping): resolve workflowId via useDependsOnGate canonical pattern
* fix(docs): restore media centering and full-width intro image
* fix(docs): drop overflow-hidden from intro media wrappers so focus ring is not clipped
* fix(docs): use inset focus ring on lightbox media so parent overflow-hidden cannot clip it
* fix(docs): drop focus ring on lightbox media to match original UI
* fix(helm): preserve STS serviceName + networkPolicy.egress back-compat
Greptile flagged two real upgrade-breaking changes vs the prior chart:
1. statefulset-postgresql spec.serviceName flipped from <name>-postgresql
to <name>-postgresql-headless. spec.serviceName is immutable, so any
existing install would hit 'Forbidden: updates to statefulset spec ...'
on helm upgrade. Revert to the original name (the headless Service in
services.yaml is added alongside, not as a swap).
2. networkPolicy.egress changed from a list to a map ({extraRules, exceptCidrs}),
silently dropping any custom egress list set by existing users. Restore
the original list semantics for networkPolicy.egress and move cloud-metadata
blocking to a sibling top-level field networkPolicy.egressExceptCidrs.
Adds NOTES.txt upgrade-notes entry covering both + the ESO v1→v1beta1 default
flip (functionally a no-op, but worth surfacing).
* docs(helm): update README egress reference to new key name
* fix(helm): revert copilot-postgresql STS serviceName too (same immutability issue)
Audit caught that the main fix in d5c2e8ef5 missed statefulset-copilot-postgres.yaml,
which had the identical immutable-field rename from -copilot-postgresql to
-copilot-postgresql-headless. Same upgrade-break vector for anyone running
copilot.enabled=true on a prior chart version. Mirrors the fix and comment
from the main postgresql STS.
* improvement(helm): postgres startupProbe + otel-collector NetworkPolicy
- add startupProbe defaults for both postgresql + copilot-postgresql STSs
to shield liveness from slow first-boot (pgvector init, WAL replay)
- render a dedicated NetworkPolicy for the otel-collector when
telemetry.enabled=true (OTLP ingress from app/realtime/copilot, DNS +
HTTPS egress for forwarding to external observability backends)
- document why copilot + copilot-postgresql intentionally do NOT ship
dedicated NetworkPolicies (Redis URL is unknowable at render time)
- regression test pins the otel-collector NP at documentIndex 3
* test(helm): assert custom egress applied to realtime NP too
The prior test claimed coverage of both app and realtime NPs but only
asserted documentIndex 0. Split into two tests so a regression that drops
custom egress from realtime would fail loudly.
* docs(helm-skill): trim narrative bloat in values-model
Cut the historical 'Layer 2 was added in chart 1.0.0' note and the
generic 'single source of truth' framing. Kept the two actionable
points: ESO requires mapping Layer 1 keys; app.env overrides
envDefaults.
* improvement(scheduler): raise per-tick claim budget to drain backlog
MAX_CRON_CLAIMS 20 -> 100; reserved workflow/job slots 10/10 -> 50/50.
Throughput was capped at 20 schedules/tick which created a 20+ hour
backlog when due work exceeded ~1 item per cron-second.
* improvement(scheduler): raise per-tick claim budget to 200
Bumps MAX_CRON_CLAIMS 100 -> 200 (workflow/job split 100/100). Pairs
with the fire-and-forget cron Lambda change so per-tick processing
time is no longer bounded by the Lambda's 50s HTTP timeout.