mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
badfbc3bdf467f98836d94ba9fc05ab68a3b3d2f
4993
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
badfbc3bdf |
fix(resource): left-align table filter/sort when there's no search (#5128)
The unconditional ml-auto from #5117 right-aligned the embedded table editor's filter/sort cluster, which has no search bar. Only push the aside + filter/sort group right when a search occupies the left; without a search it stays left-aligned as before. |
||
|
|
63a3e6d2cb |
feat(files): stream large CSV previews and add import-as-table (#5125)
* feat(files): stream large CSV previews and add import-as-table * fix(files): validate fileId in csv-preview route, guard double-import, fix sniff perf and toggle flash * fix(files): scope mothership preview-toggle loading guard to CSV files only |
||
|
|
a028d07e7b |
improvement(mothership): user_table speed parity — limit bounds, background import/delete/update jobs (#5012)
* improvement(mothership): user_table speed parity — limit bounds, async import/delete/update jobs
- query_rows / filter ops clamp limit to the contract maxes; query_rows
skips execution metadata.
- import_file / create_from_file (large CSV/TSV) and delete_rows_by_filter
(>1000 unbounded matches) dispatch background table jobs, claiming the
per-table job slot; inline paths claim the slot too.
- update_rows_by_filter now escalates the same way: >1000 unbounded matches
run as a background table job (new 'update' job type + runTableUpdate worker
+ tableUpdateTask), so a broad update on a huge table no longer loads every
row into the request. Best-effort/non-atomic and skips workflow recompute
(documented); unique-column patches stay inline. Pagination is limit/offset.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(mothership): trim user_table catalog copy to the essentials
Drop the verbose doomedCount/affectedCount, delete-mask, workflow-recompute,
and unique-column asides from the bulk-op descriptions. The model only needs:
large ops return { jobId }, limit maxes at 1000, one job per table.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* improvement(mothership): make user_table limit cap internal, not model-facing
The model can now pass any limit — no "cannot exceed 1000" rejection. 1000
becomes an internal threshold: query_rows clamps the page to MAX_QUERY_LIMIT
(totalCount signals truncation; the model pages with offset), and bulk filter
ops above the cap run as background jobs.
update_rows_by_filter loads full row data inline, so an explicit limit above
the cap escalates to the background worker with a new maxRows budget (the worker
stops after maxRows; update has no read mask so the cap is exact). delete only
loads ids inline, so an explicit limit (any size) stays inline — only unbounded
deletes use the masked background path, which would over-hide a bounded delete.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* improvement(mothership): bounded delete above the cap runs async, not inline
An explicit delete limit now mirrors update: ≤1000 runs inline, above the cap it
escalates to the background worker honoring the limit via maxRows — instead of
always staying inline. The worker stops after maxRows (per-page fetch capped to
the remaining budget).
Bounded background deletes skip pendingDeleteMask: the filter-based mask hides
every match, which would over-hide the rows beyond the cap the job never deletes.
Unmasked, a bounded delete is eventually consistent like a bounded update (rows
disappear as deleted), and doomedCount is omitted from the payload so the count
isn't double-subtracted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs(mothership): tidy user_table limit/offset param copy
Drop "Any value is allowed" from the limit description and restore the original
offset description.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(tables): skip pendingDeleteMask for bounded background deletes
The bounded-delete commit (
|
||
|
|
7d46103d09 |
chore(deps): remove unused dependencies and harden CI supply chain (#5119)
* chore(deps): remove unused dependencies and harden CI supply chain Dependency cleanup: - Remove unused deps: papaparse, unified, and 6 unused Radix primitives (alert-dialog, radio-group, scroll-area, separator, toggle, visually-hidden) plus @tanstack/react-query-devtools (all verified zero imports repo-wide) - Consolidate jwt-decode into the existing jose dependency (decodeJwt) - Migrate react-window to @tanstack/react-virtual to drop a redundant virtualization library (terminal, structured-output, code viewer) - Remove the better-auth-harmony plugin and its gating env flag Supply-chain hardening: - SHA-pin every GitHub Action to a full commit SHA with a version comment - Pin CI bun-version to 1.3.13 (was "latest" in the release job) - Raise bun minimumReleaseAge cooldown from 3 to 7 days - Add a non-blocking `bun audit` step in CI - Add a CODEOWNERS gate routing dependency-manifest changes to @simstudioai/deps * chore(deps): remove unused apps/docs dependencies (@tabler/icons-react, dotenv-cli) * style(search-modal): use Send icon for Invite teammates action * feat(search-modal): surface New chat as the top action above Create workflow * feat(search-modal): add Secrets to the pages list |
||
|
|
08bcacd74e | fix(copilot): mount input tables with display-name CSV headers, not column IDs (#5121) | ||
|
|
fcfa41cd42 | fix(azure): replace Azure DevOps icon with Azure icon and remove AzureDevOpsIcon (#5118) | ||
|
|
cae1769117 |
improvement(knowledge): align connected-sources rows and move source chip left of filter/sort (#5117)
* improvement(knowledge): align connected-sources rows and move source chip left of filter/sort - Drop the -mx-2 on the connectors list so rows respect the ChipModalBody gutter: the row hover no longer bleeds to the modal edges and row content lines up with the px-4 header. - Add a 'leading' slot to ResourceOptions (left of the filter/sort cluster) and render the knowledge connected-source chip there instead of the far-right 'aside', so it reads as part of the control row. 'aside' stays right-aligned for the table editor's run/stop control. * improvement(resource): render options aside left of filter/sort The options-bar aside has a single other consumer (the table editor's embedded run/stop control), so instead of adding a separate slot, render aside itself to the left of the filter/sort cluster. Drops the extra slot and keeps one canonical control position; the run/stop control moves left too, which is fine for a status widget. * fix(resource): keep options aside grouped with filter/sort without a search bar Group aside + the filter/sort cluster in one ml-auto right-aligned container instead of relying on the search's flex-1 to anchor them. Without this, an options bar with no search (the embedded mothership table editor) split aside to the far left and filter/sort to the far right via justify-between. * docs(resource): clarify aside groups with filter/sort regardless of search |
||
|
|
4d39b0cbf8 |
feat(connectors): use resource selectors for KB connector config (#5116)
* feat(connectors): use resource selectors for KB connector config Replace raw ID text inputs with selector pickers (canonical selector + manual-input pairs) across Google Drive/Docs/Forms/Sheets, Notion, Monday, and Webflow KB connectors, so users pick folders/spreadsheets/pages/boards/ collections instead of pasting IDs — matching the workflow blocks. - Add multi-select where the sync handler supports it (Drive/Docs/Forms folders, Monday boards, Webflow collections) via parseMultiValue - Add shared escapeDriveQueryValue/buildDriveParentsClause helpers for safe multi-folder Drive queries - Add ConnectorConfigField.mimeType, plumbed into the selector context - Fix Webflow listingCapped not set on maxItems truncation (deletion- reconciliation data-loss safety) Fully backward compatible: legacy single-string IDs and CSV both normalize via parseMultiValue; resolved canonical keys are unchanged. * fix(webflow): set listingCapped on within-page maxItems truncation When a collection's items fit in a single API page but maxItems cuts the list within that page, neither hasMoreInCollection nor hasMoreCollections is true, so listingCapped was not set and the sync engine could hard-delete still-existing documents. Add the within-page drop signal to the guard. |
||
|
|
ea505f0388 |
improvement(tables): versioned CSV snapshot cache for table mounts + parallel multipart uploader (#5108)
* improvement(tables): versioned CSV snapshot cache for table mounts + parallel multipart uploader * chore(db): drop colliding 0239 migration (renumber pending) * chore(db): renumber rows_version migration to 0240 (off staging's 0239) * improvement(tables): mount snapshots by presigned URL so the sandbox fetches directly (raise cap to 500MB) * fix(tables): allow url sandbox entries in the function-execute contract; key snapshot by column shape so schema edits invalidate it * chore(e2b): log sandbox inputs split by url-fetch vs inline write * improvement(tables): order export + snapshot rows by order_key so the CSV matches the grid under fractional ordering |
||
|
|
c907b1194b |
improvement(supabase): add Edge Functions tool; correct storage output shapes + harden tools (#5112)
* improvement(supabase): add Edge Functions tool; correct storage output shapes + harden tools
- Add supabase_invoke_function tool (POST/GET/PUT/PATCH/DELETE /functions/v1/{name})
- upsert: support on_conflict; storage upload: support cache-control
- Fix storage copy/move/upload/delete-bucket output properties to match live API
- get_public_url: build URL via directExecution (no spurious network call)
- text_search: validate column identifier
- Strip non-TSDoc section-label comments
* fix(supabase): harden rpc/text_search identifiers; drop unused get_public_url apiKey
- rpc: validate + encode functionName (SSRF/injection parity with vector_search)
- text_search: validate language config interpolated into the PostgREST operator
- get_public_url: remove unused apiKey param + dead auth headers (public endpoint needs no auth)
- create_bucket: tighten output description to match the {name}-only response
* improvement(tavily): mark optional params advanced; fix empty content output
- Mark 29 optional search/extract/crawl/map subBlocks as mode: advanced (keep query/urls/url/apiKey basic)
- Fix search transformResponse: populate content from result.content (was result.snippet, always empty)
- Guard data.results with ?? []; correct country placeholder to a lowercase name
- Rewrite stale TavilySearch/Extract response interfaces; drop dead duplicate interfaces
* fix(supabase): valid Cache-Control directive on upload; clarify functionName description
- storage upload: expand a bare numeric cache-control to `max-age=<n>` (a raw number is not a valid Cache-Control header)
- block: functionName input description now covers RPC, vector search, and Edge Function invoke
* fix(supabase): edge-function error handling + reject array headers
- invoke_function: drop unreachable !response.ok branch (executor throws on non-OK before transformResponse runs and surfaces the error body); document the success-only contract
- invoke_function: ignore non-object (array) headers so JSON arrays can't produce numeric-index header names
- block: reject array/non-object Edge Function headers with a clear error in config.params
* fix(supabase): scope Edge Function method/body/headers to invoke_function
Prevents a stale `method` value (e.g. from the Edge Function field) from leaking
into other operations' params. The tool executor lets `params.method` override a
tool's static verb (tools/utils.ts), so an unscoped value could turn a read into
DELETE/POST against PostgREST. Now method/body/headers are only passed for the
invoke_function operation. Adds a block-level regression test.
* fix(supabase): only parse Edge Function body/headers for invoke_function
Stale or invalid functionBody/functionHeaders left in the block (common when
switching operations) were parsed and validated for every operation, so they
could throw before unrelated tools ran even while hidden. Moved parsing and
validation inside the invoke_function guard; added a regression test.
* fix(supabase): include last_accessed_at as a storage list sort option
The Storage list API accepts last_accessed_at for sortBy; add it to the tool
description and the block dropdown so the surfaced options match the API.
|
||
|
|
11e23131fe |
feat(google): Maps Pollen/Solar, Custom Search expansion, and live-API fixes across Google integrations (#5113)
* feat(google): add Maps Pollen/Solar, expand Custom Search, fix Ads/Groups/Contacts/Slides New capability: - Google Maps: add Pollen Forecast and Solar Potential tools (API-key, google_cloud BYOK) - Google Custom Search: add start/dateRestrict/fileType/safe/searchType/siteSearch/ siteSearchFilter/lr/gl/sort params, htmlTitle/htmlSnippet/formattedUrl/mime/fileFormat/ cacheId/image result fields, and nextPageStartIndex pagination Fixes (validated against live API docs): - Google Ads: bump all tools from sunset v19 to v24 - Google Groups: forward OAuth credential under oauthCredential (was dropping token in 11 ops), forward all update_settings fields, JSON.stringify update_settings/add_alias bodies - Google Contacts: include required metadata.sources[].etag in updateContact body (fixed 400) - Google Slides: remove unsupported GIF thumbnail mimeType (API only allows PNG) - Google Sheets: wire delete_rows/delete_sheet/delete_spreadsheet into the V2 block - Google Custom Search: throw on API error responses instead of returning empty success; num optional + Number-coerced; pagemap typed unknown * docs(google): regenerate integration docs for new and updated operations * fix(google_maps): correct Solar requiredQuality enum to BASE The Solar API ImageryQuality enum is HIGH/MEDIUM/BASE (+ UNSPECIFIED) per the live docs; there is no LOW. Selecting "Low" sent requiredQuality=LOW which the API rejects as INVALID_ARGUMENT, and the valid BASE tier was unreachable. Replace LOW with BASE in the tool param/output descriptions, the type union, and the block dropdown. * fix(google_maps): guard !response.ok in Pollen/Solar; use ?? for color channels Address Greptile review: - Pollen and Solar transformResponse now check !response.ok || data.error (matches the Custom Search fix); a gateway error without an error key in the body no longer returns empty/zeroed output silently. - Pollen color channels use ?? instead of || so a legitimate 0 isn't treated as missing (consistent with the other numeric fields in the file). * fix(google_maps): guard against NaN days in Pollen forecast Address Cursor Bugbot: a non-numeric `days` input parsed to NaN and was forwarded as `days=NaN` (the tool's `?? 1` only catches undefined, not NaN), breaking the forecast call. The block now coerces invalid input to undefined, and the tool defaults to 1 unless `days` is a finite number. * fix(google): clamp Pollen days to 1-5; stop forwarding stale group settings fields Address Cursor Bugbot: - Pollen: clamp days to the documented 1-5 range (truncating fractionals) so 0, negatives, or >5 can't be sent to the API. - Google Groups update_settings: the block has no dedicated settings subblocks, so forwarding name/description from params could leak stale values from create_group/update_group and unintentionally rename the group. Forward only oauthCredential + groupEmail from the block (the tool's own param schema still exposes the settings fields for the agent path). * fix(google_sheets): fail fast on non-numeric delete indices Address Cursor Bugbot: delete_sheet/delete_rows parsed deleteSheetId/startIndex/ endIndex with Number.parseInt but didn't validate, so non-numeric UI input became NaN and was forwarded (the v2 delete tools only reject null/undefined), breaking the batchUpdate. The block now throws a clear error when any of these is not a valid number. * fix(google_search): clamp num to 1-10 and normalize start Address Cursor Bugbot: num was coerced with Number() but not bounded, so values like 11 or fractionals reached the API and failed. The tool now truncates and clamps num to the documented 1-10 range and only sends a positive integer start, ignoring non-numeric/out-of-range input. |
||
|
|
d7fd0405f9 |
improvement(search): align cmd+k action icons + highlight with the design system (#5114)
* improvement(search): align cmd+k action icons + highlight with the design system - Each Actions verb now uses the exact icon from its real location: Fit to view -> Scan (workflow-controls), Copy workflow link -> Duplicate (nav context menu), Invite teammates -> User (settings teammates nav). Run/Create/Import already matched. - Remove the Toggle theme action and its now-dead useTheme wiring. - Matched-text highlight now uses the design-system search tokens (--highlight-match-bg / --highlight-match-text), matching the SearchHighlight component used in knowledge-base and code search, instead of an ad-hoc font-semibold. * improvement(search): use font-medium for matched-text emphasis in cmd+k Drop the colored background highlight in favor of the design system's standard emphasis weight (font-medium, used by Button/Label/Input/Table). Lighter than the previous semibold and avoids a background, keeping the palette's clean, undecorated text style. |
||
|
|
9e9f2b9e1f |
fix(realtime): debounce the reconnecting toast to stop transient-blip flashes (#5111)
* fix(realtime): debounce the reconnecting toast to stop transient-blip flashes The "Reconnecting..." persistent toast fired the instant isReconnecting flipped true, so sub-second transport blips that self-heal on the first retry flashed a scary alert. Add useStableFlag, an anti-flicker boolean that delays the rising edge (2s, so brief blips never surface) and holds the falling edge (1.5s min visible, so a drop just past the delay does not flash-and-vanish). The socket flag stays accurate; only the user-facing alarm is smoothed. State machine extracted into a framework-agnostic controller with unit coverage for both flicker modes. * fix(realtime): reset stable-flag React state on options change; de-vacuous blip test Address Greptile review: - useStableFlag: reset React state to the fresh controller's baseline when the controller is recreated on an options change, so a dynamic consumer changing delayMs/minVisibleMs while active with value already false can no longer strand the flag at true. - test: read the live probe.active getter in the blip test instead of a destructured snapshot, which was bound to false at destructure time and made the assertion vacuous. |
||
|
|
05cd7d92a8 |
feat(search): actions, fuzzy matching, and highlighting in cmd+k palette (#5110)
* feat(search): actions, fuzzy matching, and highlighting in cmd+k palette Add a context-aware actions layer to the cmd+k search palette (Run workflow, Create workflow/folder, Import workflow, Fit to view, Copy link, Invite teammates, Toggle theme), replace the substring matcher with a boundary-anchored fuzzy matcher (initialisms, typos, multi-word) that is a strict superset of the old behavior, highlight matched characters, and rank against clean human text instead of structural id/uuid tokens. Expose invoke() on the global commands provider so the palette runs real registered commands. * fix(search): highlight the matched substring, not an earlier scattered occurrence Contiguous substring matches (exact/prefix/contains) now report the substring's own indices instead of the greedy subsequence scan positions, so HighlightedText bolds the characters the user actually matched. Restructures fuzzyMatch to handle the substring tier first; scores are unchanged for these cases. * fix(search): log clipboard copy failures and make fuzzy positions read-only - Copy workflow link now logs on clipboard write failure instead of silently swallowing the error, matching the sidebar's copy-link convention. - FuzzyResult.positions is now readonly and the NO_MATCH singleton's array is frozen, so the shared instance can never be mutated by a caller. |
||
|
|
8b93e43037 |
improvement(integrations): validate BigQuery/Forms/PageSpeed + regenerate integration docs (#5109)
* improvement(integrations): validate BigQuery/Forms/PageSpeed + regenerate integration docs - BigQuery: mark null-defaulted outputs optional (get_table type/numRows/numBytes/creationTime/lastModifiedTime/location, list_datasets location, list_tables type, query totalBytesProcessed) - Google Forms: add response pagination (pageToken + filter params, nextPageToken output), fix pageSize visibility, advanced-mode pagination subBlocks + filter wandConfig - PageSpeed: add a 7th BlockMeta template (competitor benchmark) - Regenerate integration docs; add manual intro sections to new datagma/dropcontact/enrow/icypeas/leadmagic pages * fix(docs-gen): preserve apostrophes in tool descriptions when generating docs The doc generator extracted tool descriptions with a character class that excluded both quote types (['"]([^'"]...)['"]), so a double-quoted description containing an apostrophe (e.g. "Find someone's email") was truncated at the apostrophe — the generated docs/catalog showed stubs like "Find someone". Anchor extraction on the actual opening quote (single/double/backtick), matching the existing extractDescription helper, in both buildToolDescriptionMap and extractToolInfo. Regenerated docs restore full descriptions across all affected integrations (Apollo, Ahrefs, LeadMagic, Findymail, OpenAI, Slack, etc.). * fix(docs-gen): resolve tools defined in a sibling file + scope params per tool The doc generator located a tool's definition only by filename convention (decompress.ts / index.ts), so file_decompress — which lives in compress.ts alongside file_compress — fell back to index.ts and rendered an empty Input table. It also read the params block from the first tool in a multi-tool file, so every tool in such a file inherited the first tool's inputs/outputs. - getToolInfo: when no candidate file declares the exact tool ID, scan the whole tool-prefix directory for the file that does. - extractToolInfo: read the params block scoped to the specific tool, falling back to the full file for tools that inherit params via spread. Regenerated docs eliminate ~50 empty/incorrect input tables across integrations (clickhouse, rb2b, reddit, file, etc.); param-less OAuth-only tools correctly keep an empty input table. |
||
|
|
80735b424b |
fix(locks): enforce workflow/folder locks on the agent + close manual-UI create gaps (#5107)
* fix(locks): enforce workflow/folder locks on the agent + close manual-UI create gaps The copilot/agent workflow & folder mutation tools and the edit_workflow tool bypassed lock enforcement, so the agent could edit a locked workflow and move/create workflows into a locked folder. Add assertWorkflowMutable/ assertFolderMutable guards (from @sim/workflow-authz) to every agent mutation path, mirroring the REST API. Also close two parity gaps on the manual-UI REST side: creating a workflow into a locked folder and creating a subfolder under a locked parent were previously unguarded. The realtime collaborative canvas already enforced workflow-level locks server-side. * fix(locks): normalize optional folderId to null for assertFolderMutable * refactor(locks): hoist constant folder-lock check out of move loop; scope test mocks with Once * refactor(locks): drop redundant ensureWorkflowAccess fetch in rename |
||
|
|
83531452e5 |
fix(sidebar): prefetch chats + workflows so cold loads don't flash skeletons (#5104)
* fix(sidebar): prefetch chats + workflows so cold loads don't flash skeletons On a cold load (e.g. when the browser discards an idle tab and reloads), the persistent sidebar started with an empty React Query cache and client-fetched its chat + workflow lists, flashing loading skeletons. Prefetch both lists server-side in the workspace layout and hydrate them via HydrationBoundary, under the same query keys and mappers the client hooks use, so the sidebar paints populated on the first render. The prefetch runs concurrently with the existing org-settings fetch and never throws, so it adds no blocking work in the common case and falls back to client fetching on error. * refactor(prefetch): call data layer directly instead of internal HTTP self-fetch The sidebar and settings prefetches fetched their data by making internal HTTP requests to our own API routes. Replace those self-fetches with direct calls to shared server-side data functions, so each route handler and its prefetch read from one source with no extra network hop, serialization, or re-auth. - Extract listWorkflowsForUser (lib/workflows/queries) and listMothershipChats (lib/copilot/chat) from their routes; both routes and the sidebar prefetch now call them. - Extract getUserSettings/getUserProfile (lib/users/queries) shared by the settings/profile routes and their prefetches. - Subscription prefetch calls the existing getSimplifiedBillingSummary + getEffectiveBillingStatus directly. - Sidebar prefetch checks workspace access once via checkWorkspaceAccess and skips silently when denied. * refactor(prefetch): share mothership chat list staleTime constant Export MOTHERSHIP_CHAT_LIST_STALE_TIME from the chats hook and use it in both useMothershipChats and the sidebar prefetch, mirroring WORKFLOW_LIST_STALE_TIME so the prefetch and client hook can't drift. * fix(prefetch): keep subscription prefetch on the wire shape via internal billing API The billing summary returns Date fields (and an untyped metadata blob) that the JSON API serializes to strings. Calling the data layer directly would cache Date objects (App Router preserves them through RSC serialization), mismatching the string wire shape the client useSubscriptionData hook caches. Route the subscription prefetch through the internal billing API so server-hydrated and client-fetched data share the exact same shape. The date-free general-settings and profile prefetches keep calling the data layer directly. |
||
|
|
a82b44d36d |
perf(db): logs-list index, drop redundant indexes, replica routing, hot-path write cleanups (#5105)
* perf(db): logs-list index, drop redundant indexes, replica routing, hot-path write cleanups * fix(logs): keep /api/v1/logs on primary db — its permissions join is the auth gate, not replica-safe |
||
|
|
8fe090a3a1 |
fix(input-format): field not editable race condition (#5102)
* fix(input-format): field not editable race condition * remove dead code * simplify |
||
|
|
2ffc004a0a | improvement(models): add DeepSeek V4 + Mistral Medium 3.5, fix Codestral context window (#5103) | ||
|
|
15a970d805 |
feat(integrations): hosted email-enrichment providers + cascade wiring (#5087)
* feat(integrations): hosted email-enrichment providers + cascade wiring Add Datagma, Dropcontact, LeadMagic, Icypeas, and Enrow integrations — tools, blocks, brand icons, and BYOK + metered hosted-key support — and register each in the tool/block registries and BYOK provider list. Wire the new finders/verifiers into the enrichment cascades: - work-email: Datagma, LeadMagic, Dropcontact, Icypeas, Enrow - phone-number: LeadMagic, Datagma, Dropcontact - email-verification: Icypeas, Enrow - company-info: Datagma, LeadMagic - company-domain: Datagma Add hosting tests for all five providers and cascade tests covering the new providers (incl. new test files for email-verification, company-info, and company-domain). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(enrichment): address PR review on Icypeas success + Datagma billing - Icypeas find_email/verify_email postProcess return success:true for all terminal statuses (NOT_FOUND/DEBITED_NOT_FOUND included) so the cascade runner calls mapOutput and records invalid/not-found verdicts instead of throwing and inflating the error count - Bill Icypeas verify FOUND (not just DEBITED*) per the documented 0.1-credit charge - Datagma enrich_person only applies the 30-credit phone surcharge when a phone lookup (phoneFull) was requested - Note Datagma's URL-param (apiId) auth in the hosted-key doc comment - Update hosting tests to match Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(enrichment): only bill Enrow verify on a completed verification getCost returned a flat 0.25 credits regardless of output, so a job that fell back to the initial submit response (poll never completed, no qualification) was still metered. Charge 0.25 only when a qualification is present; 0 otherwise. Add a no-qualification test case. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(enrichment): peg hosted credit cost to each provider's lowest paid plan Align *_CREDIT_USD to the entry tier Sim will provision: - Datagma: Regular $49/3,000 emails → $0.0163 (was Popular $0.0132) - LeadMagic: Basic $49/2,000 → $0.0245 (was Growth $0.0104) Icypeas (Basic $0.019), Enrow (Starter $0.012), and Dropcontact (Starter ~$0.17) already reflect their lowest plan. Tests derive from the constants, so values stay consistent. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(enrichment): address PR review on mononyms + Icypeas verify email map - work-email LeadMagic: pass full_name + domain so single-token (mononym) names are no longer skipped - work-email Icypeas: firstname/lastname are optional on the API, so run a mononym with firstname alone instead of self-skipping - icypeas_verify_email mapItem reads item.email (verify payload shape) with a fallback to the nested results.emails[0].email Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(enrichment): case-insensitive Enrow find billing getCost compared qualification to exactly 'valid' while the cascade normalizes with toLowerCase(), so a differently-cased API qualifier could zero out billing on a valid email. Lowercase before comparing; add a test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(enrichment): drop Dropcontact from the phone-number cascade Dropcontact is an email/company-data enrichment service, not a phone-discovery provider — its phone/mobile_phone fields are unreliable and were surfacing firmographic data (an employee-count range like "5000-20077") as the phone. Keep the two purpose-built phone finders (LeadMagic find_mobile, Datagma find_phone); Dropcontact stays in the work-email and company cascades where its data is reliable. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(enrichment): only accept valid-qualified Enrow emails in work-email Enrow's finder qualifies each email valid/invalid. The work-email mapOutput accepted any non-empty email, so an invalid-qualified address could fill the cell while hosted billing (which only charges on valid) charged zero. Gate the cell on qualification === 'valid', consistent with billing. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
feca5fa6f6 |
improvement(execution, connectors): offload large function inputs, increase connector limits + better error propagation (#5089)
* fix(execution,connectors): offload large function inputs; harden KB connector size limits Addresses a class of 10 MB limit failures: - executor/variables: offload over-budget function block-output context values to durable large-value refs (lazy `sim.values.read`) so JS function blocks can merge medium files without exceeding the 10 MB inter-block request-body cap. - connectors: stream downloads via `readBodyWithLimit` (memory-safe), and surface oversized files as visible `failed` KB documents instead of silently dropping them — listing-time for github/s3/dropbox/onedrive/sharepoint, fetch-time for gitlab/azure/google-drive via a shared `ConnectorFileTooLargeError`. Raise the per-file cap from a hardcoded 10 MB to the canonical 100 MB KB document limit (`CONNECTOR_MAX_FILE_BYTES`), except Google Drive's export path (Google's hard 10 MB export-API limit). - sync-engine: `classifyExternalDoc` + bulk `skipDocuments` (failed rows with a reason, excluded from retry), byte-bounded batch concurrency to cap peak worker memory at the raised cap, and a `metadata.fileSize ?? size` fallback. * fix zoom * update skill * address comments + fix terminal event in sse stream * fix accounting issue |
||
|
|
cc56408be3 |
perf(execution): parallelize preflight gates, cache deployed state, memoize Anthropic client (#5098)
* perf(execution): parallelize preflight gates, cache deployed state, memoize Anthropic client
- Memoize Anthropic + Azure-Anthropic SDK clients (new client-cache.ts) keyed
by apiKey (+beta header; +baseURL/version/pinnedIP for Azure) so HTTP
keep-alive connections are reused instead of a fresh TLS handshake per call.
apiKey is the tenant boundary.
- Parallelize the read-only preflight gates in preprocessing.ts (ban +
subscription, then usage + org-member + rate-limit) while preserving exact
error precedence (ban 403 -> usage 402 -> rate 429) and keeping the sole
write (admission reservation) last.
- Parallelize the independent workflow-state and env-var loads in execution-core.
- Cache deployed workflow state by immutable deploymentVersionId with
deep-clone-on-read, oldest-first eviction, and a 5-min TTL bounding the
credential-mapping edge across ECS tasks.
- Parallelize the independent personal-subscription + membership queries in
getHighestPrioritySubscription.
- BYOK: drop the redundant getWorkspaceById existence check (auth already
validates the workspace); read the key list fresh every call for zero
cross-instance staleness.
Billing/usage/ban/permission reads stay fresh on the primary (no cache, no
replica). Adds tests for every new mechanism and fixes a pre-existing vitest
class-mock incompatibility that had execution-core.test.ts fully red on staging.
* fix(execution): run rate-limit gate only after ban/usage pass
The rate-limit gate is not read-only — checkRateLimitWithSubscription consumes
a token — so running it in parallel with the read-only gates debited rate-limit
quota for requests that the ban (403) or usage (402) gates reject, which the
original sequential flow never did.
Move the rate-limit gate to run sequentially after the ban and usage gates pass,
preserving the read-only gates' parallelism (ban + subscription + usage) and the
exact ban -> usage -> rate precedence. Add regression tests asserting the rate
limiter is not consumed when an earlier gate rejects, and is consumed once when
they pass.
Caught by Cursor Bugbot review.
* chore(execution): trim redundant preflight comments
Tighten the gate overview to match the sequential rate-limit gate and drop
inline notes that duplicated it or the runRateLimitGate doc.
* refactor(cache): address review — idle TTL for client cache, LRUCache for deployed state
- client-cache: add updateAgeOnGet so the TTL is genuinely idle-based (active
clients keep their warm keep-alive connections; the JSDoc now matches behavior).
- deployed-state: replace the hand-rolled Map + manual FIFO eviction/TTL with
LRUCache (real LRU eviction, built-in TTL), matching the effectiveDecryptedEnv
and integration-tool-schema caches. TTL stays absolute (not reset on read) so
the credential-migration remap still propagates across ECS tasks.
Both per review feedback from Greptile.
* test(execution): isolate rate-limit gate test from STEP 7 reservation
The 'consumes the rate-limit gate once' test reached the STEP 7 admission
reservation, which depends on Redis — it passed locally (reserve throws and is
swallowed) but failed in CI (reserve returns not-reserved -> 429). Pass
skipConcurrencyReservation so the test isolates the rate gate deterministically.
* perf(providers): memoize SDK clients where the pool is per-client (bedrock, vllm)
Generalize the Anthropic client cache into one shared memoizer
(providers/client-cache.ts) and apply it only where each new client owns its own
connection pool — so reuse actually keeps connections warm:
- bedrock: AWS SDK clients hold a per-client connection pool (reuse is the AWS
best practice). Keyed by region + credential identity.
- vllm: a pinned endpoint creates its own undici Agent per call; key by the
resolved IP so DNS re-validation still runs each request.
- anthropic + azure-anthropic: migrated onto the shared memoizer.
Deliberately NOT applied to the OpenAI-compatible providers, groq, cerebras, or
google: their SDKs share a process-global keep-alive pool (Node openai-sdk module
singleton agent; anthropic/global undici), so a fresh client per request already
reuses connections and memoization would add complexity with ~no benefit. litellm
uses a plain shared-agent client (no pinning) and is likewise skipped.
Bounded LRU (max 1000, 30m idle TTL) with no close-on-eviction, avoiding the
unbounded-growth and eviction-closes-in-use-client failure modes seen in similar
client caches.
* chore(perf): trim verbose comments to terse why-notes
* chore(perf): drop obvious inline comments, keep nuance as TSDoc
* fix(bedrock): key client cache on full credential, not just access key id
A corrected secret under the same access key id would otherwise keep serving the
stale cached client until TTL/eviction. Caught by Cursor Bugbot.
* test(execution,providers): fix preflight mock reset + isolate provider client cache in tests
- preprocessing.test: re-establish the checkOrgMemberUsageLimit mock in beforeEach
(the only gate mock not re-set). In the full suite its implementation was reset
so the success-path test got undefined -> threw -> 500 -> success:false. Mirrors
how checkServerSideUsageLimits is handled.
- client-cache: add clearProviderClientCacheForTests; call it in the bedrock and
vllm test beforeEach so construction assertions always start from a cache miss
now that those providers memoize their client.
* test(execution): make RateLimiter mock constructable under vitest 4.x
The RateLimiter mock used an arrow factory (vi.fn(() => ({...}))). vitest 4.x
(CI) rejects `new` on an arrow-implemented mock ("not a constructor"); 3.2.4
allowed it. The new rate-gate test is the first to actually `new RateLimiter()`,
so it surfaced the failure only in CI. Switch the mock to a regular function and
drop the speculative beforeEach re-establishments that didn't address it.
|
||
|
|
f238184fa0 |
feat(file): add Compress and Decompress operations to the File block (#5100)
* feat(file): add Compress operation to bundle files into a .zip archive * feat(file): add Decompress operation to extract .zip archives Adds the inbound half of the archive pair: extracts a .zip back into the workspace with zip-slip path sanitization, symlink skipping, and entry/ size caps to bound zip-bomb expansion. Extracted files are returned in the files output, ready to chain downstream. * fix(file): align archive ops with v5 output surface and zip mime - Drop the single 'file' output reintroduced for compress/decompress; v5 intentionally exposes only 'files' (plus id/name/size/url scalars), so compress/decompress reuse the existing surface with no new block output - Add zip/gz to EXTENSION_TO_MIME (previously only in the reverse map), so archive extensions resolve to a real mime instead of octet-stream - Update File v5 block test for the two new operations * fix(file): harden compress naming per review - Flatten zip entry names to a safe basename so untrusted fileInput names with .. or / cannot produce zip-slip entry paths (cursor) - Treat archiveName as a flat name landing at the workspace root instead of passing it through splitWorkspaceFilePath, which silently created folders for names with separators (greptile) - Add the upfront empty-input guard before any DB calls, matching the read and content operations (greptile) * fix(file): make decompress extraction atomic and bound per-entry size - Read and validate every entry before writing any file, so hitting a size cap no longer leaves partially-extracted files in the workspace (cursor) - Enforce the per-entry cap on the materialized buffer in addition to the declared size, covering entries that omit an uncompressed size (cursor) - Pre-check declared sizes up front to reject standard zip bombs before materializing, and return 422 when no files could be extracted (cursor) * fix(file): exclude skipped entries from caps and reject multi-archive decompress - Resolve safe (sanitized) zip entries up front so unsafe/skipped entries no longer count toward the per-entry and total uncompressed-size caps (cursor) - Reject decompress input that resolves to more than one archive with a clear error instead of silently extracting only the first (cursor) * fix(file): enforce single-archive decompress at the API boundary The block already rejects multiple archives, but the manage route is the real boundary (callable directly and by the LLM tool) and still took the first of multiple resolved inputs. Add the empty-input and >1-archive guards in the route so extra archives are rejected with a clear error rather than silently ignored (cursor). * docs(file): correct compress description and stale file-output references - Drop the misleading 'under provider upload limits' claim from the compress tool description (models cannot read zip archives) - Fix bestPractices to reference the 'files' output, not a non-existent 'file' - Remove the stale 'file' property from the compress test fixture so it matches the real API response (greptile) |
||
|
|
c864a928bf |
improvement(models): sort model dropdown by latest release date within each provider (#5099)
* improvement(models): sort model dropdown by latest release date within each provider * fix(models): preserve input provider order and build catalog index once |
||
|
|
73c73ff0f2 |
fix(kb): canonicalize knowledge-base upload keys (#5096)
* fix(kb): canonicalize knowledge-base upload keys * fix tests |
||
|
|
d14bc78d03 |
fix(realtime): re-check workspace role on mutating socket events (#5080)
* fix(realtime): re-check workspace role on mutating socket events * address comments |
||
|
|
2b7f57a788 |
improvement(providers): tighten Gemini and vLLM agent-attachment ceilings (#5095)
A live-doc audit of the merged large-file feature found two ceilings that were higher than the provider actually accepts: - Gemini: 100MB -> 50MB. Gemini hard-caps PDFs at 50MB, so a 50-100MB PDF passed our gate, got uploaded + polled, then failed at generateContent. 50MB respects the documented limit and is more memory-safe. - vLLM: 50MB -> 25MB. vLLM's default image-fetch timeout is 5s; a 50MB remote fetch routinely exceeds it. 25MB aligns with that reality and matches Baseten (the other vLLM-backed provider). |
||
|
|
2c1392e3a0 |
fix(chat): autoscroll follow-ups — re-engage threshold + keep end-of-turn options in view (#5094)
* fix(chat): align scrollbar/keyboard detach with wheel/touch re-engage threshold The onScroll detach branch set only stickyRef.current = false, leaving userDetachedRef false, so a scrollbar-drag or keyboard detach kept the lenient 30px (STICK_THRESHOLD) re-engage threshold instead of the strict 5px (REATTACH_THRESHOLD) used after wheel/touch. A programmatic virtualizer re-pin landing within 30px could then snap autoscroll back on right after the user deliberately scrolled away. Reuse the detach() helper so all detach paths set userDetachedRef consistently. * fix(chat): keep end-of-turn options in view after streaming When a stream ends, the suggested-follow-up options and the actions row (gated on !isStreaming) mount, but the virtualizer's getTotalSize — which drives the scroll container's scrollHeight — only catches up a frame or two later via its ResizeObserver. The single scrollToBottom() on effect teardown therefore landed on a stale, too-short bottom and the options were clipped behind the input. (Pre-virtualization this worked because scrollHeight reflected the new rows immediately.) Extract the rAF follow loop already used for CSS height animations into a shared followToBottom(window) helper and run it for a short settle window on teardown, so the bottom is chased until the virtualizer re-measures. The follow is self-interrupting — height growth leaves scrollTop where we put it, while a user scroll moves it up, so it bails the instant the user scrolls and never fights a real gesture even with listeners torn down. |
||
|
|
e1f22bdac9 |
feat(providers): support large agent-block attachments via Files APIs and remote URLs (#5092)
* feat(providers): support large agent-block attachments via Files APIs and remote URLs
Agent-block file uploads were inlined as base64 with a hard 10MB cap. Files
above the threshold now use each provider's native large-file path:
- OpenAI / Gemini: upload to the provider Files API, reference by file_id/uri
- Anthropic: GA url content-block source (no Files API beta, no upload)
- OpenRouter/Groq/Together/Baseten/xAI/vLLM: remote signed URL in image_url/file
- Limits live per-provider in models.ts; the agent block + /models page reflect them
Files <=10MB keep the identical base64 path (zero regression). Server-only file
handles are stripped from untrusted input to prevent SSRF.
* fix(providers): clear forged file handles for inline providers too
attachLargeFileRemoteUrls early-returned for inline-strategy providers before
clearing server-only handle fields, so a forged remoteUrl on an inline-provider
file could still reach a builder (e.g. buildOpenAICompatibleChatContent for
mistral/ollama). Clear the handles for every provider before the strategy check.
* fix(providers): correct OpenAI expiry serialization and Anthropic large-text-doc handling
- OpenAI upload now uses the SDK (client.files.create) so expires_after is
serialized as a real nested object; the prior expires_after[anchor] bracket
FormData keys were ignored by OpenAI's server, leaving files un-expiring.
- Anthropic url document source only supports PDFs/images; large non-PDF text
docs now throw a clear error instead of emitting an unsupported url source.
- Warn when an oversized file can't be sent because cloud storage is unavailable.
* fix(providers): harden large-file path (SSRF fetch, ceiling gate, per-file UI limit)
- Download files for OpenAI/Gemini uploads via validateUrlWithDNS + IP-pinned
fetch so a forged URL can't reach internal addresses (covers all callers).
- Reject files above the provider ceiling before downloading/uploading.
- UI now validates each file against the provider's per-file ceiling instead of
summing all files against it, matching server-side per-file validation.
- Lower Anthropic ceiling to 50MB (documented 32MB request cap / page limits).
* refactor(providers): read files-api upload bytes via storage SDK
Read OpenAI/Gemini upload bytes through downloadFileFromStorage instead of
HTTP-fetching the presigned URL. Removes any server-side URL fetch (no SSRF
vector) and works with internal object storage (e.g. self-hosted MinIO), which
an IP-pinned URL fetch would have blocked.
* docs(providers): clarify files-api bytes are read from storage at upload time
* fix(providers): enforce access checks and strip forged ids in the upload path
uploadLargeFilesToProvider runs on raw request messages for every caller (incl.
the internal providers passthrough), so harden it independently of the agent path:
- verifyFileAccess on each file's storage key before reading its bytes, so a forged
key can't exfiltrate another user's file.
- clear any inbound providerFileId/providerFileUri up front (legit ids are only set
by the upload itself), so a forged id can't reference a file in a hosted account.
* fix(providers): resolve UI attachment limit with the same model->provider helper as execution
The file-upload control imported getProviderFromModel from @/providers/models, but
the execution path and every other consumer use the one in @/providers/utils (runtime
registry + reseller patterns). Align the UI so its size cap can't disagree with
server-side validation for reseller or dynamically-listed models.
* test(providers): add new models.ts exports to provider mocks
attachments.ts now reads getProviderFileAttachment / INLINE_ATTACHMENT_MAX_BYTES
from @/providers/models; the provider unit tests that fully mock that module need
both exports or attachments.ts fails to load.
* fix(providers): guard Gemini upload response name before polling
ai.files.upload returns name as string | undefined; guard it (instead of an
as-string cast) so a missing name surfaces a clear error at the upload site
rather than an opaque files.get failure on the first poll.
* fix(uploads): type the file-handle key list so omit preserves UserFile fields
The 'as const' readonly tuple widened omit's K to all keys, collapsing
Omit<UserFile, K> to {} and failing the production build's type check. Declare
the array as Array<keyof handle fields> so K is the precise literal union.
* refactor(providers): run handle-clear + URL-mint in executeProviderRequest for all callers
Move attachLargeFileRemoteUrls out of the agent handler and into
executeProviderRequest (right before uploadLargeFilesToProvider), so every entry
point — including the internal providers passthrough — clears forged handles and
mints/access-checks large-file URLs uniformly. The agent handler now only hydrates
base64; its missing-file guard exempts large files (resolved downstream).
* fix(azure-openai): guard optional attachment dataUrl in inline image part
PreparedProviderAttachment.dataUrl is now optional (large files carry a handle
instead); azure-openai builds chat content inline and assigned it directly to a
required url field, failing the production build's type check.
* fix(providers): upload OpenAI files via multipart and fix Buffer Blob part
The installed openai SDK (4.104) does not type expires_after on files.create, so
upload via POST /v1/files directly with the documented expires_after[...] form
fields (gives the file an auto-expiry). Also wrap the storage Buffer in a
Uint8Array for the Blob, which the production build's stricter lib types require.
These two type errors were masked locally because tsc was OOMing silently without
the type-check script's --max-old-space-size flag.
* fix(providers): forward userId from the providers API to executeProviderRequest
Large-attachment prep now needs request.userId for presigned URLs and access
checks; the authenticated providers proxy has auth.userId but wasn't passing it,
so oversized attachments failed for logged-in callers. Forwarding it makes large
files work there and keeps the access check (verifyFileAccess) intact.
* fix(providers): fail clearly when a large attachment has no cloud storage
The doc claimed a base64 fallback that doesn't exist — above the inline cap there
is no base64, so without cloud storage the file previously reached the builder and
died with a generic read error. Throw a clear 'requires cloud file storage' error
at the point of detection and correct the doc.
|
||
|
|
e48c960fe8 |
fix(chat): keep autoscroll pinned when the virtualizer re-scrolls during streaming (#5093)
* fix(chat): keep autoscroll pinned when the virtualizer re-scrolls during streaming The sticky-scroll detach heuristic (scrollTop drops while scrollHeight doesn't grow) could not distinguish a user scrollbar drag from a programmatic scroll. react-virtual re-pins content by moving scrollTop whenever a measured row's size changes — including the transient height shrinks streamdown emits as it re-parses each streaming token — so the hook misread those upward programmatic scrolls as the user scrolling away and detached mid-stream. Gate the scroll-delta detach branch behind a genuine recent user gesture (pointerdown/up tracking + wheel/touch/keydown stamp). Programmatic scrolls have no preceding gesture, so they no longer detach; scrollbar drag, wheel, and keyboard detach are preserved. * fix(chat): address review — reset pointer ref on teardown, stop wheel/touch opening detach window - Reset pointerDownRef in effect cleanup so a pointer held through teardown (e.g. dragging the scrollbar as a stream finishes) can't leak a stuck-true ref into the next session and detach on the first programmatic re-pin. - Wheel-up and touch-drag already detach directly, so the onScroll delta heuristic only needs to authorize scrollbar drag (pointerDownRef) and keyboard. Stop stamping the gesture window on wheel/touch, which otherwise let a harmless downward wheel open a 250ms window where a virtualizer shrink could falsely detach. * fix(chat): scope detach authorization to real scroll gestures; TSDoc comments - onPointerDown only marks an active drag when the press targets the scroll container itself (the scrollbar), not its content, so a text-selection drag on a message can't authorize a detach during a programmatic re-pin. - Reset lastUserGestureAtRef on teardown alongside pointerDownRef so neither a held pointer nor a late keydown can leak across streaming sessions. - Convert the hook's inline comments to TSDoc on the relevant declarations per codebase conventions. * fix(chat): only upward scroll keys authorize a keyboard detach onKeyDown stamped the gesture window on any bubbling key, so an unrelated keypress within USER_GESTURE_WINDOW of a programmatic virtualizer re-pin could satisfy userDriven and detach mid-stream. Filter to the upward scroll keys (ArrowUp, PageUp, Home, Shift+Space), mirroring the wheel handler's upward-only rule, so only a genuine upward keyboard scroll authorizes detach. |
||
|
|
6cbaf4282d |
fix(scheduled-tasks): fix scheduled tasks schema validation (#5091)
* fix(scheduled-tasks): fix schema rejection * fix(db): fix duplicate db query |
||
|
|
91666b585f | feat(scheduled-tasks): migrate jobs agent to scheduled tasks agent (#5090) | ||
|
|
7b4626e547 |
improvement(perm-groups): allow workspace filter for permission groups (#5070)
* improvement(perm-groups): allow workspace filter for permission groups * show errors correctly * address comments * address concurrent edit concern * address locks * address comments" * index migration safety * address at route level |
||
|
|
b7d30c89f9 |
feat(google-calendar): wire freebusy, align tools with API v3, add calendar + sharing tools (#5084)
* feat(google-calendar): wire freebusy, align tools with API v3, add calendar + sharing tools * fix(google-calendar): address review — trust offset timezones, make list_acl showDeleted usable, harden unshare error parse, clarify update attendees * fix(google-calendar): wire list q/pageToken into block, harden invite PUT error parse * fix(google-calendar): make list orderBy user-selectable, clarify update timeZone applies to start/end * fix(google-calendar): require timeZone for recurring timed events, clarify recurrence replace semantics * fix(google-calendar): validate grantee before building share ACL body Throw a clear error when scopeType is user/group/domain but scopeValue is missing or blank, instead of POSTing a scope-type-only body that the Calendar ACL API rejects with an opaque error. |
||
|
|
06f1e72d6a |
feat(feature-flags): migrate 3 env-flags to AppConfig-backed runtime flags (#5086)
* feat(feature-flags): migrate 3 env-flags to AppConfig-backed runtime flags * fix(feature-flags): hardcode workflow-columns on, fix feature-flags tests * chore(feature-flags): document mothership-beta userId targeting limitation |
||
|
|
05e8c7cd71 |
refactor(connectors): split client metadata from server runtime (#5076)
* refactor(connectors): split client metadata from server runtime + cover node:net in client bundle The browser build broke with `Cannot find module 'node:net'`. Server-only SSRF code in `input-validation.server.ts` (`dns/promises`, and since PR #5060 `undici` → `node:net`/`node:tls`) is statically reachable from the client bundle via the tool/connector registries, which the workflow editor imports for metadata. Node networking builtins have no browser shim, so Turbopack cannot compile them for the client. Two changes: 1. Split each connector's client-safe declarative metadata into a sibling `meta.ts` (`<name>ConnectorMeta`), mirroring the `XBlockMeta` / `BLOCK_META_REGISTRY` pattern. `connectors/registry.ts` is now the client-safe `CONNECTOR_META_REGISTRY` (+ `getConnectorMeta` / `getAllConnectorMeta`); the full registry with runtime fns moves to `connectors/registry.server.ts`. Client components consume the meta registry; the sync engine and knowledge API routes consume the server registry. This removes connectors from the client's server-only graph. Connector metadata is byte-for-byte identical before/after; runtime fns are untouched. 2. Extend the existing #4899 `turbopack.resolveAlias` browser stub — which already mapped `dns`/`dns/promises` to an empty module for the browser — to also cover `net`/`tls` (+ `node:` variants), since `undici` now pulls those in. The remaining tool/provider definitions still reach `input-validation.server` server-side; the browser-only stub keeps those Node builtins out of the client bundle while the real modules stay on the server, so SSRF validation and IP pinning are unaffected. Connector authoring/validation skills updated to teach the meta.ts split. * fix(icons): use Square logo glyph only, drop wordmark * fix(connectors): share Discord max-messages default across meta and runtime Discord defined DEFAULT_MAX_MESSAGES separately in meta.ts (config placeholder) and discord.ts (sync behavior), which could drift. Export it from meta.ts and import it in the runtime, matching the single-source pattern used by the other connectors (e.g. gmail, intercom). * refactor(tools): route grafana/agiloft egress server-side, drop SSRF browser shim Move the server-only SSRF-pinned fetch out of the grafana (update_dashboard, update_alert_rule) and agiloft (11 record/search tools) definitions and into internal API routes, the same pattern the rest of the server-side tools (and agiloft's own attach/retrieve) already use. The tool definitions are now purely declarative (request → internal route), so they no longer import `input-validation.server` and the tools registry is fully client-safe. With connectors (meta split) and these tools no longer reaching server-only code from the client bundle, the browser no longer pulls in `dns`/`net`/`tls`: - Add `import 'server-only'` to `input-validation.server.ts` so any future client import fails loudly at build time instead of silently bloating the bundle. - Remove the `turbopack.resolveAlias` browser stub and delete `empty-node-fallback.browser.ts` — the root cause is fixed, the shim is gone. Behavior is unchanged: each route runs the exact merge/validation/fetch logic the tool ran before (every header, param branch, JSON-parse guard, error string, and SSRF pinning preserved); only the location of execution moved from the client- bundled definition to a server route. * fix(connectors): move onedrive tagDefinitions into meta; drop server-only guard - onedrive's tagDefinitions lived in the runtime file, so the client meta registry returned undefined for it and the add-connector tag opt-out section stopped rendering for onedrive. Move it into meta.ts like the other connectors so client and server see identical metadata (verified across all 50). - Remove the 'server-only' import from input-validation.server.ts: the meta/route split already keeps it out of the client bundle, and blocks/tools registries don't use the guard either. * fix(grafana): surface upstream error when the prefetch GET fails Check response.ok on the existing-resource GET in both update routes and return the upstream status/body, matching how the tool framework surfaced GET errors before the move to internal routes (the framework checks response.ok before transformResponse). Without this, a failed prefetch produced a generic 'Failed to fetch existing ...' message and dropped Grafana's error detail. * fix(grafana): reject invalid panels JSON instead of silently ignoring it Grafana's dashboard API treats panels as a required array and returns 400 on invalid JSON; this route already errors on every other JSON param. Return 'Invalid JSON for panels parameter' instead of swallowing the parse error and proceeding with a misleading success. * fix(grafana): trim dashboard/alert-rule UID in route URLs (carry over #5082) PR #5082 added .trim() on dashboardUid/alertRuleUid in the original tool URL builders. Those tools now build their URLs in the internal routes, so apply the same trim there to preserve that behavior. * fix(grafana): route update_folder egress server-side (carry over #5082) #5082 added a grafana update_folder tool that does SSRF-pinned fetch in its postProcess, re-introducing the client-bundle leak. Convert it to the internal API route pattern like the other update tools so the def is declarative and input-validation.server stays out of the client bundle. * fix(grafana): surface route failures in transformResponse instead of masking them The grafana update tools' transformResponse hardcoded success: true and dropped the route's error, so an upstream/validation failure (HTTP 200 with { success: false, error }) was reported to the workflow as a success. Forward data.success and data.error (matching the agiloft tools) so failures propagate as before the move to internal routes. |
||
|
|
324189ed39 |
feat(grafana): validate integration and add folder, health, and contact-point tools (#5082)
* feat(grafana): validate integration and add folder, health, and contact-point tools - Require alert-rule title/ruleGroup/data in the block (create would 400 without them) - Trim UID path params across dashboard and alert-rule tools to avoid copy-paste 404s - Use Grafana brand color for the block background - Surface previously-unsettable list params (limit, starred, annotation type) - Add get/update/delete folder, check data source health, get health, and create contact point tools - Strip non-TSDoc comments; regenerate docs and integrations.json * fix(grafana): scope block param remaps per operation to prevent cross-operation leaks |
||
|
|
ed31edf068 | improvement(ci): fix companion regex (#5083) | ||
|
|
3fe061e3b3 |
feat(feature-flags): AppConfig-backed gated feature flags (#5059)
* feat(feature-flags): AppConfig-backed gated feature flags * fix(ci): repoint 'Validate feature flags' step to env-flags.ts after rename * improvement(feature-flags): drop in-code defaults; fallback resolves a per-flag secret, gating is AppConfig-only * improvement(feature-flags): make flag names a closed set so every flag requires a fallback secret * improvement(feature-flags): single FEATURE_FLAGS registry — each entry defines name, description, and fallback in one place * improvement(feature-flags): fallback is the env secret key (keyof typeof env), resolved to a boolean |
||
|
|
aaeef82e22 | improvement(ci): rename companion tags to be more descriptive (#5081) | ||
|
|
0c226b9cbe |
fix(providers): allow HTTP for self-hosted vLLM endpoints (#5078)
Pass allowHttp to validateUrlWithDNS so plain-HTTP self-hosted vLLM endpoints are permitted. This only relaxes the protocol check; the private/reserved-IP blocklist and blocked-port checks still apply, so SSRF protection is unchanged. |
||
|
|
cbd3d2220f | feat(ci): mship companion pr check (#5079) | ||
|
|
4a7c2ef656 |
fix(providers): pin vLLM provider endpoint to validated IP (#5077)
Validate the user-supplied vLLM endpoint (request.azureEndpoint) against the central SSRF guard and pin the connection to the resolved IP before issuing any request, mirroring the Azure OpenAI/Anthropic providers. The operator-configured VLLM_BASE_URL stays trusted and unvalidated. |
||
|
|
a09e3939f2 |
fix(webhooks): cap request body size on public webhook receivers (#5075)
* fix(webhooks): cap request body size on public webhook receivers Public, unauthenticated webhook endpoints read the entire request body into memory before any lookup or signature verification, letting a caller exhaust pod memory with arbitrarily large bodies. Bound the body via the existing size-limited stream reader (content-length guard + streamed cap) and return 413 on oversize. Applies to parseWebhookBody (trigger receiver) and the agentmail route. Cap defaults to 10 MB, overridable via WEBHOOK_MAX_REQUEST_BYTES. * refactor(webhooks): extract shared body-size cap to constants module Address review feedback: hoist WEBHOOK_MAX_BODY_BYTES into a single lib/webhooks/constants.ts so the trigger receiver and AgentMail route share one source of truth instead of recomputing the env-derived cap (prevents drift). Also drop the redundant request clone when the body stream is null. * refactor(webhooks): drop redundant null-body branch in capped readers Both capped body readers had an `if (!stream)` fallback to an uncapped `.text()`/empty string. `readStreamToBufferWithLimit` already returns an empty buffer for a null stream, so the branch is redundant and the `.text()` fallback was a theoretical bypass (chunked request, no content-length, null body). Collapse both to a single capped read. * chore(webhooks): drop inline comments from capped body readers |
||
|
|
f277f5fba5 |
feat(db): zero-downtime migration safety lint + db-migrate skill (#5041)
* feat(db): zero-downtime migration safety lint + db-migrate skill Add scripts/check-migrations-safety.ts (check:migrations), a CI gate that classifies statements in newly-added migrations into hard errors (rewrite), annotate-to-acknowledge contract ops (`-- migration-safe: <reason>`), and backfill warnings. Wire it into test-build.yml. Add the /db-migrate skill as the judgment half (expand/contract phasing, app-code cross-ref, annotation authoring). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(skills): run cleanup and db-migrate safety checks in /ship * fix(db): address review — DROP INDEX lock symmetry, RENAME CONSTRAINT false-positive, alter-type literal match - Non-concurrent DROP INDEX is now a hard error (ACCESS EXCLUSIVE lock), symmetric with CREATE INDEX; DROP INDEX CONCURRENTLY after a COMMIT passes clean. Removes the false-confidence annotate path. - RENAME rule narrowed to RENAME COLUMN / table RENAME TO; RENAME CONSTRAINT and ALTER INDEX ... RENAME (metadata-only) no longer flagged. - alter-type regex now requires TYPE to follow the column identifier, so it no longer matches TYPE inside a string default. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(db): enforce IF EXISTS on DROP INDEX CONCURRENTLY for replay idempotency Symmetric with the CREATE INDEX CONCURRENTLY rule: a post-COMMIT DROP INDEX CONCURRENTLY replays from the top on failure, so without IF EXISTS it aborts re-dropping an already-gone index. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * improvement(skills): gate /ship cleanup on UI changes; default migration base to staging - /ship runs /cleanup only when the diff touches UI code (.tsx or apps/sim/components|hooks|stores); the six passes are React-only. - /ship runs check:migrations against origin/staging (the PR base). - check:migrations default baseRef is now origin/staging instead of origin/main. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(db-migrate): add contract-pending TODO convention for deferred drops Establishes a durable, greppable marker (`contract-pending(<precondition>): ...`) left on the legacy column in schema.ts when an expand defers a drop, so the contract phase doesn't rot. The outstanding-work list is `grep -rn contract-pending`; the contract PR's `-- migration-safe:` annotation references the expand and deletes the marker. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
a49e755418 |
feat(auth): OAuth-only signup with Microsoft provider (#5073)
* feat(auth): OAuth-only signup with Microsoft provider - Remove email/password form from /signup — Google, Microsoft, GitHub OAuth only - Add Microsoft as a social provider (MICROSOFT_CLIENT_ID / MICROSOFT_CLIENT_SECRET / DISABLE_MICROSOFT_AUTH) - Wire microsoftAvailable through provider checker, API contract, providers route, and all auth UI - Hide "Continue with email" in auth modal signup view; login view unchanged - Fix MicrosoftIcon SVG to use official brand colors and proportions * fix(auth): remove unused useSession, guard invalid-callback warn with ref * feat(auth): gate email signup via DISABLE_EMAIL_SIGNUP flag * feat(auth): restore signup email form gated by NEXT_PUBLIC_DISABLE_EMAIL_SIGNUP * refactor(auth): single DISABLE_EMAIL_SIGNUP env var controls both ui and backend * fix(config): restore isHosted hostname check |
||
|
|
4f4ff53411 |
fix(uploads): authorize internal file URLs before download (#5049)
* fix(uploads): authorize internal file URLs before download downloadFileFromUrl treated any URL containing /api/files/serve/ as trusted-internal and read the object straight from storage by key with no access check, while every other resolution path in the file calls verifyFileAccess. Reachable during workflow execution via file[] inputs (type: 'url'), letting an authenticated user read arbitrary storage objects across tenants by supplying a storage key. Thread the caller's userId into downloadFileFromUrl and run verifyFileAccess(key, userId, undefined, context, false) on the resolved key before downloadFile; fail closed when no userId is present. Update all callers (execution file inputs, tool file outputs, KB ingestion); webhook and chat inputs already thread userId via processExecutionFiles. * chore(uploads): log denied internal file downloads for rollout telemetry * fix(uploads): derive internal file context from key, not query param Cursor Bugbot flagged a context-spoofing bypass: downloadFileFromUrl resolved context via parseInternalFileUrl, which honors a caller-controlled ?context= query param. An attacker could label a private storage key with a world-readable context (profile-pictures/og-images/workspace-logos) so verifyFileAccess short-circuits to granted while downloadFile still reads the private object. Infer context from the key only (inferContextFromKey), mirroring how /api/files/serve resolves it; ignore the query param. Also move the userId guard ahead of key resolution so auth failures surface first. * docs(uploads): move context-derivation rationale into TSDoc * fix(uploads): match internal file marker in URL path only isInternalFileUrl matched the /api/files/serve/ substring anywhere in the string, so a crafted URL could carry it in a query string or fragment and skip DNS/SSRF validation. Match it in the path component only. The raw path is checked without URL normalization on purpose: the files parse route relies on traversal sequences surviving this check (an absolute https://host/api/files/serve/../.. URL must classify as internal so the '..' check rejects it, rather than being normalized to /etc/... and waved through as external). Host is intentionally not gated — cross-tenant reads are prevented at the storage sink by verifyFileAccess, and host-allowlisting would break self-hosted/multi-domain deployments. Adds unit tests. * consolidate access, billing principals --------- Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai> |
||
|
|
dd32abef1e |
feat(jsm): add Atlassian Assets (Insight/CMDB) tools for asset management (#5072)
* feat(jsm): add Atlassian Assets (Insight/CMDB) tools for asset management
Add nine JSM Assets tools so workflows can read and write Atlassian Assets
(Insight/CMDB) objects — the foundation for keeping JSM asset tables in sync
for software/hardware asset management.
Tools (wired into the Jira Service Management block):
- jsm_list_object_schemas, jsm_get_object_schema
- jsm_list_object_types, jsm_get_object_type_attributes
- jsm_search_objects_aql (AQL search with pagination)
- jsm_get_object, jsm_create_object, jsm_update_object, jsm_delete_object
Each tool proxies through an internal route that resolves the Jira cloudId and
the Assets workspaceId, then calls the Assets API via the OAuth 2.0 (3LO)
gateway form (/ex/jira/{cloudId}/jsm/assets/workspace/{workspaceId}/v1).
Adds the CMDB OAuth scopes to the jira provider (read/write/delete cmdb-object,
read cmdb-schema/type/attribute) with descriptions, contract schemas for each
route, and block operations/subBlocks/outputs. Bumps the API-validation route
baseline for the nine new routes.
* refactor(jsm): harden Assets param coercion and response typing
- Add toOptionalInt helper so non-numeric pagination inputs never emit NaN
into the Assets query string (startAt/maxResults/page/resultsPerPage)
- Replace Record<string, any> in mapAssetObject with typed Raw* interfaces
* fix(jsm): validate Assets workspaceId and honor `last` pagination flag
Address review findings on the Assets tools:
- Add validateAssetsWorkspaceId and guard the workspaceId in every Assets
route before it is interpolated into the API path (mirrors the existing
cloudId guard) — prevents a crafted workspaceId from escaping the
workspace-scoped path
- Object schema list now falls back to the `last` flag when `isLast` is
absent, so pagination doesn't stop early
* feat(jsm): allow overriding the auto-resolved Assets workspace
Atlassian provisions one Assets workspace per site, so workspace discovery
uses values[0] by design. For the rare multi-workspace site, expose an
advanced "Assets Workspace ID" override on the block that flows through to
every Assets operation, and document the single-workspace assumption.
* refactor(jsm): include Assets responses in the JsmResponse union
Append the nine Assets tool response types to JsmResponse for completeness
and consistency with the rest of the JSM tool surface.
|
||
|
|
d538b76eda | feat(copilot): server-side mothership tool/vfs/file metrics (#5071) |