mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
cb63ecaefbfd14c7fc4daa26cb7c8cd9db5f4d8f
99
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
35fd4ef42f |
improvement(self-host): simplify capability setup configuration (#6230)
* feat(self-host): add capability-aware setup * fix(self-host): preserve capability compatibility * fix(copilot): honor preview availability server-side * improvement(self-host): centralize capability resolution * fix(self-host): preserve integration availability paths * fix(testing): align capability-aware config mocks * improvement(self-host): simplify capability setup configuration * fix(setup): preserve unowned storage overrides * fix(self-host): reconcile storage and allowlists * fix(integrations): preserve connect deep links |
||
|
|
3de63c94e3 |
feat(self-host): align Docker Compose with Helm and overhaul self-hosting docs (#6225)
* feat(self-host): align Docker Compose with Helm and overhaul self-hosting docs Docker Compose shipped no scheduler, so scheduled workflows, every polling trigger, connector syncs, the outbox, and data drains silently never ran. Adds a cron service running the same 18 jobs the Helm chart schedules as CronJobs, and closes the remaining behavioral gaps between the two paths: bundled Redis in the chart, no hosted plan caps in chart defaults, pinned image tags, and fail-fast secrets. A CI check keeps the schedulers in sync. Also rewrites the self-hosting docs: 14 new pages, 8 updated, reorganized into Install / Configure / Operate. * fix(self-host): drop bun install from chart CI, remove air-gapped and backup docs The scheduler-parity check pulled a full dependency install into the chart-validation job, which fails building isolated-vm on that runner. Rewritten to use only node builtins so the job installs nothing. Also removes the air-gapped and backup/restore pages, and stops pinning a concrete release in the docs so the examples do not go stale each release. * fix(helm): bundle Redis in secret-manager modes unless the URL is supplied Suppressing Redis whenever a secret mode was active left those deployments with no Redis at all — REDIS_URL is optional there and both shipped examples omit it. The chart now steps aside only on a detectable signal: an explicit app.env.REDIS_URL, an ESO remoteRefs.app.REDIS_URL mapping, or the new redis.provideUrl=false opt-out for a pre-created Secret it cannot read. * fix(compose): derive realtime BETTER_AUTH_URL from NEXT_PUBLIC_APP_URL realtime read BETTER_AUTH_URL directly and fell back to localhost while simstudio derived it from NEXT_PUBLIC_APP_URL, so setting only the public origin left realtime authenticating against http://localhost:3000. * fix(helm): deliver bundled REDIS_URL via ConfigMap so an operator value always wins Injecting REDIS_URL as an inline container env made it beat every envFrom source, so a REDIS_URL held in a pre-created Secret or synced by External Secrets was silently shadowed and traffic moved to a fresh in-cluster Redis. Kubernetes resolves duplicate envFrom keys by letting the last source win, so the bundled URL now ships as a ConfigMap listed before the app Secret. Any operator-supplied value overrides it without the chart needing to read it, which also removes the redis.provideUrl flag the previous attempt required. * docs(helm): spell out the egress rule external datastores need The default NetworkPolicy allows 443 plus the bundled Postgres and Redis by pod selector. Anything you run outside the chart on another port needs its own rule, which is easiest to miss when REDIS_URL arrives via a Secret the chart cannot inspect. Adds a copyable example to the production checklist and the security guide. * feat(helm): add networkPolicy.allowExternalEgress for managed datastores The default policy allows 443 plus the bundled Postgres and Redis by pod selector, so a managed datastore on another port needs a hand-written CIDR rule — awkward when REDIS_URL arrives via a Secret the chart cannot inspect. Adds an opt-in switch that drops the port restriction while still blocking the cloud metadata endpoints. Defaults to false, keeping this chart stricter than the common chart default of unrestricted egress. |
||
|
|
7798e83489 |
feat(function): custom sandboxes (#6071)
* feat(sandboxes): workspace dependency sets for Function blocks Named package sets a Function block can import from. The server canonicalizes and hashes the list; E2B prebuilds a content-addressed template per set, Daytona installs per execution. Create/edit is gated to Max or Enterprise via the shared workspace entitlement check; execution is deliberately ungated, so a downgraded workspace keeps running what it already built. Also on this branch: - Extract the duplicated dropdown/combobox option-fetch lifecycle into use-fetched-options. Only combobox had the dependency-change reset, so every dropdown with dependsOn + fetchOptions cleared its list and never repopulated until reopened. - Collapse the repeated Max-tier entitlement check onto one hasMaxTierWorkspaceAccess, shared by inbox, live sync, and sandboxes. - Resolve a personal payer's block state through getEffectiveBillingStatus in getBillingEntityBlockStatus, so the client-side Max gates agree with the server-side ones when blockOrgMembers' fan-out is stale. - Carve the Daytona dependency install out of the caller's execution budget instead of stacking on top of it. Co-Authored-By: Claude <noreply@anthropic.com> * chore(db): regenerate the sandboxes migration as 0273 Staging claimed 0271 and 0272 while this branch was out, so the hand-authored 0271_workspace_sandboxes was dropped before the merge and regenerated on top of the merged schema. Same DDL; drizzle emits plain CREATE TABLE/INDEX rather than the hand-added IF NOT EXISTS, which matches the repo default — that idempotent form is only needed for files with CONCURRENTLY ops below an embedded COMMIT. Regenerating also restores the meta snapshot the hand-authored migration never had. Co-Authored-By: Claude <noreply@anthropic.com> * chore(db): drop the sandboxes migration ahead of the staging merge Staging independently claims idx 0273, so remove ours before merging to avoid an add/add conflict on the drizzle migration index. Regenerated at the next free index once the merge lands. Co-Authored-By: Claude <noreply@anthropic.com> * chore(db): regenerate the sandboxes migration as 0275 Staging took 0273 and 0274, so the sandboxes DDL lands at the next free index. The emitted SQL is byte-identical to the dropped 0273. Co-Authored-By: Claude <noreply@anthropic.com> * fix(billing): consolidate the Max-tier entitlement onto one predicate The Max tier was spelled five ways. The odd one out — `isMax`, defined as `isPro(plan) && credits >= 25000` — excluded both `team_25000` and `enterprise`, and it was the sole input to the personal-workspace cap. A delinquent Max-for-Teams org admin got 1 personal workspace while a delinquent Max individual got 10. Only free/pro_6000/pro_25000 were tested, so the two broken tiers were unpinned. Separately, the server gate and the client `hasUsableMaxAccess` were independent copies of the same rule. The settings sidebar renders Sandboxes and Sim Mailer from the client one while the API answers 403 from the server one, so any drift renders a feature unlocked that the API refuses. - `MAX_TIER_CREDITS` is derived from the `CREDIT_TIERS` table; `isMaxTier` in plan-helpers is now the single definition, shared by the server gates, the client derivation, `getPlanTypeForLimits`, `plan-view`, and the cap - `hasWorkspaceTierAccess(id, predicate, { intent, onMissingWorkspace })` becomes the one org-vs-personal payer fork. `intent: 'active-use'` means active and not billing-blocked; `'retention'` means active/past_due with block state ignored, so the inbox teardown guard keeps its fail-open semantics instead of implying them through a duplicated fork - `isWorkspaceOnEnterprisePlan`'s personal branch now applies the status and block checks its own org branch always had, and its TSDoc names its real consumer (copilot BYOK, not Access Control) - the client live-sync gate gained the server's `isHosted` branch, so a self-hosted deploy with billing on no longer locks an interval the API accepts. It reads both flags directly rather than taking one as a parameter the callers sourced from the same module - `sqlIsPro`/`sqlIsTeam` escape the `_` LIKE wildcard, matching the already correct hand-rolled filter in seat-drift - deletes the `TERMINAL_SUBSCRIPTION_STATUSES` and `ENTITLED_STATUSES` shadow constants, and corrects three test mocks that asserted `trialing` was entitled or usable `max-tier-parity.test.ts` asserts the client and server answers match for every plan name. Both new guards were checked against the old code: the parity test fails 3 assertions with the previous predicate, and the self-hosted test fails without the `isHosted` branch. Co-Authored-By: Claude <noreply@anthropic.com> * chore(db): drop the sandboxes migration ahead of the staging merge Staging has claimed 0275 (table_views) and 0276 (drop_legacy_folder_tables) since the last merge, so our 0275_workspace_sandboxes collides on the index. Dropping ours first — the .sql, meta/0275_snapshot.json, and the journal entry — leaves packages/db/migrations byte-identical to the merge-base, so the merge sees no add/add conflict at all. Regenerated on the far side. Ours is the droppable side: plain additive DDL with no hand edits, which drizzle reproduces exactly. Staging's migrations are hand-written and must survive. Co-Authored-By: Claude <noreply@anthropic.com> * chore(db): regenerate the sandboxes migration as 0277 Staging claimed 0275 (table_views) and 0276 (drop_legacy_folder_tables), so the sandboxes migration dropped before the merge comes back on top as 0277. The emitted SQL is byte-identical to what was dropped — the original had no hand edits, so there is nothing to reapply. It is purely additive: two enums, sandbox_image and workspace_sandbox, their two FKs and six indexes. That it regenerated unchanged also confirms the schema.ts auto-merge was correct — had it lost staging's legacy-folder-table drops, drizzle would have emitted CREATE TABLE for them here. Snapshot chain is continuous (0273 -> 0277, each prevId matching the previous id) and the table counts track the DDL: 100 -> 101 (table_views) -> 99 (legacy folder tables dropped) -> 101 (the two sandbox tables). Co-Authored-By: Claude <noreply@anthropic.com> * feat(sandboxes): gate on the enterprise feature flags, drop the rollout switch Sandboxes shipped behind `custom-sandboxes`, an AppConfig rollout flag falling back to a `CUSTOM_SANDBOXES` secret. That made it the only Max-gated surface with no self-hosted path: `INBOX_ENABLED` can force Sim Mailer on for an operator running their own billing, and `ENTERPRISE_ENABLED` turns on the other nine features at once, but neither reached sandboxes. A self-hoster had to find a separately-named variable that was not part of that family, and one running with billing enabled could not enable it at all. Sandboxes now joins the enterprise feature set and the rollout flag is gone: - `sandboxes` is an `EnterpriseFeature` with `SANDBOXES_ENABLED` and its `NEXT_PUBLIC_` twin, so the master switch and the per-feature override both reach it like every sibling - `hasWorkspaceSandboxAccess` takes the inbox's shape exactly — the override wins, then a deployment without billing is unrestricted, then the workspace payer needs usable Max or Enterprise - the settings nav gains `selfHostedOverride`, so the section resolves through the same path as Sim Mailer instead of a second entitlement AND-ed in - `custom-sandboxes`, the `CUSTOM_SANDBOXES` secret, the now-unreachable `SANDBOXES_UNAVAILABLE` 403 copy, and the route's kill-switch branch are deleted Its legacy default is `true`, matching `inbox`: the gate already returns true whenever billing is off, so `false` would leave the nav override disagreeing with the gate that answers the request. Self-hosted builds run on the operator's own E2B/Daytona credentials, so there is no Sim-side cost to withhold — the docs now say so, since enabling the feature without a provider configured is the obvious trap. The new gate tests run with billing enabled on purpose; the `!isBillingEnabled` bail would otherwise answer every case and hide whether the override is wired. Verified by deleting the override line — exactly the one assertion fails. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): let the language menu match its trigger width `matchTriggerWidth={false}` exists for the opposite case — a narrow trigger whose option labels would truncate, letting the menu grow past it. The language field is a full-width form control with two short labels, so the override shrank the menu to "JavaScript" and pinned it to the right edge instead. The default (`true`) is correct here. Every other consumer passing `false` is a genuinely narrow trigger — a role picker in a member row, a table filter chip. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): re-queue a build when resolution finds the image unusable `ensureSandboxImage` only ran when a sandbox was saved, so resolution treated an unusable image as terminal and told the user to go fix a definition that was never wrong. Three states stuck permanently until someone re-saved in Settings: - a build that failed - a build whose worker died mid-flight, stranding the row in `building` - every sandbox created while the deployment ran a `runtime` provider, after a switch to a `prebuilt` one — `runtime` writes no image rows at all, so the whole fleet resolved to "no completed build" with nothing to repair it Resolution now re-queues through the registry's existing idempotent entry point before failing, and says a build is on its way instead of pointing at Settings. The conflict guard already claims only a `failed` row or a stale `pending`/ `building` one, so executions arriving during a healthy build enqueue nothing — no thundering herd from a hot workflow. The registry is imported dynamically for the same reason `sandboxDb` is: it pulls `@sim/db` into the static graph, which this module keeps out of the executor bundle. That also avoids a cycle, since the registry imports `invalidateSandboxResolution` from here. A repair that itself fails is logged and swallowed — it must never replace the build error naming the sandbox. Verified by deleting the repair call: exactly the three new assertions fail. Co-Authored-By: Claude <noreply@anthropic.com> * improvement(sandboxes): let the picker show just the sandbox name The label read "Test · Python · 1 package". The block's own list is already scoped to the language its sibling `language` subblock selects, so the language repeated on every row said nothing, and the package count is decoration next to the name that identifies the sandbox. The language stays for the one caller that cannot filter — agent tool-input renders this field under a synthetic id where the sibling `language` value is unreachable, so its list spans both languages and the name alone is ambiguous. That is the same missing value which disables filtering, so `showLanguage` is derived from it directly rather than passed independently and left to drift. A failed build is still marked: that suffix is the difference between a selection that runs and one that does not. Passing the flag also means dropping `.map(toSandboxOption)` for an explicit arrow — `Array.map` hands the index to the second parameter. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): show the sandbox name on the block card, not its uuid The card printed "443f4934-26ab-44ab-8...". `resolveDropdownLabel` only reads a subblock's static `options` array, and the sandbox picker is a `combobox` whose options load asynchronously, so its array is empty and the raw stored id fell through to the label. Resolved the same way skills and tools already are: a `resolveSandboxLabel` in the display layer, fed from the shared sandbox list query — the same cache entry the picker reads, so this adds no request. Two deliberate scopings: - the query is subscribed only for the sandbox row. `SubBlockRow` is memoized per subblock, and the list query polls while a build is in flight, so an unconditional hook would re-render every row on the canvas on each poll tick - the resolver matches the field id, not just the type. There is no dedicated subblock type for it, and matching `combobox` alone would relabel unrelated pickers An id with no matching sandbox resolves to null rather than a guess, so a deleted sandbox falls through to the caller's placeholder. The template preview surface is left alone: it is explicitly hook-free and passes empty lists for tools and skills too. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): hide the Sandboxes section with no provider configured Entitlement decides whether a workspace may author sandboxes; nothing decided whether anything could run one. A self-hosted deployment with SANDBOXES_ENABLED but no E2B or Daytona credentials got a fully functional tab whose output no Function block could select — the picker is gated on the provider vars, the tab was not. Both navigation planes now drop the section when neither NEXT_PUBLIC_SANDBOX_ENABLED nor the pre-Daytona NEXT_PUBLIC_E2B_ENABLED is set — the same pair the picker's `showWhenEnvSet` reads, so the two cannot disagree. Dropped rather than locked: an upgrade does not conjure a provider. The unified plane drops it in `buildUnifiedSettingsNavigation` rather than in the sidebar's filter, because the sidebar's `selfHostedOverride` short-circuit runs before its `requiresMax` check and would have revealed the tab anyway. It reads the browser twins, not the server's `isRemoteSandboxEnabled`, since this module renders on both sides. The predicate is a function, not a module constant, because the constant form was untestable and ambient: the env mock falls through to `process.env`, and `apps/sim/.env` (gitignored, so absent on CI) sets NEXT_PUBLIC_E2B_ENABLED=true. The nav tests passed locally and failed 6 assertions with the flag cleared. They now pin both flags, so the suite is identical with and without a local env file — verified by running it both ways. Co-Authored-By: Claude <noreply@anthropic.com> * docs(sandboxes): correct three claims the code no longer makes The Sandboxes section described behavior two commits on this branch changed, and led with an internal detail no reader needs. - entitlement is no longer Max/Enterprise only: self-hosted deployments unlock sandboxes with SANDBOXES_ENABLED, and the section is hidden outright when a deployment has no sandbox provider, which is the state a self-hoster is most likely to hit and least likely to diagnose - a build that is not Ready is no longer terminal. It is queued again on the next run, so the advice is to wait and re-run, not to go edit a package list that was never wrong - deleting a sandbox frees its build once nothing else references it. Builds are shared by content, so this is the one place a reader could reasonably assume deletion is immediate Dropped the `ModuleNotFoundError` aside: what the old code did instead is not something a reader needs to know to use the feature. The page is hand-written — `function` has category 'blocks' and is absent from `NATIVE_RESOURCE_BLOCK_TYPES`, so generate-docs skips it and these edits will not be overwritten. Co-Authored-By: Claude <noreply@anthropic.com> * feat(sandboxes): release the provider image when nothing references it Deleting a sandbox only removed its row, leaving the built template in E2B until the 30-day retention sweep — up to a month of paying to store an image nothing could select. Editing a package list had the same effect on the old content address, which is the more common case since every edit re-points the sandbox. `releaseSandboxImage(specHash)` now deletes the provider image and its row from both paths. It reuses the sweep's provider call and its ordering: image first, row second, so a refused delete leaves the row for the sweep to retry rather than orphaning a remote template nothing points at. Two guards make eager deletion safe: - builds are keyed by content, not by workspace, so two workspaces declaring the same package list share one image. The release no-ops while any sandbox still references the hash — otherwise one workspace's delete would break the other's - an in-flight build is left alone rather than raced; the sweep collects it once it settles Called detached from both routes. The row is already committed by then, so the user's action has succeeded whatever the provider says, and awaiting would hold a UI delete open on a remote call the sweep would retry anyway. Every failure inside is logged and swallowed for the same reason. E2B's delete verified against their API reference: DELETE /templates/{templateID} with X-API-Key, 204 on success. The existing implementation already matched, so this commit only adds the call sites and the guards. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): rate-limit the automatic rebuild, drop the one-off status dot Two follow-ups to the resolution repair. The repair had no rate limit. `ensureSandboxImage` re-claims a `failed` row on sight, and a bad package name fails in seconds, so the in-flight guard never closed the window: a workflow on a one-minute schedule would enqueue a build a minute against a package list that will never resolve, each one real provider build compute. Before the repair existed resolution simply threw, so this was introduced with it. The two callers want different things, so the cooldown is opt-in. A save is a person explicitly asking for another attempt and still retries immediately; resolution passes `FAILED_BUILD_RETRY_COOLDOWN_MS` and gets at most one attempt per window no matter how often the workflow runs. Ten minutes: long enough that per-minute runs cannot drive per-minute builds, short enough that a transient registry outage clears within the hour. The status line loses its colour dot. `size-[6px] rounded-full` appeared in exactly one file in the repo, so it was a new primitive rather than a pattern, and it duplicated state the text colour already carries — the label now turns `--text-error` on a failed build, which is what every other status row in settings does. `ChipTag` was the wrong home for this: its variants are `mono`/`invite`, with no semantic tone, so a status version would have meant overriding its chrome from the consumer. Also corrects the docs line this changes: a failed build is retried periodically, and saving is the way to retry now, so "wait a moment and run again" no longer describes it. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): claim the image row and its reference check in one statement Greptile P1. Reading references in one statement and deleting in another left a window — a wide one, since a provider delete is a network call — where a second workspace could declare the same package list, inherit the `ready` row, and have its next run fail against a template already on its way out. Content addressing is what makes that reachable: the image is shared, so one workspace's delete can strand another's sandbox. The reference check now lives in the conditional DELETE itself, so winning the delete is the proof that nothing referenced the hash. A workspace that adopts the hash first makes the delete match nothing and the release becomes a no-op. Claiming the row before the provider call would otherwise strand a template nothing points at if the provider then refused, so that path puts the row back and the retention sweep inherits the retry — the same property the previous ordering had. The sweep is deliberately left as it is: its equivalent window needs a hash unreferenced AND unused for 30 days, and its provider-first ordering encodes the documented retry-on-refusal behaviour this path now reproduces explicitly. No transaction is opened. The provider call sits between discrete statements rather than inside one, so no pooled connection is held across it — which is why this uses a conditional delete instead of the repo's `pg_advisory_xact_lock` pattern, whose lock only releases at commit. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): route the retention sweep through the same image claim Cursor and Greptile both flagged the sweep as still carrying the interleaving just fixed in releaseSandboxImage, and they are right — the reason given for leaving it alone last round does not survive scrutiny. That reason was that provider-first ordering encodes retry-on-refusal, so making the claim atomic would trade a race for an orphaned template. The release path already answers that: claim the row, and put it back if the provider refuses. The sweep can have both properties too. The rarity argument was also weaker than stated. The sweep nominates up to 200 candidates and then works through them eight network deletes at a time, so its check-to-delete gap is seconds to minutes — wider than the window that was just closed, not narrower. Both callers now share `claimAndDeleteImage`, which owns the whole contract: the unreferenced check lives inside the DELETE, the provider call runs only after the claim succeeds, and a refusal restores the row. Having written that ordering twice is what let the two paths drift, so it exists once now. The sweep's query becomes a nomination step only. Its retention cutoff is passed into the claim rather than trusted from the earlier read, so a candidate that stops qualifying mid-sweep fails its claim and is skipped instead of losing its image. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): rebuild a hash adopted while its image was being deleted Greptile's third pass on this path, and a case the previous two did not cover: the adopter starting a *fresh build* rather than inheriting a ready row. Claiming removes the registry row, so between that and the provider delete finishing, a workspace can declare the same package list, get a new row, and start a build under the same content-derived imageRef — which the in-flight delete then removes. The window itself is inherent. The registry row and the provider template are two systems with no shared transaction, so it can be narrowed but not closed. A Redis lock would not close it either: acquireLock returns true when Redis is absent, so it cannot be a correctness guarantee for self-hosted. Holding a Postgres advisory lock would, but only by pinning a pooled connection for the length of a provider call, which is a worse trade. What was avoidable is the adopter finding out the slow way. Its row is new and healthy-looking, so nothing noticed: resolution only repairs a row that is missing or failed, and a failed one waits out the retry cooldown first. The release path now re-checks after the delete and re-enqueues, so the rebuild starts immediately instead of one failed run plus a cooldown later. A build already in flight is left to the conflict guard, since it may still outlive the delete. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): reclaim a ready row whose image was deleted underneath it Greptile found the hole the previous commit left, and it is the case that made the claim in that commit's message wrong: this one is permanent, not transient. If a re-adopted hash reaches `ready` before the in-flight provider delete lands — plausible, since E2B layer caching can rebuild an identical spec in seconds — the row looks healthy while its imageRef points at nothing. Resolution repairs a row that is missing or failed, never one claiming to be ready, so nothing recovers it. The sandbox stays broken until someone re-saves it by hand. `rebuildIfReadopted` called `ensureSandboxImage` with no options, whose conflict guard reclaims only a failed or stale in-flight row, so it silently did nothing in exactly that case. The release path now passes `imageKnownGone`, which widens the re-claim to any settled row rather than only a failed one. It is the one caller that knows the image is gone regardless of what the row says. An in-flight build is still left alone: it either recreates the template it was building or fails into the normal repair path, and resetting it would only add a duplicate build. The three ways a settled row may be re-claimed now sit in one `settledRebuildBranch` helper — any settled row when the image is known gone, a failed one after the cooldown for an automatic caller, a failed one immediately for a person — because inlining the third case is what hid the gap. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): let a same-spec save retry a failed build Cursor Bugbot. `scheduleSandboxBuild` sat inside the changed-hash branch, so a save that did not alter the package list never reached the registry. The comment above it described the opposite — that an unchanged spec finds a ready row and enqueues nothing — which is what `ensureSandboxImage` does, but only if it is called. That made the docs wrong too. They tell a reader to save the sandbox again to retry a failed build immediately, and this branch is exactly why that did nothing: the only way to retry was to edit the package list into a different hash, which is not what someone recovering from a transient registry failure wants to do. The call is now unconditional and the registry decides what a save costs, which is what its conflict guard is for: a ready or in-flight row is left alone, a failed one is re-claimed at once. Releasing the previous image stays behind the hash check, since only a changed hash orphans one. Cache invalidation is unchanged — `scheduleSandboxBuild` already does it, which is why the else branch existed. Co-Authored-By: Claude <noreply@anthropic.com> * docs(sandboxes): correct the image cache's staleness invariant Cursor Bugbot found that a released image can still be served from another replica's cache. The finding is real, and the reason it went unnoticed is that the cache documented an invariant which eager release quietly broke. It claimed a `ready` row is terminal for its spec hash, so a cached hit could not go stale in a way that matters. That held while the only ways a row changed were an edit (new hash) or a delete (caught by the `workspace_sandbox` read). Releasing an image eagerly made a `ready` row disappear with the hash unchanged, so the premise no longer holds and the comment was actively misleading to the next reader. No behaviour change here — the exposure is bounded at IMAGE_TTL_MS on replicas other than the one that ran the release, and it self-heals once the entry expires and the row read finds nothing. Closing it properly needs cross-replica invalidation or a provider-error path that invalidates on "template not found", both of which are larger than a review fix; the comment now says so instead of implying the problem cannot exist. Co-Authored-By: Claude <noreply@anthropic.com> * docs(sandboxes): note that a JavaScript sandbox needs an import to apply Cursor Bugbot pointed out that `useRemoteSandbox` keys on detected static import/require and never on the selected sandbox, so JavaScript without one runs locally and the selection has no effect. Keeping the behaviour: honouring the selection would force those blocks remote, and the large-value-ref guard immediately below would then reject code that runs fine today. Documenting it instead, next to the picker, since a selection that silently does nothing is only surprising if nothing says so. Python is unaffected — it always runs remotely, so its sandbox always applies. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): stop create mode surviving a return to an open sandbox Cursor Bugbot. Create mode and having a sandbox open are mutually exclusive, but nothing enforced it, so both could be set at once — and the screen then lied about which sandbox its Delete pointed at. With `isCreating` true and `selectedId` restored, `baseline` is null, so the editor renders an empty "New sandbox" form, while the Delete action is built from `selected` and still targets the restored sandbox. An admin looking at a blank create form could delete a sandbox it never named. Two ways in, both closed: - Browser Forward after starting a new sandbox restores `selectedId` without going through `closeEditor`. The render-time sync that already drops a stale draft now also leaves create mode, which is the same class of correction and the reason that block exists. - "New sandbox" set `isCreating` without clearing `selectedId`, so the same contradiction was reachable without touching history at all. It now clears the selection, with `history: 'replace'` because switching mode is not a destination. Co-Authored-By: Claude <noreply@anthropic.com> * fix(ci): pin the sandbox flag in the second nav catalog test, bump the chart Two CI failures, both mine. `app/workspace/[workspaceId]/settings/navigation.test.ts` asserts the unified catalog and was left on ambient env. Dropping the Sandboxes section without a sandbox provider made it 26 items instead of 27 on CI, which has no `apps/sim/.env` — the same trap already fixed in the sibling `components/settings/navigation.test.ts`, in the one file that was missed. Fixing it needs `vi.hoisted` rather than the sibling's `beforeEach`, because this file reads `allNavigationItems`, built once at module load; a hook would run after the value it is trying to influence already exists. The chart gate is separate: this branch adds sandbox settings to `helm/sim/values.yaml`, and the workflow requires a Chart.yaml bump whenever `helm/sim/**` changes. Additive config, so 1.3.0 -> 1.4.0 by SemVer. Verified by running the whole suite with the flags forced off, not just the two navigation files — no other test depends on a local env file. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): keep the row restore to a refused delete only Cursor and Greptile, independently, on the same code. `deleteImage` and `rebuildIfReadopted` shared one try/catch, so a rebuild failure after a *successful* provider delete was handled as if the provider had refused: the catch put the claimed row back, `ready` status and all, pointing at a template that no longer exists. That is the one state resolution cannot repair — it fixes a row that is missing or failed, never one claiming to be ready — so it reintroduced the permanent breakage an earlier commit had just closed, through the error path rather than the happy one. Restoring now belongs strictly to a refused delete. Once the template is gone the row stays gone, and the rebuild runs past that catch. The rebuild also swallows its own failures: it follows a delete that already succeeded, so it must not be reported as a failed release, and inside the sweep it must not reject the rest of its chunk. The adopter's next run still reaches the normal repair path. The regression test drives a rebuild failure and asserts no row is restored. It fails against the original shape — rebuild inside the shared try, no inner catch — which is what the two reviewers were describing. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): drop the dead row when a re-adopt rebuild cannot be scheduled Greptile, one layer under the previous fix. Making the post-delete rebuild swallow its own failures kept it from being reported as a failed release, but left the adopter's row claiming a `ready` image whose template is already deleted — the one state resolution cannot repair, since it rebuilds a row that is missing or failed and never one that says ready. So the row is now dropped when the rebuild does not take. That turns the adopter into the missing-row case, which the next execution repairs on its own, instead of a sandbox that stays broken until someone re-saves it by hand. A failure to drop it is logged at error, because at that point two writes in a row have failed and there is nothing further this path can do. Also gives the release tests a default "nothing re-adopted" select. Without it the rebuild threw on an unstubbed mock and the cleanup delete overwrote the predicate the claim assertions read, so two of them were passing on the wrong statement. Co-Authored-By: Claude <noreply@anthropic.com> * feat(sandboxes): repair a missing image at create, where the truth is observable Six review rounds narrowed the window between deleting a shared template and another workspace adopting its content hash, and each fix exposed the next facet. They all share a cause: the registry row and the provider template are two systems with no shared transaction, so any scheme that keeps them in step is guessing. Create is the one step that does not have to guess. It either gets a sandbox or it does not, so a `ready` row pointing at a deleted template now corrects itself the first time it is used, rather than needing someone to re-save the sandbox. - `SandboxImageBuilder.isMissingImage` asks the provider to classify its own failure. Prebuilt-only, because a runtime provider has no image to miss - E2B answers it off `NotFoundError`, which the SDK maps from a 404. The only resource a create names is the template, and the two subclasses that describe other calls — a missing file, an exited sandbox — are excluded. The classifier stays deliberately narrow: treating auth or rate-limit failures as a missing image would turn a provider outage into a build storm - `repairMissingSandboxImage` invalidates the cache, rebuilds with `imageKnownGone` (no cooldown, since this observed the image is gone rather than inferring it), and returns copy telling the author to run again - `ResolvedSandbox` carries `specHash` so the failing execution can name what to rebuild This subsumes the open facets rather than adding another guard beside them: the stale per-replica cache, an adopter left `ready` against a deleted ref, and a rebuild that never took all end at the same place — the next run repairs itself. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): key the build trigger by attempt, not by spec Cursor Bugbot. The Trigger.dev idempotency key was the content address alone, so a second attempt at the same spec was deduped against the first: the SDK returns the finished run instead of starting one, and the row that `ensureSandboxImage` just flipped to `pending` sits there with no worker. Nothing can re-claim a `pending` row until it goes stale, so a retry inside the 5-minute TTL did nothing for the next half hour. That silently disabled every repair path — save-to-retry, which the docs name explicitly, and both the resolution and create-time rebuilds. The key's own comment already said it exists "to collapse concurrent saves of the same spec into one build, not to suppress a retry after one failed". The conditional update above it is what actually collapses concurrent saves: only one caller gets a row back, so only one ever reaches the trigger. Keying by the claim's `updatedAt` keeps that property and makes each genuine attempt distinct, while a duplicate delivery of one attempt still collapses. Co-Authored-By: Claude <noreply@anthropic.com> * feat(sandboxes): create a sandbox from the picker, and fix three UI papercuts The Function block's sandbox field now pins a "Create Sandbox" row above its options, matching the "Create Skill" / "Create Tool" rows it sits beside, so authoring a package list no longer means leaving the workflow for Settings. The row is declared by the field (`createAction`) rather than hardcoded by id; block configs are read by the serializer and executor, so the name maps to a modal in the picker rather than carrying a component. Two things the modal has to get right. It seeds the new sandbox's language from the sibling the list is scoped by, or a sandbox created off a JavaScript block would land in the Python list and vanish. And the created option is held locally until a real fetch carries it, or the field would sit on a raw uuid until hydration answered. Also: - The Sandboxes icon was the Logs block's icon (`blocks/blocks/logs.ts`), in both the settings nav and the list rows. It is the Function block's now. - "Default image (no extra packages)" claimed something untrue: E2B and Daytona base images both ship with packages installed. - A new sandbox opened in Python while the Function block defaults to JavaScript. The test pins the two together rather than the literal. Draft shape and helpers moved out of the editor component into `utils.ts` — three consumers now, and it makes the defaults testable without a DOM. * feat(settings): one Max-plan wall, and give the create modal the same one The create-sandbox modal answered a non-Max workspace with a red line under a form it could never submit, and no way to act on it. It now renders the same wall the Settings > Sandboxes tab does — heading, one sentence on what the plan unlocks, and an Upgrade to Max chip — instead of the fields. That wall existed twice already (sandboxes and Sim Mailer), so this extracts it rather than adding a third copy. `SettingsUpgradeNotice` owns the copy rhythm and the route, and `compact` trades the page's full-height centering for a modal's. Both settings consumers now compose it; neither keeps its own markup. The action lands on billing, which `resolveSettingsHref` already redirects to the plan-comparison page for a member who cannot manage billing — so it is a route to explore plans, never a dead end. The chip stays hidden for non-admins, exactly as the settings pages had it. A non-admin on an entitled workspace gets the muted reason rather than the upgrade wall: buying a plan is not what is in their way. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
c809845b99 |
improvement(self-host): enterprise features enabling (#6028)
* improvement(self-host): enterprise features enabling * chore(helm): bump chart to 1.3.0 for the enterprise self-host values values.yaml gained the ENTERPRISE_ENABLED switch and INSTANCE_ORG_* keys, and the feature-flag envDefaults moved from "false" to empty so the master switch can resolve them. Additive and backward compatible, so a minor bump. * fix(self-host): address review findings on instance org and org delete Drop the per-process instance-org id cache. It went stale once the organization was deleted through the Admin API, and clearing it from the delete handler would only heal the replica that served that request. The lookup runs on the signup path against a single-row table, so re-reading costs nothing and keeps every replica self-correcting. Scope the org-delete subscription conflict to entitled statuses. Matching any row regardless of status let a canceled subscription — which bills nobody — permanently block deletion. * fix(admin): block org delete on any live subscription, not just entitled ones ENTITLED_SUBSCRIPTION_STATUSES excludes trialing, so a trial — which grants no entitlement but is a live Stripe subscription that will convert — slipped past the delete guard and could be stranded against a removed organization id. Adds TERMINAL_SUBSCRIPTION_STATUSES and inverts the predicate: block unless the row is finished. Expressed as the terminal set so a status Stripe adds later defaults to blocking, which is the safe direction for a destructive operation. * fix(self-host): resolve SSO and access-control in the UI, not the raw env var Nine client consumers still read NEXT_PUBLIC_SSO_ENABLED / NEXT_PUBLIC_ACCESS_CONTROL_ENABLED directly while the server gates and settings nav had moved to the resolver. With only ENTERPRISE_ENABLED set that produced dead ends: the SSO settings section appeared but ssoClient() was never registered and no login button rendered, and the Access Control section appeared but its page reported "not entitled". Points every consumer at isSsoEnabled / isAccessControlEnabled so visibility and capability come from one place. * fix(admin): validate retention workspace targets on the Admin API too retentionOverrides and per-workspace PII rules both name a workspace, and neither field is a foreign key. The settings UI rejected ids belonging to another organization; the Admin API did not, so the two paths could persist different data for the same org. Extracts the check as getForeignWorkspaceTargetsReason and points both routes at it, so they cannot drift apart again. * fix(self-host): close three review findings on admin routes and cleanup Make org delete atomic. detachOrganizationWorkspaces committed on its own, so a failed delete left workspaces detached and re-billed while the organization, its members, and its settings survived. Adds a Tx variant so both commit together. Gate the admin session-policy PATCH on entitlement, matching the settings UI. Without it the stored policy was inert — getSessionPolicy resolves to no-op when the feature is off, so the one eager clamp would be undone on the next refresh. Stop emitting plan-wide housekeeping when billing is off. It is keyed to the hosted free-tier 30-day window, the same default the per-workspace pass deliberately refuses to apply off-hosted. * fix(admin): gate whitelabel on entitlement and emit detach audits post-commit The Admin whitelabel PATCH skipped the entitlement check the settings UI runs, so an admin key could set branding the product had not granted the org. detachOrganizationWorkspacesTx also wrote its audit rows inside the caller's transaction, contradicting its own doc comment — a rolled-back delete would have left audit history describing detachments that never happened. It now returns the rows and callers emit them after commit. * fix(self-host): refuse instance-org resolution when the slug is ambiguous organization.slug has no unique constraint, and the lookup took the first of however many matched. The choice is unordered, so two replicas could resolve different organizations and split new signups between them. Resolution is now three-state. Ambiguity is distinct from absence, so it both declines to adopt an arbitrary organization and declines to provision another one on top of the duplicates. |
||
|
|
90f6708270 |
improvement(helm): hygiene pass — CI gating, strict values schema, ESO v1 default (#5939)
* improvement(helm): hygiene pass — CI gating, strict values schema, ESO v1 default, ci values Closes the gaps from a best-practices audit of the chart (template-level conformance was already clean: full label set, 82 unit tests, kubeconform- valid renders): - new Helm Chart workflow gates every chart change: helm lint, the 82 helm-unittest cases, kubeconform validation of default + all-components renders (k8s 1.29 strict, CRD catalog), and a render of all 10 example values files - helm/sim/ci/ values files (chart-testing convention) so the chart lints and templates cleanly out of the box with dummy secrets - values.schema.json declares all 30 top-level keys (16 were invisible) and sets root additionalProperties: false, so top-level typos fail fast - externalSecrets.apiVersion defaults to v1: current ESO releases removed the v1beta1 compatibility path in 2026, so the old default produced rejected manifests on new installs; NOTES/values comments updated - wait-for-postgres init container gets requests/limits (the only container in the chart without them; broke ResourceQuota'd namespaces) - drop the telemetry Prometheus scrape config for app/realtime — neither exposes /metrics, so it was dead config that also rendered a realtime target with realtime disabled - chart 1.2.0 with README upgrade notes * fix(helm): version-agnostic schema-error assertions in the secret-length suite Helm v4 phrases schema rejections as 'minLength: got N, want 32' while v3 says 'String length must be greater than or equal to 32'. The suite grepped for 'minLength' only, so it passed on local helm v4 and failed on CI's v3.16.4 — the enforcement itself works on both. Patterns now assert the key name plus either wording. Caught by the new Helm Chart workflow on its very first run. * improvement(helm): kind install test + chart version-bump gate in CI Benchmarked against the flagship OSS charts (ingress-nginx, argo-cd, kube-prometheus-stack, grafana, bitnami, cert-manager): sim already exceeds most of them on validation rigor (strict schema — 4 of 6 ship none; 82 unit tests vs argo-cd's zero; kubeconform manifest validation none of them run), but every top community chart repo actually installs the chart on a kind cluster in chart CI — the one majority practice we lacked. Adds: - install job: kind cluster, helm install with ci/default + a new small-footprint ci/kind-values.yaml overlay (default app requests of 4Gi can't schedule on a CI node), --wait, then the chart's helm test hook, with pod/event/log diagnostics on failure - version-bump job (PR-only): fails when helm/sim/** changes without a Chart.yaml version increment — argo/kps/grafana all enforce this * fix(helm): declare naming overrides in the strict schema; SHA-pin CI actions - nameOverride/fullnameOverride are consumed by sim.name/sim.fullname but were never in values.yaml, so root additionalProperties: false rejected Helm's standard naming overrides — both now declared (a helper-wide sweep confirmed they were the only template-read keys missing), with a smoke regression test that installs under the strict schema and asserts the override lands in resource names - CI supply-chain hardening: checkout/setup-helm/kind-action pinned to full commit SHAs (tag comments retained) and the kubeconform archive verified against its published sha256 * fix(helm): immutable unittest runner and fail-closed version gate in CI - helm-unittest runs via the project's official docker image pinned by immutable sha256 digest instead of a plugin install from a mutable git tag - the version-bump gate fetches the full base ref (a --depth=1 fetch could leave no merge base), computes the merge-base and diff outside the if so any git failure fails the job instead of falling into the skip branch * fix(ci): run the helm-unittest container as the runner UID The digest-pinned image runs as a non-root user that cannot write into the runner-owned bind mount (it creates tests/__snapshot__, absent from the checkout since empty dirs aren't tracked). Standard bind-mount pattern: --user "$(id -u):$(id -g)" with HOME=/tmp for helm's cache. * fix(helm): appVersion points at a real GHCR tag (v0.7.44) The kind install job caught this on its first full run: the default image tag (Chart.AppVersion 0.6.73) returns 404 on GHCR for all three images — the registry's tags are v-prefixed — so an unpinned default install could never pull. Updated appVersion to v0.7.44 (verified 200 for simstudio, realtime, and migrations manifests), stale values comment refreshed, and an upgrade note added. The CI kind values deliberately stay tag-free so the job keeps exercising the true default path. |
||
|
|
d48722a04e |
fix(helm): correct chart docs, examples, and dead config across the board (#5907)
* fix(helm): correct chart docs, examples, and dead config across the board Audit-driven accuracy pass over the chart's entire documentation surface, verified by rendering every example against the templates: - migrations run as an init container on the app pod, not a Job — fix the README component list, troubleshooting commands, and sim-helm skill refs; drop the dead migrations-job NetworkPolicy ingress rule - referenced-but-never-created resources: document the GKE ManagedCertificate creation (values-gcp), comment out the key-file Secret mount that stuck all pods in ContainerCreating (values-gcp), enable certManager for the postgres TLS issuerRef (values-production), add the cert-manager cluster-issuer annotation nginx needs (values-azure) - values-external-db: networkPolicy.egress is a list, not a map (the map rendered an invalid manifest); fill schema-failing placeholder host/username - realtime >1 replica requires REDIS_URL (Socket.IO Redis adapter) — default examples to 1 replica with the scaling note, and warn where autoscaling HPAs override replicaCount - pod anti-affinity selectors matched nothing (simstudio vs sim name label) - kubernetes.mdx: install commands were missing required CRON_SECRET and postgresql password (failed at template time), wrong deployment name in port-forward, stale version requirements, unsupported key-remapping claim - remove unimplemented app.secrets.existingSecret.keys from values + schema; fix README PDB default, cronjob list, /metrics caveat, NOTES secret count, Azure-only StorageClass in generic examples, dead SOCKET_SERVER_URL and GOOGLE_CLOUD_* env, ESO apiVersion mismatch, and skill-reference drift - bump chart to 1.0.1 * fix(helm): review round 1 — scoped example egress, in-tab secret note, copilot Job wording - external-db example egress scopes to a placeholder database CIDR instead of to: [] (which allowed every destination on 5432, defeating the isolation the example teaches) - kubernetes.mdx cloud tabs state explicitly that they reuse the variables generated in the Installation block - Copilot migrations really do run as a Helm-hook Job — restore Job wording there (only the app migrations are an init container) * fix(helm): template-sweep fixes — telemetry validity, ESO rollout checksums, dead passwordKey knob - telemetry: memory_limiter gets the required check_interval (collector failed startup validation whenever telemetry.enabled=true); the jaeger exporter was removed from collector-contrib in v0.86 — export to Jaeger via its native OTLP endpoint instead (otlp/jaeger, default port 4317) - app/realtime rollout checksums now hash the ExternalSecret manifest too, mirroring the copilot pattern — with ESO enabled the inline Secret renders empty, so remoteRefs changes never rolled the pods - remove the unimplemented existingSecret.passwordKey knob (values, schema, README, dead helpers): nothing consumed it, and a non-default value silently produced a DATABASE_URL with an unexpandable placeholder; secrets must use the standard POSTGRES_PASSWORD / EXTERNAL_DB_PASSWORD keys - drop the orphaned sim.migrations.labels helper (its only consumer was the dead NetworkPolicy rule removed earlier) - helm test pod image resolves through sim.image so global.imageRegistry mirroring applies; NetworkPolicy realtime-ingress comment reflects actual traffic direction; smoke unittest suite loads the newly referenced external-secret template * chore(helm): bump chart to 1.1.0 with upgrade notes Removing (inert) documented values keys and changing the rollout-checksum inputs is a values-surface change — per SemVer chart conventions that is more than a patch. Adds an Upgrading section documenting the one-time pod roll, the removed no-op keys, and the Jaeger-over-OTLP change. * fix(helm): review round — external-db NP opt-in with real-CIDR-first flow, prod Jaeger OTLP endpoint - external-db example ships networkPolicy disabled so a verbatim install always reaches the database; the scoped egress rule stays as the documented opt-in (set your CIDR first, then enable) - values-production still pointed telemetry.jaeger at the legacy 14250 collector port — now Jaeger's OTLP gRPC endpoint to match the otlp/jaeger exporter * feat(helm): autoscaling.realtime.enabled toggle so examples can scale the app without unsafe realtime replicas Cursor correctly flagged that comment-level warnings didn't stop a verbatim production/external-db install from running the realtime HPA at minReplicas 2 without REDIS_URL (silent cross-pod event loss). Adds an opt-out toggle (default true — existing deployments unchanged): the realtime HPA renders only when autoscaling.realtime.enabled, and the realtime Deployment keeps spec.replicas under its control when the HPA is excluded. The three autoscaling examples set it false with the Redis rationale; README and upgrade notes document the toggle. * fix(helm): review round — whitelabeled realtime HPA opt-out, external-db isolation on with required CIDR in install flow - values-whitelabeled now actually sets autoscaling.realtime.enabled: false (the earlier batch aborted before reaching this file — Cursor caught it) - external-db keeps networkPolicy enabled (no isolation regression); the DB egress CIDR is marked REQUIRED and wired into both documented install commands via --set, so the copy-paste flow sets the real subnet in the same breath as the DB host * fix(helm): use a syntactically valid example CIDR in the external-db install commands An unreplaced <YOUR_DB_CIDR> literal fails Kubernetes CIDR validation and aborts the install; every sibling placeholder in the same command (host, username) installs fine and simply doesn't connect until replaced. The CIDR now behaves the same way: valid example value (10.20.0.0/24), explicitly marked as the operator's database subnet in both commands and the values comment. * fix(helm): remove text after continuation backslashes in external-db install commands A trailing comment after the line-continuation backslash (and equally a comment line spliced mid-command) breaks the shell command when copied — the remaining --set overrides run as separate commands and required-value validation fails. Both documented commands now reconstruct to bash-clean multi-line invocations (verified with bash -n); the CIDR guidance lives in the networkPolicy section comment. * fix(docs): cloud-tab installs are alternatives via helm upgrade --install Following Installation and then a cloud tab ran helm install twice for the same release and failed on the second. The tabs now state they replace the generic install and use helm upgrade --install, which is idempotent and also converts an existing generic install to the cloud values. * fix(docs): honest conversion caveat for cloud-tab upgrades over an existing install The cloud values rename the bundled Postgres database to simstudio, which Postgres only applies at first initialization — an in-place conversion of a generic install would point DATABASE_URL at a nonexistent database. Document the two safe paths: keep the original name via --set, or uninstall + delete PVCs and install fresh. * fix(helm): external-db header install command declares its secrets and includes CRON_SECRET The primary documented command failed the chart's required-value validation (missing app.env.CRON_SECRET with cronjobs default-on) and referenced an undeclared DB_PASSWORD. It now shows the export lines for every variable it uses and sets CRON_SECRET; both commands verified with bash -n. * chore(helm): restore values.schema.json formatting — surgical deletions only The earlier programmatic edit reformatted the whole file (~390 lines of whitespace churn hiding the 12 real deleted lines). Re-applied the removal of the dead keys/passwordKey properties as text-level deletions preserving the original style. * fix(helm+docs): explicit secret exports in external-db header; conversion must reuse original secrets - all five export lines are written out (three were only named in a trailing comment, so a verbatim copy passed empty required values) - the cloud-conversion caveat now leads with reusing the original secret values (helm get values) — a regenerated ENCRYPTION_KEY makes previously encrypted credentials undecryptable |
||
|
|
c083be9def |
improvement(auth): bump better-auth to 1.6.23 and add trusted-proxy client IP resolution (#5857)
* improvement(auth): bump better-auth to 1.6.23 and add trusted-proxy client IP resolution * chore(billing): record checkout-scope mirror re-verification against @better-auth/stripe 1.6.23 * chore(deploy): expose AUTH_TRUSTED_PROXIES in docker-compose.prod and Helm chart |
||
|
|
5eafa86992 |
feat(email): native Gmail API mail provider for GCP self-hosting (#5736)
* feat(email): native Gmail API mail provider for GCP self-hosting Adds Gmail as a fifth transactional mail provider (Resend → SES → SMTP → ACS → Gmail). GCP has no first-party SES/ACS equivalent, so the native Google path is the Gmail API: a service account with domain-wide delegation impersonates a Workspace sender (GMAIL_SENDER) and posts the raw RFC 822 message (built via nodemailer's MailComposer — full parity incl. attachments, replyTo, unsubscribe headers) to the media-upload messages.send endpoint. The Workspace SMTP relay alternative is documented against the existing SMTP provider. * fix(email): review round 1 + audit hardening for Gmail provider - a 2xx from Gmail with an empty/malformed body no longer surfaces as a send failure (the mailer's fallback chain would deliver the same email twice); covered by a regression test - normalize bare-LF line endings in html/text bodies to CRLF (RFC 822) before composing the raw message - add multi-recipient and text-only test cases |
||
|
|
f6bb8e6d3e |
feat(storage): native Google Cloud Storage support for self-hosting (#5728)
* feat(storage): native Google Cloud Storage support for self-hosting
Adds GCS as a third object-storage backend with full parity with S3 and
Azure Blob: uploads, streaming downloads, deletes, head, V4 signed URLs
(single + batch), and browser/server multipart uploads via the GCS XML
API. Selection precedence is Azure Blob > S3 > GCS > local disk.
- new provider client at lib/uploads/providers/gcs (cached singleton,
ADC/Workload Identity or inline GCS_CREDENTIALS_JSON auth)
- per-context GCS_*_BUCKET_NAME config wired through getStorageConfig
- shared getServeStoragePrefix() replaces hardcoded blob/s3 serve paths
- docs (object-storage, environment-variables), .env.example, helm
values.yaml + values-gcp.yaml storage section
* fix(storage): review round 1 — GCS per-context bucket fallback + ETag quote normalization
- getGcsConfig falls back to the general bucket for every context (GCS bucket
names are globally unique, so the S3-style sim-execution-files literal default
would point at an unowned bucket; empty per-context buckets previously made
uploads and downloads disagree)
- completeGcsMultipartUpload restores quotes on ETags stripped by the shared
browser upload client before building the completion XML
- docs/.env.example/helm updated for the fallback behavior
* fix(storage): review round 2 — route chat authz and execution-URL detection through getStorageConfig
- getChatStorageConfig delegates to getStorageConfig('chat') (identical for
S3/Azure, picks up the GCS general-bucket fallback instead of reading the
raw chat config and rejecting valid chat files)
- parse route resolves the execution bucket via getStorageConfig('execution')
for all providers, so GCS execution files in the fallback bucket are still
recognized as our own objects
* fix(storage): validation pass — gcs serve-prefix parity in key parsers + CORS doc fix
- extractStorageKey, extractFilename, and extractEmbeddedFileRef now strip the
gcs/ serve prefix like s3/ and blob/, so direct-uploaded files on GCS parse,
delete, download, and embed correctly (previously only the serve route knew
the prefix)
- file-download storageProvider union includes 'gcs'
- completeGcsMultipartUpload defensively rejects a 200 response carrying an
XML error document
- docs: CORS example lists concrete x-goog-meta-* header names (GCS matches
responseHeader entries exactly; wildcards are only supported for origin)
|
||
|
|
4d6301c900 |
feat(pii): regex-only block-output redaction + drop GLiNER/GPU image (#5697)
* chore(pii): remove GLiNER/GPU image + add spaCy-skip fast path to CPU server * feat(pii): restrict block-output redaction to regex-only entities * fix(pii): derive spaCy-NER set from registry + skip fast path when score_threshold set * fix(pii): include ORGANIZATION in app-side NER set (align with server) |
||
|
|
2e5b33c2db | feat(community): replace Discord community links with Slack across app, docs, emails, and readme (#5653) | ||
|
|
2d41360807 |
improvement(concurrency): limits configurable, docs updates (#5640)
* improvement(concurrency): limits configurable, docs updates * remove dead tests * limits self hosted vars |
||
|
|
896d15b666 |
fix(helm): close blocker/real-gap findings from Helm chart best-practices audit (#5555)
* fix(helm): close blocker/real-gap findings from Helm chart best-practices audit Verified every finding against the official Helm docs and Kubernetes Pod Security Standards docs before fixing, and validated each fix with helm lint/template plus the chart's own helm-unittest suite (65 -> 79 tests, all new tests confirmed to fail on the pre-fix code): - Blocker: values.schema.json documented "minimum 32/8 characters" on BETTER_AUTH_SECRET/ENCRYPTION_KEY/postgresql.auth.password but never enforced it. Added anyOf minLength-or-empty constraints (empty stays legal for existingSecret/ESO modes) — verified negative/positive cases live, no regression for any secret-delivery mode. - Real gap: copilot didn't support the External Secrets Operator mode the rest of the chart offers (app/postgresql/externalDatabase). Added external-secret-copilot.yaml, remoteRefs.copilot, and extended sim.copilot.validate with the same "map it or remove it" fail-fast guard app.env/realtime.env already have. Verified byte-identical rendering for the existing non-ESO path. - Real gap: the OpenTelemetry Collector was the only workload missing the shared Restricted-profile securityContext helpers (no container-level hardening at all). Wired sim.podSecurityContext/containerSecurityContext in, preserving the collector's original UID/GID/fsGroup. - Real gap: copilot templates hand-rolled label/selector blocks instead of using the chart's established sim.<component>.labels/selectorLabels pattern. Added sim.copilot.*/sim.copilotPostgresql.* helpers and refactored every consumer — confirmed byte-identical helm template output before/after (selector labels are immutable on upgrade, so this was verified, not assumed). - Documented (README): the ingressFrom default and readOnlyRootFilesystem posture, both real but intentional tradeoffs the audit flagged as underdocumented. Added extraVolumes/extraVolumeMounts to copilot's Deployment (realtime/pii already had it) so the readOnlyRootFilesystem guidance is actually actionable for all three stateless services. Deferred (nice-to-have, not blocking): pinning the two floating Postgres image tags, values.schema.json stubs for ~13 uncovered top-level sections, and an OTel collector image version bump — none are correctness issues. * fix(helm): move copilot's static config out of the ESO-required Secret Greptile caught a real bug: copilot.server.env shipped with non-empty static defaults (PORT, SERVICE_NAME, ENVIRONMENT, LOG_LEVEL), unlike app.env/realtime.env which ship fully empty. The new ESO validation correctly required every non-empty env key to be mapped in externalSecrets.remoteRefs.copilot — but that meant a default install with copilot + ESO enabled failed demanding secret-store paths for values that were never secrets. Fixed by applying the chart's own existing pattern for this exact problem: moved the 4 static keys into copilot.server.envDefaults (mirroring app.envDefaults) and inlined them as plain container env, bypassing the Secret/ExternalSecret system entirely — same rationale already documented for app.envDefaults. Verified live that Greptile's exact repro (default copilot env + ESO enabled, only the 7 real secrets mapped) now renders cleanly and the four values still reach the container. Added a regression test that fails on the pre-fix code. * fix(helm): don't shadow copilot's existingSecret with envDefaults Greptile and Cursor Bugbot both independently caught this: in copilot.server.secret.create=false (existingSecret) mode, the chart still unconditionally inlined copilot.server.envDefaults as explicit container env. Kubernetes gives explicit env precedence over envFrom, so a pre-existing Secret's PORT/LOG_LEVEL/etc values were silently overridden by the chart defaults — the exact shadowing bug app.envDefaults already guards against via its own $useExistingSecret skip, which I forgot to mirror when copying the pattern to copilot. Skip envDefaults entirely in existingSecret mode (matching app's existing behavior — the pre-created Secret is the sole source of truth), while still rendering extraEnv. Verified live: existingSecret mode now renders no env: block at all when extraEnv is unset, and still renders extraEnv without envDefaults leaking in when it is set. Added two regression tests, confirmed both fail on the pre-fix code. * fix(helm): key copilot's existingSecret check off its own secret.create, not the global ESO flag Round 2's fix (which I copied nearly verbatim from Greptile's own suggested diff) used $useExistingSecret := and (not externalSecrets.enabled) (not copilot.server.secret.create) — Greptile caught its own suggestion's remaining bug on round 3: when externalSecrets.enabled=true globally (for app/postgresql) but copilot itself uses copilot.server.secret.create=false with its own pre-created Secret, that condition evaluated to non-existingSecret mode, so envDefaults still inlined and shadowed the user's Secret values — same bug, different trigger condition. envFrom always points at the user-provided Secret name whenever secret.create=false, independent of what other components do with ESO, so the check should key on that alone. Verified live: global ESO enabled + copilot's own existingSecret now renders no env: block and envFrom correctly points at the pre-created secret name; the two scenarios that should still inline (copilot itself on ESO, plain inline mode) still work. Added a regression test, confirmed it fails against round 2's guard. * fix(helm): checksum/secret annotation on copilot ignores ESO-sourced secret Cursor Bugbot caught a real bug: checksum/secret only hashed secrets-copilot.yaml's rendered output, but under externalSecrets.enabled=true that template renders nothing (env credentials come from external-secret-copilot.yaml instead). Result: changing externalSecrets.remoteRefs.copilot mappings wouldn't change the pod template hash, so Kubernetes would never restart the copilot pod to pick up the new mapping — stale envFrom values until a manual restart. Fixed by hashing the concatenation of both templates' rendered output: whichever mode is active, only one renders non-empty content, but the concatenated hash still changes on a mode switch or a remoteRefs change. This can't reach into the live secret store value ESO syncs (Helm only sees the ExternalSecret manifest at render time) — that's an inherent ESO limitation, not something a checksum annotation can close; documented as such in the template comment. Note: deployment-app.yaml and deployment-realtime.yaml have this same latent limitation in ESO mode (checksum/secret only hashes secrets-app.yaml), but that's pre-existing code outside this PR's diff — not fixed here to stay scoped to what Cursor actually flagged. Verified live: the checksum differs across two different remoteRefs.copilot.LICENSE_KEY mappings, and still changes correctly in plain inline mode. Added a regression test. |
||
|
|
cbe5e0f3f4 |
fix(uploads): fix Azure Blob connection-string-only auth and document self-host parity (#5553)
* fix(uploads): fix Azure Blob connection-string-only auth and document self-host parity Every Azure Blob operation (upload/download/delete/head/presigned URLs) threw when only AZURE_CONNECTION_STRING was set, despite that being the documented alternative to AZURE_ACCOUNT_NAME/KEY across .env.example, env.ts, and the Helm chart. createBlobConfig required accountName unconditionally, and the upload-SAS path had no fallback to derive credentials from the connection string. Fixed both, verified end-to-end against a live Azurite emulator (upload, download, head, delete, multipart/block-blob upload, and real HTTP PUT/GET through generated SAS URLs), and added regression tests. Also closes the remaining self-host Azure documentation gaps: AZURE_ACS_CONNECTION_STRING, OCR_AZURE_*, KB_OPENAI_MODEL_NAME, and WAND_OPENAI_MODEL_NAME are now documented in the Helm chart (values.yaml, values.schema.json, values-azure.yaml) and .env.example alongside their AWS/S3 counterparts. * fix(uploads): map new Azure keys in the ESO remoteRefs example Greptile flagged that the values-azure.yaml External Secrets example didn't map KB_OPENAI_MODEL_NAME, WAND_OPENAI_MODEL_NAME, OCR_AZURE_ENDPOINT, and OCR_AZURE_MODEL_NAME, so a user who fills those in and switches to ESO would hit a Helm template render failure. Added the remoteRefs entries and verified with an isolated helm template render. |
||
|
|
4e6594dc54 |
feat(pii): add opt-in GLiNER NER engine (PII_ENGINE), device-agnostic (#5495)
* feat(pii): add opt-in GLiNER NER engine (PII_ENGINE), device-agnostic Swap the 4 NER entity types (PERSON/LOCATION/NRP/DATE_TIME) to a single multilingual GLiNER zero-shot model when PII_ENGINE=gliner; spaCy stays the default and all ~36 regex/checksum recognizers are identical on both engines. Device-agnostic via PII_DEVICE / cuda auto-detect — same code on Fargate CPU now and EC2-GPU later. - engines.py: side-effect-free builders; SharedModelGLiNERRecognizer loads ONE model shared across the 5 per-language instances and restricts labels to the entities it owns; small spaCy models keep tokenization/lemmas for the regex recognizers; fail-fast on the lean image - pii.Dockerfile: multi-stage — default target unchanged (lean spaCy); --target gliner is a superset (torch CPU + gliner + baked model) where both engines work; gliner-gpu scaffold for the GPU fleet - CI publishes the gliner variant (:staging-gliner/:latest-gliner, amd64) - Helm: pii.engine / pii.device values wired to PII_ENGINE/PII_DEVICE - scripts/bench_engines.py: throughput + NER-parity diff harness - tests: unit (mocked GLiNER) + in-image integration for both engines Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Up3F97mjCH9HCj1pX4J8VJ * refactor(pii): ship both engines in one image — engine is a pure env flip Collapse the gliner build target into the single pii image: spaCy lg models, torch (CPU), gliner, and the baked GLiNER weights all ship in it, so PII_ENGINE switches engines with no image swap and no tag matrix. CI reverts to the single pii build (no -gliner tags). The GPU variant becomes the same Dockerfile built with --build-arg TORCH_INDEX_URL=.../cu128. Image grows ~6.1GB -> ~9.6GB. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Up3F97mjCH9HCj1pX4J8VJ --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
24ebba9acc | chore(credential-sets): cleanup feature (#5460) | ||
|
|
69b81a679b |
feat(data-retention): granular PII redaction stages (input + block outputs) (#5272)
* feat(data-retention): granular PII redaction stages (input + block outputs) * fix(data-retention): propagate block-output redaction into child workflows * fix(data-retention): close block-output redaction gaps on streaming + resume * fix(data-retention): drain+mask streamed output, resolve PII policy unconditionally (no fail-open) * test(testing): support leftJoin().where().limit() in shared db mock * fix(data-retention): mask agent/Pi memory writes under block-output redaction * fix(data-retention): guard partial PII stages in GET normalize * fix(data-retention): mask seeded memory messages under block-output redaction * fix(guardrails): fail closed on misaligned Presidio batch responses * fix(data-retention): enabled stage with no entity types redacts all (no fail-open) * fix(data-retention): reject enabled stage with no entity types; empty = off everywhere * docs(data-retention): note resume remask covers inline values only * fix(data-retention): scrub offloaded large-value refs from logs when block-output redaction is off * fix(data-retention): hydrate, mask, and re-store large-value refs in logs (preserve redacted content) * fix(data-retention): always apply logs policy to large-value refs when logs stage is on * perf(data-retention): drop redaction byte ceiling, parallelize chunks (env-tunable), remove request timeouts, sync large-value walk * feat(data-retention): gate granular PII stages behind pii-granular-redaction flag - New pii-granular-redaction feature flag (fallback PII_GRANULAR_REDACTION), layered on pii-redaction, gating the execution-altering input + block-output stages - Route returns piiGranularRedactionEnabled and rejects enabling granular stages when off - UI shows only the Logs stage tab unless the flag is on; clamps active stage - Drop the per-search Select all toggle; add a Deselect all action to the PII section header * docs(pii): describe Presidio as a standalone service, not a sidecar Presidio now runs as its own ECS service (and, in Helm, its own Deployment + Service) reached over the network via PII_URL — not a sidecar in the app task. Update README, code comments, env docs, Dockerfiles, and the Helm chart docs to match, and note the deploy requirement that PII_URL must be reachable. * fix(data-retention): re-mask offloaded large-value refs on resume + don't lock out granular saves - Resume/run-from-block restore now hydrates → masks → re-stores large-value refs in restored blockStates (not just inline strings), so a value offloaded before the block-output stage was enabled can't warm raw PII into downstream blocks. Fails fast. - pii-large-values: add onFailure mode (throw on the execution path, scrub for logs) and redactLargeValueRefsInValue for arbitrary (non-RedactablePayload) values - Granular flag gate now rejects only NEW off→on granular enablement, so orgs that already configured granular stages can still save retention settings when the flag is off |
||
|
|
ca34301d7f |
fix(mailer): permissions entitlements for enabling/disabling (#5312)
* v0.6.29: login improvements, posthog telemetry (#4026) * feat(posthog): Add tracking on mothership abort (#4023) Co-authored-by: Theodore Li <theo@sim.ai> * fix(login): fix captcha headers for manual login (#4025) * fix(signup): fix turnstile key loading * fix(login): fix captcha header passing * Catch user already exists, remove login form captcha * fix(mailer): permissions entitlements for enabling/disabling * fix lifecycle for agentmail infra --------- Co-authored-by: Waleed <walif6@gmail.com> Co-authored-by: Theodore Li <theodoreqili@gmail.com> Co-authored-by: Siddharth Ganesan <33737564+Sg312@users.noreply.github.com> Co-authored-by: Theodore Li <theo@sim.ai> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5a8134119a |
feat(workspaces): fork + push/pull (#5210)
* feat(workspaces): fork + push/pull * type fix * fix tests * progress on ux * remove modal section * improve UI of modal * update more ui * make rollback part of the footer * track skipped count correctly * address comments * make it workspace admin level * update skipped count * address more comments * deal with unbounded memory possibility * fix deleted kb article bug * no deployed workflow case * UI/UX cleanup * fix oauth dropdown case * fix oauth selector issue * infra work + activity log * consolidate migration * update modal state * more UI simplification * grammar * update audit report ui * perf improvements * fix tool input scenarios and add dependsOn UI handling * minor comments * fix webhook stability issues + drift detection removal * make dependsOn subblock mapping cleanly stored * fix: harden fork dependent-value mapping (clear stale rows, identity-guard first-sync fallback, perf + cleanup) * address comments * update comment * enforce admin perms for activity api * fix required + dependsOn combo |
||
|
|
76867062e5 |
feat(pii): publish PII image to GHCR and add Presidio sidecar to Helm chart (#5188)
* feat(pii): publish PII image to GHCR and add Presidio sidecar to Helm chart * fix(pii): allow app→PII NetworkPolicy egress, global tolerations, topology spread |
||
|
|
7d46103d09 |
chore(deps): remove unused dependencies and harden CI supply chain (#5119)
* chore(deps): remove unused dependencies and harden CI supply chain Dependency cleanup: - Remove unused deps: papaparse, unified, and 6 unused Radix primitives (alert-dialog, radio-group, scroll-area, separator, toggle, visually-hidden) plus @tanstack/react-query-devtools (all verified zero imports repo-wide) - Consolidate jwt-decode into the existing jose dependency (decodeJwt) - Migrate react-window to @tanstack/react-virtual to drop a redundant virtualization library (terminal, structured-output, code viewer) - Remove the better-auth-harmony plugin and its gating env flag Supply-chain hardening: - SHA-pin every GitHub Action to a full commit SHA with a version comment - Pin CI bun-version to 1.3.13 (was "latest" in the release job) - Raise bun minimumReleaseAge cooldown from 3 to 7 days - Add a non-blocking `bun audit` step in CI - Add a CODEOWNERS gate routing dependency-manifest changes to @simstudioai/deps * chore(deps): remove unused apps/docs dependencies (@tabler/icons-react, dotenv-cli) * style(search-modal): use Send icon for Invite teammates action * feat(search-modal): surface New chat as the top action above Create workflow * feat(search-modal): add Secrets to the pages list |
||
|
|
6abcf82db2 |
feat(workflows): sim trigger, logs v2 block, toolbar renaming (#4941)
* feat(workflows): sim trigger, logs v2 block, toolbar renaming * fix(review): bound rule queries, canonical logs params, watched-workflow SQL scoping Code-review fixes: read the canonical workflowIds param in logs_v2 (the serializer deletes the source pair ids), aggregate failure-rate in the DB and switch rule windows to the indexed startedAt column, clamp rule config to the legacy contract bounds, push no_activity watch scoping into SQL before the LIMIT, fix the generated sim icon-map key, normalize docs wording, and drop dead exports. Co-authored-by: Cursor <cursoragent@cursor.com> * address comments * fix(review): integer rule rounding, success-gated workflow labels, display module hygiene Second-pass review fixes: round integer rule fields so fractional input never reaches SQL LIMIT, gate workflow-name readiness on a successful non-placeholder load in both editor and preview (errored loads mislabeled valid workflows as deleted), lazily read the variables store in preview rows, move the filter-field JSON preview into the shared display module and unexport its single-consumer helpers, and align >= boundary copy (failure rate, error count, cooldown window) with implementation. Co-authored-by: Cursor <cursoragent@cursor.com> * chore: sync lockfile after staging merge Co-authored-by: Cursor <cursoragent@cursor.com> * fix(workspace-events): keyset-paginate the no_activity subscription scan A fixed LIMIT 500 with no ORDER BY silently starved subscriptions beyond the cap once the global count exceeded it. The poll now pages by webhook id so every subscription is visited each cycle; pagination bounds memory, not total work. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(workspace-events): keyset-paginate the watched-workflow scan The 500-row LIMIT silently and deterministically excluded high-id workflows from no_activity coverage in watch-everything subscriptions on large workspaces. The scan now pages by workflow id, mirroring the subscription scan; per-workflow checks move into a helper so the pagination loop stays flat. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(workspace-events): skip no_activity subscriptions on the execution-completion path no_activity is poller-owned and can never fire from a completed execution, but it passed into the rule branch and cost a pointless cooldown point-read per subscription on the hottest path. Early-continue alongside the workflow_deployed guard. Co-authored-by: Cursor <cursoragent@cursor.com> * docs(sim-trigger): note failure-based alert conditions evaluate on failed runs Co-authored-by: Cursor <cursoragent@cursor.com> * fix(blocks): recategorize Data Enrichment as a core block It's a Sim-native capability (registry enrichments over a managed provider cascade, like Search), not a third-party integration. Moves it to Core Blocks in the toolbar, out of the integrations catalog, and relocates its docs page to blocks/ with the icon-map allowlist keeping the docs card icon. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(blocks): recategorize MySQL, PostgreSQL, SFTP, SMTP, SSH as integrations External-system connectors with host/credential auth belong under Integrations, not Core Blocks — consistent with MongoDB, Redis, ClickHouse, and the other datastore integrations. They already carried integrationType and /tools docsLinks; the regenerated docs pages turn those previously-dangling links into real pages, and the blocks join the integrations catalog and icon maps. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
0075ab9cf6 |
improvement(platform): remove tour, simplify sidebar/header, drop loading skeletons (#4354)
* improvement(platform): workspace UI/UX overhaul + integrations catalog Rework the workspace around the AI-workspace model: a Mothership home, a top-level Skills route, connected-credential and integration-detail pages, and a polished sidebar/settings surface. Replace the notifications store with a unified toast system (provider-level dismiss/pause, countdown ring). Integrations & catalog: - Add a BlockMeta layer (tags + catalog templates) scoped to catalog-visible integrations; every catalog integration carries >=7 grounded templates. - Rework the taxonomy: each block declares category tools|blocks|triggers. 3rd-party services are 'tools'; first-party primitives (postgres, mysql, knowledge, file, search, stt/tts, image/video generators, thinking, etc.) are 'blocks'. Versioned blocks follow the upgrade paradigm (old hidden, latest in toolbar/docs). - Generate integrations.json + tool docs canonically from block configs. Architecture & cleanup: - Consolidate block data extraction behind a single latest-version strategy (getCanonicalBlocksByCategory; version-consistent getBlockMeta). - Unify version-suffix handling in @sim/utils/string (stripVersionSuffix / isVersionedType, with tests); registry, generate-docs, tools/utils, and integrations all route through it. - Repair latent broken barrels, remove dead code, fix BlockMeta-related type errors and 5 broken docs links. Behavior-preserving for block execution and the toolbar's tool/block listing. * refactor(platform): remove forms, templates, and creators features Remove three standalone features and their supporting code: - Forms: form-deployment pages, API routes, execution path, and docs. - Templates: the template gallery (landing + workspace) and template APIs. - Creators: creator-profile routes and contracts. Add a super-user permissions module (lib/permissions/super-user) and an organizations API contract; update the audit/db/testing packages, billing, and the session/theme providers accordingly. * test(workflows): update archiveWorkflow update count after forms removal The forms feature was removed, dropping the form-table update from archiveWorkflow. Update the stale assertion from 8 to 7 tx.update calls. * upgrade * improvement(knowledge): polish tag filter dropdowns (#4816) * improvement(logs): object storage backed tracespans (#4787) * improvement(logs): obj storage backed tracespans * fix storage write context * fix tests * address comments * address comments * chore(db): remove migration 0219 to regenerate after staging merge Drops the 0219_robust_shard SQL, its snapshot, and the journal entry so the trace-spans/cost schema migration can be regenerated on top of the latest staging migration chain (avoids a number collision with staging's migrations). Co-authored-by: Cursor <cursoragent@cursor.com> * improvement(billing): accurate per-member usage via shared ledger helper Per-member/per-user usage in the org-member routes now adds the usage_log ledger to the currentPeriodCost baseline (which is no longer incremented), via a shared getOrgMemberLedgerByUser helper to avoid repeating the subscription→period→ledger lookup across the admin and member-facing routes. Co-authored-by: Cursor <cursoragent@cursor.com> * regen migrations * update migration * address comments * more code cleanup * incorrect type cast --------- Co-authored-by: Cursor <cursoragent@cursor.com> * improvement(providers): harden OpenAI-compatible providers + add tests (#4796) * improvement(providers): harden OpenAI-compatible providers + add tests * fix(vllm): let tool-loop errors propagate instead of returning silent partial success * fix(litellm): force tool_choice 'none' on final structured-output call The deferred final call used tool_choice 'auto', so the model could emit another tool_calls round instead of the structured answer, leaving content stale. Use 'none' (matching vLLM/Fireworks) on both the streaming and non-streaming final calls so the model must return the structured response. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(providers/ollama): drop tools from post-tool streaming call Ollama ignores tool_choice (not in its supported fields), so vLLM/Fireworks' tool_choice:'none' guard is a no-op here. Omit tools from the final streaming payload instead so the summarization turn can't emit dropped tool calls. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(litellm): spread payload into deferred final call so reasoning_effort carries over The non-streaming deferred finalPayload hand-picked fields and dropped reasoning_effort (and any future payload field), diverging from the streaming path which spreads ...payload. Spread payload here too for consistency. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(providers/ollama): restore enrichment TSDoc block Keeps parity with sibling Chat Completions providers (cerebras/mistral/xai). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(fireworks): restore TSDoc on utils helpers Restore the TSDoc blocks on supportsNativeStructuredOutputs, createReadableStreamFromOpenAIStream, and checkForForcedToolUsage — TSDoc is the codebase documentation standard and should not have been stripped. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(litellm): remove inline rationale comments (codebase uses TSDoc) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(providers/ollama): drop orphaned enrichment TSDoc The block documented a function that now lives in trace-enrichment.ts, so it documents nothing in this file. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> * chore(copilot): deprecate mcp server (#4797) * chore(copilot): deprecate mcp * update error codes * deprecate copilot api v1 route * feat(integrations): hosted API keys for Findymail, Prospeo, and Wiza (#4777) * feat(integrations): hosted API keys for Findymail, Prospeo, and Wiza Add hosted-key support across all credit-consuming Findymail, Prospeo, and Wiza operations so Sim provides the key when a workspace has not brought its own. Register the three BYOK providers, consolidate Wiza's two-step reveal into a single polling wiza_individual_reveal op, and hide the API key field on hosted Sim for hosted operations. * fix(integrations): harden Wiza reveal polling, soften enrichment getCost guards Address Greptile + Cursor Bugbot review on #4777: return explicit failures from the Wiza individual_reveal poller instead of throwing (thrown errors were swallowed into a false queued success), short-circuit when the initial reveal is already terminal, tolerate transient 5xx/429 during polling, and return 0 (not throw) from Findymail getCost when the contacts/employees array is absent. * chore(integrations): biome formatting after wiza merge resolution * fix(wiza): type isTerminalReveal param structurally for next build typecheck * feat(enrichments): add Findymail, Prospeo, Wiza to work-email waterfall * feat(enrichments): add Wiza + Prospeo phone reveal to phone-number waterfall * feat(enrichments): opportunistic identifiers + LinkedIn URL input across work-email & phone cascades * fix(tables): reduce column header chevron size and fix sidebar shadow bleed (#4800) * feat(slack): add install + privacy section to integration landing page (#4799) * feat(slack): add install + privacy section to integration landing page Adds a hand-authored, slug-keyed landing-content module (separate from the generated integrations.json so it survives regeneration) and renders an install walkthrough + privacy-policy link on integration pages when present. Also refreshes generated docs (data-enrichment entry, icon mappings, tool mdx). * fix(landing): render privacy section independently, align CTA analytics label * docs(landing): clarify the Slack install button is behind sign-in * refactor(landing): bake integration landing content into generated json via docs-gen Moves landing content (install walkthrough + privacy) out of a render-time augment and into the generation pipeline: generate-docs reads the pure-data content map and writes landingContent into integrations.json, so the page reads a single source (integration.landingContent). Canonical types live in integrations/data/types.ts. * improvement(enrichments): align enrichments sidebar with design system (#4801) * improvement(enrichments): align enrichments sidebar with design system * fix(enrichments): consistent close button pattern and fix url link hover * fix(misc): upgrade path change for new better-auth version, billing issue for workflow block agent usage (#4803) * fix(misc): upgrade path change for new better-auth version, double-billing for workflow block agent usage * fail loudly if stripe sub id missing * fix(copilot): seq migration (#4804) * chore(db): drop redundant idx_webhook_on_workflow_id_block_id index (#4809) Removed because (workflow_id, block_id) is a left-prefix of idx_webhook_on_workflow_id_block_id_updated_at_desc, which fully covers it. The dropped index was non-unique and enforced no constraint. * perf(copilot): read chat transcripts from copilot_messages (R+1 cutover) (#4808) * perf(copilot): read chat transcripts from copilot_messages, not JSONB Flip user-facing chat reads from the legacy copilot_chats.messages JSONB array (5.7GB, 99% TOAST) to the normalized copilot_messages table via a new loadCopilotChatMessages helper ordered by seq NULLS LAST, created_at, id — the verified canonical order. Both chat-detail getters (getAccessibleCopilotChat, getAccessibleCopilotChatWithMessages) now drop the messages column from their metadata select (no more whole-array detoast on every load) and assemble the transcript from the table after authorization. This cascades to the copilot + mothership GET endpoints and to resolveOrCreateChat's conversationHistory (the LLM payload). The normalize/effective-transcript pipeline is source-agnostic (copilot_messages.content == a JSONB array element), so transcripts are byte-identical. Dual-write and the JSONB column stay in place as the internal-logic source and fallback; removing JSONB writes is a later step. Prod integrity verified before cutover: 0 messages missing, 0 NULL-seq, 0 dup keys/seq, 0 orphans, order-parity vs JSONB = 0 mismatches. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(copilot): cover auth-deny on a found row skips the messages query Address PR review: exercise the `if (!authorized) return null` contract — when the chat row exists but authorization fails, the getter returns null and never issues the copilot_messages read. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> * fix(tables): right-align run/stop in embedded toolbar; workflow cells format like normal cells (#4806) * fix(tables): right-align run/stop in the embedded table toolbar Add a right-aligned `trailing` slot to ResourceOptionsBar and move the embedded mothership table's run/stop control into it, so Filter + Sort stay left-aligned and run/stop sits opposite on the right. No-op for the search-bearing consumers (logs, resource list), which don't pass `trailing`. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(tables): workflow-output cells format values like normal cells Workflow-output columns short-circuited in resolveCellRender and rendered their value as plain text, so a sim-resource URL / external URL / JSON / date produced by a workflow never got the chip, favicon link, or typed formatting a normal cell gets. Factor value formatting into a shared `resolveValueKind` helper used by both the workflow-value branch and the plain-cell branch; the workflow branch keeps the typewriter reveal for plain streaming text via a `typewriter` flag. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(tables): detect resource/URL links on workflow output regardless of column type Workflow output columns default to `json` (columnTypeForLeaf), so routing their values through the type-based formatter (a) gated chip/URL promotion behind `column.type === 'string'` — a URL produced by a json-typed output never became a chip — and (b) JSON.stringify'd plain string values, adding quotes and losing the typewriter reveal. Detect links (sim-resource chip / favicon URL) on the value string directly for workflow outputs, falling back to the plain `value` kind; plain cells keep the type-based formatting. Addresses Greptile P2 on #4806. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(icons): repair broken integration icon rendering (#4810) * fix(icons): repair broken integration icon rendering Two distinct bugs left integration icons broken on the /integrations page (visible at 32-40px, hidden at the toolbar's 16px): 1. Corrupted SVG paths (Notion, Greptile, Granola, Calendly, Grafana, Bedrock): over-minified data dropped elliptical-arc flag digits (e.g. `A1 1 0 5.9 7` instead of `A1 1 0 0 0 5.9 7`); Granola's cubic stream was truncated. Browsers abort path parsing at the first invalid arc flag, so each rendered as a fragment or blank. Replaced with correct path data from canonical sources, preserving each icon's existing fill/gradient and bgColor. 2. Invisible glyph (Bright Data): its icon uses fill='currentColor' but bgColor was '#FFFFFF', and every surface forces text-white on the glyph - white-on-white. Changed bgColor to Bright Data's brand blue (#3d7ffc) so the white glyph reads, matching the white-glyph-on-brand-chip convention. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(icons): restore Calendly dual-tone brand colors Addresses review feedback: the previous fix replaced the broken Calendly icon with a monochrome #006BFF path, dropping the cyan #0ae8f0 accent from the original dual-tone mark. Restored the two-tone logo (blue + cyan) using clean, valid path data, cropped to a tight square viewBox so it fills the chip. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * improvement(icons): enlarge icons, fix Zoom contrast and Quiver chip - Zoom: glyph was blue-on-blue (#0B5CFF on #2D8CFF chip); switched to currentColor so it renders as a white glyph on the blue chip. - Quiver: chip bgColor #000000 -> #FFFFFF to match the icon's near-white box, and enlarged the mark slightly (viewBox crop). - Enlarged (tightened viewBox, verified no clipping): RevenueCat, Prospeo, Granola, Firecrawl, Enrich.so, and the AWS icons (RDS, DynamoDB, SQS, CloudFormation, Athena, CloudWatch, SES, Bedrock, S3). - ZoomInfo left unchanged: it is a full red rounded-square logo that already fills its frame, so a crop would clip it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(icons): use Bright Data wordmark on white chip; repair Circleback - Bright Data: replaced the flame glyph with the official two-tone 'bright data' wordmark (provided asset), centered in a symmetric viewBox. Reverted the chip bgColor from #3d7ffc to #FFFFFF since the blue wordmark is invisible on a blue chip (the wordmark is designed for a light background). - Circleback: a minifier had rounded the pattern's image scale to scale(0), collapsing the embedded logo to zero size (invisible). Restored the correct scale (1/280 = 0.00357142857) so the C. mark renders. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(docs): sync Quiver block color card to white chip Reflects the Quiver bgColor change (#000000 -> #FFFFFF) in the docs block info card. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * improvement(icons): enlarge AWS/Cloudflare/Dagster icons, fully white Zoom - Enlarged (tighter viewBox, render-verified, no clipping): Cloudflare, Dagster, and the red AWS icons AWS IAM, Identity Center, Secrets Manager, SES, STS. Identity Center was anomalously small (filled ~32% of its frame); the group is now sized consistently (~80% fill). - Zoom: the camera lens triangle was still #0B5CFF (blue-on-blue); switched it to currentColor so the whole camera renders white on the blue chip. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs(wiza): consolidate individual reveal into a single operation Merges the separate Start/Get Individual Reveal operations into one Individual Reveal operation in the Wiza docs and integrations data (operationCount 5 -> 4). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * improvement(icons): size remaining AWS icons to match the set (~80% fill) Bring RDS, DynamoDB, SQS, CloudFormation, Athena, CloudWatch and S3 up to the same ~80% fill as the AWS IAM/Identity Center/Secrets Manager/SES/STS group, so all AWS icons are visually consistent. Bedrock left as-is (already ~92% fill). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(icons): use Bright Data flame mark, enlarge ZoomInfo - Bright Data: the full 'bright data' wordmark was illegible at chip size. Replaced with just the flame-'i' brand mark (blue #4280f6 on the white chip), centered. - ZoomInfo: cropped the viewBox toward the white 'Zi' so it's larger; the red rounded-square background still fills the chip. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * improvement(icons): enlarge CrowdStrike icon The falcon mark sat small in its chip because the icon used a wide 768x500 viewBox (letterboxed in the square chip). Switched to a square viewBox centered on the mark so it fills ~80%, consistent with the other icons. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> * fix(tables): serialize schema mutations to prevent parallel column clobber (#4812) * Make workflow description nullable * fix(tables): serialize schema mutations to prevent parallel column clobber * fix(tables): load workflow outside schema lock; use DbOrTx for getTableById * fix(tables): scale idle timeout in updateColumnType to avoid aborting large type changes * fix(tables): skip stale remap types when workflowId changes concurrently * fix(tables): scale idle timeout in updateColumnConstraints for large tables * fix(wait): resume live/draft async waits and preserve cell context on chained waits (#4814) * Make workflow description nullable * fix(wait): resume live/draft async waits and preserve cell context on chained waits * improvement(knowledge): polish tag filter dropdowns * improvement(knowledge): soften filter section labels * improvement(knowledge): soften list filter labels * fix(security): harden SSO domain registration, webhook path isolation, and CSV export (#4813) * fix(security): harden KB file access, SSO domain registration, webhook path isolation, env secrets, and CSV export * fix(sso): scope domain conflict query with indexed lower(domain) filter Address PR review: avoid a full-table scan on every SSO provider registration by filtering candidate rows in SQL with lower(domain) = <normalized>, keeping the in-memory ownership check. Also tighten the normalizeSSODomain TSDoc. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: condense env route security comments Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * icons update * chore(security): tighten inline comments in CSV export and KB file authorization Condense verbose comment blocks to concise TSDoc/single-line form; no behavior change. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(security): validate internal serve origin in KB file authorization Replace the bypassable isInternalFileUrl substring check in resolveInternalKbKey with an origin allow-list (base URL, internal API base URL, TRUSTED_ORIGINS). A crafted external host whose path is /api/files/serve/<victim-key> no longer resolves to the victim key. Relative same-origin URLs are unaffected. * style(sso): use idiomatic sql lower() comparison for domain conflict query Match the repo's prevailing `sql`lower(col) = value`` idiom for the case-insensitive SSO domain conflict lookup. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(security): align workspace env admin gate with hasWorkspaceAdminAccess Use the same admin check the secrets UI uses (owner, admin permission, or org-admin) so owners and org-admins are not wrongly denied their own decrypted workspace secrets, while read-only members remain restricted to names only. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(sso): rely on lower(domain) match for conflict detection, drop dead in-memory recheck Address PR review: the SQL `lower(domain) = <normalized>` predicate already excludes rows that the in-memory `normalizeSSODomain(...) === domain` recheck claimed to catch, making that recheck dead/misleading code. Match on the canonical lower-cased domain and filter purely by ownership. Malformed legacy values (wildcards, schemes, ports) never match an email domain at sign-in, so excluding them is not a gap. Test DB mock now applies the lower() predicate so the casing-variant case is genuinely exercised. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(security): scope webhook deploy path conflict to active webhooks findConflictingWebhookPathOwner omitted the isActive filter that the runtime dispatcher (findAllWebhooksForPath) applies, so an inactive but non-archived webhook from another workflow (e.g. after undeploy or failure auto-disable) would permanently block any new deployment on that path even though it never receives deliveries. Align the guard with the runtime isActive + archivedAt filter; the earliest-owner runtime check remains the authoritative cross-tenant protection. Also trims verbose TSDoc on the webhook path-isolation helpers. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(security): exclude archived workflows from webhook deploy path conflict findConflictingWebhookPathOwner now joins workflow and filters isNull(workflow.archivedAt), matching the runtime dispatcher (findAllWebhooksForPath). A webhook on an archived workflow can never receive deliveries at runtime, so it must not block legitimate path reuse with a 409. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(security): anchor KB file ownership to earliest document in any state A KB file's owner is now the earliest document referencing its key regardless of state (active/archived/deleted/excluded); access is granted only when that owning document is still active. Closes the residual where an attacker could plant an active document to claim a file whose original document was archived or deleted. * updated greptile icon * revert(security): drop KB file authorization changes Reverts the knowledge-base file-access work (origin-pinning / owner-pinning / origin allow-list in verifyKBFileAccess) and its test. The other hardening fixes (SSO domain registration, webhook path isolation, workspace env secrets, CSV export) are unchanged. apps/sim/app/api/files/authorization.ts is restored to its origin/staging baseline. * fix(sso): treat caller's own user-scoped provider as owned during conflict check Self-hosters often register SSO user-scoped via the CLI script (no SSO_ORGANIZATION_ID). If they later enable organizations and reconfigure the same domain org-scoped through the UI, the conflict check previously treated their own user-scoped row as another tenant's and returned a misleading 409. Recognize the caller's own user-scoped provider as owned so that migration is allowed, while still blocking another user's or another org's domain. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * revert(security): remove workspace-env admin gate Defer to a credential-based access model (separate change). Restores GET /api/workspaces/[id]/environment to main behavior and removes the test. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * refactor(security): consolidate webhook path-collision check into one helper Extract findConflictingWebhookPathOwner to lib/webhooks/utils.server.ts as the single source of truth for cross-tenant path-collision detection, used by both webhook creation paths (deploy sync and the manual /api/webhooks route). This also repairs two latent issues in the manual route's previous inline check, which queried with limit(1) and only webhook.archivedAt: - limit(1) inspected one arbitrary row, so a same-workflow row could mask a foreign collision (false negative). The shared helper scans all matching rows. - It omitted isActive/workflow.archivedAt, so inactive or archived-workflow webhooks (which never receive deliveries) permanently blocked path reuse. The helper mirrors the runtime dispatcher's filter. Same-workflow webhook reuse for upsert is now a separate, explicit lookup. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> * fix(security): block private/reserved IPs for hosted 1Password Connect SSRF (#4818) * fix(security): block private/reserved IPs for hosted 1Password Connect SSRF * test(security): use real isPrivateOrReservedIP and cover IPv6 edge cases * improvement(integrations): validate and expand devin, cursor, and greptile (#4820) * improvement(integrations): validate and expand devin, cursor, and greptile - devin: fix missing org_id path segment on all session endpoints, add 7 session sub-resource tools (list messages/attachments, get/append/replace tags, archive, terminate), pagination, and is_archived output - cursor: add get_api_key_info, list_models, list_repositories tools - greptile: align block and docs - normalize array outputs to default [] and tighten types * refactor(cursor): simplify list_repositories v2 array normalization Collapse the redundant `?? []` + `Array.isArray` double-guard into a single Array.isArray check, per PR review feedback. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(devin): scope session-tag mapping to tag ops and normalize array tag inputs - Only map sessionTags into the tools tags param for append/replace operations, preventing stale sessionTags state from clobbering create_session tags - Fall back to a wired tags value when sessionTags is empty for tag operations - Normalize tag inputs (string or wired string[]) via normalizeTags so array values from other blocks no longer throw on .split Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(cursor): restore base64 file data in legacy download_artifact metadata The legacy CursorBlock exposes only content + metadata (no v2 file output), so metadata.data was the only way legacy-block workflows could access downloaded artifact bytes. Restore the base64 data field and document it in the outputs/type instead of dropping it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(devin): coerce terminateArchive to archive flag for boolean-wired input * docs(integrations): regenerate tool docs for new devin and cursor operations --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> * fix(search-replace): don't auto-navigate when content edits invalidate the active match (#4819) * fix(search-replace): don't auto-navigate when content edits invalidate the active match * fix(search-replace): clear afterReplaceIndexRef on apply failure and zero matches * fix(search-replace): remove duplicate setActiveSearchTarget(null) on close * fix(search-replace): move afterReplaceIndexRef write inside handleApply past the guard * fix(search-replace): auto-navigate when hydration resolves with no prior active match * chore(search-replace): remove inline comments * fix(search-replace): revert !activeMatchId guard that caused immediate re-navigation after deselect * improvement(enrichments): limit company-info to fields both providers return (#4817) Hunter's company dataset returns null industry/foundedYear for many large companies (verified against the live API for Microsoft, Amazon, Google), so under the first-non-empty-wins cascade those columns appeared inconsistently across rows. Limit company-info outputs to employee count and description — the fields Hunter and PDL both reliably return — so every row is consistent. employeeCount is a string so Hunter's range bucket and PDL's exact count share the column. * fix(files): don't reject external URLs containing '..' in file parse validation (#4821) * fix(files): don't reject external URLs containing '..' in file parse validation The file block's file_fetch operation rejected any external URL whose path contained '..' (e.g. Slack files-pri slugs with a literal '...') with 'Access denied: path traversal detected'. Traversal checks only apply to local paths — external http(s) URLs are fetched with SSRF protection downstream and are never resolved against the filesystem, so they now short-circuit as valid. Internal /api/files/serve/ URLs keep full traversal protection. * test(files): fix external-URL assertion to handle undefined error * test(files): assert success explicitly in external-URL traversal test * fix(files): keep traversal protection for https URLs matching internal serve paths * feat(google-sheets): add row filtering to read with numeric operators (#4822) * feat(google-sheets): add row filtering to read with numeric operators Adds client-side row filtering to the Google Sheets read (v2) operation. Filter the returned rows by a header column using text operators (contains, not_contains, exact, not_equals, starts_with, ends_with) and numeric/ordering operators (gt, gte, lt, lte). Filtering lives in a pure, unit-tested helper (filterSheetRows) and runs over the fetched read range; an optional `filter` output reports whether the column was found and how many rows matched. Also hardens the surrounding tools: - trim spreadsheetId in write/update/append URL builders (matches read) - URL-encode the v1 read default range - expose valueInputOption for the update operation in the block Backwards compatible: with no filter requested, read output is byte- identical and the `filter` field is omitted. The filterMatchType union is widened additively (4 -> 10 values). * fix(google-sheets): correct filter metadata for missing column and header-only sheets - matchedRows is now 0 (not totalRows) when the filter column is not found, so it no longer contradicts applied=false / columnFound=false - columnFound now reflects an actual header lookup for empty/header-only sheets instead of being hardcoded true - add tests covering header-only and empty sheets with present/absent columns * fix(selectors): fetch all pages for paginated dropdown list routes (#4823) * fix(selectors): fetch all pages for paginated dropdown list routes Dropdown selectors fetched only the first page of paginated provider APIs, silently hiding results past page one. Add bounded server-side draining to the list routes across Microsoft Graph, Google, Notion, Atlassian, Linear, AWS CloudWatch, and offset/token REST APIs, plus a shared client-side drain cap in the selector hook. Response shapes, stored values, and tool execution are unchanged; CloudWatch list tools still honor a caller-supplied limit. Also fixes the Word file picker that was searching for .xlsx files. * fix(selectors): harden JSM and Monday pagination draining - JSM service-desk/request-type drains advance `start` by the actual row count returned (not the fixed page size) and stop on an empty page, so a short non-final page can't skip items. - Monday boards drain now checks `response.ok` per page, surfacing a mid-drain HTTP failure instead of treating it as an empty final page and returning a partial 200. * docs(selectors): clarify JSM drain advances start by actual row count The offset-advancement fix (advance `start` by the rows returned, not the fixed page size) landed in 7b19788a8; update the TSDoc to match so it no longer reads as advancing by `limit`. * fix(selectors): drain fetchPage in direct fetchList callers Making `fetchList` optional left three direct callers (outside the useSelectorOptions hook) calling it unguarded, which broke the build's type check. Route them through a shared `loadAllSelectorOptions` helper that uses `fetchList` when present and otherwise drains `fetchPage`. This also prevents a regression: `confluence.spaces` / `knowledge.documents` now paginate via `fetchPage` only, and these callers (search/replace, value resolution) would otherwise have silently returned no options. * chore(selectors): rename MAX_PAGE_PAGES to MAX_NOTION_PAGES for readability * fix(sso): re-check domain conflict before write and reject IP-address domains (#4825) * improvement(copilot): make copilot_messages the sole transcript store, remove JSONB dual-write (#4826) Stop writing/reading the legacy copilot_chats.messages JSONB column now that reads are cut over to copilot_messages. Make appendCopilotChatMessages the primary write (throws on failure instead of swallowing), repoint peripheral readers (workspace VFS, chat cleanup, data drains, fork, superuser import) to copilot_messages, and persist the assistant turn inside finalizeAssistantTurn's transaction so it commits atomically with the stream-marker clear. The column itself is dropped in a follow-up migration after this bakes. * feat(tables): expand filter operators (not-contains, starts/ends-with, not-in, empty) (#4827) Add does-not-contain ($ncontains), starts-with ($startsWith), ends-with ($endsWith), not-in-array ($nin, previously executed server-side but unexposed in the UI), and is-empty/is-not-empty ($empty) filter operators end-to-end — SQL builder, condition types, query-builder converters/constants, the filter UI, the Table tools/block descriptions, and docs. Also fix correctness bugs in the filter builder surfaced by the wider operator set: - Same-column AND rules (e.g. age > 18 AND age < 65, or name startsWith 'A' AND name endsWith 'Z') silently overwrote each other because the AND group was keyed by column name. They now merge into one operator object, which also makes Filter -> rules -> Filter round-trip losslessly for multi-operator columns. - $nin values were not split into an array like $in, and textual-match values like "123" were numeric-coerced (breaking the ILIKE path). - A non-boolean $empty operand from the raw API silently inverted the check; it now coerces 'true'/'false' strings and otherwise returns a 400. * improvement(copilot): stop persisting tool-call result outputs in transcripts (#4829) Opening a Mothership task could take many seconds because a single persisted assistant message in copilot_messages.content can reach hundreds of MB, almost entirely inside contentBlocks[].toolCall.result.output (e.g. a get_workflow_logs or run_workflow result). The DB query is ~2ms; the cost is detoasting that payload, shipping it to the browser, and parsing it. These outputs are dead weight on the Sim side: they are never rendered (the thread shows only tool name/title/status) and never replayed to the model (the upstream copilot service owns conversation memory). So drop result.output before it is persisted, keeping result.success/error plus the tool metadata. - add stripToolResultOutput() in persisted-message.ts - apply it in messages-store toRow (covers every write path) and in loadCopilotChatMessages (existing rows render fast on read) Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> * feat(providers): add Together AI, Baseten, and Ollama Cloud model providers (#4830) * feat(providers): add Together AI, Baseten, and Ollama Cloud model providers * fix(providers): guard Ollama streaming fast-path with hasActiveTools Match Together/Baseten/Fireworks: when tools are supplied but all are filtered out (usageControl 'none'), take the single streaming call instead of an extra non-streaming round-trip. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(providers): filter non-chat model types from Together model list * refactor(providers): dedupe Ollama Cloud upstream schema ollamaCloudUpstreamResponseSchema was byte-for-byte identical to ollamaUpstreamResponseSchema (both /api/tags endpoints return the same { models: [{ name }] } shape). Drop the duplicate and reuse the shared schema. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> * fix(knowledge): calendar view sync, deduplicate popover animation classes, type-safe filter cast * cleanup(knowledge): remove TRIGGER_BORDER_CLASS duplication, inline displayLabel, drop enabledFilterParam alias --------- Co-authored-by: Vikhyath Mondreti <vikhyathvikku@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Waleed <walif6@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Theodore Li <theo@sim.ai> Co-authored-by: andresdjasso <andresdjasso@users.noreply.github.com> * feat(blocks): add BlockMeta to Quiver and Linq; fix invalid block config fields; update skills Block fixes: - Add QuiverBlockMeta (tags + 3 templates: icon generator, diagram creator, vectorizer) - Fix QuiverBlock: remove invalid tags field from BlockConfig, IntegrationType.Design → IntegrationType.AI (Design doesn't exist in the enum) - Fix GreptileBlock: remove invalid tags field from BlockConfig, IntegrationType.DeveloperTools → IntegrationType.DevOps - Fix LinqBlock: remove invalid tags field from BlockConfig (tags belong only in BlockMeta) Skills: - add-block: add dedicated BlockMeta section with structure, rules, and registration pattern; add BlockMeta checklist items - add-integration: add BlockMeta to block structure template, add rules clarifying that tags must NOT appear on BlockConfig and integrationType must be a valid enum value; update registry snippet to include blocksMeta; add checklist items * fix(integrations): fix category dropdown by defining missing LANDING_INTEGRATIONS_DATA_PATH and regenerating integrations.json The staging merge introduced landing-content.ts but forgot to define LANDING_INTEGRATIONS_DATA_PATH in generate-docs.ts, causing the script to crash before writing integrations.json. The stale JSON had integrationTypes (plural array) from an older script version, while the Integration type and workspace UI both read integrationType (singular string) — so ALL_CATEGORY_SECTIONS bucketed to undefined and the category filters never appeared in the dropdown. Fixed by adding the missing path constant and re-running the generator. integrations.json now has 192 entries with the correct integrationType field. * fix(sidebar): restore resize handle on all pages commit |
||
|
|
b329c36b1a |
fix(auth): link SSO sign-in to existing same-email accounts (#4866)
* fix(auth): link SSO sign-in to existing same-email accounts SSO sign-ins failed with "account not linked" (then a cascading "Invalid callbackURL") when an account with the same email already existed. Better Auth's `@better-auth/sso` plugin hardcodes the provisioned user's `emailVerified: options?.trustEmailVerified ? <claim> : false`, so with the option unset every SSO login arrived unverified and tripped the account linking gate `(!isTrustedProvider && !userInfo.emailVerified)` whenever the provider was not in `accountLinking.trustedProviders`. - Set `trustEmailVerified: true` on the SSO plugin so the IdP's verified-email claim is honored (Okta, Entra ID, Google Workspace, Auth0 all assert it). - Trust the operator's configured provider for linking: merge `SSO_PROVIDER_ID` (when present in the app env) plus a new `SSO_TRUSTED_PROVIDER_IDS` list into `trustedProviders`. Empty/unset => no-op, so existing deployments are unchanged. - Invite callback URL: return a clean `/invite/<id>` (token already persists in sessionStorage) so an appended `?error=` cannot produce a malformed URL. - Document `SSO_TRUSTED_PROVIDER_IDS` in SSO docs, Helm values, and schema. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(auth): address review — guard trusted SSO providers, revert invite callback - Only compute additionalTrustedSsoProviders when SSO_ENABLED, so trustedProviders is exactly unchanged for non-SSO deployments. - Revert the invite getCallbackUrl change: keep the token in the callback URL (with sessionStorage/searchParams fallback) so the token survives when sessionStorage is unavailable. The account-linking fix removes the "account not linked" error that caused the malformed callback URL, so the callback cleanup is unnecessary. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(auth): guard trusted SSO providers with isSsoEnabled (isTruthy) env.SSO_ENABLED can be the string "false" (t3-env returns strings for booleans), which is truthy in JS. Use the canonical isSsoEnabled flag (isTruthy(env.SSO_ENABLED)) so SSO_ENABLED="false"/"0" correctly yields an empty trusted-provider list, matching how SSO is gated elsewhere. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
648a5a117d |
feat(storage): support S3-compatible endpoints (R2, MinIO, B2) for file storage (#4865)
* feat(storage): support S3-compatible endpoints (R2, MinIO, B2) for file storage Add S3_ENDPOINT and S3_FORCE_PATH_STYLE env vars, wired into the shared upload S3 client so Cloudflare R2, MinIO, Backblaze B2, and other S3-compatible stores work for self-hosted file storage. The endpoint is trusted operator config (no SSRF/HTTPS gate). Makes the multipart Location fallback endpoint-aware, extends the S3 client unit tests, and documents the new vars in Helm values, .env.example, and the English self-hosting docs (incl. browser-reachability + CORS guidance). * docs(storage): add RustFS as an S3-compatible provider example * fix(storage): address review feedback and fix env mock for CI - Add envBoolean to the shared env test mock (createEnvMock) so config.ts's forcePathStyle coercion resolves — fixes failing knowledge/utils.test.ts - Declare S3_FORCE_PATH_STYLE as z.string() (every other env var's pattern); it's coerced via envBoolean at the consumption site, avoiding a boolean type that never matches the string process.env value - Log path-style from S3_CONFIG.forcePathStyle (envBoolean) instead of a separate isTruthy call, so the startup log can't disagree with the client - Make buildObjectFallbackUrl honor forcePathStyle: virtual-hosted-style URL (bucket as subdomain) for R2, path-style only when forcePathStyle is set * docs(storage): add backlinks to S3-compatible providers (R2, MinIO, Ceph, B2, RustFS) and backends |
||
|
|
952eb1216f |
feat(mailer): add AWS SES and SMTP providers with auto-detect fallback (#4710)
* feat(mailer): add AWS SES and SMTP providers with auto-detect fallback * fix(mailer): cast SES options to bridge duplicate @aws-sdk type identities * fix(mailer): dedupe aws-sdk-sesv2, address review feedback - Force a single @aws-sdk/client-sesv2 install via root package.json overrides; @types/nodemailer pulled in a nested copy whose nominal class brand made the two SDK type identities incompatible, breaking the CI build. With one install the cast disappears. - Batch result message now reports successCount instead of sendable.length when entries are skipped, so "5 emails sent" no longer overstates delivery on partial failures. - SMTP provider now warns when SMTP_HOST is set without SMTP_PORT, and when only one of SMTP_USER/SMTP_PASS is set — both previously silent misconfigurations. - SMTP_SECURE schema is z.boolean() to match every other boolean in env.ts; runtime parsing is still handled by envBoolean. - Strip the verbose TSDoc comments I had added. * fix(mailer): exact sent counts in batch results, restore SES type cast - mergeBatchResults: data.count and the message now report only emails that were actually delivered, not skipped-unsubscribed ones (they returned success: true and inflated the count). Empty-sendable branch distinguishes "all unsubscribed" from "mixed skip/failure" so the message stops lying when some entries fail validation. - ses.ts: revert the package.json override approach (bun honors it locally but CI still installs a nested @types/nodemailer copy). Reinstate the `as unknown as` cast with a single-line WHY comment. * fix(mailer): annotate double-cast in ses provider for strict api-validation * fix(mailer): batch degrades isUnsubscribed errors to per-entry failures A transient DB error in isUnsubscribed used to abort the whole batch because the call sat outside the per-email try/catch in prepareBatch. Wrap the unsubscribe check inside the same catch so a rejection becomes a per-recipient failure, matching sendEmail's behavior. Lock it in with a regression test. |
||
|
|
21c956cf97 |
improvement(hubspot): OAuth-native polling trigger replacing webhook flow (#4705)
* improvement(hubspot): OAuth-native polling trigger replacing webhook flow * feat(hubspot): property autocomplete, multi-filter, property-changed, list-membership, pipeline/owner dropdowns * fix(hubspot): freeze cursor on failure + request full OAuth scopes * chore(api): bump API route baseline from 749 to 753 for HubSpot selector routes * fix(hubspot): make eventType required conditional on visibility * improvement(hubspot): align trigger name and longDescription with poll-trigger conventions * fix(hubspot): encodeURIComponent on search path segment for defense-in-depth * fix(hubspot): cursor-based seed for list_membership polling * fix(hubspot): Map-backed property snapshot + drop redundant filter parse |
||
|
|
48cf200ccd |
fix(helm): allow host[:port][/path] form in global.imageRegistry schema (#4686)
The values.schema.json constrained global.imageRegistry to JSON Schema hostname format (RFC 1123), which forbids '/'. That rejected the host+path form required by Artifactory virtual repos, Harbor projects, GCR (gcr.io/project-id), and ECR-with-namespace — all of which the chart's image-rendering helper already supports (it prints '%s/%s:%s'). Drop the format constraint and document the supported shapes. Matches the bitnami common-chart convention of validating image registry as a plain string and deferring to Docker for the actual reference parse. |
||
|
|
d1eb79ecd3 |
fix(helm): preserve STS serviceName + networkPolicy.egress back-compat (#4569)
* fix(helm): preserve STS serviceName + networkPolicy.egress back-compat
Greptile flagged two real upgrade-breaking changes vs the prior chart:
1. statefulset-postgresql spec.serviceName flipped from <name>-postgresql
to <name>-postgresql-headless. spec.serviceName is immutable, so any
existing install would hit 'Forbidden: updates to statefulset spec ...'
on helm upgrade. Revert to the original name (the headless Service in
services.yaml is added alongside, not as a swap).
2. networkPolicy.egress changed from a list to a map ({extraRules, exceptCidrs}),
silently dropping any custom egress list set by existing users. Restore
the original list semantics for networkPolicy.egress and move cloud-metadata
blocking to a sibling top-level field networkPolicy.egressExceptCidrs.
Adds NOTES.txt upgrade-notes entry covering both + the ESO v1→v1beta1 default
flip (functionally a no-op, but worth surfacing).
* docs(helm): update README egress reference to new key name
* fix(helm): revert copilot-postgresql STS serviceName too (same immutability issue)
Audit caught that the main fix in
|
||
|
|
9d2dd8f550 |
improvement(helm): helm chart updates with security, ESO, and docs overhaul (#4565)
* improvement(helm): production-ready chart with security, ESO, and docs overhaul
Comprehensive Helm chart improvements bringing the chart up to industry
standards for security, secret management, and documentation.
Security
- Pod Security Standards "restricted" defaults on every pod and container
(runAsNonRoot, allowPrivilegeEscalation=false, capabilities.drop=[ALL],
seccompProfile=RuntimeDefault)
- automountServiceAccountToken=false on ServiceAccount and every pod
- NetworkPolicy egress blocks cloud metadata endpoints by default
- Sensitive app/realtime env keys auto-partitioned into chart-managed Secret
via envFrom; no more plaintext secrets on container specs
Secret management
- Three modes: inline, existingSecret, ExternalSecrets Operator (ESO)
- ESO sync supports arbitrary sensitive keys
- Fail-fast template rendering when ESO enabled but sensitive key unmapped
- AWS/Azure/GCP example files document all three modes
Reliability
- Headless Services for both Postgres StatefulSets
- HPA-aware replicas (omits spec.replicas when autoscaling.enabled)
- PodDisruptionBudget auto-activates when replicaCount > 1
- Startup / liveness / readiness probes with distinct timings
- CronJob ttlSecondsAfterFinished for automatic cleanup
Chart hygiene
- Image tags default to Chart.AppVersion; pullPolicy IfNotPresent
- Optional image.digest pin for content-addressed deploys
- kubeVersion >=1.25.0-0 enforced
- Ollama pinned to 0.23.2; mount moved to /data
Documentation
- README rewritten in cert-manager / Bitnami style
- NOTES.txt with post-install guidance
- Example values files annotated with usage and secret-strategy guidance
* fix(helm): correct resource names in README (sim-sim-* → sim-*)
The sim.fullname helper collapses to the release name when the release
name contains the chart name. With the documented release name 'sim',
actual resources are 'sim-app', 'sim-postgresql', etc. — not the
'sim-sim-*' form previously documented. Fixes copy-paste commands in the
pre-1.0.0 upgrade walkthrough and several troubleshooting snippets.
Also expands the cronjobs component description to reflect the full set
of 13 scheduled jobs (was understated as just Gmail/Outlook polling).
* improvement(helm): split app/realtime env into Secret-bound + inline defaults
- Add app.envDefaults / realtime.envDefaults for chart-shipped operational
tunables (rate limits, timeouts, IVM, feature-flag defaults, localhost URL
fallbacks). Rendered inline on the container, not into the Secret
- Remove operational defaults from app.env / realtime.env so the chart-managed
Secret stays minimal and External Secrets Operator users only map keys they
actually set, not every chart default
- Skip an envDefaults key when the user explicitly sets it in env (K8s `env`
overrides `envFrom`, so an inline default would otherwise mask a Secret
value at runtime)
- Relax values.schema.json to allow empty strings on NEXT_PUBLIC_APP_URL,
BETTER_AUTH_URL, NEXT_PUBLIC_SUPPORT_EMAIL (defaults supplied via envDefaults)
* fix(helm): address PR review — cronjob validation, ESO apiVersion, secret merge order, image guard
- CronJobs reference CRON_SECRET via secretKeyRef; fail-fast at template
time when cronjobs.enabled=true and app.env.CRON_SECRET is empty so users
get a clear error instead of a CreateContainerConfigError loop
- Default externalSecrets.apiVersion to "v1beta1" (supported by every ESO
release since v0.7). The previous "v1" default targets only ESO v0.17+
- Swap merge order in secrets-app.yaml so app.env wins over realtime.env
for shared keys (BETTER_AUTH_SECRET, BETTER_AUTH_URL, …) — both pods
consume the same Secret via envFrom, so the app value must be canonical
- Add `required` guard on sim.image so an empty tag + empty digest +
empty Chart.AppVersion surfaces as a clear template-time error instead
of rendering an invalid `repo:` reference
* fix(helm): require critical secrets to be mapped when ESO is enabled
Previously, enabling externalSecrets without mapping BETTER_AUTH_SECRET /
ENCRYPTION_KEY / INTERNAL_API_SECRET (and CRON_SECRET when cronjobs are
on) rendered cleanly but produced CrashLoopBackOff at runtime with
cryptic missing-env errors. Fail at template time instead.
* fix(helm): auto-enable PDB when HPA minReplicas > 1
Previously the auto-enable predicate only checked the static
app.replicaCount, which defaults to 1 even when autoscaling is on
(HPA owns spec.replicas). PDB now also activates when
autoscaling.enabled=true and minReplicas > 1.
* fix(helm): prevent realtime envDefaults from masking app.env Secret values; add StatefulSet upgrade NOTES
- Realtime override-skip now considers keys set in either app.env or
realtime.env. The shared app Secret is mounted via envFrom on the
realtime pod, so a key set in app.env (e.g. NEXT_PUBLIC_APP_URL) would
previously be masked by the realtime envDefault (inline env overrides
envFrom in K8s).
- NOTES.txt now prints a StatefulSet orphan-delete reminder on upgrade,
surfacing the immutable serviceName issue documented in the README.
* feat(helm): add Claude Skill for chart deployment
Adds a skill at helm/sim/.claude/skills/sim-helm/ that teaches agents how
to deploy and troubleshoot the Sim Helm chart: install path selection
(inline / existingSecret / ESO), secret generation, the values.yaml
four-layer mental model, common-failure troubleshooting, and the
pre-1.0.0 StatefulSet orphan-delete upgrade procedure.
Skill is loadable by Claude Code, Codex, and OpenCode via the standard
skills convention (directory name matches frontmatter name).
* docs(helm): add CRON_SECRET to TL;DR, dry-run, and example install headers
The validateSecrets guard requires CRON_SECRET when cronjobs.enabled=true
(the default), but the quickstart and example file install commands
omitted it — users following the docs hit a hard template-render failure.
Adds CRON_SECRET to README TL;DR, validate-the-install dry-run snippet,
and the install command headers in all example values files.
* fix(helm): require INTERNAL_API_SECRET in inline secret mode
The ESO coverage validator already required INTERNAL_API_SECRET, but the
inline validateSecrets path only checked BETTER_AUTH_SECRET, ENCRYPTION_KEY,
and CRON_SECRET — letting inline installs render successfully and then
crash at runtime when the realtime↔app shared auth secret was missing.
Adds the same fail-fast check to the inline path.
* docs(helm): surface INTERNAL_API_SECRET upgrade requirement in NOTES.txt
The new validateSecrets check makes app.env.INTERNAL_API_SECRET mandatory
on upgrade. Existing installs that never set it would hit a template
render failure with no in-context guidance. Adds an upgrade-only note
with the generation snippet and storage guidance alongside the existing
StatefulSet orphan-delete instructions.
* fix(helm): NetworkPolicy egress to OTEL collector + external-db example format
- Add app/realtime NetworkPolicy egress rules for the OpenTelemetry
collector pod on ports 4317 (OTLP gRPC) and 4318 (OTLP HTTP) when
telemetry.enabled=true. Without these, traces and metrics were silently
dropped with connection-refused errors when both telemetry and
networkPolicy were enabled.
- Migrate values-external-db.yaml from the legacy list-shaped egress
format to the new {exceptCidrs, extraRules} object. The list form would
replace the default object on merge and crash template rendering when
the chart tried to access .exceptCidrs on a list.
* fix(helm): NOTES.txt no longer prints false secret warning for ESO users
The secrets-empty warning only checked app.secrets.existingSecret.enabled
before scanning app.env. ESO users intentionally leave app.env empty —
secrets come from the ESO-synced Secret — so every ESO install/upgrade
printed a misleading 'pods will fail to start' warning.
Reorders the branches so externalSecrets.enabled takes precedence: ESO
users now see a confirmation message with kubectl commands to verify the
ExternalSecret has synced. The empty-app.env warning only fires when
both ESO and existingSecret are disabled.
* fix(helm): existingSecret mode no longer drops app.env / realtime.env values
In existingSecret mode the chart-managed Secret is not rendered, so non-empty
values in app.env / realtime.env had nowhere to land — yet the envDefaults
skip logic still suppressed the matching defaults. Result: keys like
NEXT_PUBLIC_APP_URL, BETTER_AUTH_URL, and NODE_ENV silently went missing
on both pods (the example values-existing-secret.yaml hit this directly).
Both app and realtime deployments now inline non-empty values from app.env
(plus realtime.env on the realtime container) when existingSecret is enabled
and ESO is not. Inline / ESO modes are unchanged: inline still flows through
the chart-managed Secret, ESO still owns the synced Secret.
* fix(helm): correct realtime env overlay + filter chart-computed keys in existingSecret mode
Realtime: Sprig merge gives the first source precedence and treats "" as a
real value, so realtime.env empty defaults for shared keys shadowed
non-empty app.env values. Replace with deepCopy($appEnv) base + manual
non-empty overlay of $rtEnv.
Both deployments: exclude DATABASE_URL/SOCKET_SERVER_URL/OLLAMA_URL from
the existingSecret inline path so user-supplied values can't override
chart-computed ones via last-wins env semantics.
* fix(helm): skip envDefaults in existingSecret mode + document egress rename
In existingSecret mode the user's pre-existing Secret is the source of
truth (loaded via envFrom). Inlining localhost envDefaults for URL keys
(BETTER_AUTH_URL, NEXT_PUBLIC_APP_URL, ALLOWED_ORIGINS) silently shadowed
the Secret-bound values because K8s env always wins over envFrom. Skip
envDefaults entirely on both deployments when existingSecret is enabled.
Also call out the networkPolicy.egress shape change (list -> map with
exceptCidrs + extraRules) in the NOTES.txt upgrade block so operators
migrate their custom rules rather than silently losing them.
* fix(helm): copy-pasteable install commands in copilot + ESO examples
values-copilot.yaml: the install header was missing every required
copilot.server.env.* secret (AGENT_API_DB_ENCRYPTION_KEY, INTERNAL_API_SECRET,
LICENSE_KEY, SIM_BASE_URL, SIM_AGENT_API_KEY, REDIS_URL, one model key) plus
copilot.postgresql.auth.password. Pasting it as-is failed at template render.
values-external-secrets.yaml: NEXT_PUBLIC_APP_URL, BETTER_AUTH_URL, etc. were
declared under app.env / realtime.env. In ESO mode the chart-managed Secret
isn't rendered, so the validator (rightly) rejects keys in app.env that
aren't mapped under externalSecrets.remoteRefs. Moved non-secret URL/config
to envDefaults, which is inlined and not subject to the ESO mapping rule.
* polish(helm): configurable NetworkPolicy ingress peers + clearer API_ENCRYPTION_KEY comment
- networkPolicy.ingressFrom lets operators scope the ingress-controller
rule to a specific namespace/podSelector. Defaults to a single empty
peer (`- {}`), which is the explicit form of "any source" — same
effective behavior as the old `from: []` but unambiguous across CNIs.
To restrict, override with e.g.:
networkPolicy:
ingressFrom:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: ingress-nginx
- API_ENCRYPTION_KEY comment: drop the "must be exactly 64 hex
characters" phrasing that sat awkwardly next to `openssl rand -hex 32`.
The generation command already produces the required length.
* test(helm): add helm-unittest suites + CI workflow + ci values matrix
- 7 helm-unittest suites covering smoke, validators, secret modes,
envDefaults secret-mode-aware inlining (round-9 regression net),
chart-computed env keys (round-8 regression net), NetworkPolicy
shape, and PDB/HPA conditional rendering (38 tests, ~265ms).
- ci/*.yaml render fixtures for default, production, existingSecret,
ESO, and external-db install modes.
- GitHub Actions workflow runs helm lint --strict, helm unittest,
helm template across the ci matrix, and kubeconform validation
against Kubernetes 1.30 schemas.
- CONTRIBUTING.md documents how to run the same gates locally.
* test(helm): add helm test hook + kind apiserver dry-run in CI
- New templates/tests/test-connection.yaml renders a Pod with
helm.sh/hook=test that wgets the app Service (and realtime when
enabled). Lets users run `helm test <release>` after install for
a real in-cluster connectivity check. Restricted PSS context.
- tests.* values block (image, timeoutSeconds, resources) is the
knob to disable or tune the probe; documented in values.schema.json.
- 3 helm-unittest tests cover the hook annotations, PSS context,
and tests.enabled=false skip path (41 tests total).
- New CI job spins up a kind v1.30 cluster and runs
`kubectl apply --dry-run=server` against the rendered manifests
for the CRD-free ci fixtures (default / existing-secret /
external-db). Catches admission and validation issues the static
kubeconform schema check can't see.
* chore(helm): remove pre-1.0.0 upgrade fluff + tighten .helmignore
This is the 1.0.0 release of the chart — there is no pre-1.0.0
predecessor for users to upgrade from, so all of the dedicated upgrade
narration was hypothetical.
- Drop the 'Upgrading from a pre-1.0.0 build' README section and the
matching troubleshooting entry.
- Drop the .Release.IsUpgrade block from NOTES.txt: items 5 (StatefulSet
orphan-delete), 6 (INTERNAL_API_SECRET 'new in 1.0.0'), 7
(networkPolicy.egress shape change). Each described a migration off a
chart version that never shipped.
- Delete references/upgrade-pre-1.0.0.md and remove the corresponding
pointers from SKILL.md.
- Anchor .helmignore patterns to chart root so /tests/ (unit suites)
and /examples/ are dropped from the packaged tarball without also
dropping templates/tests/ (the helm test hook).
* chore(helm): drop CI workflow + ci/ fixtures + CONTRIBUTING.md
The helm-unittest suites in helm/sim/tests/ and the helm test hook
in helm/sim/templates/tests/ stay — those are chart-internal quality
scaffolding, not CI. Removed:
- .github/workflows/helm-chart.yml
- helm/sim/ci/*.yaml (5 render fixtures used only by the workflow)
- helm/sim/CONTRIBUTING.md (mostly documented those gates)
- dead /ci/ and /CONTRIBUTING.md entries in .helmignore
* feat(helm): pod rollout on Secret change + topologySpreadConstraints
- Add checksum/secret pod annotations on app, realtime, and copilot
Deployments (plus checksum/config on app when branding ConfigMap is
enabled). Closes the long-standing footgun where 'helm upgrade' with
a changed Secret would silently leave pods running the old values
until a manual rollout restart.
- New top-level topologySpreadConstraints value (and sim.topologySpreadConstraints
helper) applied to app and realtime Deployments. Mirrors how affinity
and tolerations are plumbed; users supply their own labelSelector
to mirror Bitnami convention.
- 5 helm-unittest cases cover the checksum annotations and topology
spread rendering (46 tests total).
* fix(helm): drop empty-string shadowing in app/realtime env merge
Sprig 'merge' treats "" as a real value, so a default-empty
app.env.BETTER_AUTH_URL would shadow a non-empty realtime.env override
and the URL would never reach the rendered Secret. Replace 'merge'
with an explicit two-pass overlay that filters empties before writing,
mirroring the same pattern already used in deployment-realtime.yaml's
existingSecret block.
Adds two regression tests: realtime.env-only value reaches the Secret
when app.env is empty, and app.env still wins on collision when both
are non-empty (48 tests total).
* fix(helm): make topologySpreadConstraints per-component to match docstring
Greptile flagged that sim.topologySpreadConstraints helper docstring promised
per-component config (.Values.app, .Values.realtime, ...) but call sites
passed .Values, so any app.topologySpreadConstraints / realtime.topologySpreadConstraints
set by the user was silently dropped. The single global key also prevented
distinct app-vs-realtime spread rules.
Pass .Values.app / .Values.realtime to the helper at each call site; move
the top-level topologySpreadConstraints key into both component sections in
values.yaml. Adds a regression test that app constraints don't leak onto
the realtime pod.
* fix(helm): allow cron pods through app NetworkPolicy
Cursor flagged that when networkPolicy.enabled=true and cronjobs.enabled=true
(the recommended production config), the app NetworkPolicy only allowed
ingress from realtime and the ingress controller — silently blocking every
cron pod's HTTP call to /api/schedules/execute, webhook polls, etc. All 13
default cronjobs would fail.
Tag cron pods with a stable simstudio.ai/component-group: cronjob label so
the app NetworkPolicy can allow them with a single rule (no per-job
enumeration). Rule is conditional on cronjobs.enabled. Adds positive and
negative regression tests.
|
||
|
|
d721dc3358 |
feat(enterprise): add data drains for continuous export to S3 / webhook (#4440)
* feat(enterprise): add data drains for continuous export to S3 / webhook * chore(data-drains): regenerate migration on top of staging + bump route baseline Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * docs(data-drains): clarify retention pairing is user-coupled, not enforced Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(data-drains): preserve explicit forcePathStyle=false + reserve x-sim-signature Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(data-drains): drift guard ensures every webhook header is reserved Asserts that any header buildHeaders writes is rejected when reused as a custom signatureHeader. Adding a new metadata header without mirroring it into RESERVED_SIGNATURE_HEADER_NAMES now fails CI. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
09f4c94b3c |
feat(block): Allow wait block to wait up to 30 days (#4331)
* v0.6.29: login improvements, posthog telemetry (#4026) * feat(posthog): Add tracking on mothership abort (#4023) Co-authored-by: Theodore Li <theo@sim.ai> * fix(login): fix captcha headers for manual login (#4025) * fix(signup): fix turnstile key loading * fix(login): fix captcha header passing * Catch user already exists, remove login form captcha * feat(block): Allow wait block to wait up to 30 days * restore ff * Filter out waits from hitl endpoints * Use correct count, filtering out wait blocks * improvement(wait): tighten poll route and pause-manager helpers - Parallelize per-row dispatch with Promise.all - Add status='paused' guard on nextResumeAt rewrite to prevent clobbering concurrent resumes - Extract computeEarliestResumeAt + PauseResumeManager.setNextResumeAt helpers - Use canonical PausePoint type in poll route (drop StoredPausePoint) - Narrow UNIT_TO_MS via as const + WaitUnit guard - Bump LOCK_TTL_SECONDS above route maxDuration - Clearer error when allowedPauseKinds rejects a resume --------- Co-authored-by: Waleed <walif6@gmail.com> Co-authored-by: Siddharth Ganesan <33737564+Sg312@users.noreply.github.com> Co-authored-by: Vikhyath Mondreti <vikhyathvikku@gmail.com> |
||
|
|
57dc745bab |
feat(knowledge): expose Cohere reranker controls (#4429)
* feat(knowledge): expose Cohere reranker controls on knowledge block Add a self-hosted Cohere API key field (mirroring the agent block's hosted-key pattern), a configurable reranker input pool size (1-100), and surface meta.warnings from Cohere rerank responses via logger.warn. All new contract fields are optional and nullable for full backwards compatibility. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(knowledge): address PR feedback on Cohere reranker controls - Drop required:true on apiKey field — server has BYOK→env→rotation fallback chain, so self-hosted users with COHERE_API_KEY env should not be blocked - Drop .min(1) on rerankerApiKey contract field so empty strings coerce to undefined via the transform (matches the existing query field pattern) - Log a warning when rerankerInputCount is clamped up to topK so users notice their setting was overridden Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * feat(knowledge): mirror agent block API key visibility for Cohere reranker Restore required:true on the Cohere API Key field and hide it server-side via a new NEXT_PUBLIC_COHERE_CONFIGURED public env flag — same pattern the Agent block uses for Azure (NEXT_PUBLIC_AZURE_CONFIGURED). Self-hosters who set COHERE_API_KEY in their environment also set NEXT_PUBLIC_COHERE_CONFIGURED=true, which removes the field from the UI; everyone else sees a required field. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(knowledge): treat empty rerankerInputCount as unset An empty string from the Documents Sent to Reranker input passed the undefined/null guard, so Number('') = 0 → clamped to 1, sending only 1 document to the reranker instead of falling back to the 4× topK auto default. Add the empty-string check to the guard. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
879dab9f19 |
feat(table): make plan table limits configurable via env vars (#4406)
* feat(table): make plan table limits configurable via env vars * fix(table): coerce env table limits to number for skipValidation env * improvement(env): extract envNumber helper for numeric env coercion * improvement(knowledge): use envNumber helper for KB_CONFIG_* env reads * fix(testing): add envNumber to env mock factory * fix(env): allow zero in envNumber for max-throughput configs * fix(env): add min option to envNumber for strict-positive configs |
||
|
|
0c25fc4ee1 |
fix(auth): resolve CORS errors for self-hosted deployments behind reverse proxies (#4369)
* fix(auth): resolve CORS errors for self-hosted deployments behind reverse proxies
- auth client now uses browser origin first, falling back to NEXT_PUBLIC_APP_URL
- socket client falls back to page origin when served from non-localhost (assumes /socket.io is proxied)
- add TRUSTED_ORIGINS env var to extend Better Auth trustedOrigins (apex+www, alias hostnames)
- warn at startup when NEXT_PUBLIC_APP_URL is localhost in production
- preprocess empty NEXT_PUBLIC_SOCKET_URL so docker-compose ${VAR:-} works
- migrate remaining uuid/nanoid/randomUUID usages to @sim/utils generateId/generateShortId
- extend generateShortId with optional alphabet param (rejection sampling)
- document TRUSTED_ORIGINS in .env.example, docker-compose.prod.yml, and helm values.yaml
Fixes simstudioai/sim#1243
* fix(auth): address PR review comments
* chore(env): drop unnecessary NEXT_PUBLIC_SOCKET_URL preprocess (skipValidation is true)
* fix(docker): include @sim/utils in migrations image
Migration scripts now import generateId from @sim/utils/id; without copying packages/utils into the image, bun install fails to resolve the workspace dep at build time and the import fails at runtime.
* fix(helm): remove unused NEXT_PUBLIC_SOCKET_URL from realtime sections
The realtime service never reads NEXT_PUBLIC_SOCKET_URL — its env schema
only includes BETTER_AUTH_URL, NEXT_PUBLIC_APP_URL, ALLOWED_ORIGINS,
BETTER_AUTH_SECRET, INTERNAL_API_SECRET, DATABASE_URL, and REDIS_URL.
Remove the dead config from all helm values files and the values schema.
* fix(helm): allow empty NEXT_PUBLIC_SOCKET_URL in values schema
The default in values.yaml is now "" (empty string), which falls back to
the page origin at runtime. The schema previously required a valid URI,
which would reject the default. Mirror the INTERNAL_API_BASE_URL pattern
using anyOf with const "". Also add TRUSTED_ORIGINS to the schema.
* docs(self-hosting): mark NEXT_PUBLIC_SOCKET_URL as optional
The page-origin fallback in getSocketUrl() means self-hosters no longer
need to set NEXT_PUBLIC_SOCKET_URL when realtime is on the same origin
as the app. Update docs to reflect this:
- Remove NEXT_PUBLIC_SOCKET_URL from .env scaffolding examples in
docker.mdx, platforms.mdx, environment-variables.mdx
- Mark the variable as Optional in the env vars table with the new
default behavior described
- Update troubleshooting to point at reverse-proxy /socket.io routing
rather than the env var
- Flip dev docker-compose defaults (local, ollama, devcontainer) from
http://localhost:3002 to empty for consistency with prod.yml; the
in-code localhost fallback handles the dev case identically
Applied across all 6 documentation languages (en/fr/de/ja/es/zh).
* chore: untrack and ignore .claude/scheduled_tasks.lock
|
||
|
|
0abcc6e813 |
improvement(mothership): restructured stream, tool structures, code typing, file write/patch/append tools, timing issues (#4090)
* fix build error * improvement(mothership): new agent loop (#3920) * feat(transport): replace shared chat transport with mothership-stream module * improvement(contracts): regenerate contracts from go * feat(tools): add tool catalog codegen from go tool contracts * feat(tools): add tool-executor dispatch framework for sim side tool routing * feat(orchestrator): rewrite tool dispatch with catalog-driven executor and simplified resume loop * feat(orchestrator): checkpoint resume flow * refactor(copilot): consolidate orchestrator into request/ layer * refactor(mothership): reorganize lib/copilot into structured subdirectories * refactor(mothership): canonical transcript layer, dead code cleanup, type consolidation * refactor(mothership): rebase onto latest staging * refactor(mothership): rename request continue to lifecycle * feat(trace): add initial version of request traces * improvement(stream): batch stream from redis * fix(resume): fix the resume checkpoint * fix(resume): fix resume client tool * fix(subagents): subagent resume should join on existing subagent text block * improvement(reconnect): harden reconnect logic * fix(superagent): fix superagent integration tools * improvement(stream): improve stream perf * Rebase with origin dev * fix(tests): fix failing test * fix(build): fix type errors * fix(build): fix build errors * fix(build): fix type errors * feat(mothership): add cli execution * fix(mothership): fix function execute tests * Force redeploy * feat(motheship): add docx support * feat(mothership): append * Add deps * improvement(mothership): docs * File types * Add client retry logic * Fix stream reconnect * Eager tool streaming * Fix client side tools * Security * Fix shell var injection * Remove auto injected tasks * Fix 10mb tool response limit * Fix trailing leak * Remove dead tools * file/folder tools * Folder tools * Hide function code inline * Dont show internal tool result reads * Fix spacing * Auth vfs * Empty folders should show in vfs * Fix run workflow * change to node runtime * revert back to bun runtime * Fix * Appends * Remove debug logs * Patch * Fix patch tool * Temp * Checkpoint * File writes * Fix * Remove tool truncation limits * Bad hook * replace react markdown with streamdown * Checkpoitn * fix code block * fix stream persistence * temp * Fix file tools * tool joining * cleanup subagent + streaming issues * streamed text change * Tool display intetns * Fix dev * Fix tests * Fix dev * Speed up dev ci * Add req id * Fix persistence * Tool call names * fix payload accesses * Fix name * fix snapshot crash bug * fix * Fix * remove worker code * Clickable resources * Options ordering * Folder vfs * Restore and mass delete tools * Fix * lint * Update request tracing and skills and handlers * Fix editable * fix type error * Html code * fix(chat): make inline code inherit parent font size in markdown headers Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * improved autolayout * durable stream for files * one more fix * POSSIBLE BREAKAGE: SCROLLING * Fixes * Fixes * Lint fix * fix(resource): fix resource view disappearing on ats (#4103) Co-authored-by: Theodore Li <theo@sim.ai> * Fixes * feat(mothership): add execution logs as a resource type Adds `log` as a first-class mothership resource type so copilot can open and display workflow execution logs as tabs alongside workflows, tables, files, and knowledge bases. - Add `log` to MothershipResourceType, all Zod enums, and VALID_RESOURCE_TYPES - Register log in RESOURCE_REGISTRY (Library icon) and RESOURCE_INVALIDATORS - Add EmbeddedLog and EmbeddedLogActions components in resource-content - Export WorkflowOutputSection from log-details for reuse in EmbeddedLog - Add log resolution branch in open_resource handler via new getLogById service - Include log id in get_workflow_logs response and extract resources from output - Exclude log from manual add-resource dropdown (enters via copilot tools only) - Regenerate copilot contracts after adding log to open_resource Go enum * Fix perf and message queueing * Fix abort * fix(ui): dont delete resource on clearing from context, set resource closed on new task (#4113) Co-authored-by: Theodore Li <theo@sim.ai> * improvement(mothership): structure sim side typing * address comments * reactive text editor tweaks * Fix file read and tool call name persistence bug * Fix code stream + create file opening resource * fix use chat race + headless trace issues * Fix type issue * Fix mothership block req lifecycle * Fix build * Move copy reqid * Fix * fix(ui): fix resource tag transition from home to task (#4132) Co-authored-by: Theodore Li <theo@sim.ai> * Fix persistence --------- Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai> Co-authored-by: Waleed Latif <walif6@gmail.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Theodore Li <theo@sim.ai> Co-authored-by: Theodore Li <theodoreqili@gmail.com> |
||
|
|
7491d70a67 |
feat(workspaces): add workspace logo upload (#4136)
* feat(workspaces): add workspace logo upload * feat(workspaces): add workspace logo upload * fix(workspaces): validate logoUrl accepts only paths or HTTPS URLs * fix(workspaces): add admin authorization, audit log, and posthog event for workspace logo uploads * lint * fix: add WebP support and use refs pattern in useProfilePictureUpload - Add image/webp to ACCEPTED_IMAGE_TYPES in useProfilePictureUpload - Add image/webp to file input accept attributes in whitelabeling settings - Refactor useProfilePictureUpload to use refs for onUpload, onError, and currentImage callbacks, matching the established codebase pattern Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: restore cloudwatch/cloudformation files from staging These files were accidentally regressed during rebase conflict resolution, reverting changes from #4027. Restoring to staging versions. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add workspace_logo_uploaded to PostHogEventMap Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: separate workspaceId ref sync to prevent overwrite on re-render Split the ref sync useEffect so workspaceIdRef only updates when the workspaceId prop changes, not when onUpload/onError callbacks get new references. Prevents setTargetWorkspaceId from being overwritten by a re-render before the file upload completes. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: use Pick type for workspace dropdown in knowledge header The shared Workspace type requires ownerId and other fields that aren't available from the workspaces API response mapping. Use a Pick type to accurately represent the subset of fields actually constructed. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor: replace raw fetch with useWorkspacesQuery in knowledge header Remove useState + useEffect + fetch anti-pattern for loading workspaces. Use useWorkspacesQuery from React Query with inline filter for write/admin permissions. Eliminates ~30 lines of manual state management, any casts, and the Pick type workaround. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
85f1d96859 |
feat(ee): enterprise feature flags, permission group platform controls, audit logs ui, delete account (#4115)
* feat(ee): enterprise feature flags, permission group platform controls, audit logs ui, delete account * fix(settings): improve sidebar skeleton fidelity and fix credit purchase org cache invalidation - Bump skeleton icon and text from 16/14px to 24px to better match real nav item visual weight - Add orgId support to usePurchaseCredits so org billing/subscription caches are invalidated on credit purchase, matching the pattern used by useUpgradeSubscription - Polish ColorInput in whitelabeling settings with auto-prefix and select-on-focus UX * revert(settings): remove delete account feature * fix(settings): address pr review — atomic autoAddNewMembers, extract query hook, fix types and signal forwarding * chore(helm): add CREDENTIAL_SETS_ENABLED to values.yaml * fix(access-control): dynamic platform category columns, atomic permission group delete * fix(access-control): restore triggers section in blocks tab * fix(access-control): merge triggers into tools section in blocks tab * upgrade tubro * fix(access-control): fix Select All state when config has stale blacklisted provider IDs * fix(access-control): derive platform Select All from features list; revert turbo schema version * fix(access-control): fix blocks Select All check, filter empty platform columns * revert(settings): restore original skeleton icon and text sizes |
||
|
|
6099683e5a |
feat(trigger): add Google Sheets, Drive, and Calendar polling triggers (#4081)
* feat(trigger): add Google Sheets, Drive, and Calendar polling triggers Add polling triggers for Google Sheets (new rows), Google Drive (file changes via changes.list API), and Google Calendar (event updates via updatedMin). Each includes OAuth credential support, configurable filters (event type, MIME type, folder, search term, render options), idempotency, and first-poll seeding. Wire triggers into block configs and regenerate integrations.json. Update add-trigger skill with polling instructions and versioned block wiring guidance. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(polling): address PR review feedback for Google polling triggers - Fix Drive cursor stall: use nextPageToken as resume point when breaking early from pagination instead of re-using the original token - Eliminate redundant Drive API call in Sheets poller by returning modifiedTime from the pre-check function - Add 403/429 rate-limit handling to Sheets API calls matching the Calendar handler pattern - Remove unused changeType field from DriveChangeEntry interface - Rename triggers/google_drive to triggers/google-drive for consistency Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(polling): fix Drive pre-check never activating in Sheets poller isDriveFileUnchanged short-circuited when lastModifiedTime was undefined, never calling the Drive API — so currentModifiedTime was never populated, creating a permanent chicken-and-egg loop. Now always calls the Drive API and returns the modifiedTime regardless of whether there's a previous value to compare against. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * chore(lint): fix import ordering in triggers registry Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(polling): address PR review feedback for Google polling handlers - Fix fetchHeaderRow to throw on 403/429 rate limits instead of silently returning empty headers (prevents rows from being processed without headers and lastKnownRowCount from advancing past them permanently) - Fix Drive pagination to avoid advancing resume cursor past sliced changes (prevents permanent change loss when allChanges > maxFiles) - Remove unused logger import from Google Drive trigger config * fix(polling): prevent data loss on partial row failures and harden idempotency key - Sheets: only advance lastKnownRowCount by processedCount when there are failures, so failed rows are retried on the next poll cycle (idempotency deduplicates already-processed rows on re-fetch) - Drive: add fallback for change.time in idempotency key to prevent key collisions if the field is ever absent from the API response * fix(polling): remove unused variable and preserve lastModifiedTime on Drive API failure - Remove unused `now` variable from Google Drive polling handler - Preserve stored lastModifiedTime when Drive API pre-check fails (previously wrote undefined, disabling the optimization until the next successful Drive API call) * fix(polling): don't advance state when all events fail across sheets, calendar, drive handlers * fix(polling): retry failed idempotency keys, fix drive cursor overshoot, fix calendar inclusive updatedMin * fix(polling): revert calendar timestamp on any failure, not just all-fail * fix(polling): revert drive cursor on any failure, not just all-fail * feat(triggers): add canonical selector toggle to google polling triggers - Add 'trigger-advanced' mode to SubBlockConfig so canonical pairs work in trigger mode - Fix buildCanonicalIndex: trigger-mode subblocks don't overwrite non-trigger basicId, deduplicate advancedIds from block spreads - Update editor, subblock layout, and trigger config aggregation to include trigger-advanced subblocks - Replace dropdown+fetchOptions in Calendar/Sheets/Drive pollers with file-selector (basic) + short-input (advanced) canonical pairs - Add canonicalParamId: 'oauthCredential' to triggerCredentials for selector context resolution - Update polling handlers to read canonical fallbacks (calendarId||manualCalendarId, etc.) * test(blocks): handle trigger-advanced mode in canonical validation tests * fix(triggers): handle trigger-advanced mode in deploy, preview, params, and copilot * fix(polling): use position-only idempotency key for sheets rows * fix(polling): don't advance calendar timestamp to client clock on empty poll * fix(polling): remove extraneous comment from calendar poller * fix(polling): drive cursor stall on full page, calendar latestUpdated past filtered events * fix(polling): advance calendar cursor past fully-filtered event batches --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
1189400167 |
feat(enterprise): cloud whitelabeling for enterprise orgs (#4047)
* feat(enterprise): cloud whitelabeling for enterprise orgs * fix(enterprise): scope enterprise plan check to target org in whitelabel PUT * fix(enterprise): use isOrganizationOnEnterprisePlan for org-scoped enterprise check * fix(enterprise): allow clearing whitelabel fields and guard against empty update result * fix(enterprise): remove webp from logo accept attribute to match upload hook validation * improvement(billing): use isBillingEnabled instead of isProd for plan gate bypasses * fix(enterprise): show whitelabeling nav item when billing is enabled on non-hosted environments * fix(enterprise): accept relative paths for logoUrl since upload API returns /api/files/serve/ paths * fix(whitelabeling): prevent logo flash on refresh by hiding logo while branding loads * fix(whitelabeling): wire hover color through CSS token on tertiary buttons * fix(whitelabeling): show sim logo by default, only replace when org logo loads * fix(whitelabeling): cache org logo url in localstorage to eliminate flash on repeat visits * feat(whitelabeling): add wordmark support with drag/drop upload * updated turbo * fix(whitelabeling): defer localstorage read to effect to prevent hydration mismatch * fix(whitelabeling): use layout effect for cache read to eliminate logo flash before paint * fix(whitelabeling): cache theme css to eliminate color flash before org settings resolve * fix(whitelabeling): deduplicate HEX_COLOR_REGEX into lib/branding and remove mutation from useCallback deps * fix(whitelabeling): use cookie-based SSR cache to eliminate brand flash on all page loads * fix(whitelabeling): use !orgSettings condition to fix SSR brand cache injection React Query returns isLoading: false with data: undefined during SSR, so the previous brandingLoading condition was always false on the server — initialCache was never injected into brandConfig. Changing to !orgSettings correctly applies the cookie cache both during SSR and while the client-side query loads, eliminating the logo flash on hard refresh. |
||
|
|
3c7bfa797a |
improvement(kb): deferred content fetching and metadata-based hashes for connectors (#4044)
* improvement(kb): deferred content fetching and metadata-based hashes for connectors * fix(kb): remove message count from outlook contentHash to prevent list/get divergence * fix(kb): increase outlook getDocument message limit from 50 to 250 * fix(kb): skip outlook messages without conversationId to prevent broken stubs * fix(kb): scope outlook getDocument to same folder as listDocuments to prevent hash divergence * fix(kb): add missing connector sync cron job to Helm values The connector sync endpoint existed but had no cron job configured to trigger it, meaning scheduled syncs would never fire. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
c89a95d606 |
feat(auth): add DISABLE_GOOGLE_AUTH and DISABLE_GITHUB_AUTH env vars (#4019)
* feat(auth): add DISABLE_GOOGLE_AUTH and DISABLE_GITHUB_AUTH env vars * fix(auth): also disable server-side OAuth provider registration when flags are set * lint |
||
|
|
8527ae5d3b |
feat(providers): server-side credential hiding for Azure and Bedrock (#3884)
* fix: allow Bedrock provider to use AWS SDK default credential chain Remove hard requirement for explicit AWS credentials in Bedrock provider. When access key and secret key are not provided, the AWS SDK automatically falls back to its default credential chain (env vars, instance profile, ECS task role, EKS IRSA, SSO). Closes #3694 Signed-off-by: majiayu000 <1835304752@qq.com> * fix: add partial credential guard for Bedrock provider Reject configurations where only one of bedrockAccessKeyId or bedrockSecretKey is provided, preventing silent fallback to the default credential chain with a potentially different identity. Add tests covering all credential configuration scenarios. Signed-off-by: majiayu000 <1835304752@qq.com> * fix: clean up bedrock test lint and dead code Remove unused config parameter and dead _lastConfig assignment from mock factory. Break long mockReturnValue chain to satisfy biome line-length rule. Signed-off-by: majiayu000 <1835304752@qq.com> * fix: address greptile review feedback on PR #3708 Use BedrockRuntimeClientConfig from SDK instead of inline type. Add default return value for prepareToolsWithUsageControl mock. Signed-off-by: majiayu000 <1835304752@qq.com> * feat(providers): server-side credential hiding for Azure and Bedrock * fix(providers): revert Bedrock credential fields to required with original placeholders * fix(blocks): add hideWhenEnvSet to getProviderCredentialSubBlocks for Azure and Bedrock * fix(agent): use getProviderCredentialSubBlocks() instead of duplicating credential subblocks * fix(blocks): consolidate Vertex credential into shared factory with basic/advanced mode * fix(types): resolve pre-existing TypeScript errors across auth, secrets, and copilot * lint * improvement(blocks): make Vertex AI project ID a password field * fix(blocks): preserve vertexCredential subblock ID for backwards compatibility * fix(blocks): follow canonicalParamId pattern correctly for vertex credential subblocks * fix(blocks): keep vertexCredential subblock ID stable to preserve saved workflow state * fix(blocks): add canonicalParamId to vertexCredential basic subblock to complete the swap pair * fix types * more types --------- Signed-off-by: majiayu000 <1835304752@qq.com> Co-authored-by: majiayu000 <1835304752@qq.com> Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai> |
||
|
|
f1ead2ed55 | fix docker image build | ||
|
|
d2c3c1c39e |
improvement(worker): configuration defaults (#3821)
* improvement(worker): configuration defaults * update readmes * realtime curl import |
||
|
|
21156dd54a |
fix(worker): dockerfile + helm updates (#3818)
* fix(worker): dockerfile + helm updates * address comments |
||
|
|
dda012eae9 |
feat(concurrency): bullmq based concurrency control system (#3605)
* feat(concurrency): bullmq based queueing system * fix bun lock * remove manual execs off queues * address comments * fix legacy team limits * cleanup enterprise typing code * inline child triggers * fix status check * address more comments * optimize reconciler scan * remove dead code * add to landing page * Add load testing framework * update bullmq * fix * fix headless path --------- Co-authored-by: Theodore Li <teddy@zenobiapay.com> |
||
|
|
4a34ac3015 |
feat(auth): add Turnstile captcha + harmony disposable email blocking (#3699)
* feat(turnstile): conditionally added CF turnstile to signup * feat(auth): add execute-on-submit Turnstile, conditional harmony, and feature flag - Switch Turnstile to execution: 'execute' mode so challenge runs on form submit (fresh token every time, no expiry issues) - Make emailHarmony conditional via SIGNUP_EMAIL_VALIDATION_ENABLED feature flag so self-hosted users can opt out - Add isSignupEmailValidationEnabled to feature-flags.ts following existing pattern - Add better-auth-harmony to Next.js transpilePackages (required for validator.js ESM compatibility) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(validation): remove dead validateEmail and checkMXRecord Server-side disposable email blocking is now handled by better-auth-harmony. The async validateEmail (with MX check) had no remaining callers. Only quickValidateEmail remains for client-side form feedback. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(auth): add 15s timeout to Turnstile captcha promise Prevents form from hanging indefinitely if Turnstile never fires onSuccess/onError (e.g. script fails to load, network drop). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(helm): add Turnstile and harmony env vars to values.yaml Adds TURNSTILE_SECRET_KEY, NEXT_PUBLIC_TURNSTILE_SITE_KEY, and SIGNUP_EMAIL_VALIDATION_ENABLED to the helm chart so self-hosted deployments can configure captcha and disposable email blocking. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(auth): reject captcha promise on token expiry onExpire now rejects the pending promise so the form doesn't hang if the Turnstile token expires mid-challenge. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(login): replace useEffect keydown listener with form onSubmit The forgot-password modal used a global window keydown listener in a useEffect to handle Enter key — a "you might not need an effect" anti-pattern with a stale closure risk. Replaced with a native <form onSubmit> wrapper which handles Enter natively, eliminating the useEffect, the global listener, and the stale closure. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(auth): clear dangling timeout after captcha promise settles Use .finally(() => clearTimeout(timeoutId)) to clean up the 15s timeout timer when the captcha resolves before the deadline. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * refactor(auth): use getResponsePromise() for Turnstile token retrieval Replace the manual Promise + refs + timeout pattern with the documented getResponsePromise(timeout) API from @marsidev/react-turnstile. This eliminates captchaToken state, captchaResolveRef, captchaRejectRef, and all callback wiring on the Turnstile component. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(auth): show captcha errors as form-level message, not password error Captcha failures were misleadingly displayed under the password field. Added a dedicated formError state that renders above the submit button, making it clear the issue is with verification, not the password. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
d4a014f423 |
feat(public-api): add env var and permission group controls to disable public API access (#3317)
Add DISABLE_PUBLIC_API / NEXT_PUBLIC_DISABLE_PUBLIC_API environment variables and disablePublicApi permission group config option to allow self-hosted deployments and enterprise admins to globally disable the public API toggle. When disabled: the Access toggle is hidden in the Edit API Info modal, the execute route blocks unauthenticated public access (401), and the public-api PATCH route rejects enabling public API (403). Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
bbcef7ce5c |
feat(access-control): add ALLOWED_INTEGRATIONS env var for self-hosted block restrictions (#3238)
* feat(access-control): add ALLOWED_INTEGRATIONS env var for self-hosted block restrictions * fix(tests): add getAllowedIntegrationsFromEnv mock to agent-handler tests * fix(access-control): add auth to allowlist endpoint, fix loading state race, use accurate error message * fix(access-control): remove auth from allowed-integrations endpoint to match models endpoint pattern * fix(access-control): normalize blockType to lowercase before env allowlist check * fix(access-control): expose merged allowedIntegrations on config to prevent bypass via direct access * consolidate merging of allowed blocks so all callers have it by default * normalize to lower case * added tests * added tests, normalize to lower case * added safety incase userId is missing * fix failing tests |