* feat(sandboxes): workspace dependency sets for Function blocks Named package sets a Function block can import from. The server canonicalizes and hashes the list; E2B prebuilds a content-addressed template per set, Daytona installs per execution. Create/edit is gated to Max or Enterprise via the shared workspace entitlement check; execution is deliberately ungated, so a downgraded workspace keeps running what it already built. Also on this branch: - Extract the duplicated dropdown/combobox option-fetch lifecycle into use-fetched-options. Only combobox had the dependency-change reset, so every dropdown with dependsOn + fetchOptions cleared its list and never repopulated until reopened. - Collapse the repeated Max-tier entitlement check onto one hasMaxTierWorkspaceAccess, shared by inbox, live sync, and sandboxes. - Resolve a personal payer's block state through getEffectiveBillingStatus in getBillingEntityBlockStatus, so the client-side Max gates agree with the server-side ones when blockOrgMembers' fan-out is stale. - Carve the Daytona dependency install out of the caller's execution budget instead of stacking on top of it. Co-Authored-By: Claude <noreply@anthropic.com> * chore(db): regenerate the sandboxes migration as 0273 Staging claimed 0271 and 0272 while this branch was out, so the hand-authored 0271_workspace_sandboxes was dropped before the merge and regenerated on top of the merged schema. Same DDL; drizzle emits plain CREATE TABLE/INDEX rather than the hand-added IF NOT EXISTS, which matches the repo default — that idempotent form is only needed for files with CONCURRENTLY ops below an embedded COMMIT. Regenerating also restores the meta snapshot the hand-authored migration never had. Co-Authored-By: Claude <noreply@anthropic.com> * chore(db): drop the sandboxes migration ahead of the staging merge Staging independently claims idx 0273, so remove ours before merging to avoid an add/add conflict on the drizzle migration index. Regenerated at the next free index once the merge lands. Co-Authored-By: Claude <noreply@anthropic.com> * chore(db): regenerate the sandboxes migration as 0275 Staging took 0273 and 0274, so the sandboxes DDL lands at the next free index. The emitted SQL is byte-identical to the dropped 0273. Co-Authored-By: Claude <noreply@anthropic.com> * fix(billing): consolidate the Max-tier entitlement onto one predicate The Max tier was spelled five ways. The odd one out — `isMax`, defined as `isPro(plan) && credits >= 25000` — excluded both `team_25000` and `enterprise`, and it was the sole input to the personal-workspace cap. A delinquent Max-for-Teams org admin got 1 personal workspace while a delinquent Max individual got 10. Only free/pro_6000/pro_25000 were tested, so the two broken tiers were unpinned. Separately, the server gate and the client `hasUsableMaxAccess` were independent copies of the same rule. The settings sidebar renders Sandboxes and Sim Mailer from the client one while the API answers 403 from the server one, so any drift renders a feature unlocked that the API refuses. - `MAX_TIER_CREDITS` is derived from the `CREDIT_TIERS` table; `isMaxTier` in plan-helpers is now the single definition, shared by the server gates, the client derivation, `getPlanTypeForLimits`, `plan-view`, and the cap - `hasWorkspaceTierAccess(id, predicate, { intent, onMissingWorkspace })` becomes the one org-vs-personal payer fork. `intent: 'active-use'` means active and not billing-blocked; `'retention'` means active/past_due with block state ignored, so the inbox teardown guard keeps its fail-open semantics instead of implying them through a duplicated fork - `isWorkspaceOnEnterprisePlan`'s personal branch now applies the status and block checks its own org branch always had, and its TSDoc names its real consumer (copilot BYOK, not Access Control) - the client live-sync gate gained the server's `isHosted` branch, so a self-hosted deploy with billing on no longer locks an interval the API accepts. It reads both flags directly rather than taking one as a parameter the callers sourced from the same module - `sqlIsPro`/`sqlIsTeam` escape the `_` LIKE wildcard, matching the already correct hand-rolled filter in seat-drift - deletes the `TERMINAL_SUBSCRIPTION_STATUSES` and `ENTITLED_STATUSES` shadow constants, and corrects three test mocks that asserted `trialing` was entitled or usable `max-tier-parity.test.ts` asserts the client and server answers match for every plan name. Both new guards were checked against the old code: the parity test fails 3 assertions with the previous predicate, and the self-hosted test fails without the `isHosted` branch. Co-Authored-By: Claude <noreply@anthropic.com> * chore(db): drop the sandboxes migration ahead of the staging merge Staging has claimed 0275 (table_views) and 0276 (drop_legacy_folder_tables) since the last merge, so our 0275_workspace_sandboxes collides on the index. Dropping ours first — the .sql, meta/0275_snapshot.json, and the journal entry — leaves packages/db/migrations byte-identical to the merge-base, so the merge sees no add/add conflict at all. Regenerated on the far side. Ours is the droppable side: plain additive DDL with no hand edits, which drizzle reproduces exactly. Staging's migrations are hand-written and must survive. Co-Authored-By: Claude <noreply@anthropic.com> * chore(db): regenerate the sandboxes migration as 0277 Staging claimed 0275 (table_views) and 0276 (drop_legacy_folder_tables), so the sandboxes migration dropped before the merge comes back on top as 0277. The emitted SQL is byte-identical to what was dropped — the original had no hand edits, so there is nothing to reapply. It is purely additive: two enums, sandbox_image and workspace_sandbox, their two FKs and six indexes. That it regenerated unchanged also confirms the schema.ts auto-merge was correct — had it lost staging's legacy-folder-table drops, drizzle would have emitted CREATE TABLE for them here. Snapshot chain is continuous (0273 -> 0277, each prevId matching the previous id) and the table counts track the DDL: 100 -> 101 (table_views) -> 99 (legacy folder tables dropped) -> 101 (the two sandbox tables). Co-Authored-By: Claude <noreply@anthropic.com> * feat(sandboxes): gate on the enterprise feature flags, drop the rollout switch Sandboxes shipped behind `custom-sandboxes`, an AppConfig rollout flag falling back to a `CUSTOM_SANDBOXES` secret. That made it the only Max-gated surface with no self-hosted path: `INBOX_ENABLED` can force Sim Mailer on for an operator running their own billing, and `ENTERPRISE_ENABLED` turns on the other nine features at once, but neither reached sandboxes. A self-hoster had to find a separately-named variable that was not part of that family, and one running with billing enabled could not enable it at all. Sandboxes now joins the enterprise feature set and the rollout flag is gone: - `sandboxes` is an `EnterpriseFeature` with `SANDBOXES_ENABLED` and its `NEXT_PUBLIC_` twin, so the master switch and the per-feature override both reach it like every sibling - `hasWorkspaceSandboxAccess` takes the inbox's shape exactly — the override wins, then a deployment without billing is unrestricted, then the workspace payer needs usable Max or Enterprise - the settings nav gains `selfHostedOverride`, so the section resolves through the same path as Sim Mailer instead of a second entitlement AND-ed in - `custom-sandboxes`, the `CUSTOM_SANDBOXES` secret, the now-unreachable `SANDBOXES_UNAVAILABLE` 403 copy, and the route's kill-switch branch are deleted Its legacy default is `true`, matching `inbox`: the gate already returns true whenever billing is off, so `false` would leave the nav override disagreeing with the gate that answers the request. Self-hosted builds run on the operator's own E2B/Daytona credentials, so there is no Sim-side cost to withhold — the docs now say so, since enabling the feature without a provider configured is the obvious trap. The new gate tests run with billing enabled on purpose; the `!isBillingEnabled` bail would otherwise answer every case and hide whether the override is wired. Verified by deleting the override line — exactly the one assertion fails. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): let the language menu match its trigger width `matchTriggerWidth={false}` exists for the opposite case — a narrow trigger whose option labels would truncate, letting the menu grow past it. The language field is a full-width form control with two short labels, so the override shrank the menu to "JavaScript" and pinned it to the right edge instead. The default (`true`) is correct here. Every other consumer passing `false` is a genuinely narrow trigger — a role picker in a member row, a table filter chip. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): re-queue a build when resolution finds the image unusable `ensureSandboxImage` only ran when a sandbox was saved, so resolution treated an unusable image as terminal and told the user to go fix a definition that was never wrong. Three states stuck permanently until someone re-saved in Settings: - a build that failed - a build whose worker died mid-flight, stranding the row in `building` - every sandbox created while the deployment ran a `runtime` provider, after a switch to a `prebuilt` one — `runtime` writes no image rows at all, so the whole fleet resolved to "no completed build" with nothing to repair it Resolution now re-queues through the registry's existing idempotent entry point before failing, and says a build is on its way instead of pointing at Settings. The conflict guard already claims only a `failed` row or a stale `pending`/ `building` one, so executions arriving during a healthy build enqueue nothing — no thundering herd from a hot workflow. The registry is imported dynamically for the same reason `sandboxDb` is: it pulls `@sim/db` into the static graph, which this module keeps out of the executor bundle. That also avoids a cycle, since the registry imports `invalidateSandboxResolution` from here. A repair that itself fails is logged and swallowed — it must never replace the build error naming the sandbox. Verified by deleting the repair call: exactly the three new assertions fail. Co-Authored-By: Claude <noreply@anthropic.com> * improvement(sandboxes): let the picker show just the sandbox name The label read "Test · Python · 1 package". The block's own list is already scoped to the language its sibling `language` subblock selects, so the language repeated on every row said nothing, and the package count is decoration next to the name that identifies the sandbox. The language stays for the one caller that cannot filter — agent tool-input renders this field under a synthetic id where the sibling `language` value is unreachable, so its list spans both languages and the name alone is ambiguous. That is the same missing value which disables filtering, so `showLanguage` is derived from it directly rather than passed independently and left to drift. A failed build is still marked: that suffix is the difference between a selection that runs and one that does not. Passing the flag also means dropping `.map(toSandboxOption)` for an explicit arrow — `Array.map` hands the index to the second parameter. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): show the sandbox name on the block card, not its uuid The card printed "443f4934-26ab-44ab-8...". `resolveDropdownLabel` only reads a subblock's static `options` array, and the sandbox picker is a `combobox` whose options load asynchronously, so its array is empty and the raw stored id fell through to the label. Resolved the same way skills and tools already are: a `resolveSandboxLabel` in the display layer, fed from the shared sandbox list query — the same cache entry the picker reads, so this adds no request. Two deliberate scopings: - the query is subscribed only for the sandbox row. `SubBlockRow` is memoized per subblock, and the list query polls while a build is in flight, so an unconditional hook would re-render every row on the canvas on each poll tick - the resolver matches the field id, not just the type. There is no dedicated subblock type for it, and matching `combobox` alone would relabel unrelated pickers An id with no matching sandbox resolves to null rather than a guess, so a deleted sandbox falls through to the caller's placeholder. The template preview surface is left alone: it is explicitly hook-free and passes empty lists for tools and skills too. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): hide the Sandboxes section with no provider configured Entitlement decides whether a workspace may author sandboxes; nothing decided whether anything could run one. A self-hosted deployment with SANDBOXES_ENABLED but no E2B or Daytona credentials got a fully functional tab whose output no Function block could select — the picker is gated on the provider vars, the tab was not. Both navigation planes now drop the section when neither NEXT_PUBLIC_SANDBOX_ENABLED nor the pre-Daytona NEXT_PUBLIC_E2B_ENABLED is set — the same pair the picker's `showWhenEnvSet` reads, so the two cannot disagree. Dropped rather than locked: an upgrade does not conjure a provider. The unified plane drops it in `buildUnifiedSettingsNavigation` rather than in the sidebar's filter, because the sidebar's `selfHostedOverride` short-circuit runs before its `requiresMax` check and would have revealed the tab anyway. It reads the browser twins, not the server's `isRemoteSandboxEnabled`, since this module renders on both sides. The predicate is a function, not a module constant, because the constant form was untestable and ambient: the env mock falls through to `process.env`, and `apps/sim/.env` (gitignored, so absent on CI) sets NEXT_PUBLIC_E2B_ENABLED=true. The nav tests passed locally and failed 6 assertions with the flag cleared. They now pin both flags, so the suite is identical with and without a local env file — verified by running it both ways. Co-Authored-By: Claude <noreply@anthropic.com> * docs(sandboxes): correct three claims the code no longer makes The Sandboxes section described behavior two commits on this branch changed, and led with an internal detail no reader needs. - entitlement is no longer Max/Enterprise only: self-hosted deployments unlock sandboxes with SANDBOXES_ENABLED, and the section is hidden outright when a deployment has no sandbox provider, which is the state a self-hoster is most likely to hit and least likely to diagnose - a build that is not Ready is no longer terminal. It is queued again on the next run, so the advice is to wait and re-run, not to go edit a package list that was never wrong - deleting a sandbox frees its build once nothing else references it. Builds are shared by content, so this is the one place a reader could reasonably assume deletion is immediate Dropped the `ModuleNotFoundError` aside: what the old code did instead is not something a reader needs to know to use the feature. The page is hand-written — `function` has category 'blocks' and is absent from `NATIVE_RESOURCE_BLOCK_TYPES`, so generate-docs skips it and these edits will not be overwritten. Co-Authored-By: Claude <noreply@anthropic.com> * feat(sandboxes): release the provider image when nothing references it Deleting a sandbox only removed its row, leaving the built template in E2B until the 30-day retention sweep — up to a month of paying to store an image nothing could select. Editing a package list had the same effect on the old content address, which is the more common case since every edit re-points the sandbox. `releaseSandboxImage(specHash)` now deletes the provider image and its row from both paths. It reuses the sweep's provider call and its ordering: image first, row second, so a refused delete leaves the row for the sweep to retry rather than orphaning a remote template nothing points at. Two guards make eager deletion safe: - builds are keyed by content, not by workspace, so two workspaces declaring the same package list share one image. The release no-ops while any sandbox still references the hash — otherwise one workspace's delete would break the other's - an in-flight build is left alone rather than raced; the sweep collects it once it settles Called detached from both routes. The row is already committed by then, so the user's action has succeeded whatever the provider says, and awaiting would hold a UI delete open on a remote call the sweep would retry anyway. Every failure inside is logged and swallowed for the same reason. E2B's delete verified against their API reference: DELETE /templates/{templateID} with X-API-Key, 204 on success. The existing implementation already matched, so this commit only adds the call sites and the guards. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): rate-limit the automatic rebuild, drop the one-off status dot Two follow-ups to the resolution repair. The repair had no rate limit. `ensureSandboxImage` re-claims a `failed` row on sight, and a bad package name fails in seconds, so the in-flight guard never closed the window: a workflow on a one-minute schedule would enqueue a build a minute against a package list that will never resolve, each one real provider build compute. Before the repair existed resolution simply threw, so this was introduced with it. The two callers want different things, so the cooldown is opt-in. A save is a person explicitly asking for another attempt and still retries immediately; resolution passes `FAILED_BUILD_RETRY_COOLDOWN_MS` and gets at most one attempt per window no matter how often the workflow runs. Ten minutes: long enough that per-minute runs cannot drive per-minute builds, short enough that a transient registry outage clears within the hour. The status line loses its colour dot. `size-[6px] rounded-full` appeared in exactly one file in the repo, so it was a new primitive rather than a pattern, and it duplicated state the text colour already carries — the label now turns `--text-error` on a failed build, which is what every other status row in settings does. `ChipTag` was the wrong home for this: its variants are `mono`/`invite`, with no semantic tone, so a status version would have meant overriding its chrome from the consumer. Also corrects the docs line this changes: a failed build is retried periodically, and saving is the way to retry now, so "wait a moment and run again" no longer describes it. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): claim the image row and its reference check in one statement Greptile P1. Reading references in one statement and deleting in another left a window — a wide one, since a provider delete is a network call — where a second workspace could declare the same package list, inherit the `ready` row, and have its next run fail against a template already on its way out. Content addressing is what makes that reachable: the image is shared, so one workspace's delete can strand another's sandbox. The reference check now lives in the conditional DELETE itself, so winning the delete is the proof that nothing referenced the hash. A workspace that adopts the hash first makes the delete match nothing and the release becomes a no-op. Claiming the row before the provider call would otherwise strand a template nothing points at if the provider then refused, so that path puts the row back and the retention sweep inherits the retry — the same property the previous ordering had. The sweep is deliberately left as it is: its equivalent window needs a hash unreferenced AND unused for 30 days, and its provider-first ordering encodes the documented retry-on-refusal behaviour this path now reproduces explicitly. No transaction is opened. The provider call sits between discrete statements rather than inside one, so no pooled connection is held across it — which is why this uses a conditional delete instead of the repo's `pg_advisory_xact_lock` pattern, whose lock only releases at commit. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): route the retention sweep through the same image claim Cursor and Greptile both flagged the sweep as still carrying the interleaving just fixed in releaseSandboxImage, and they are right — the reason given for leaving it alone last round does not survive scrutiny. That reason was that provider-first ordering encodes retry-on-refusal, so making the claim atomic would trade a race for an orphaned template. The release path already answers that: claim the row, and put it back if the provider refuses. The sweep can have both properties too. The rarity argument was also weaker than stated. The sweep nominates up to 200 candidates and then works through them eight network deletes at a time, so its check-to-delete gap is seconds to minutes — wider than the window that was just closed, not narrower. Both callers now share `claimAndDeleteImage`, which owns the whole contract: the unreferenced check lives inside the DELETE, the provider call runs only after the claim succeeds, and a refusal restores the row. Having written that ordering twice is what let the two paths drift, so it exists once now. The sweep's query becomes a nomination step only. Its retention cutoff is passed into the claim rather than trusted from the earlier read, so a candidate that stops qualifying mid-sweep fails its claim and is skipped instead of losing its image. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): rebuild a hash adopted while its image was being deleted Greptile's third pass on this path, and a case the previous two did not cover: the adopter starting a *fresh build* rather than inheriting a ready row. Claiming removes the registry row, so between that and the provider delete finishing, a workspace can declare the same package list, get a new row, and start a build under the same content-derived imageRef — which the in-flight delete then removes. The window itself is inherent. The registry row and the provider template are two systems with no shared transaction, so it can be narrowed but not closed. A Redis lock would not close it either: acquireLock returns true when Redis is absent, so it cannot be a correctness guarantee for self-hosted. Holding a Postgres advisory lock would, but only by pinning a pooled connection for the length of a provider call, which is a worse trade. What was avoidable is the adopter finding out the slow way. Its row is new and healthy-looking, so nothing noticed: resolution only repairs a row that is missing or failed, and a failed one waits out the retry cooldown first. The release path now re-checks after the delete and re-enqueues, so the rebuild starts immediately instead of one failed run plus a cooldown later. A build already in flight is left to the conflict guard, since it may still outlive the delete. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): reclaim a ready row whose image was deleted underneath it Greptile found the hole the previous commit left, and it is the case that made the claim in that commit's message wrong: this one is permanent, not transient. If a re-adopted hash reaches `ready` before the in-flight provider delete lands — plausible, since E2B layer caching can rebuild an identical spec in seconds — the row looks healthy while its imageRef points at nothing. Resolution repairs a row that is missing or failed, never one claiming to be ready, so nothing recovers it. The sandbox stays broken until someone re-saves it by hand. `rebuildIfReadopted` called `ensureSandboxImage` with no options, whose conflict guard reclaims only a failed or stale in-flight row, so it silently did nothing in exactly that case. The release path now passes `imageKnownGone`, which widens the re-claim to any settled row rather than only a failed one. It is the one caller that knows the image is gone regardless of what the row says. An in-flight build is still left alone: it either recreates the template it was building or fails into the normal repair path, and resetting it would only add a duplicate build. The three ways a settled row may be re-claimed now sit in one `settledRebuildBranch` helper — any settled row when the image is known gone, a failed one after the cooldown for an automatic caller, a failed one immediately for a person — because inlining the third case is what hid the gap. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): let a same-spec save retry a failed build Cursor Bugbot. `scheduleSandboxBuild` sat inside the changed-hash branch, so a save that did not alter the package list never reached the registry. The comment above it described the opposite — that an unchanged spec finds a ready row and enqueues nothing — which is what `ensureSandboxImage` does, but only if it is called. That made the docs wrong too. They tell a reader to save the sandbox again to retry a failed build immediately, and this branch is exactly why that did nothing: the only way to retry was to edit the package list into a different hash, which is not what someone recovering from a transient registry failure wants to do. The call is now unconditional and the registry decides what a save costs, which is what its conflict guard is for: a ready or in-flight row is left alone, a failed one is re-claimed at once. Releasing the previous image stays behind the hash check, since only a changed hash orphans one. Cache invalidation is unchanged — `scheduleSandboxBuild` already does it, which is why the else branch existed. Co-Authored-By: Claude <noreply@anthropic.com> * docs(sandboxes): correct the image cache's staleness invariant Cursor Bugbot found that a released image can still be served from another replica's cache. The finding is real, and the reason it went unnoticed is that the cache documented an invariant which eager release quietly broke. It claimed a `ready` row is terminal for its spec hash, so a cached hit could not go stale in a way that matters. That held while the only ways a row changed were an edit (new hash) or a delete (caught by the `workspace_sandbox` read). Releasing an image eagerly made a `ready` row disappear with the hash unchanged, so the premise no longer holds and the comment was actively misleading to the next reader. No behaviour change here — the exposure is bounded at IMAGE_TTL_MS on replicas other than the one that ran the release, and it self-heals once the entry expires and the row read finds nothing. Closing it properly needs cross-replica invalidation or a provider-error path that invalidates on "template not found", both of which are larger than a review fix; the comment now says so instead of implying the problem cannot exist. Co-Authored-By: Claude <noreply@anthropic.com> * docs(sandboxes): note that a JavaScript sandbox needs an import to apply Cursor Bugbot pointed out that `useRemoteSandbox` keys on detected static import/require and never on the selected sandbox, so JavaScript without one runs locally and the selection has no effect. Keeping the behaviour: honouring the selection would force those blocks remote, and the large-value-ref guard immediately below would then reject code that runs fine today. Documenting it instead, next to the picker, since a selection that silently does nothing is only surprising if nothing says so. Python is unaffected — it always runs remotely, so its sandbox always applies. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): stop create mode surviving a return to an open sandbox Cursor Bugbot. Create mode and having a sandbox open are mutually exclusive, but nothing enforced it, so both could be set at once — and the screen then lied about which sandbox its Delete pointed at. With `isCreating` true and `selectedId` restored, `baseline` is null, so the editor renders an empty "New sandbox" form, while the Delete action is built from `selected` and still targets the restored sandbox. An admin looking at a blank create form could delete a sandbox it never named. Two ways in, both closed: - Browser Forward after starting a new sandbox restores `selectedId` without going through `closeEditor`. The render-time sync that already drops a stale draft now also leaves create mode, which is the same class of correction and the reason that block exists. - "New sandbox" set `isCreating` without clearing `selectedId`, so the same contradiction was reachable without touching history at all. It now clears the selection, with `history: 'replace'` because switching mode is not a destination. Co-Authored-By: Claude <noreply@anthropic.com> * fix(ci): pin the sandbox flag in the second nav catalog test, bump the chart Two CI failures, both mine. `app/workspace/[workspaceId]/settings/navigation.test.ts` asserts the unified catalog and was left on ambient env. Dropping the Sandboxes section without a sandbox provider made it 26 items instead of 27 on CI, which has no `apps/sim/.env` — the same trap already fixed in the sibling `components/settings/navigation.test.ts`, in the one file that was missed. Fixing it needs `vi.hoisted` rather than the sibling's `beforeEach`, because this file reads `allNavigationItems`, built once at module load; a hook would run after the value it is trying to influence already exists. The chart gate is separate: this branch adds sandbox settings to `helm/sim/values.yaml`, and the workflow requires a Chart.yaml bump whenever `helm/sim/**` changes. Additive config, so 1.3.0 -> 1.4.0 by SemVer. Verified by running the whole suite with the flags forced off, not just the two navigation files — no other test depends on a local env file. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): keep the row restore to a refused delete only Cursor and Greptile, independently, on the same code. `deleteImage` and `rebuildIfReadopted` shared one try/catch, so a rebuild failure after a *successful* provider delete was handled as if the provider had refused: the catch put the claimed row back, `ready` status and all, pointing at a template that no longer exists. That is the one state resolution cannot repair — it fixes a row that is missing or failed, never one claiming to be ready — so it reintroduced the permanent breakage an earlier commit had just closed, through the error path rather than the happy one. Restoring now belongs strictly to a refused delete. Once the template is gone the row stays gone, and the rebuild runs past that catch. The rebuild also swallows its own failures: it follows a delete that already succeeded, so it must not be reported as a failed release, and inside the sweep it must not reject the rest of its chunk. The adopter's next run still reaches the normal repair path. The regression test drives a rebuild failure and asserts no row is restored. It fails against the original shape — rebuild inside the shared try, no inner catch — which is what the two reviewers were describing. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): drop the dead row when a re-adopt rebuild cannot be scheduled Greptile, one layer under the previous fix. Making the post-delete rebuild swallow its own failures kept it from being reported as a failed release, but left the adopter's row claiming a `ready` image whose template is already deleted — the one state resolution cannot repair, since it rebuilds a row that is missing or failed and never one that says ready. So the row is now dropped when the rebuild does not take. That turns the adopter into the missing-row case, which the next execution repairs on its own, instead of a sandbox that stays broken until someone re-saves it by hand. A failure to drop it is logged at error, because at that point two writes in a row have failed and there is nothing further this path can do. Also gives the release tests a default "nothing re-adopted" select. Without it the rebuild threw on an unstubbed mock and the cleanup delete overwrote the predicate the claim assertions read, so two of them were passing on the wrong statement. Co-Authored-By: Claude <noreply@anthropic.com> * feat(sandboxes): repair a missing image at create, where the truth is observable Six review rounds narrowed the window between deleting a shared template and another workspace adopting its content hash, and each fix exposed the next facet. They all share a cause: the registry row and the provider template are two systems with no shared transaction, so any scheme that keeps them in step is guessing. Create is the one step that does not have to guess. It either gets a sandbox or it does not, so a `ready` row pointing at a deleted template now corrects itself the first time it is used, rather than needing someone to re-save the sandbox. - `SandboxImageBuilder.isMissingImage` asks the provider to classify its own failure. Prebuilt-only, because a runtime provider has no image to miss - E2B answers it off `NotFoundError`, which the SDK maps from a 404. The only resource a create names is the template, and the two subclasses that describe other calls — a missing file, an exited sandbox — are excluded. The classifier stays deliberately narrow: treating auth or rate-limit failures as a missing image would turn a provider outage into a build storm - `repairMissingSandboxImage` invalidates the cache, rebuilds with `imageKnownGone` (no cooldown, since this observed the image is gone rather than inferring it), and returns copy telling the author to run again - `ResolvedSandbox` carries `specHash` so the failing execution can name what to rebuild This subsumes the open facets rather than adding another guard beside them: the stale per-replica cache, an adopter left `ready` against a deleted ref, and a rebuild that never took all end at the same place — the next run repairs itself. Co-Authored-By: Claude <noreply@anthropic.com> * fix(sandboxes): key the build trigger by attempt, not by spec Cursor Bugbot. The Trigger.dev idempotency key was the content address alone, so a second attempt at the same spec was deduped against the first: the SDK returns the finished run instead of starting one, and the row that `ensureSandboxImage` just flipped to `pending` sits there with no worker. Nothing can re-claim a `pending` row until it goes stale, so a retry inside the 5-minute TTL did nothing for the next half hour. That silently disabled every repair path — save-to-retry, which the docs name explicitly, and both the resolution and create-time rebuilds. The key's own comment already said it exists "to collapse concurrent saves of the same spec into one build, not to suppress a retry after one failed". The conditional update above it is what actually collapses concurrent saves: only one caller gets a row back, so only one ever reaches the trigger. Keying by the claim's `updatedAt` keeps that property and makes each genuine attempt distinct, while a duplicate delivery of one attempt still collapses. Co-Authored-By: Claude <noreply@anthropic.com> * feat(sandboxes): create a sandbox from the picker, and fix three UI papercuts The Function block's sandbox field now pins a "Create Sandbox" row above its options, matching the "Create Skill" / "Create Tool" rows it sits beside, so authoring a package list no longer means leaving the workflow for Settings. The row is declared by the field (`createAction`) rather than hardcoded by id; block configs are read by the serializer and executor, so the name maps to a modal in the picker rather than carrying a component. Two things the modal has to get right. It seeds the new sandbox's language from the sibling the list is scoped by, or a sandbox created off a JavaScript block would land in the Python list and vanish. And the created option is held locally until a real fetch carries it, or the field would sit on a raw uuid until hydration answered. Also: - The Sandboxes icon was the Logs block's icon (`blocks/blocks/logs.ts`), in both the settings nav and the list rows. It is the Function block's now. - "Default image (no extra packages)" claimed something untrue: E2B and Daytona base images both ship with packages installed. - A new sandbox opened in Python while the Function block defaults to JavaScript. The test pins the two together rather than the literal. Draft shape and helpers moved out of the editor component into `utils.ts` — three consumers now, and it makes the defaults testable without a DOM. * feat(settings): one Max-plan wall, and give the create modal the same one The create-sandbox modal answered a non-Max workspace with a red line under a form it could never submit, and no way to act on it. It now renders the same wall the Settings > Sandboxes tab does — heading, one sentence on what the plan unlocks, and an Upgrade to Max chip — instead of the fields. That wall existed twice already (sandboxes and Sim Mailer), so this extracts it rather than adding a third copy. `SettingsUpgradeNotice` owns the copy rhythm and the route, and `compact` trades the page's full-height centering for a modal's. Both settings consumers now compose it; neither keeps its own markup. The action lands on billing, which `resolveSettingsHref` already redirects to the plan-comparison page for a member who cannot manage billing — so it is a route to explore plans, never a dead end. The chip stays hidden for non-admins, exactly as the settings pages had it. A non-admin on an entitled workspace gets the muted reason rather than the upgrade wall: buying a plan is not what is in their way. --------- Co-authored-by: Claude <noreply@anthropic.com>
Sim Helm Chart
Deploy Sim — the open-source AI workspace where teams build, deploy, and manage AI agents — on Kubernetes.
- Chart version: see
Chart.yaml - App version: tracks the upstream Sim release
- Kubernetes: 1.25+
- License: Apache-2.0
TL;DR
# Generate required secrets
export BETTER_AUTH_SECRET=$(openssl rand -hex 32)
export ENCRYPTION_KEY=$(openssl rand -hex 32)
export INTERNAL_API_SECRET=$(openssl rand -hex 32)
export CRON_SECRET=$(openssl rand -hex 32)
export POSTGRES_PASSWORD=$(openssl rand -base64 24 | tr -d '/+=')
# Install from this repository
helm install sim ./helm/sim \
--namespace sim --create-namespace \
--set app.env.BETTER_AUTH_SECRET="$BETTER_AUTH_SECRET" \
--set app.env.ENCRYPTION_KEY="$ENCRYPTION_KEY" \
--set app.env.INTERNAL_API_SECRET="$INTERNAL_API_SECRET" \
--set app.env.CRON_SECRET="$CRON_SECRET" \
--set postgresql.auth.password="$POSTGRES_PASSWORD"
After install, follow the on-screen NOTES.txt to reach the app.
Introduction
This chart deploys the Sim platform on a Kubernetes cluster using the Helm package manager. A default install includes:
app— the Sim Next.js web application (Deployment).realtime— the WebSocket service for live workflow updates (Deployment).postgresql— an in-clusterpgvector/pgvectorPostgres (StatefulSet, with a headless Service for stable per-pod DNS).migrations— an init container on the app Deployment that applies database migrations before each app pod starts.cronjobs— scheduled jobs for workflow schedule execution, inbox/calendar/drive polling (Gmail, Outlook, Calendar, Drive, Sheets, IMAP, RSS), workspace event and HubSpot webhook polling, outbox processing, subscription renewal, billing-seat and inbox-entitlement reconciliation, time-pause/resume polling, data drains, and connector syncs.serviceaccount— a dedicated ServiceAccount withautomountServiceAccountToken: false.
Optional components (off by default):
copilot— the Sim Copilot service plus its own Postgres StatefulSet.ollama— local LLM inference, with optional NVIDIA GPU support.pii— Presidio PII redaction service (analyzer + anonymizer) for the Guardrails PII block and log redaction. See PII redaction.telemetry— OpenTelemetry Collector wired to Jaeger / Prometheus / OTLP backends.ingress— NGINX-style Ingress for the app and realtime services.networkPolicy— east-west and egress isolation (blocks cloud metadata endpoints by default).hpa— HorizontalPodAutoscaler forappandrealtime.podDisruptionBudget— auto-activates whenreplicaCount > 1.servicemonitor— Prometheus Operator integration.
Prerequisites
| Requirement | Version / Notes |
|---|---|
| Kubernetes | 1.25+ (Chart.yaml enforces kubeVersion: ">=1.25.0-0") |
| Helm | 3.8+ |
| StorageClass | A default StorageClass that supports ReadWriteOnce PVCs (for Postgres, Ollama). Set global.storageClass to pick a non-default class. |
| Ingress controller | Only if ingress.enabled=true. The chart's defaults assume nginx. |
| cert-manager | Only if you want auto-issued TLS certificates. See cert-manager docs. |
| metrics-server | Only if autoscaling.enabled=true (HPA needs metrics). |
| External Secrets Operator | Only if externalSecrets.enabled=true. See ESO docs. |
| Prometheus Operator | Only if monitoring.serviceMonitor.enabled=true. |
| Namespace PSS labels | Recommended: pod-security.kubernetes.io/enforce=restricted. The chart's pod and container security contexts are PSS-restricted by default. |
Generate required secrets
Sim will not start without these. Generate them once and feed them via --set, an existing Kubernetes Secret, or External Secrets Operator.
# Application secrets (32 bytes hex each)
openssl rand -hex 32 # BETTER_AUTH_SECRET - signs auth JWTs
openssl rand -hex 32 # ENCRYPTION_KEY - encrypts sensitive env vars
openssl rand -hex 32 # INTERNAL_API_SECRET - service-to-service auth
openssl rand -hex 32 # CRON_SECRET - required if cronjobs.enabled (default true)
openssl rand -hex 32 # API_ENCRYPTION_KEY - optional; encrypts user API keys at rest
# Postgres password
openssl rand -base64 24 | tr -d '/+='
If you set app.secrets.existingSecret.enabled=true and point at a pre-created Secret, you do not also pass these via --set — pick one path.
Installing the chart
From this repository
helm install sim ./helm/sim \
--namespace sim --create-namespace \
--set app.env.BETTER_AUTH_SECRET="$BETTER_AUTH_SECRET" \
--set app.env.ENCRYPTION_KEY="$ENCRYPTION_KEY" \
--set app.env.INTERNAL_API_SECRET="$INTERNAL_API_SECRET" \
--set app.env.CRON_SECRET="$CRON_SECRET" \
--set postgresql.auth.password="$POSTGRES_PASSWORD"
With a values file
helm install sim ./helm/sim \
--namespace sim --create-namespace \
--values my-values.yaml
Run helm template ./helm/sim --values my-values.yaml | less first to see what will be applied.
Validate the install
helm install sim ./helm/sim --dry-run --debug \
--values my-values.yaml \
--set app.env.BETTER_AUTH_SECRET=$(openssl rand -hex 16) \
--set app.env.ENCRYPTION_KEY=$(openssl rand -hex 16) \
--set app.env.INTERNAL_API_SECRET=$(openssl rand -hex 16) \
--set app.env.CRON_SECRET=$(openssl rand -hex 16) \
--set postgresql.auth.password=$(openssl rand -base64 12 | tr -d '/+=')
Upgrading
helm upgrade sim ./helm/sim --namespace sim --values my-values.yaml
Uninstalling
helm uninstall sim --namespace sim
PVCs are not deleted by helm uninstall. If you want to wipe data too:
# WARNING: this destroys all Postgres, Ollama, and shared-storage data.
kubectl delete pvc --namespace sim \
-l app.kubernetes.io/instance=sim
# Or list and delete by name
kubectl get pvc --namespace sim
kubectl delete pvc <pvc-name> --namespace sim
# Then delete the namespace if you're done with it
kubectl delete namespace sim
Examples
Pre-built values files for common scenarios live in helm/sim/examples/. Each file has a header explaining when to use it and any prerequisites.
| File | When to use |
|---|---|
values-development.yaml |
Local dev / kind / minikube. Minimal resources, no TLS. |
values-production.yaml |
Generic production: HA, network policy, autoscaling, monitoring. |
values-aws.yaml |
EKS — EBS GP3 storage, ALB ingress, IRSA-friendly. |
values-gcp.yaml |
GKE — Persistent Disk storage, GCP managed certs, Workload Identity. |
values-azure.yaml |
AKS — managed-csi storage, NGINX ingress, GPU node pools. |
values-external-db.yaml |
Production with a managed Postgres (RDS, Cloud SQL, Azure DB). |
values-external-secrets.yaml |
Sync secrets from Vault / AWS SM / Azure KV / GCP SM via External Secrets Operator. |
values-existing-secret.yaml |
GitOps / Sealed Secrets / SOPS — reference pre-created Kubernetes Secrets. |
values-copilot.yaml |
Enables the Copilot service + its Postgres StatefulSet. |
values-whitelabeled.yaml |
Custom branding (logo, name, support links). |
Use one with:
helm install sim ./helm/sim \
--namespace sim --create-namespace \
--values ./helm/sim/examples/values-production.yaml \
--set app.env.BETTER_AUTH_SECRET="$BETTER_AUTH_SECRET" \
--set app.env.ENCRYPTION_KEY="$ENCRYPTION_KEY" \
--set app.env.INTERNAL_API_SECRET="$INTERNAL_API_SECRET" \
--set postgresql.auth.password="$POSTGRES_PASSWORD"
Parameters
This chart is intentionally configurable. Rather than maintain a hand-curated parameter table (which would drift), read the canonical sources:
# Print all values with comments and defaults
helm show values ./helm/sim
# Print the JSON Schema (used by `helm install` to validate your values)
cat ./helm/sim/values.schema.json
values.yaml is heavily commented; each top-level section explains what it controls and which sub-keys are required vs optional. For per-cloud examples and idiomatic overrides, see examples/.
Production checklist
Before installing in production, confirm each of the following:
- High availability — scale
app.replicaCount > 1. The chart auto-creates aPodDisruptionBudgetwithmaxUnavailable: "25%". SetpodDisruptionBudget.minAvailableinstead for a stricter policy. - Pinned images — override
image.tag(orimage.digest) with an explicit version. Do not rely on the chart's default tag in production. - Secrets management — provide secrets via External Secrets Operator (ESO) or pre-created Kubernetes Secrets. Never commit secrets to
values.yaml. - TLS / Ingress — set the
cert-manager.io/cluster-issuerannotation on the ingress and tuneproxy-body-size/proxy-read-timeoutfor your workload. See commented examples invalues.yaml. - Network policy egress — review
networkPolicy.egressExceptCidrs. Defaults block cloud metadata endpoints (169.254.169.254/32,169.254.170.2/32); add your cluster's API server CIDR for stronger isolation. Custom egress rules go innetworkPolicy.egress(a list). - Network policy ingress —
networkPolicy.ingressFromdefaults to[{}](an empty peer selector), which allows ingress traffic from any pod in the cluster, not just your ingress controller. This is a deliberate simple default, not a locked-down one. On a shared or multi-tenant cluster, scope it down, e.g. to the ingress-nginx namespace:networkPolicy: ingressFrom: - namespaceSelector: matchLabels: kubernetes.io/metadata.name: ingress-nginx - Namespace hardening — label the install namespace with Pod Security Standards
restrictedenforcement (pod-security.kubernetes.io/enforce=restricted). All workloads setrunAsNonRoot, drop all Linux capabilities, disable privilege escalation, and setseccompProfile: RuntimeDefault— the four controls the Restricted profile requires.readOnlyRootFilesystemis intentionally not defaulted anywhere (Postgres/Ollama genuinely need a writable root; the stateless services —realtime,pii,copilot— could tolerate it but aren't pre-wired with a/tmpemptyDir). If your policy requires it, set<component>.securityContext.readOnlyRootFilesystem: trueand mount anemptyDirat/tmpyourself viaextraVolumes/extraVolumeMounts. - Env validation — keys under
app.env,realtime.env, andcopilot.envare passed through to the application and validated at startup. The JSON Schema intentionally does not enforceadditionalProperties: false(would break custom user envs), so typos likeOPENA_API_KEY(instead ofOPENAI_API_KEY) surface as missing-key errors at runtime, not athelm installtime. Review your env block carefully. - Set public URLs —
app.env.NEXT_PUBLIC_APP_URLandapp.env.BETTER_AUTH_URLmust match your public origin (e.g.https://sim.example.com). Leaving them aslocalhostbreaks sign-in.
Secrets
The chart supports three ways to provide secrets, in increasing order of production-readiness:
1. Inline --set (dev / dry-run only)
helm install sim ./helm/sim --set app.env.BETTER_AUTH_SECRET=...
Discouraged for production — values land in helm get values output.
2. Pre-existing Kubernetes Secret
Create the Secret first, then reference it:
kubectl create secret generic sim-app-secrets --namespace sim \
--from-literal=BETTER_AUTH_SECRET=$(openssl rand -hex 32) \
--from-literal=ENCRYPTION_KEY=$(openssl rand -hex 32) \
--from-literal=INTERNAL_API_SECRET=$(openssl rand -hex 32) \
--from-literal=CRON_SECRET=$(openssl rand -hex 32)
kubectl create secret generic sim-postgres-secret --namespace sim \
--from-literal=POSTGRES_PASSWORD=$(openssl rand -base64 24 | tr -d '/+=')
app:
secrets:
existingSecret:
enabled: true
name: sim-app-secrets
postgresql:
auth:
existingSecret:
enabled: true
name: sim-postgres-secret # must contain the password under the key POSTGRES_PASSWORD
See examples/values-existing-secret.yaml.
3. External Secrets Operator (recommended)
Sync from Azure Key Vault, AWS Secrets Manager, HashiCorp Vault, or GCP Secret Manager. Install ESO once, create a ClusterSecretStore, then:
externalSecrets:
enabled: true
refreshInterval: 1h
secretStoreRef:
name: my-secret-store
kind: ClusterSecretStore
remoteRefs:
app:
BETTER_AUTH_SECRET: sim/app/better-auth-secret
ENCRYPTION_KEY: sim/app/encryption-key
INTERNAL_API_SECRET: sim/app/internal-api-secret
postgresql:
password: sim/postgresql/password
# Only needed when copilot.enabled=true and copilot.server.secret.create=true.
# Every non-empty copilot.server.env key must have a matching entry here —
# template rendering fails with a clear message naming the missing key otherwise.
copilot:
AGENT_API_DB_ENCRYPTION_KEY: sim/copilot/agent-api-db-encryption-key
INTERNAL_API_SECRET: sim/copilot/internal-api-secret
LICENSE_KEY: sim/copilot/license-key
SIM_BASE_URL: sim/copilot/sim-base-url
SIM_AGENT_API_KEY: sim/copilot/sim-agent-api-key
REDIS_URL: sim/copilot/redis-url
OPENAI_API_KEY_1: sim/copilot/openai-api-key
See examples/values-external-secrets.yaml.
Persistence
Postgres, Ollama, and any configured sharedStorage.volumes[] use PersistentVolumeClaims. PVCs survive helm uninstall — see Uninstalling for full cleanup.
| Component | Default size | Access mode | Storage class |
|---|---|---|---|
postgresql |
10Gi | ReadWriteOnce |
global.storageClass |
copilot.postgresql |
10Gi | ReadWriteOnce |
global.storageClass |
ollama |
100Gi | ReadWriteOnce |
global.storageClass |
sharedStorage.volumes[] |
user-defined | ReadWriteMany recommended |
sharedStorage.storageClass |
For production, use a StorageClass with reclaimPolicy: Retain on database volumes.
Security
The chart applies Pod Security Standards restricted defaults to every workload:
runAsNonRoot: trueallowPrivilegeEscalation: falsecapabilities.drop: [ALL]seccompProfile.type: RuntimeDefault
User-supplied securityContext values are merged with the defaults — your values win, but you don't have to repeat the defaults.
Other security features:
automountServiceAccountToken: falseon the ServiceAccount and every pod.- Every value in
app.envandrealtime.envis written to a chart-managed Secret and mounted viaenvFrom: secretRef— no values are inlined on the container spec. This eliminates a sensitivity classifier (no static list of "secret" keys to maintain) and ensures new provider keys can never accidentally leak into pod manifests. Two categories are inlined on the container instead: chart-computed values (DATABASE_URL,SOCKET_SERVER_URL,OLLAMA_URL,PII_URL) and operational defaults underapp.envDefaults/realtime.envDefaults(rate limits, timeouts, IVM tunables, feature-flag defaults, branding defaults,http://localhost:3000URL fallbacks). Operational defaults are non-sensitive by design — moving them out ofapp.envkeeps the Secret small and means External Secrets Operator users only have to map the keys they actually set, not every chart default. A value placed inapp.envalways wins over the same key inapp.envDefaults(the template skips the inline default when an override exists). - Optional
networkPolicy.enabled=trueenforces east-west isolation and blocks cloud metadata endpoints in egress.
Autoscaling
autoscaling:
enabled: true
minReplicas: 2
maxReplicas: 20
targetCPUUtilizationPercentage: 70
targetMemoryUtilizationPercentage: 80
When autoscaling.enabled=true, the chart omits spec.replicas from the Deployment so the HPA owns replica count. Requires metrics-server in the cluster. The realtime Deployment gets the same HPA unless autoscaling.realtime.enabled=false — scale realtime past one replica only with REDIS_URL set (Socket.IO Redis adapter), or cross-pod collaboration events are dropped.
Monitoring
monitoring:
serviceMonitor:
enabled: true
interval: 30s
Requires the Prometheus Operator CRDs. Scrapes /metrics on the app and realtime services — note the default images do not currently expose a /metrics endpoint, so enable this only with a build that does.
PII redaction
Sim can redact personally identifiable information using a Presidio service (analyzer + anonymizer combined into one image listening on port 5001). Enable it with:
pii:
enabled: true
When enabled, the chart deploys it as a standalone <release>-pii Deployment + Service and auto-wires PII_URL on the app to the in-cluster service. The service bundles five large spaCy models (en/es/it/pl/fi, ~2.2GB), so the first start takes ~3 minutes while models load — the startupProbe allows for this. Size the pii.resources for at least ~4Gi memory.
This alone powers the Guardrails PII block and on-demand masking. To additionally turn on automatic log redaction (the org/workspace data-retention scrub), you must:
app:
env:
PII_REDACTION: "true"
# The log-redaction path calls the app's own /api/guardrails/mask-batch,
# which must be reachable from inside the cluster. Set this to the in-cluster
# app Service URL (NOT the public ingress, which usually isn't hairpin-reachable).
INTERNAL_API_BASE_URL: "http://<release>-app.<namespace>.svc.cluster.local:3000"
Without a cluster-reachable INTERNAL_API_BASE_URL (it falls back to NEXT_PUBLIC_APP_URL), the redaction path fails closed — it scrubs affected fields to [REDACTION_FAILED] rather than leaking, but redaction won't actually run.
The PII image is published at
ghcr.io/simstudioai/pii(multi-arch). If you mirror images into a private registry, retag it alongside the app/realtime/migrations images.
Troubleshooting
Error: execution error at (sim/templates/...): app.env.BETTER_AUTH_SECRET is required for production deployment
You ran helm install without setting required secrets. Generate them and pass with --set:
helm install sim ./helm/sim \
--set app.env.BETTER_AUTH_SECRET=$(openssl rand -hex 32) \
--set app.env.ENCRYPTION_KEY=$(openssl rand -hex 32) \
--set app.env.INTERNAL_API_SECRET=$(openssl rand -hex 32) \
--set postgresql.auth.password=$(openssl rand -base64 24 | tr -d '/+=')
App pods stuck in CrashLoopBackOff
kubectl logs --namespace sim deploy/sim-app --tail 200
Common causes:
NEXT_PUBLIC_APP_URLstill set tohttp://localhost:3000in a clustered deploy → set it to your public origin.DATABASE_URLnot reachable → check the Postgres pod is running andpostgresql.auth.passwordmatches.- Missing migration → check
kubectl logs deploy/sim-app -c migrations(migrations run as an init container on the app pod).
Image pull errors (ErrImagePull / ImagePullBackOff)
- You pushed Sim to a private registry but haven't configured pull secrets. Set
global.imagePullSecretsandglobal.imageRegistry. - You overrode
image.tagto a tag that doesn't exist in the registry.helm get values simand verify.
Postgres pod Pending
kubectl describe pvc --namespace sim
Almost always one of:
- No default
StorageClass→ setglobal.storageClass. - No PV provisioner → install one (e.g. EBS CSI on EKS,
local-path-provisionerfor dev). - StorageClass exists but doesn't support
ReadWriteOnce→ pick another class.
Ingress not routing
kubectl get ingress --namespace sim
kubectl describe ingress --namespace sim
- Ingress controller not installed → install
ingress-nginxor similar. ingress.classNamedoesn't match your controller → set it to your installed class.- DNS not pointed at the ingress's external IP / LoadBalancer.
Get logs from each component
kubectl --namespace sim logs -f deployment/sim-app
kubectl --namespace sim logs -f deployment/sim-realtime
kubectl --namespace sim logs -f statefulset/sim-postgresql
kubectl --namespace sim logs deploy/sim-app -c migrations
Upgrading to 1.2.0
appVersion(the default image tag whenimage.tagis unset) is nowv0.7.44— the previous0.6.73referenced a tag that does not exist on GHCR, so an unpinned default install could not pull images. Production installs should still pinimage.tagexplicitly.externalSecrets.apiVersionnow defaults to"v1"— current External Secrets Operator releases no longer servev1beta1(removed upstream in 2026). SetexternalSecrets.apiVersion: "v1beta1"only if you still run ESO < 0.17.values.schema.jsonnow declares every top-level key and rejects unknown top-level keys, so a typo likenetworkPolciy:fails fast at install time instead of being silently ignored. If an upgrade suddenly fails schema validation, check your values file for stray top-level keys.- The opt-in telemetry collector no longer ships a Prometheus scrape config for the app/realtime services (they expose no
/metricsendpoint); OTLP ingestion is unchanged.
Upgrading to 1.1.0
No action is required for working configurations. Notes:
- Pods for
appandrealtimeroll once on upgrade (their rollout checksum now also covers the ExternalSecret manifest, fixing missed rollouts in ESO mode). - Two values keys that were never consumed by any template were removed:
app.secrets.existingSecret.keysand*.existingSecret.passwordKey. Existing secrets must use the standard key names (BETTER_AUTH_SECRET, ...,POSTGRES_PASSWORD,EXTERNAL_DB_PASSWORD); leftover keys in your values file are ignored, not rejected. telemetry.jaegernow exports over OTLP (otlp/jaeger) — pointtelemetry.jaeger.endpointat Jaeger's OTLP gRPC port (4317). The previousjaegerexporter did not exist in the pinned collector image, so any prior jaeger-enabled config was already failing at collector startup.
Support
- Docs: https://docs.sim.ai
- GitHub: https://github.com/simstudioai/sim
- Issues: https://github.com/simstudioai/sim/issues
- Slack: https://join.slack.com/t/sim-ott9864/shared_invite/zt-43lp8tc5v-0qrrqHGBKUsvQlpoouH~TA
License
Apache-2.0 © Sim. See LICENSE.