* fix(db): encode Date binds in raw sql, document the fetch_types contract
* fix(testing): drop template placeholders from mock error message for biome
* feat(sso): DNS domain verification gating org SSO registration
Add org-scoped domain ownership verification (DNS TXT challenge) as the
security precondition for configuring SSO. Closes the first-come domain-claim
vuln where any org could wire another company's domain to its own IdP.
- New sso_domain table + migration 0266; existing org SSO domains are
grandfathered as verified so live tenants are unaffected
- Verified-domains settings UI (enterprise-gated) with add/verify/remove
- Register route now requires a verified domain for org-scoped registration;
personal SSO and already-grandfathered domains are unaffected
- Self-host register script writes the verified sso_domain row directly, so
script-driven registration stays backwards compatible
* fix(sso): harden domain verification against concurrency + fix CI lint
Addresses review findings on state invariants under concurrent/failed writes:
- Add unique index on (organization_id, domain) so concurrent claims can't
create duplicate pending rows; POST re-reads and stays idempotent on conflict
- Verify flips the row only if it's still the exact pending challenge checked
(guards deletion/token-rotation mid-DNS-lookup) and maps the partial unique
index violation to 409 instead of an unhandled 500
- Wrap the self-host script's provider write + verified-domain upsert in a
transaction so a failed ownership write can't leave a provider committed
- Format 0266 snapshot/journal with biome (fixes @sim/db lint:check)
* fix(sso): re-check domain verification before provider write (TOCTOU)
The register gate checked the verified sso_domain row only at handler entry,
then ran OIDC discovery before writing the provider. A verified row removed
during that window could still complete registration. Extract the check into a
closure and call it both as an entry fast-fail and authoritatively right before
registerSSOProvider, alongside the existing domain-conflict re-check.
* fix(sso): stop rotating verification token on idempotent re-add
Re-adding a pending domain rotated its verification token, which invalidated a
TXT record the admin may have already published and — under two concurrent
re-adds — could return a token the racing write had already superseded, so the
admin's DNS record would never verify. Return the existing row unchanged
instead; the pending token is always shown in the UI, so it is never lost.
* fix(sso): close register TOCTOU with compensating delete + harden edges
Audit-driven hardening:
- Close the residual register TOCTOU: registerSSOProvider is create-only (throws
if the providerId exists), so a compensating delete after the write is provably
safe — it can only remove the just-created row. If verification was revoked
during the write, roll the provider back and 403.
- Verify is now idempotent under concurrency: a same-org row already flipped to
verified by a racing request returns 200, not a confusing 409.
- Grandfather backfill + self-host script now match normalizeSSODomain's dominant
transforms (lower + trim + strip leading wildcard) so a non-canonical legacy
domain can't miss the runtime gate's lookup. Prod backfill result is unchanged.
- Cleanup: drop dead default export, align card radius to sibling convention.
* fix(sso): redact domain tokens from non-admins + fix script stale-update
Round-5 review findings:
- GET /domains redacted the pending TXT verification token (a management
secret) to any org member. Now only owner/admins read it; members see the
list and status without it. Non-Enterprise orgs get an empty list (entitlement
flag only), never the domains/tokens.
- Self-host script decided update-vs-insert from a read taken OUTSIDE the
transaction; a provider deleted mid-flight made the UPDATE match zero rows
silently while the verified-domain upsert still committed (orphaned domain).
The decision now happens inside the transaction from the UPDATE's row count.
* docs(sso): drop unshipped enforce-SSO / auto-join copy from verified domains
Verified domains currently only gate SSO configuration. Remove the
forward-looking references to enforcing SSO and auto-joining members (deferred
to a later release) from the docs, settings copy, nav description, and schema
comment so we don't promise unshipped features.
* fix(sso): guard rollback to new providers only + Enterprise-gate domain removal
Round-6 review findings:
- The compensating provider rollback now only fires when the provider did not
exist before this request (providerExistedBefore). registerSSOProvider is
create-only today so reaching the rollback already implies a fresh create, but
this makes the safety local and future-proof: if Better Auth ever allowed
updating an existing provider, a revoked-verification rollback must not delete
that pre-existing row.
- DELETE /domains now requires an Enterprise plan like add/list/verify, so all
domain mutations share one entitlement (the UI already hides removal from
non-Enterprise orgs). Adds a delete-route test.
* fix(sso): roll back the SSO provider by row id, not logical keys
The compensating rollback deleted by (providerId, orgId). providerId is unique,
so if this request's row were deleted and recreated by a concurrent registration
in the narrow window before the rollback, the logical-key delete would remove
that other request's provider. Delete by the primary-key id registerSSOProvider
returns instead, so only the exact row this request created is ever removed.
* chore(sso): final-review polish — trim script read, unify copy, doc migration edge
Cosmetic cleanup from a final 4-track adversarial review (no bugs found in the
new logic):
- Self-host script: narrow the pre-transaction existence read to select({ id })
instead of SELECT * (it only feeds a log line now).
- Unify invalid-domain copy ("for example acme.com") and the verified-elsewhere
409 wording ("is already verified by another organization") across routes.
- p-3 shorthand on the domain row card.
- Document the migration's rare two-orgs-share-a-domain grandfather behavior
(login unaffected; validated no such duplicates in prod).
* fix(sso): apply attribute mapping + make SSO edit work; drop dead guard
Two pre-existing SSO bugs the final review surfaced (prod has one SSO org, RVW,
script-registered with the default mapping, so neither change affects it):
- Attribute mapping was passed at the top level of the register payload, which
Better Auth ignores — it reads oidcConfig.mapping / samlConfig.mapping. Nest it
so custom mappings actually apply. (Default mapping is unchanged, so existing
logins are unaffected.)
- Editing an SSO provider was broken: registerSSOProvider is create-only and
threw on the existing providerId → generic 500. Route now detects a provider
the caller already owns and updates it via Better Auth's updateSSOProvider, and
surfaces Better Auth's own error status/message instead of a blanket 500.
Also drops the now-unnecessary providerExistedBefore guard (the rollback deletes
by the created row's primary-key id and register is create-only) and the earlier
final-review polish (script read, unified copy, migration edge note).
Smoke-test SSO login + edit on staging before merge (auth-path change).
* fix(sso): require null org on personal-mode provider lookups (gate bypass)
The personal branch of both provider-ownership lookups keyed on
(providerId, userId) without requiring organizationId IS NULL. Because org
providers store userId = their creator and providerId is globally unique, an org
admin could send a personal-mode request (no orgId) — which skips the membership
check and the domain-verification gate — yet still match, and then via the new
update path move, their org's provider to an unverified domain. Add
isNull(organizationId) to the personal branch of both clauses so it can only
match a genuinely personal provider, matching the route's own isOwnedByCaller.
Found by an adversarial review of the update path added in 394bda9f7.
* fix(sso): script updates the observed provider by id, not providerId
Inside the registration transaction the script updated WHERE providerId — the
logical key. If the observed provider was deregistered and a replacement created
with the same providerId before the transaction ran, that update would clobber
the replacement's config and ownership. Update the specific observed row by its
primary-key id instead; if it's gone we insert, which fails cleanly on the
providerId unique constraint rather than overwriting the replacement.
* fix(sso): script upserts provider via delete-then-insert (no unique constraint)
sso_provider.provider_id is a plain (non-unique) index and prod holds legitimate
duplicates, so the previous "update by id, else insert" could create a duplicate
provider when the observed row was deregistered and replaced before the
transaction — the fallback insert would succeed. Delete every row for the
providerId then insert exactly one, inside the transaction, so the providerId
ends up as exactly this config atomically. Linked accounts key on the providerId
string (not the row id), so existing logins are unaffected.
* fix(sso): guard compensating-delete row id so rollback can't silently no-op
* chore(sso): regenerate migration as 0268 after merging staging
Staging landed migrations 0266/0267, colliding with our 0266. Removed our
migration, merged staging, and regenerated cleanly with drizzle-kit as
0268_sso_domain_verification (identical sso_domain table + indexes), then
re-appended the grandfather backfill. api-validation baseline reconciled to 973
(staging 970 + our 3 domain routes). Also make the register-route test's
registerSSOProvider mock return an id so the guarded compensating delete runs.
* refactor(sso): share normalizeSSODomain via @sim/utils so script matches gate
The self-host script canonicalized SSO domains with a minimal inline transform
(lower+trim+wildcard) that diverged from the app's full normalizeSSODomain
(protocol, port, path, trailing dot, email local part) — equivalent spellings
could store a different ownership key than the runtime gate looks up. Move
normalizeSSODomain into @sim/utils/sso-domain (a pure function) so the register
route, the domain-claim route, and the script all use the identical canonicalizer.
The script now skips the verified-domain record when SSO_DOMAIN isn't a valid
registrable domain instead of storing a malformed key.
* feat(byok): migrate agent/router/evaluator LLM keys to BYOK by model prefix
* fix(script): address review — cover translate/guardrails, trim provider ids, reject mixed mapping, add key-prefix guard
* fix(script): include pi block in LLM-family set, trim key-prefix value
* fix(script): add router_v2 to LLM-family set, canonicalize provider ids to lowercase
* fix(script): resolve overlapping model prefixes deterministically (longest match wins)
* fix(invites): preserve active organization for external access
Keep organization activation server-owned so failed membership checks cannot clear a valid session context.
* feat(admin, billing, settings): cleanup settings visibility, billing actor resolution, new admin routes
* address comments
* chore(db): reset pending migrations before staging merge
Remove locally generated migrations so they can be regenerated against the latest staging schema without preserving stale snapshots or numbering.
* regen migrations
* address comments
* chore(db): reset generated migrations before staging merge
Remove this branch's generated migrations so they can be regenerated against the latest staging schema with fresh numbering.
* upgrade global work
* fix lint
* address comments
* legacy callbacks correctness
* address comments
* update
* guardrail attribution
* fix(sso): support skipping the OIDC UserInfo endpoint at registration
* fix(sso): cap OIDC discovery fetch at 10s to avoid stalling registration
* test(sso): default-mock discovery fetch so intent is explicit
* fix(sso): prefer client_secret_post and surface discovery failure reasons
* fix(sso): always resolve token auth method and skip SSRF-checking a discarded userInfoEndpoint
* ci(migrations): skip db:migrate on merges that change no migration files
Every push to main/staging ran db:migrate against the production/staging
database even when the merge changed no schema, so a no-op migration would dial
the DB and fail whenever it was at its connection limit (53300, slots reserved
for SUPERUSER) — red-X'ing UI-only merges.
Add a detect-migrations job (dorny/paths-filter on packages/db/migrations/**)
and pass the result into the reusable migrations workflow, which now skips the
apply step when no migration files changed. The migrate job still runs so
downstream build/deploy jobs that need it are never skipped, and the flag
defaults to 'true' so manual dispatch and any unknown value always apply
migrations — the gate only ever skips a provably-empty change.
* fix(db): retry the migration connection on transient slot exhaustion
The migration opens its session on the first query (the advisory-lock
acquire). When the deploy database briefly exhausts every non-superuser
connection slot at peak, that connect fails with 53300 ("remaining connection
slots are reserved for roles with the SUPERUSER attribute") and the whole
deploy's migrate step errors out — even when the spike clears within seconds.
Add a bounded connectWithRetry() before acquiring the lock that retries 53300,
the 08xxx connection_exception class, and the driver's transport errors with
backoff (10 attempts, ~90s ceiling). Non-transient errors (auth, bad config)
still fail fast. The migration is a single short-lived session, so waiting out
a transient spike is far safer than failing the deploy.
* ci: drop the migration paths-filter gate (out of scope)
Revert the detect-migrations gate carried over from the closed CI PR; we are
fixing the connection failure at its source (migrate.ts connection retry)
rather than gating db:migrate, which the reviewers correctly noted could leave
a previously-merged migration unapplied after a failed deploy.
* feat(db): attribute Postgres connections by runtime via application_name
* improvement(db): label migration-runner connection sim-migrate; trim DB_APP_NAME comment
* fix(db): label realtime's shared @sim/db connections sim-realtime too
The realtime process uses both its own socketDb pool and the shared @sim/db
client (handlers, preflight, permissions). Only socketDb was labeled, so the
shared client defaulted to sim-app, mislabeling much of realtime's DB traffic.
Set DB_APP_NAME=sim-realtime at the process level (bootstrap before the dynamic
@/index import for prod; dev/start scripts for local) so both clients report it.
* feat(byok): support multiple keys per provider with round-robin rotation
* fix(byok): address review feedback — serialize cap check, defer encryption, explicit delete guard
* fix(byok): drop ON CONFLICT on removed unique index in legacy key migration script
* improvement(byok): guard double-submit, Enter-to-save on key field, cap hint in manage modal
* fix(db): serialize concurrent migrations with a Postgres advisory lock
Deployments start N app replicas at once, each with a migration sidecar.
drizzle migrate() has no cross-process lock, so all N read
__drizzle_migrations, all see the same migration pending, and all apply it
concurrently — one wins, the losers run the same DDL against already-mutated
state and exit 1 (e.g. DROP TABLE "form" -> table does not exist /
TaskFailedToStart). Wrap migrate() in a session-level pg_advisory_lock so
runners serialize: the winner migrates, the losers block, then re-read and
find nothing pending. Session locks auto-release on disconnect, so a crashed
runner never wedges the lock.
* fix(db): guard pg_advisory_unlock so it cannot mask a successful migration
If the explicit unlock throws (e.g. connection drops in the window after
migrate() commits), the exception bubbled to the outer catch and exited 1 —
falsely reporting a failed migration to the deploy orchestrator. The session
lock auto-releases on disconnect anyway, so swallow and log instead.
* refactor(db): move unlock-guard rationale to TSDoc helper
* improvement(logs): obj storage backed tracespans
* fix storage write context
* fix tests
* address comments
* address comments
* chore(db): remove migration 0219 to regenerate after staging merge
Drops the 0219_robust_shard SQL, its snapshot, and the journal entry so the
trace-spans/cost schema migration can be regenerated on top of the latest
staging migration chain (avoids a number collision with staging's migrations).
Co-authored-by: Cursor <cursoragent@cursor.com>
* improvement(billing): accurate per-member usage via shared ledger helper
Per-member/per-user usage in the org-member routes now adds the usage_log
ledger to the currentPeriodCost baseline (which is no longer incremented),
via a shared getOrgMemberLedgerByUser helper to avoid repeating the
subscription→period→ledger lookup across the admin and member-facing routes.
Co-authored-by: Cursor <cursoragent@cursor.com>
* regen migrations
* update migration
* address comments
* more code cleanup
* incorrect type cast
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore(utils): migrate to shared random/ID utilities and add enforcement linting
- Replace all Math.random(), crypto.randomUUID(), crypto.randomBytes(), nanoid, and uuid usages with shared @sim/utils/random and @sim/utils/id helpers across 72 files
- Add new @sim/utils exports: deepClone, omit, filterUndefined (object), truncate (string), backoffWithJitter, parseRetryAfter (retry), getErrorMessage (errors)
- Sweep all getErrorMessage, sleep, deepClone callsites across 500+ files to use shared utilities
- Add Biome noRestrictedImports rule to catch nanoid, uuid, and crypto named imports at lint time
- Add scripts/check-utils-enforcement.ts to catch Math.random and crypto.* global property access
- Add check:utils script to package.json
* chore(utils): replace deepClone wrapper with structuredClone built-in
deepClone() was a one-line wrapper around structuredClone(), which is
universally available in Node 17+ and all modern browsers. Removing the
abstraction reduces indirection and means contributors don't need to
learn a project-specific name for a well-known built-in.
- Remove deepClone from packages/utils/src/object.ts and index.ts
- Replace all 17 call sites with structuredClone() directly
- Update check:utils script suggestion text
- Update CLAUDE.md and global.md docs
* fix(utils): add missing biome noRestrictedImports rule and correct truncate docs
- Add noRestrictedImports to biome.json under style — bans nanoid and uuid
package imports at lint time (crypto.randomUUID/randomBytes are caught by
the check:utils grep script which handles global property access)
- Correct truncate() TSDoc and parameter name: sliceLength makes it clear
that total output length is sliceLength + suffix.length, matching the
behavior all callers were already written to expect
* fix(utils): add missing getErrorMessage imports at 4 call sites
The sweep agents added getErrorMessage calls without the corresponding
import in 4 files, causing test failures. Added the missing imports.
* fix(utils): fix build errors from getErrorMessage sweep and retry.ts Turbopack issue
- Fix retry.ts cross-file import: Turbopack cannot resolve './random.js' for
internal package imports; inline the jitter crypto call directly
- Add missing getErrorMessage imports to 32 files where the sweep added calls
without the corresponding import (caught by type-check and test runs)
- Remove accidental getErrorMessage import from crowdstrike/query/route.ts
which has its own domain-specific getErrorMessage for parsing CrowdStrike's
JSON error format
- Fix use-sub-block-value.ts type error from structuredClone narrowing:
add 'as T' cast at emitValue callsite (safe — valueCopy is always a
structural copy of newValue)
* fix(tools): use toError in crowdstrike catch block instead of local getErrorMessage
The catch block was calling the local getErrorMessage function which
parses CrowdStrike API JSON responses, not JavaScript Error objects.
Use toError(error).message to correctly extract the message from a
caught value in this context.
* fix(auth): resolve CORS errors for self-hosted deployments behind reverse proxies
- auth client now uses browser origin first, falling back to NEXT_PUBLIC_APP_URL
- socket client falls back to page origin when served from non-localhost (assumes /socket.io is proxied)
- add TRUSTED_ORIGINS env var to extend Better Auth trustedOrigins (apex+www, alias hostnames)
- warn at startup when NEXT_PUBLIC_APP_URL is localhost in production
- preprocess empty NEXT_PUBLIC_SOCKET_URL so docker-compose ${VAR:-} works
- migrate remaining uuid/nanoid/randomUUID usages to @sim/utils generateId/generateShortId
- extend generateShortId with optional alphabet param (rejection sampling)
- document TRUSTED_ORIGINS in .env.example, docker-compose.prod.yml, and helm values.yaml
Fixessimstudioai/sim#1243
* fix(auth): address PR review comments
* chore(env): drop unnecessary NEXT_PUBLIC_SOCKET_URL preprocess (skipValidation is true)
* fix(docker): include @sim/utils in migrations image
Migration scripts now import generateId from @sim/utils/id; without copying packages/utils into the image, bun install fails to resolve the workspace dep at build time and the import fails at runtime.
* fix(helm): remove unused NEXT_PUBLIC_SOCKET_URL from realtime sections
The realtime service never reads NEXT_PUBLIC_SOCKET_URL — its env schema
only includes BETTER_AUTH_URL, NEXT_PUBLIC_APP_URL, ALLOWED_ORIGINS,
BETTER_AUTH_SECRET, INTERNAL_API_SECRET, DATABASE_URL, and REDIS_URL.
Remove the dead config from all helm values files and the values schema.
* fix(helm): allow empty NEXT_PUBLIC_SOCKET_URL in values schema
The default in values.yaml is now "" (empty string), which falls back to
the page origin at runtime. The schema previously required a valid URI,
which would reject the default. Mirror the INTERNAL_API_BASE_URL pattern
using anyOf with const "". Also add TRUSTED_ORIGINS to the schema.
* docs(self-hosting): mark NEXT_PUBLIC_SOCKET_URL as optional
The page-origin fallback in getSocketUrl() means self-hosters no longer
need to set NEXT_PUBLIC_SOCKET_URL when realtime is on the same origin
as the app. Update docs to reflect this:
- Remove NEXT_PUBLIC_SOCKET_URL from .env scaffolding examples in
docker.mdx, platforms.mdx, environment-variables.mdx
- Mark the variable as Optional in the env vars table with the new
default behavior described
- Update troubleshooting to point at reverse-proxy /socket.io routing
rather than the env var
- Flip dev docker-compose defaults (local, ollama, devcontainer) from
http://localhost:3002 to empty for consistency with prod.yml; the
in-code localhost fallback handles the dev case identically
Applied across all 6 documentation languages (en/fr/de/ja/es/zh).
* chore: untrack and ignore .claude/scheduled_tasks.lock
* fix(sso): default tokenEndpointAuthentication to client_secret_post
better-auth's SSO plugin does not URL-encode credentials before Base64
encoding in client_secret_basic mode (RFC 6749 §2.3.1). When the client
secret contains special characters (+, =, /), OIDC providers decode them
incorrectly, causing invalid_client errors.
Default to client_secret_post when tokenEndpointAuthentication is not
explicitly set to avoid this upstream encoding issue.
Fixes#3626
* fix(sso): use nullish coalescing and add env var for tokenEndpointAuthentication
- Use ?? instead of || for semantic correctness
- Add SSO_OIDC_TOKEN_ENDPOINT_AUTH env var so users can explicitly
set client_secret_basic when their provider requires it
* docs(sso): add SSO_OIDC_TOKEN_ENDPOINT_AUTH to script usage comment
Signed-off-by: Mini Jeong <mini.jeong@navercorp.com>
* fix(sso): validate SSO_OIDC_TOKEN_ENDPOINT_AUTH env var value
Replace unsafe `as` type cast with runtime validation to ensure only
'client_secret_post' or 'client_secret_basic' are accepted. Invalid
values (typos, empty strings) now fall back to undefined, letting the
downstream ?? fallback apply correctly.
Signed-off-by: Mini Jeong <mini.jeong@navercorp.com>
---------
Signed-off-by: Mini Jeong <mini.jeong@navercorp.com>
* fix: specify authTagLength in AES-GCM decipheriv calls
Fixes missing authTagLength parameter in createDecipheriv calls using
AES-256-GCM mode. Without explicit tag length specification, the
application may be tricked into accepting shorter authentication tags,
potentially allowing ciphertext spoofing.
CWE-310: Cryptographic Issues (gcm-no-tag-length)
* fix: specify authTagLength on createCipheriv calls for AES-GCM consistency
Complements #3881 by adding explicit authTagLength: 16 to the encrypt
side as well, ensuring both cipher and decipher specify the tag length.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* refactor: clean up crypto modules
- Fix error: any → error: unknown with proper type guard in encryption.ts
- Eliminate duplicate iv.toString('hex') calls in both encrypt functions
- Remove redundant string split in decryptApiKey (was splitting twice)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* new turborepo version
---------
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Lakee Sivaraya <71339072+lakeesiv@users.noreply.github.com>
Co-authored-by: Vikhyath Mondreti <vikhyath@simstudio.ai>
Co-authored-by: Vikhyath Mondreti <vikhyathvikku@gmail.com>
Co-authored-by: Siddharth Ganesan <33737564+Sg312@users.noreply.github.com>
Co-authored-by: NLmejiro <kuroda.k1021@gmail.com>
* improvement(code-structure): move db into separate package
* make db separate package
* remake bun lock
* update imports to not maintain two separate ones
* fix CI for tests by adding dummy url
* vercel build fix attempt
* update bun lock
* regenerate bun lock
* fix mocks
* remove db commands from apps/sim package json