mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
feat(setup): setup wizard with browser-based Chat key handoff (#5911)
* feat(setup): setup wizard with browser-based Chat key handoff
Adds `bun run setup` and `bun run doctor` for local installs, and replaces
the wizard's paste-your-Chat-key step with a browser handoff that never puts
the key in a URL.
* improvement(setup): drop the paste-a-key fallback, simplify consent copy
The browser handoff is now the only path — the wizard waits on a spinner
instead of racing a paste prompt. Consent card leads with "Connect your
terminal" and moves the match-the-code disclaimer into the description.
* fix(setup): pin kube context, keep secrets out of argv, validate reused keys
Review findings from #5911:
- helm/kubectl now run against the validated context instead of the ambient one
- helm values are piped on stdin rather than passed as --set arguments
- ENCRYPTION_KEY/API_ENCRYPTION_KEY are checked for the 64-hex format the app
requires, not just length, so an unusable key is replaced rather than kept
- the managed Redis container's published port is read back instead of assumed
* refactor(copilot): one module for Chat API key operations
list/generate/delete each repeated the same /api/validate-key envelope in
their route. They now share callValidateKey in lib/copilot/server/api-keys.ts,
which also keeps the display masking server-side so the full key can only ever
leave at creation.
* improvement(setup): reuse shared helpers, parallelize probes, drop dead code
- PKCE verifier/state/pairing code now use generateSecureToken, generateRandomHex
and generateShortId instead of hand-rolled randomBytes; the pairing loop's
modulo was unbiased only because 256 % 32 == 0
- new sha256Base64Url in @sim/security/hash so both sides of the PKCE exchange
derive the challenge from one implementation
- isUsableSecret moved beside SECRET_KEYS so setup and doctor apply the same
rule; doctor previously passed a key setup would replace
- isTruthy narrowed to true/1, matching the app it claims to mirror — it accepted
yes/on, so a flag could read on in doctor and off in the app
- checkLive runs its five probes concurrently (~17s serial worst case)
- detection overlaps the banner animation instead of queueing behind it
- glyph.fail/glyph.warn at 13 sites that bypassed the constant; removed unused
prompter exports, a dead ENV_PATHS re-export, and an unused export keyword
* fix(setup): make doctor understand the compose env layout
Compose writes a single root .env (what docker-compose reads via env_file) but
the checks required the three per-app files, so a successful compose install was
followed by doctor printing three failures and exiting 1 — and the whole
coherence catalog was skipped because it keyed off apps/sim/.env existing.
Layout is now derived from what's on disk and every check consults it: file and
schema checks iterate the layout's targets, consistency reports skip when
there's only one file to mirror, and coherence/live read the layout's primary
file. The wizard's existing-config detection counts root for the same reason —
a compose install used to read as unconfigured and re-run from scratch.
* feat(cli-auth): device-authorization poll flow, drop the loopback listener
The CLI no longer binds a local port. It generates a request id + poll secret,
opens /cli/auth, and polls /api/cli/auth/poll over TLS while the user approves
in the browser — so the flow works over SSH and inside containers, where the
browser and terminal don't share a machine.
- approve stores the approval keyed by request id (session-authed, userId from
the session only); poll verifies the secret before an atomic claim, so an
observer of the semi-public request id can neither mint nor cancel it
- pairing code stays as the anti-phishing compare; no key ever crosses the
browser; done page just confirms
- removes the loopback listener, /token exchange, buildCliHandoffUrl, and
validateCliCallbackUrl (+ its tests) — nothing hands a key to a URL anymore
* fix(setup): reuse an existing managed Postgres container instead of colliding
A running sim-postgres fell through to `docker run --name sim-postgres` and died
on the name conflict; a stopped one failed with "no DATABASE_URL to reach it"
because the generated password only lived in the env files a fresh clone lacks.
Both facts are recoverable from Docker: the ladder now reads the published port
and password back via `docker inspect` and reuses the container (starting it if
stopped). A container that won't answer prompts before recreating, and never
drops the data volume silently.
* improvement(setup): audience-first run-mode hints
Each run mode now names who it's for — compose for self-hosting/evaluating, dev
for contributing to Sim, k8s for rehearsing a production deploy — with the live
detection state (Docker/kube/VM) appended.
* fix(cli-auth): retry a failed mint, port container/port fixes to Redis + k8s
Review findings from #5911:
- poll now reserves the mint with an atomic NX lock instead of deleting the
approval up front, so a failed mint (e.g. mothership blip) is retried by the
next poll instead of forcing a fresh browser approval; the lock still prevents
a double-mint and its TTL frees the slot if the caller dies
- setup reuses/recreates an unhealthy managed sim-redis instead of colliding on
the name (Redis has no data volume, so it removes and recreates without a prompt)
- k8s failure-path hints carry --context, matching the success-path hints, so a
changed ambient context can't send diagnostics to the wrong cluster
- compose port-free waits for a killed port to actually release before
re-checking; SIGKILL is async, so the immediate re-check re-saw the port
* fix(setup): harden mint cleanup, Windows browser, container detection, helm cwd
Review findings from #5911:
- a post-mint completeApproval failure no longer routes into releaseMint — the
mint lock now outlives the approval (shared TTL), so a cleanup blip can't leave
a re-mintable window and orphan a key; cleanup is best-effort after the key ships
- compose doctor --fix writes the feature-flag twin to the layout's primary env
(root .env on a compose install), not always apps/sim/.env
- Windows opens the browser via `cmd /c start "" <url>` — `start` is a shell
builtin, so spawning it directly ENOENT'd and the handoff never opened
- managed-container detection filters loosely and pins the exact name in code;
Docker's `name=^x$` anchor matches the internal `/x` form and often missed,
skipping the reuse branch
- the shared helm/kind run helper pins cwd to the repo root, matching helm test,
so `helm upgrade --install ./helm/sim` works from any working directory
* feat(chat-keys): standalone manage page, drop from settings nav, refresh README
- Add /account/settings/chat-keys — a linkable page to view, create, and revoke Chat API keys
- Remove Chat keys from the settings sidebar (account + unified nav) and its render branches
- README: replace Docker Compose + Manual Setup with the bun run setup wizard; drop the manual COPILOT_API_KEY step, point to the manage page
* fix(setup): per-key reason in the secret-replacement warning
Cursor: the warn hardcoded '64-character hex key', but only ENCRYPTION_KEY/API_ENCRYPTION_KEY require that — BETTER_AUTH_SECRET/INTERNAL_API_SECRET only need length >= 32. Use the existing secretRequirement(key) helper so each replaced key reports its actual requirement.
* fix(setup): compose doctor schema, cross-platform binary detection, quoted context hints
- Doctor: for the compose (root) env layout, require only the secrets compose has no interpolation default for (BETTER_AUTH_SECRET/ENCRYPTION_KEY/INTERNAL_API_SECRET). DATABASE_URL/BETTER_AUTH_URL/NEXT_PUBLIC_APP_URL come from docker-compose ${VAR:-default}, so a healthy compose install no longer fails doctor.
- Binary detection: use Bun.which instead of which (which is absent on Windows), so kubectl/helm/kind/docker resolve cross-platform.
- k8s diagnostic hints: POSIX-quote the kube-context so a context with whitespace/metacharacters can't break or inject into a copied command.
* fix(setup): quote kube-context in the helm uninstall tear-down hint too
The tear-down hint used --kube-context ${context} raw while the sibling kubectl hints already used shq(); a context with whitespace/metacharacters could break or inject into the copied command. All copyable k8s hints now go through shq(context).
* feat(setup): sim lifecycle CLI — start/stop/status/logs/down/reset
Turn the setup entry into a 'sim' command umbrella so there's one place to run everything, not scattered docker/bun commands. Adds a global bin (bun link) + a bun run sim fallback.
- Detects how you're running (compose file / managed dev containers / helm release) from disk + docker/helm state — no persisted mode. Ambiguous installs prompt.
- start/stop/restart/logs work per mode; down removes containers (volumes kept); reset archives .env + wipes managed data; both destructive verbs confirm first.
- status shows detected mode, container states, and app/realtime health.
- Wizard outro + README now point at the sim commands and the one-time bun link.
* feat(setup): 'bun run sim' is the primary entry; bare invocation prints help
- Lead usage/wizard-outro/README with 'bun run sim <cmd>' (works with zero PATH setup); global bare 'sim' via bun link is an optional upgrade, with the ~/.bun/bin PATH caveat spelled out (Homebrew's bun omits it).
- Bare 'sim' now prints help instead of launching the wizard; the wizard is 'sim setup'. The 'setup' npm script passes the keyword so 'bun run setup' is unchanged.
* fix(setup): quote the auth URL for cmd /c start on Windows
Cursor (High): cmd re-parses the command line and treats & in the query string as a command separator, so cmd /c start opened a URL truncated at the first &, breaking the key flow on win32 (the handoff URL always has request/challenge/pairing). Quote the URL and pass args verbatim so & stays literal.
* fix(setup): verify kube-context is really local; lengthen CLI handoff wait
- k8s: a context named like a local cluster (kind-*, docker-desktop) can actually point at a remote API server. Verify the server host is loopback/docker-internal before defaulting the 'use this context?' confirm to yes; otherwise warn and default to no, so generated secrets can't ship to a remote cluster on a blind Enter.
- cli-auth: bump the device-flow wait from 3 to 15 minutes so first-time users have time to sign up, wait for the email OTP, and approve before the terminal stops polling. The server-side approval record keeps its own short TTL, so a longer client wait only costs cheap rate-limited polls.
* fix(setup): only manage k8s lifecycle on a verified-local context
Greptile: sim down/reset used the ambient kube-context, so switching context after setup could uninstall a same-named sim-dev release from the wrong cluster. Gate k8sInstall on the same locality check the wizard uses (API server is loopback/docker-internal) via a shared isLocalKubeContext helper — the wizard only ever deploys locally, so a remote current-context is never treated as a Sim install.
* fix(setup): doctor skips placeholder secrets when seeding; reset names its target
- checks: the missing-file autofix copied shared keys from apps/sim/.env whenever truthy, including .env.example placeholders — doctor --fix could seed unusable secrets into realtime/db env files. Skip placeholders, matching autofixForMissing.
- lifecycle: reset now names the exact install (k8s context / compose file / dev containers) in its confirm, so a destructive reset can't silently hit the wrong same-named install after a context switch (down already names the context).
* fix(cli-auth): size the poll rate limit to the poll cadence; honor Retry-After
The poll route used the default public-IP bucket (10 burst, 5/min) but the CLI polls every 2s (30/min), so it 429'd within ~20s — worse behind a slow dev cold-compile. Give the endpoint a bucket matched to its cadence (60 burst, 60/min); it's not a brute-force surface (unknown request id returns pending, minting needs the 256-bit verifier). Also make the CLI honor Retry-After and back off on 429 so a shared-NAT per-IP limit degrades gracefully instead of hammering.
* fix(setup): check ports before starting the dev server, not just compose
Local dev auto-start spawned bun run dev:full with no port check, so it silently started a server that couldn't bind when 3000/3002 were already taken (e.g. another worktree's dev server). Extract compose's port-conflict resolver into a shared ensurePortsFree(ports) and run it before the dev start too — kill/recheck/leave, same as compose. Leaving the ports skips the auto-start with guidance instead of failing; compose still treats it as fatal.
* fix(setup): verify the kube cluster is reachable, not just local
A kubeconfig context can outlive its cluster — a kind cluster gets deleted or its Docker container stops (Docker/machine restart), but the context entry remains, pointing at a dead API-server port. The wizard checked the context looked local and handed it to helm, which failed with 'cluster unreachable'.
Add a clusterReachable() liveness probe: only offer the current context when it actually answers; if a local context is dead, fall through to the kind path. There, if kind still knows 'sim' but it's stopped, start its node containers and wait for the API; if it's gone, create fresh. Either way the user gets a working cluster instead of a cryptic helm failure.
* fix(helm): point appVersion at published image tags (v-prefixed, current)
The chart's appVersion was "0.6.73", but CI publishes GHCR tags with a v prefix (its release-commit regex captures v0.7.45). Since sim.image defaults every image tag to Chart.AppVersion, a default helm install requested ghcr.io/simstudioai/{simstudio,realtime,migrations}:0.6.73 — a tag that has never existed — so app and realtime sat in ImagePullBackOff and helm --wait failed with 'progress deadline exceeded'. Any self-hoster installing with default values hit this, not just the setup wizard.
Set appVersion to v0.7.45 (latest release on main; all three images verified present on ghcr) and bump the chart version to 1.1.1. Verified with helm lint, helm template (all images render as v0.7.45), and a live helm upgrade on a kind cluster where the new pods pull successfully while the old 0.6.73 pods remain in ImagePullBackOff.
* Revert "fix(helm): point appVersion at published image tags (v-prefixed, current)"
This reverts commit 28b6047d1d.
* chore(api-validation): rebaseline route count to 977 after staging merge
Staging moved the baseline to 975; this branch's two CLI-auth routes (approve, poll) make 977. The clean merge absorbed the earlier +2 adjustment.
* fix(settings): don't highlight a sibling nav item on nested settings pages
/account/settings/chat-keys is a real page but deliberately not a nav item, so the sidebar's parseSettingsPathSection fell through to defaultSection ('general') and highlighted General — the page read as though it lived inside General.
Resolve the sidebar's active item with a null default so an unmatched nested route highlights nothing, and widen SettingsSidebar's activeSection to string | null. The section feeding the title/description provider keeps its default (pages override title/description anyway), and /account/settings/billing/credit-usage still correctly highlights Billing.
* fix(setup,auth): manage explicitly-confirmed k8s contexts, fail loudly on reset, clear stale post-auth redirect
- lifecycle: detection is now factual — a sim-dev release either exists on the current context or it doesn't. Gating on locality stranded a release the user explicitly confirmed during setup (status/start/stop/down/reset all claimed no k8s install). Locality is recorded instead and surfaced through describeInstall, which every destructive confirm renders, so acting on a non-local cluster is named and defaulted to no rather than silently blocked or silently allowed.
- lifecycle: reset no longer discards helm uninstall's exit status. Env files are archived by that point, so claiming 'Reset complete' while the release still runs is the worst outcome — it now throws with retry/inspect commands.
- auth: signup clears POST_AUTH_REDIRECT_STORAGE_KEY when it has no callbackUrl, and the verification-disabled path consumes it, so a stale CLI/invite destination can't leak into a later flow in the same tab.
This commit is contained in:
@@ -0,0 +1,626 @@
|
||||
import { portOpen } from './detect.ts'
|
||||
import {
|
||||
type EnvFile,
|
||||
type EnvTarget,
|
||||
generateSecret,
|
||||
isPlaceholder,
|
||||
isTruthy,
|
||||
isUsableSecret,
|
||||
readEnvFile,
|
||||
SECRET_KEYS,
|
||||
SHARED_KEYS,
|
||||
secretRequirement,
|
||||
writeEnvValues,
|
||||
} from './env-files.ts'
|
||||
import { httpHealth, pgProbe, redisPing } from './probes.ts'
|
||||
import { FLAG_TWINS, hasMailProvider, LOGIN_PROVIDERS } from './twins.ts'
|
||||
|
||||
export type CheckGroup = 'files' | 'schema' | 'consistency' | 'coherence' | 'live'
|
||||
export type CheckStatus = 'pass' | 'warn' | 'fail' | 'skip'
|
||||
|
||||
export interface Finding {
|
||||
group: CheckGroup
|
||||
status: CheckStatus
|
||||
message: string
|
||||
fix?: string
|
||||
autofix?: () => void
|
||||
}
|
||||
|
||||
/**
|
||||
* Which env-file topology this install uses. Compose mode writes a single root
|
||||
* `.env` (that's what `docker-compose.*.yml` reads via `env_file`), dev mode
|
||||
* writes the three per-app files. Checking for the wrong one reports a healthy
|
||||
* install as broken, so the layout is derived and every check consults it.
|
||||
*/
|
||||
export type EnvLayout = 'split' | 'root' | 'none'
|
||||
|
||||
export interface CheckContext {
|
||||
env: Record<EnvTarget, EnvFile>
|
||||
layout: EnvLayout
|
||||
/** The file holding app configuration for this layout — what coherence reads. */
|
||||
primary: EnvFile
|
||||
live: boolean
|
||||
}
|
||||
|
||||
/** Split wins when both exist: the per-app files are what a dev run actually loads. */
|
||||
function detectLayout(env: Record<EnvTarget, EnvFile>): EnvLayout {
|
||||
if (env.sim.exists || env.realtime.exists || env.db.exists) return 'split'
|
||||
return env.root.exists ? 'root' : 'none'
|
||||
}
|
||||
|
||||
/** Targets whose files this layout expects to exist. */
|
||||
function layoutTargets(layout: EnvLayout): EnvTarget[] {
|
||||
if (layout === 'split') return ['sim', 'realtime', 'db']
|
||||
return layout === 'root' ? ['root'] : []
|
||||
}
|
||||
|
||||
export function loadCheckContext(live: boolean): CheckContext {
|
||||
const env = {
|
||||
sim: readEnvFile('sim'),
|
||||
realtime: readEnvFile('realtime'),
|
||||
db: readEnvFile('db'),
|
||||
root: readEnvFile('root'),
|
||||
}
|
||||
const layout = detectLayout(env)
|
||||
return { env, layout, primary: layout === 'root' ? env.root : env.sim, live }
|
||||
}
|
||||
|
||||
const REQUIRED_KEYS: Partial<Record<EnvTarget, string[]>> = {
|
||||
sim: [
|
||||
'DATABASE_URL',
|
||||
'BETTER_AUTH_SECRET',
|
||||
'BETTER_AUTH_URL',
|
||||
'NEXT_PUBLIC_APP_URL',
|
||||
'ENCRYPTION_KEY',
|
||||
'INTERNAL_API_SECRET',
|
||||
],
|
||||
realtime: [
|
||||
'DATABASE_URL',
|
||||
'BETTER_AUTH_URL',
|
||||
'BETTER_AUTH_SECRET',
|
||||
'INTERNAL_API_SECRET',
|
||||
'NEXT_PUBLIC_APP_URL',
|
||||
],
|
||||
db: ['DATABASE_URL'],
|
||||
// Compose's single root .env only carries the secrets that have no safe
|
||||
// interpolation default in docker-compose.*.yml. DATABASE_URL, BETTER_AUTH_URL,
|
||||
// and NEXT_PUBLIC_APP_URL are supplied by `${VAR:-default}` there, so requiring
|
||||
// them here would fail a healthy compose install that never wrote them.
|
||||
root: ['BETTER_AUTH_SECRET', 'ENCRYPTION_KEY', 'INTERNAL_API_SECRET'],
|
||||
}
|
||||
|
||||
const MIN_32_KEYS = new Set<string>(SECRET_KEYS)
|
||||
const URL_KEYS = ['DATABASE_URL', 'BETTER_AUTH_URL', 'NEXT_PUBLIC_APP_URL']
|
||||
|
||||
function rel(file: EnvFile): string {
|
||||
return `${file.target === 'root' ? '' : file.target === 'db' ? 'packages/db/' : `apps/${file.target}/`}.env`
|
||||
}
|
||||
|
||||
function checkFiles(ctx: CheckContext): Finding[] {
|
||||
if (ctx.layout === 'none') {
|
||||
return [
|
||||
{
|
||||
group: 'files',
|
||||
status: 'fail',
|
||||
message: 'no env files found',
|
||||
fix: 'run: bun run setup',
|
||||
},
|
||||
]
|
||||
}
|
||||
const findings: Finding[] = []
|
||||
for (const target of layoutTargets(ctx.layout)) {
|
||||
const file = ctx.env[target]
|
||||
if (file.exists) {
|
||||
findings.push({ group: 'files', status: 'pass', message: `${rel(file)} exists` })
|
||||
continue
|
||||
}
|
||||
const canSeed = target !== 'sim' && ctx.env.sim.exists
|
||||
findings.push({
|
||||
group: 'files',
|
||||
status: 'fail',
|
||||
message: `${rel(file)} is missing`,
|
||||
fix: canSeed
|
||||
? `run doctor --fix to seed it from apps/${target === 'db' ? '../packages/db' : target}/.env.example + apps/sim/.env`
|
||||
: 'run: bun run setup',
|
||||
autofix: canSeed
|
||||
? () => {
|
||||
const keys = target === 'db' ? ['DATABASE_URL'] : [...SHARED_KEYS]
|
||||
const values: Record<string, string> = {}
|
||||
for (const key of keys) {
|
||||
const value = ctx.env.sim.vars.get(key)
|
||||
// Skip placeholders so seeding never copies an .env.example stub
|
||||
// into the new file (matches autofixForMissing).
|
||||
if (value && !isPlaceholder(value)) values[key] = value
|
||||
}
|
||||
writeEnvValues(target, values)
|
||||
}
|
||||
: undefined,
|
||||
})
|
||||
}
|
||||
return findings
|
||||
}
|
||||
|
||||
function autofixForMissing(
|
||||
ctx: CheckContext,
|
||||
target: EnvTarget,
|
||||
key: string
|
||||
): (() => void) | undefined {
|
||||
const simValue = ctx.env.sim.vars.get(key)
|
||||
if (
|
||||
target !== 'sim' &&
|
||||
(SHARED_KEYS as readonly string[]).includes(key) &&
|
||||
simValue &&
|
||||
!isPlaceholder(simValue)
|
||||
) {
|
||||
return () => writeEnvValues(target, { [key]: simValue })
|
||||
}
|
||||
if (MIN_32_KEYS.has(key)) {
|
||||
return () => writeEnvValues(target, { [key]: generateSecret() })
|
||||
}
|
||||
return undefined
|
||||
}
|
||||
|
||||
function checkSchema(ctx: CheckContext): Finding[] {
|
||||
const findings: Finding[] = []
|
||||
const production = process.env.NODE_ENV === 'production'
|
||||
for (const target of layoutTargets(ctx.layout)) {
|
||||
const file = ctx.env[target]
|
||||
if (!file.exists) continue
|
||||
const missing: string[] = []
|
||||
for (const key of REQUIRED_KEYS[target] ?? []) {
|
||||
const value = file.vars.get(key)
|
||||
if (!value) {
|
||||
missing.push(key)
|
||||
findings.push({
|
||||
group: 'schema',
|
||||
status: 'fail',
|
||||
message: `${rel(file)}: ${key} is missing or empty`,
|
||||
fix: MIN_32_KEYS.has(key) ? 'doctor --fix generates it' : `set ${key} in ${rel(file)}`,
|
||||
autofix: autofixForMissing(ctx, target, key),
|
||||
})
|
||||
continue
|
||||
}
|
||||
if (isPlaceholder(value)) {
|
||||
findings.push({
|
||||
group: 'schema',
|
||||
status: production ? 'fail' : 'warn',
|
||||
message: `${rel(file)}: ${key} still has the .env.example placeholder`,
|
||||
fix: MIN_32_KEYS.has(key)
|
||||
? 'doctor --fix generates a real value'
|
||||
: `replace the placeholder in ${rel(file)}`,
|
||||
autofix: MIN_32_KEYS.has(key)
|
||||
? () => writeEnvValues(target, { [key]: generateSecret() })
|
||||
: undefined,
|
||||
})
|
||||
continue
|
||||
}
|
||||
if (MIN_32_KEYS.has(key) && !isUsableSecret(key, value)) {
|
||||
findings.push({
|
||||
group: 'schema',
|
||||
status: 'fail',
|
||||
message: `${rel(file)}: ${key} ${secretRequirement(key)}`,
|
||||
fix: 'generate a new one with `openssl rand -hex 32` (rotating it invalidates existing sessions/encrypted data)',
|
||||
})
|
||||
continue
|
||||
}
|
||||
if (URL_KEYS.includes(key)) {
|
||||
try {
|
||||
new URL(value)
|
||||
} catch {
|
||||
findings.push({
|
||||
group: 'schema',
|
||||
status: 'fail',
|
||||
message: `${rel(file)}: ${key} is not a valid URL (${value})`,
|
||||
fix: `correct ${key} in ${rel(file)}`,
|
||||
})
|
||||
}
|
||||
}
|
||||
}
|
||||
if (missing.length === 0 && findings.every((f) => !f.message.startsWith(rel(file)))) {
|
||||
findings.push({
|
||||
group: 'schema',
|
||||
status: 'pass',
|
||||
message: `${rel(file)}: required keys valid`,
|
||||
})
|
||||
}
|
||||
}
|
||||
return findings
|
||||
}
|
||||
|
||||
function checkConsistency(ctx: CheckContext): Finding[] {
|
||||
// Consistency is about the same key agreeing across files; a single root
|
||||
// file has nothing to disagree with.
|
||||
if (ctx.layout !== 'split') {
|
||||
return ctx.layout === 'root'
|
||||
? [{ group: 'consistency', status: 'skip', message: 'single .env — nothing to mirror' }]
|
||||
: []
|
||||
}
|
||||
const findings: Finding[] = []
|
||||
const { sim, realtime, db } = ctx.env
|
||||
if (sim.exists && realtime.exists) {
|
||||
for (const key of SHARED_KEYS) {
|
||||
const simValue = sim.vars.get(key)
|
||||
const realtimeValue = realtime.vars.get(key)
|
||||
if (!simValue || !realtimeValue) continue
|
||||
if (simValue !== realtimeValue) {
|
||||
findings.push({
|
||||
group: 'consistency',
|
||||
status: 'fail',
|
||||
message: `${key} differs between apps/sim/.env and apps/realtime/.env`,
|
||||
fix: 'doctor --fix mirrors the apps/sim/.env value',
|
||||
autofix: () => writeEnvValues('realtime', { [key]: simValue }),
|
||||
})
|
||||
}
|
||||
}
|
||||
}
|
||||
if (sim.exists && db.exists) {
|
||||
const simDsn = sim.vars.get('DATABASE_URL')
|
||||
const dbDsn = db.vars.get('DATABASE_URL')
|
||||
if (simDsn && dbDsn && simDsn !== dbDsn) {
|
||||
findings.push({
|
||||
group: 'consistency',
|
||||
status: 'fail',
|
||||
message:
|
||||
'DATABASE_URL differs between apps/sim/.env and packages/db/.env — migrations would hit a different database',
|
||||
fix: 'doctor --fix mirrors the apps/sim/.env value',
|
||||
autofix: () => writeEnvValues('db', { DATABASE_URL: simDsn }),
|
||||
})
|
||||
}
|
||||
}
|
||||
if (findings.length === 0) {
|
||||
findings.push({
|
||||
group: 'consistency',
|
||||
status: 'pass',
|
||||
message: 'shared env subset is in sync across files',
|
||||
})
|
||||
}
|
||||
return findings
|
||||
}
|
||||
|
||||
function checkCoherence(ctx: CheckContext): Finding[] {
|
||||
const findings: Finding[] = []
|
||||
const sim = ctx.primary
|
||||
if (!sim.exists) return findings
|
||||
if (isTruthy(sim.vars.get('TRIGGER_DEV_ENABLED'))) {
|
||||
const missing = ['TRIGGER_SECRET_KEY', 'TRIGGER_PROJECT_ID'].filter((k) => !sim.vars.get(k))
|
||||
if (missing.length > 0) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'fail',
|
||||
message: `TRIGGER_DEV_ENABLED is on but ${missing.join(' and ')} ${missing.length > 1 ? 'are' : 'is'} not set`,
|
||||
fix: 'set the missing Trigger.dev vars or remove TRIGGER_DEV_ENABLED (jobs fall back to the DB queue)',
|
||||
})
|
||||
}
|
||||
}
|
||||
const redisUrl = sim.vars.get('REDIS_URL')
|
||||
if (redisUrl?.startsWith('rediss://')) {
|
||||
const host = new URL(redisUrl).hostname
|
||||
if (/^\d+\.\d+\.\d+\.\d+$/.test(host) && !sim.vars.get('REDIS_TLS_SERVERNAME')) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'fail',
|
||||
message:
|
||||
'rediss:// with a bare IP host requires REDIS_TLS_SERVERNAME — the redis client throws without it',
|
||||
fix: 'set REDIS_TLS_SERVERNAME to the certificate hostname',
|
||||
})
|
||||
}
|
||||
}
|
||||
const appUrl = sim.vars.get('NEXT_PUBLIC_APP_URL')
|
||||
if (appUrl) {
|
||||
try {
|
||||
const host = new URL(appUrl).hostname
|
||||
if (host === 'sim.ai' || host.endsWith('.sim.ai')) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'warn',
|
||||
message: `NEXT_PUBLIC_APP_URL points at ${host} — this flips isHosted=true and disables self-host overrides`,
|
||||
fix: 'use your own domain or http://localhost:3000',
|
||||
})
|
||||
}
|
||||
} catch {
|
||||
// schema group already reports the invalid URL
|
||||
}
|
||||
}
|
||||
const hasS3 = Boolean(sim.vars.get('AWS_REGION') && sim.vars.get('S3_BUCKET_NAME'))
|
||||
const s3Partial = Boolean(sim.vars.get('AWS_REGION')) !== Boolean(sim.vars.get('S3_BUCKET_NAME'))
|
||||
const hasAzure = Boolean(
|
||||
sim.vars.get('AZURE_CONNECTION_STRING') || sim.vars.get('AZURE_ACCOUNT_NAME')
|
||||
)
|
||||
const azurePartial =
|
||||
Boolean(sim.vars.get('AZURE_ACCOUNT_NAME')) &&
|
||||
!sim.vars.get('AZURE_ACCOUNT_KEY') &&
|
||||
!sim.vars.get('AZURE_CONNECTION_STRING')
|
||||
const hasGcs = Boolean(sim.vars.get('GCS_BUCKET_NAME'))
|
||||
if (s3Partial) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'fail',
|
||||
message:
|
||||
'S3 is half-configured (need BOTH AWS_REGION and S3_BUCKET_NAME) — storage silently falls back to local disk',
|
||||
fix: 'set the missing var, or remove both to use local disk intentionally',
|
||||
})
|
||||
}
|
||||
if (azurePartial) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'fail',
|
||||
message:
|
||||
'Azure storage is half-configured — AZURE_ACCOUNT_NAME needs AZURE_ACCOUNT_KEY (or use AZURE_CONNECTION_STRING)',
|
||||
fix: 'set the missing credential, or remove the Azure vars',
|
||||
})
|
||||
}
|
||||
if (hasAzure && hasS3) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'warn',
|
||||
message:
|
||||
'both Azure Blob and S3 are configured — Azure takes precedence, the S3 vars are ignored',
|
||||
fix: 'remove the backend you are not using',
|
||||
})
|
||||
}
|
||||
if (hasGcs && (hasAzure || hasS3)) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'warn',
|
||||
message:
|
||||
'GCS is configured alongside Azure/S3 — GCS is only used when neither of those is set',
|
||||
fix: 'remove the backend you are not using',
|
||||
})
|
||||
}
|
||||
for (const { server, client } of FLAG_TWINS) {
|
||||
const serverValue = sim.vars.get(server)
|
||||
const clientValue = sim.vars.get(client)
|
||||
const bothUnset = serverValue === undefined && clientValue === undefined
|
||||
if (bothUnset || isTruthy(serverValue) === isTruthy(clientValue)) continue
|
||||
const setSide = serverValue !== undefined ? server : client
|
||||
const missingSide = serverValue !== undefined ? client : server
|
||||
const value = serverValue ?? clientValue ?? ''
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'fail',
|
||||
message: `${setSide} is set but its twin ${missingSide} disagrees — server and browser will render different features`,
|
||||
fix: `doctor --fix sets ${missingSide}=${value}`,
|
||||
// Write to the layout's primary env (root on a compose install), not always sim.
|
||||
autofix: () => writeEnvValues(sim.target, { [missingSide]: value }),
|
||||
})
|
||||
}
|
||||
|
||||
const disableAuth = sim.vars.get('DISABLE_AUTH')
|
||||
if (
|
||||
isTruthy(disableAuth) &&
|
||||
ctx.env.realtime.exists &&
|
||||
!isTruthy(ctx.env.realtime.vars.get('DISABLE_AUTH'))
|
||||
) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'fail',
|
||||
message:
|
||||
'DISABLE_AUTH is on in apps/sim/.env but not apps/realtime/.env — the socket server still enforces auth, so the canvas breaks silently',
|
||||
fix: 'doctor --fix mirrors it into apps/realtime/.env',
|
||||
autofix: () => writeEnvValues('realtime', { DISABLE_AUTH: disableAuth as string }),
|
||||
})
|
||||
}
|
||||
|
||||
if (isTruthy(sim.vars.get('EMAIL_VERIFICATION_ENABLED')) && !hasMailProvider(sim.vars)) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'fail',
|
||||
message:
|
||||
'EMAIL_VERIFICATION_ENABLED is on but no mail provider is configured — verification emails only go to the console, locking out new users',
|
||||
fix: 'configure RESEND_API_KEY / SMTP_* / AWS_SES_REGION, or turn verification off',
|
||||
})
|
||||
}
|
||||
|
||||
const featureRules: Array<{ flag: string; needs: string[]; label: string }> = [
|
||||
{ flag: 'BILLING_ENABLED', needs: ['STRIPE_SECRET_KEY'], label: 'billing' },
|
||||
{ flag: 'E2B_ENABLED', needs: ['E2B_API_KEY'], label: 'E2B code execution' },
|
||||
{ flag: 'SSO_ENABLED', needs: ['SSO_ISSUER'], label: 'SSO' },
|
||||
]
|
||||
for (const rule of featureRules) {
|
||||
if (!isTruthy(sim.vars.get(rule.flag))) continue
|
||||
const missing = rule.needs.filter((key) => !sim.vars.get(key))
|
||||
if (missing.length > 0) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'fail',
|
||||
message: `${rule.flag} is on but ${missing.join(', ')} is not set — ${rule.label} will fail at runtime`,
|
||||
fix: `set ${missing.join(', ')} or remove ${rule.flag}`,
|
||||
})
|
||||
}
|
||||
}
|
||||
if (
|
||||
isTruthy(sim.vars.get('PII_GRANULAR_REDACTION')) &&
|
||||
!isTruthy(sim.vars.get('PII_REDACTION'))
|
||||
) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'warn',
|
||||
message:
|
||||
'PII_GRANULAR_REDACTION is on but PII_REDACTION is off — the granular flag is inert without it',
|
||||
fix: 'set PII_REDACTION=true or remove PII_GRANULAR_REDACTION',
|
||||
})
|
||||
}
|
||||
if (
|
||||
Boolean(sim.vars.get('TURNSTILE_SECRET_KEY')) !==
|
||||
Boolean(sim.vars.get('NEXT_PUBLIC_TURNSTILE_SITE_KEY'))
|
||||
) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'fail',
|
||||
message:
|
||||
'Turnstile is half-configured — TURNSTILE_SECRET_KEY and NEXT_PUBLIC_TURNSTILE_SITE_KEY must both be set',
|
||||
fix: 'set the missing Turnstile var or remove both',
|
||||
})
|
||||
}
|
||||
for (const provider of LOGIN_PROVIDERS) {
|
||||
if (Boolean(sim.vars.get(provider.idKey)) !== Boolean(sim.vars.get(provider.secretKey))) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'fail',
|
||||
message: `${provider.label} login is half-configured — ${provider.idKey} and ${provider.secretKey} must both be set`,
|
||||
fix: 'set the missing credential or remove both',
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
const appUrlValue = sim.vars.get('NEXT_PUBLIC_APP_URL')
|
||||
if (
|
||||
appUrlValue &&
|
||||
!appUrlValue.includes('localhost') &&
|
||||
!appUrlValue.includes('127.0.0.1') &&
|
||||
!sim.vars.get('NEXT_PUBLIC_SOCKET_URL')
|
||||
) {
|
||||
findings.push({
|
||||
group: 'coherence',
|
||||
status: 'warn',
|
||||
message:
|
||||
'NEXT_PUBLIC_APP_URL is not localhost but NEXT_PUBLIC_SOCKET_URL is unset — the browser cannot find the realtime server',
|
||||
fix: 'set NEXT_PUBLIC_SOCKET_URL to the public URL of the realtime service (:3002)',
|
||||
})
|
||||
}
|
||||
|
||||
if (findings.length === 0) {
|
||||
findings.push({ group: 'coherence', status: 'pass', message: 'no conflicting settings' })
|
||||
}
|
||||
return findings
|
||||
}
|
||||
|
||||
async function checkDatabase(sim: EnvFile): Promise<Finding[]> {
|
||||
const findings: Finding[] = []
|
||||
const dsn = sim.vars.get('DATABASE_URL')
|
||||
const dsnPassword = (() => {
|
||||
try {
|
||||
return dsn ? new URL(dsn).password : null
|
||||
} catch {
|
||||
return null
|
||||
}
|
||||
})()
|
||||
if (dsn && dsnPassword !== null && !isPlaceholder(dsnPassword)) {
|
||||
const probe = await pgProbe(dsn)
|
||||
if (!probe.ok) {
|
||||
findings.push({
|
||||
group: 'live',
|
||||
status: 'fail',
|
||||
message: `database unreachable: ${probe.error}`,
|
||||
fix: 'start Postgres (bun run setup can manage a pgvector container) or fix DATABASE_URL',
|
||||
})
|
||||
} else {
|
||||
findings.push({ group: 'live', status: 'pass', message: 'database reachable' })
|
||||
if (!probe.pgvectorAvailable) {
|
||||
findings.push({
|
||||
group: 'live',
|
||||
status: 'fail',
|
||||
message: 'pgvector extension is not available on this Postgres',
|
||||
fix: 'use the pgvector/pgvector:pg17 image or install the extension',
|
||||
})
|
||||
}
|
||||
const { applied, journal } = probe.migrations ?? { applied: null, journal: 0 }
|
||||
if (applied === null) {
|
||||
findings.push({
|
||||
group: 'live',
|
||||
status: 'fail',
|
||||
message: 'migrations have never run on this database',
|
||||
fix: 'cd packages/db && bun run db:migrate',
|
||||
})
|
||||
} else if (applied < journal) {
|
||||
findings.push({
|
||||
group: 'live',
|
||||
status: 'warn',
|
||||
message: `database has ${applied}/${journal} migrations applied`,
|
||||
fix: 'cd packages/db && bun run db:migrate',
|
||||
})
|
||||
} else {
|
||||
findings.push({
|
||||
group: 'live',
|
||||
status: 'pass',
|
||||
message: `migrations up to date (${applied})`,
|
||||
})
|
||||
}
|
||||
}
|
||||
} else {
|
||||
findings.push({
|
||||
group: 'live',
|
||||
status: 'skip',
|
||||
message: 'database: DATABASE_URL not usable yet',
|
||||
})
|
||||
}
|
||||
|
||||
return findings
|
||||
}
|
||||
|
||||
async function checkRedis(sim: EnvFile): Promise<Finding[]> {
|
||||
const redisUrl = sim.vars.get('REDIS_URL')
|
||||
if (!redisUrl) return []
|
||||
const ping = await redisPing(redisUrl)
|
||||
return [
|
||||
ping.ok
|
||||
? { group: 'live', status: 'pass', message: 'redis reachable' }
|
||||
: {
|
||||
group: 'live',
|
||||
status: 'fail',
|
||||
message: `redis unreachable: ${ping.error}`,
|
||||
fix: 'fix REDIS_URL or remove it (optional for single-replica)',
|
||||
},
|
||||
]
|
||||
}
|
||||
|
||||
async function checkService(label: string, port: number, url: string): Promise<Finding[]> {
|
||||
if (!(await portOpen(port))) {
|
||||
return [{ group: 'live', status: 'skip', message: `${label}: not running on :${port}` }]
|
||||
}
|
||||
if (await httpHealth(url)) {
|
||||
return [{ group: 'live', status: 'pass', message: `${label} healthy on :${port}` }]
|
||||
}
|
||||
return [
|
||||
{
|
||||
group: 'live',
|
||||
status: 'fail',
|
||||
message: `${label}: something is on :${port} but ${url} is not answering`,
|
||||
fix: 'check the dev server logs',
|
||||
},
|
||||
]
|
||||
}
|
||||
|
||||
async function checkOllama(sim: EnvFile): Promise<Finding[]> {
|
||||
const ollamaUrl = sim.vars.get('OLLAMA_URL')
|
||||
if (!ollamaUrl) return []
|
||||
return [
|
||||
(await httpHealth(`${ollamaUrl.replace(/\/$/, '')}/api/tags`))
|
||||
? { group: 'live', status: 'pass', message: 'ollama reachable' }
|
||||
: {
|
||||
group: 'live',
|
||||
status: 'warn',
|
||||
message: 'OLLAMA_URL is set but Ollama is not answering',
|
||||
fix: 'start Ollama or remove OLLAMA_URL',
|
||||
},
|
||||
]
|
||||
}
|
||||
|
||||
/**
|
||||
* The five probes are independent, so they run concurrently — serially this is
|
||||
* the sum of every timeout (~17s worst case) on a command whose whole job is to
|
||||
* tell you what's broken. Results are concatenated in a fixed order so the
|
||||
* report stays deterministic regardless of which probe settles first.
|
||||
*/
|
||||
async function checkLive(ctx: CheckContext): Promise<Finding[]> {
|
||||
const sim = ctx.primary
|
||||
const [database, redis, app, realtime, ollama] = await Promise.all([
|
||||
checkDatabase(sim),
|
||||
checkRedis(sim),
|
||||
checkService('app', 3000, 'http://localhost:3000/api/health'),
|
||||
checkService('realtime', 3002, 'http://localhost:3002/health'),
|
||||
checkOllama(sim),
|
||||
])
|
||||
return [...database, ...redis, ...app, ...realtime, ...ollama]
|
||||
}
|
||||
|
||||
export async function runChecks(ctx: CheckContext, groups?: CheckGroup[]): Promise<Finding[]> {
|
||||
const findings: Finding[] = [
|
||||
...checkFiles(ctx),
|
||||
...checkSchema(ctx),
|
||||
...checkConsistency(ctx),
|
||||
...checkCoherence(ctx),
|
||||
]
|
||||
if (ctx.live) findings.push(...(await checkLive(ctx)))
|
||||
return groups ? findings.filter((f) => groups.includes(f.group)) : findings
|
||||
}
|
||||
Reference in New Issue
Block a user