mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
feat(setup): setup wizard with browser-based Chat key handoff (#5911)
* feat(setup): setup wizard with browser-based Chat key handoff
Adds `bun run setup` and `bun run doctor` for local installs, and replaces
the wizard's paste-your-Chat-key step with a browser handoff that never puts
the key in a URL.
* improvement(setup): drop the paste-a-key fallback, simplify consent copy
The browser handoff is now the only path — the wizard waits on a spinner
instead of racing a paste prompt. Consent card leads with "Connect your
terminal" and moves the match-the-code disclaimer into the description.
* fix(setup): pin kube context, keep secrets out of argv, validate reused keys
Review findings from #5911:
- helm/kubectl now run against the validated context instead of the ambient one
- helm values are piped on stdin rather than passed as --set arguments
- ENCRYPTION_KEY/API_ENCRYPTION_KEY are checked for the 64-hex format the app
requires, not just length, so an unusable key is replaced rather than kept
- the managed Redis container's published port is read back instead of assumed
* refactor(copilot): one module for Chat API key operations
list/generate/delete each repeated the same /api/validate-key envelope in
their route. They now share callValidateKey in lib/copilot/server/api-keys.ts,
which also keeps the display masking server-side so the full key can only ever
leave at creation.
* improvement(setup): reuse shared helpers, parallelize probes, drop dead code
- PKCE verifier/state/pairing code now use generateSecureToken, generateRandomHex
and generateShortId instead of hand-rolled randomBytes; the pairing loop's
modulo was unbiased only because 256 % 32 == 0
- new sha256Base64Url in @sim/security/hash so both sides of the PKCE exchange
derive the challenge from one implementation
- isUsableSecret moved beside SECRET_KEYS so setup and doctor apply the same
rule; doctor previously passed a key setup would replace
- isTruthy narrowed to true/1, matching the app it claims to mirror — it accepted
yes/on, so a flag could read on in doctor and off in the app
- checkLive runs its five probes concurrently (~17s serial worst case)
- detection overlaps the banner animation instead of queueing behind it
- glyph.fail/glyph.warn at 13 sites that bypassed the constant; removed unused
prompter exports, a dead ENV_PATHS re-export, and an unused export keyword
* fix(setup): make doctor understand the compose env layout
Compose writes a single root .env (what docker-compose reads via env_file) but
the checks required the three per-app files, so a successful compose install was
followed by doctor printing three failures and exiting 1 — and the whole
coherence catalog was skipped because it keyed off apps/sim/.env existing.
Layout is now derived from what's on disk and every check consults it: file and
schema checks iterate the layout's targets, consistency reports skip when
there's only one file to mirror, and coherence/live read the layout's primary
file. The wizard's existing-config detection counts root for the same reason —
a compose install used to read as unconfigured and re-run from scratch.
* feat(cli-auth): device-authorization poll flow, drop the loopback listener
The CLI no longer binds a local port. It generates a request id + poll secret,
opens /cli/auth, and polls /api/cli/auth/poll over TLS while the user approves
in the browser — so the flow works over SSH and inside containers, where the
browser and terminal don't share a machine.
- approve stores the approval keyed by request id (session-authed, userId from
the session only); poll verifies the secret before an atomic claim, so an
observer of the semi-public request id can neither mint nor cancel it
- pairing code stays as the anti-phishing compare; no key ever crosses the
browser; done page just confirms
- removes the loopback listener, /token exchange, buildCliHandoffUrl, and
validateCliCallbackUrl (+ its tests) — nothing hands a key to a URL anymore
* fix(setup): reuse an existing managed Postgres container instead of colliding
A running sim-postgres fell through to `docker run --name sim-postgres` and died
on the name conflict; a stopped one failed with "no DATABASE_URL to reach it"
because the generated password only lived in the env files a fresh clone lacks.
Both facts are recoverable from Docker: the ladder now reads the published port
and password back via `docker inspect` and reuses the container (starting it if
stopped). A container that won't answer prompts before recreating, and never
drops the data volume silently.
* improvement(setup): audience-first run-mode hints
Each run mode now names who it's for — compose for self-hosting/evaluating, dev
for contributing to Sim, k8s for rehearsing a production deploy — with the live
detection state (Docker/kube/VM) appended.
* fix(cli-auth): retry a failed mint, port container/port fixes to Redis + k8s
Review findings from #5911:
- poll now reserves the mint with an atomic NX lock instead of deleting the
approval up front, so a failed mint (e.g. mothership blip) is retried by the
next poll instead of forcing a fresh browser approval; the lock still prevents
a double-mint and its TTL frees the slot if the caller dies
- setup reuses/recreates an unhealthy managed sim-redis instead of colliding on
the name (Redis has no data volume, so it removes and recreates without a prompt)
- k8s failure-path hints carry --context, matching the success-path hints, so a
changed ambient context can't send diagnostics to the wrong cluster
- compose port-free waits for a killed port to actually release before
re-checking; SIGKILL is async, so the immediate re-check re-saw the port
* fix(setup): harden mint cleanup, Windows browser, container detection, helm cwd
Review findings from #5911:
- a post-mint completeApproval failure no longer routes into releaseMint — the
mint lock now outlives the approval (shared TTL), so a cleanup blip can't leave
a re-mintable window and orphan a key; cleanup is best-effort after the key ships
- compose doctor --fix writes the feature-flag twin to the layout's primary env
(root .env on a compose install), not always apps/sim/.env
- Windows opens the browser via `cmd /c start "" <url>` — `start` is a shell
builtin, so spawning it directly ENOENT'd and the handoff never opened
- managed-container detection filters loosely and pins the exact name in code;
Docker's `name=^x$` anchor matches the internal `/x` form and often missed,
skipping the reuse branch
- the shared helm/kind run helper pins cwd to the repo root, matching helm test,
so `helm upgrade --install ./helm/sim` works from any working directory
* feat(chat-keys): standalone manage page, drop from settings nav, refresh README
- Add /account/settings/chat-keys — a linkable page to view, create, and revoke Chat API keys
- Remove Chat keys from the settings sidebar (account + unified nav) and its render branches
- README: replace Docker Compose + Manual Setup with the bun run setup wizard; drop the manual COPILOT_API_KEY step, point to the manage page
* fix(setup): per-key reason in the secret-replacement warning
Cursor: the warn hardcoded '64-character hex key', but only ENCRYPTION_KEY/API_ENCRYPTION_KEY require that — BETTER_AUTH_SECRET/INTERNAL_API_SECRET only need length >= 32. Use the existing secretRequirement(key) helper so each replaced key reports its actual requirement.
* fix(setup): compose doctor schema, cross-platform binary detection, quoted context hints
- Doctor: for the compose (root) env layout, require only the secrets compose has no interpolation default for (BETTER_AUTH_SECRET/ENCRYPTION_KEY/INTERNAL_API_SECRET). DATABASE_URL/BETTER_AUTH_URL/NEXT_PUBLIC_APP_URL come from docker-compose ${VAR:-default}, so a healthy compose install no longer fails doctor.
- Binary detection: use Bun.which instead of which (which is absent on Windows), so kubectl/helm/kind/docker resolve cross-platform.
- k8s diagnostic hints: POSIX-quote the kube-context so a context with whitespace/metacharacters can't break or inject into a copied command.
* fix(setup): quote kube-context in the helm uninstall tear-down hint too
The tear-down hint used --kube-context ${context} raw while the sibling kubectl hints already used shq(); a context with whitespace/metacharacters could break or inject into the copied command. All copyable k8s hints now go through shq(context).
* feat(setup): sim lifecycle CLI — start/stop/status/logs/down/reset
Turn the setup entry into a 'sim' command umbrella so there's one place to run everything, not scattered docker/bun commands. Adds a global bin (bun link) + a bun run sim fallback.
- Detects how you're running (compose file / managed dev containers / helm release) from disk + docker/helm state — no persisted mode. Ambiguous installs prompt.
- start/stop/restart/logs work per mode; down removes containers (volumes kept); reset archives .env + wipes managed data; both destructive verbs confirm first.
- status shows detected mode, container states, and app/realtime health.
- Wizard outro + README now point at the sim commands and the one-time bun link.
* feat(setup): 'bun run sim' is the primary entry; bare invocation prints help
- Lead usage/wizard-outro/README with 'bun run sim <cmd>' (works with zero PATH setup); global bare 'sim' via bun link is an optional upgrade, with the ~/.bun/bin PATH caveat spelled out (Homebrew's bun omits it).
- Bare 'sim' now prints help instead of launching the wizard; the wizard is 'sim setup'. The 'setup' npm script passes the keyword so 'bun run setup' is unchanged.
* fix(setup): quote the auth URL for cmd /c start on Windows
Cursor (High): cmd re-parses the command line and treats & in the query string as a command separator, so cmd /c start opened a URL truncated at the first &, breaking the key flow on win32 (the handoff URL always has request/challenge/pairing). Quote the URL and pass args verbatim so & stays literal.
* fix(setup): verify kube-context is really local; lengthen CLI handoff wait
- k8s: a context named like a local cluster (kind-*, docker-desktop) can actually point at a remote API server. Verify the server host is loopback/docker-internal before defaulting the 'use this context?' confirm to yes; otherwise warn and default to no, so generated secrets can't ship to a remote cluster on a blind Enter.
- cli-auth: bump the device-flow wait from 3 to 15 minutes so first-time users have time to sign up, wait for the email OTP, and approve before the terminal stops polling. The server-side approval record keeps its own short TTL, so a longer client wait only costs cheap rate-limited polls.
* fix(setup): only manage k8s lifecycle on a verified-local context
Greptile: sim down/reset used the ambient kube-context, so switching context after setup could uninstall a same-named sim-dev release from the wrong cluster. Gate k8sInstall on the same locality check the wizard uses (API server is loopback/docker-internal) via a shared isLocalKubeContext helper — the wizard only ever deploys locally, so a remote current-context is never treated as a Sim install.
* fix(setup): doctor skips placeholder secrets when seeding; reset names its target
- checks: the missing-file autofix copied shared keys from apps/sim/.env whenever truthy, including .env.example placeholders — doctor --fix could seed unusable secrets into realtime/db env files. Skip placeholders, matching autofixForMissing.
- lifecycle: reset now names the exact install (k8s context / compose file / dev containers) in its confirm, so a destructive reset can't silently hit the wrong same-named install after a context switch (down already names the context).
* fix(cli-auth): size the poll rate limit to the poll cadence; honor Retry-After
The poll route used the default public-IP bucket (10 burst, 5/min) but the CLI polls every 2s (30/min), so it 429'd within ~20s — worse behind a slow dev cold-compile. Give the endpoint a bucket matched to its cadence (60 burst, 60/min); it's not a brute-force surface (unknown request id returns pending, minting needs the 256-bit verifier). Also make the CLI honor Retry-After and back off on 429 so a shared-NAT per-IP limit degrades gracefully instead of hammering.
* fix(setup): check ports before starting the dev server, not just compose
Local dev auto-start spawned bun run dev:full with no port check, so it silently started a server that couldn't bind when 3000/3002 were already taken (e.g. another worktree's dev server). Extract compose's port-conflict resolver into a shared ensurePortsFree(ports) and run it before the dev start too — kill/recheck/leave, same as compose. Leaving the ports skips the auto-start with guidance instead of failing; compose still treats it as fatal.
* fix(setup): verify the kube cluster is reachable, not just local
A kubeconfig context can outlive its cluster — a kind cluster gets deleted or its Docker container stops (Docker/machine restart), but the context entry remains, pointing at a dead API-server port. The wizard checked the context looked local and handed it to helm, which failed with 'cluster unreachable'.
Add a clusterReachable() liveness probe: only offer the current context when it actually answers; if a local context is dead, fall through to the kind path. There, if kind still knows 'sim' but it's stopped, start its node containers and wait for the API; if it's gone, create fresh. Either way the user gets a working cluster instead of a cryptic helm failure.
* fix(helm): point appVersion at published image tags (v-prefixed, current)
The chart's appVersion was "0.6.73", but CI publishes GHCR tags with a v prefix (its release-commit regex captures v0.7.45). Since sim.image defaults every image tag to Chart.AppVersion, a default helm install requested ghcr.io/simstudioai/{simstudio,realtime,migrations}:0.6.73 — a tag that has never existed — so app and realtime sat in ImagePullBackOff and helm --wait failed with 'progress deadline exceeded'. Any self-hoster installing with default values hit this, not just the setup wizard.
Set appVersion to v0.7.45 (latest release on main; all three images verified present on ghcr) and bump the chart version to 1.1.1. Verified with helm lint, helm template (all images render as v0.7.45), and a live helm upgrade on a kind cluster where the new pods pull successfully while the old 0.6.73 pods remain in ImagePullBackOff.
* Revert "fix(helm): point appVersion at published image tags (v-prefixed, current)"
This reverts commit 28b6047d1d.
* chore(api-validation): rebaseline route count to 977 after staging merge
Staging moved the baseline to 975; this branch's two CLI-auth routes (approve, poll) make 977. The clean merge absorbed the earlier +2 adjustment.
* fix(settings): don't highlight a sibling nav item on nested settings pages
/account/settings/chat-keys is a real page but deliberately not a nav item, so the sidebar's parseSettingsPathSection fell through to defaultSection ('general') and highlighted General — the page read as though it lived inside General.
Resolve the sidebar's active item with a null default so an unmatched nested route highlights nothing, and widen SettingsSidebar's activeSection to string | null. The section feeding the title/description provider keeps its default (pages override title/description anyway), and /account/settings/billing/credit-usage still correctly highlights Billing.
* fix(setup,auth): manage explicitly-confirmed k8s contexts, fail loudly on reset, clear stale post-auth redirect
- lifecycle: detection is now factual — a sim-dev release either exists on the current context or it doesn't. Gating on locality stranded a release the user explicitly confirmed during setup (status/start/stop/down/reset all claimed no k8s install). Locality is recorded instead and surfaced through describeInstall, which every destructive confirm renders, so acting on a non-local cluster is named and defaulted to no rather than silently blocked or silently allowed.
- lifecycle: reset no longer discards helm uninstall's exit status. Env files are archived by that point, so claiming 'Reset complete' while the release still runs is the worst outcome — it now throws with retry/inspect commands.
- auth: signup clears POST_AUTH_REDIRECT_STORAGE_KEY when it has no callbackUrl, and the verification-disabled path consumes it, so a stale CLI/invite destination can't leak into a later flow in the same tab.
This commit is contained in:
@@ -0,0 +1,118 @@
|
||||
import { spawnSync } from 'node:child_process'
|
||||
import type { Detection } from '../detect.ts'
|
||||
import { ensureDocker } from '../docker.ts'
|
||||
import { ROOT, readEnvFile, writeEnvValues } from '../env-files.ts'
|
||||
import { SetupError } from '../errors.ts'
|
||||
import { ensurePortsFree } from '../ports.ts'
|
||||
import { httpHealth, waitFor } from '../probes.ts'
|
||||
import * as p from '../prompter.ts'
|
||||
import {
|
||||
collectSecrets,
|
||||
promptCopilotKey,
|
||||
promptEmail,
|
||||
promptLlmKeys,
|
||||
promptSecurity,
|
||||
promptSignInProviders,
|
||||
promptStorage,
|
||||
promptUnlocks,
|
||||
} from '../steps.ts'
|
||||
import { glyph, theme } from '../theme.ts'
|
||||
|
||||
/**
|
||||
* Compose publishes 3000 and 3002 — resolve conflicts before touching docker,
|
||||
* instead of letting `docker compose up` die halfway through startup. Aborting
|
||||
* is fatal here: compose can't come up while the ports are held.
|
||||
*/
|
||||
async function ensureComposePortsFree(composeFile: string): Promise<void> {
|
||||
if (await ensurePortsFree([3000, 3002])) return
|
||||
throw new SetupError('ports 3000/3002 are in use', [
|
||||
`free the ports, then re-run: ${theme.command('bun run setup')}`,
|
||||
`see what holds them: ${theme.command('lsof -nP -iTCP:3000 -sTCP:LISTEN')}`,
|
||||
`stop a container publishing them: ${theme.command('docker ps')}`,
|
||||
`compose file in play: ${composeFile}`,
|
||||
])
|
||||
}
|
||||
|
||||
export async function runComposeMode(detection: Detection, quick: boolean): Promise<void> {
|
||||
await ensureDocker(true)
|
||||
const variant = quick
|
||||
? 'prod'
|
||||
: await p.select({
|
||||
message: 'Which images?',
|
||||
options: [
|
||||
{
|
||||
value: 'prod',
|
||||
label: 'Published images',
|
||||
hint: 'pulls ghcr.io/simstudioai/* — fastest',
|
||||
},
|
||||
{
|
||||
value: 'local',
|
||||
label: 'Build from source',
|
||||
hint: 'builds docker/*.Dockerfile — for testing local changes',
|
||||
},
|
||||
],
|
||||
initialValue: 'prod',
|
||||
})
|
||||
const composeFile = variant === 'prod' ? 'docker-compose.prod.yml' : 'docker-compose.local.yml'
|
||||
|
||||
const root = readEnvFile('root')
|
||||
const values = collectSecrets(root)
|
||||
const copilotKey = await promptCopilotKey(root.vars.get('COPILOT_API_KEY'))
|
||||
if (copilotKey) values.COPILOT_API_KEY = copilotKey
|
||||
Object.assign(values, await promptLlmKeys(detection, !quick))
|
||||
if (!quick) {
|
||||
const storage = await promptStorage(root.vars, true)
|
||||
if (storage) Object.assign(values, storage)
|
||||
const appUrl = root.vars.get('NEXT_PUBLIC_APP_URL') ?? 'http://localhost:3000'
|
||||
Object.assign(values, await promptSignInProviders(root.vars, appUrl))
|
||||
Object.assign(values, await promptEmail(root.vars))
|
||||
const security = await promptSecurity(root.vars)
|
||||
Object.assign(values, security.sim, security.mirrorToRealtime)
|
||||
Object.assign(values, await promptUnlocks(root.vars))
|
||||
}
|
||||
if (!root.vars.get('LOG_LEVEL')) {
|
||||
values.LOG_LEVEL = 'INFO'
|
||||
p.log.step(
|
||||
'Set LOG_LEVEL=INFO (production containers default to ERROR, which hides startup problems)'
|
||||
)
|
||||
}
|
||||
if (!root.vars.get('NEXT_TELEMETRY_DISABLED')) values.NEXT_TELEMETRY_DISABLED = '1'
|
||||
writeEnvValues('root', values)
|
||||
p.log.step('Wrote .env (compose reads it for variable substitution)')
|
||||
|
||||
await ensureComposePortsFree(composeFile)
|
||||
|
||||
p.log.step(`Running docker compose -f ${composeFile} up -d`)
|
||||
const result = spawnSync('docker', ['compose', '-f', composeFile, 'up', '-d'], {
|
||||
cwd: ROOT,
|
||||
stdio: 'inherit',
|
||||
})
|
||||
if (result.status !== 0) {
|
||||
throw new SetupError(`docker compose exited with ${result.status}.`, [
|
||||
`inspect what failed: ${theme.command(`docker compose -f ${composeFile} logs --tail 50`)}`,
|
||||
`container status: ${theme.command(`docker compose -f ${composeFile} ps`)}`,
|
||||
`clean slate: ${theme.command(`docker compose -f ${composeFile} down`)} then re-run the wizard`,
|
||||
])
|
||||
}
|
||||
|
||||
const spin = p.spinner()
|
||||
spin.start('Waiting for Sim to come up (first run pulls images and migrates)…')
|
||||
const appHealthy = await waitFor(
|
||||
() => httpHealth('http://localhost:3000/api/health'),
|
||||
300_000,
|
||||
3000
|
||||
)
|
||||
const realtimeHealthy =
|
||||
appHealthy && (await waitFor(() => httpHealth('http://localhost:3002/health'), 60_000, 2000))
|
||||
if (!appHealthy || !realtimeHealthy) {
|
||||
spin.stop(`${glyph.fail} services did not become healthy`)
|
||||
throw new SetupError(
|
||||
`${!appHealthy ? 'the app (:3000)' : 'realtime (:3002)'} never answered its health check.`,
|
||||
[
|
||||
`follow the logs: ${theme.command(`docker compose -f ${composeFile} logs -f`)}`,
|
||||
'first boots on slow disks can exceed the wait — if containers are still starting, just wait and open http://localhost:3000',
|
||||
]
|
||||
)
|
||||
}
|
||||
spin.stop('App and realtime are healthy')
|
||||
}
|
||||
@@ -0,0 +1,171 @@
|
||||
import { spawnSync } from 'node:child_process'
|
||||
import path from 'node:path'
|
||||
import { truncate } from '@sim/utils/string'
|
||||
import { resolveDatabase } from '../db.ts'
|
||||
import type { Detection } from '../detect.ts'
|
||||
import { ROOT, readEnvFile, writeEnvValues } from '../env-files.ts'
|
||||
import { SetupError } from '../errors.ts'
|
||||
import { pgProbe } from '../probes.ts'
|
||||
import * as p from '../prompter.ts'
|
||||
import { resolveRedis } from '../redis.ts'
|
||||
import {
|
||||
collectSecrets,
|
||||
promptCopilotKey,
|
||||
promptEmail,
|
||||
promptLlmKeys,
|
||||
promptSecurity,
|
||||
promptSignInProviders,
|
||||
promptStorage,
|
||||
promptUnlocks,
|
||||
} from '../steps.ts'
|
||||
import { glyph, theme } from '../theme.ts'
|
||||
|
||||
const APP_URL = 'http://localhost:3000'
|
||||
|
||||
/**
|
||||
* A migrate failure on a never-migrated database means setup failed — abort.
|
||||
* On a database that already has applied migrations (a live but drifted dev
|
||||
* DB), the failure is surfaced and the user decides whether to continue.
|
||||
*/
|
||||
async function runMigrations(dsn: string): Promise<void> {
|
||||
const spin = p.spinner()
|
||||
spin.start('Running database migrations…')
|
||||
const result = spawnSync('bun', ['run', 'db:migrate'], {
|
||||
cwd: path.join(ROOT, 'packages/db'),
|
||||
encoding: 'utf8',
|
||||
})
|
||||
if (result.status === 0) {
|
||||
spin.stop('Migrations applied')
|
||||
return
|
||||
}
|
||||
spin.stop(`${glyph.fail} migrations failed`)
|
||||
const error = truncate(`${result.stdout}\n${result.stderr}`.trim(), 2000)
|
||||
const probe = await pgProbe(dsn)
|
||||
const applied = probe.ok ? (probe.migrations?.applied ?? 0) : 0
|
||||
if (applied === 0) {
|
||||
throw new SetupError(`db:migrate failed on a fresh database:\n${error}`, [
|
||||
`run it by hand to see the full output: ${theme.command('cd packages/db && bun run db:migrate')}`,
|
||||
'check DATABASE_URL points at the database you expect',
|
||||
])
|
||||
}
|
||||
p.log.warn(
|
||||
`db:migrate failed, but this database already has ${applied} applied migrations — it may have schema drift (e.g. built with db:push).`
|
||||
)
|
||||
p.log.info(theme.muted(truncate(error, 600)))
|
||||
const proceed = await p.confirm({
|
||||
message: 'Continue setup without migrating? (doctor will keep flagging the drift)',
|
||||
initialValue: true,
|
||||
})
|
||||
if (!proceed) throw new Error(`aborted: db:migrate failed:\n${error}`)
|
||||
}
|
||||
|
||||
async function promptRedis(detection: Detection, existing?: string): Promise<string | null> {
|
||||
const wants = await p.confirm({
|
||||
message:
|
||||
'Configure Redis? (only needed for multi-replica — single instance runs fine without it)',
|
||||
initialValue: Boolean(existing),
|
||||
})
|
||||
if (!wants) return null
|
||||
return resolveRedis(detection, existing)
|
||||
}
|
||||
|
||||
async function promptTrigger(): Promise<Record<string, string> | null> {
|
||||
const wants = await p.confirm({
|
||||
message: 'Enable Trigger.dev for background jobs? (off = jobs run via the DB queue)',
|
||||
initialValue: false,
|
||||
})
|
||||
if (!wants) return null
|
||||
const secretKey = await p.password({
|
||||
message: 'TRIGGER_SECRET_KEY',
|
||||
validate: (v) => (v ? undefined : 'required'),
|
||||
})
|
||||
const projectId = await p.text({
|
||||
message: 'TRIGGER_PROJECT_ID',
|
||||
validate: (v) => (v ? undefined : 'required'),
|
||||
})
|
||||
return {
|
||||
TRIGGER_DEV_ENABLED: 'true',
|
||||
TRIGGER_SECRET_KEY: secretKey,
|
||||
TRIGGER_PROJECT_ID: projectId,
|
||||
}
|
||||
}
|
||||
|
||||
export async function runDevMode(
|
||||
detection: Detection,
|
||||
quick: boolean
|
||||
): Promise<{ startNow: boolean; script: string }> {
|
||||
const sim = readEnvFile('sim')
|
||||
const dsn = await resolveDatabase(detection, sim.vars.get('DATABASE_URL'))
|
||||
const secrets = collectSecrets(sim)
|
||||
|
||||
const shared = {
|
||||
DATABASE_URL: dsn,
|
||||
BETTER_AUTH_SECRET: secrets.BETTER_AUTH_SECRET,
|
||||
INTERNAL_API_SECRET: secrets.INTERNAL_API_SECRET,
|
||||
BETTER_AUTH_URL: APP_URL,
|
||||
NEXT_PUBLIC_APP_URL: APP_URL,
|
||||
}
|
||||
writeEnvValues('sim', {
|
||||
...shared,
|
||||
ENCRYPTION_KEY: secrets.ENCRYPTION_KEY,
|
||||
API_ENCRYPTION_KEY: secrets.API_ENCRYPTION_KEY,
|
||||
})
|
||||
writeEnvValues('realtime', shared)
|
||||
writeEnvValues('db', { DATABASE_URL: dsn })
|
||||
p.log.step('Wrote apps/sim/.env, apps/realtime/.env, packages/db/.env (shared subset mirrored)')
|
||||
await runMigrations(dsn)
|
||||
|
||||
const simAfter = readEnvFile('sim')
|
||||
const values: Record<string, string> = {}
|
||||
const copilotKey = await promptCopilotKey(simAfter.vars.get('COPILOT_API_KEY'))
|
||||
if (copilotKey) values.COPILOT_API_KEY = copilotKey
|
||||
Object.assign(values, await promptLlmKeys(detection, !quick))
|
||||
|
||||
if (!quick) {
|
||||
const redisUrl = await promptRedis(detection, simAfter.vars.get('REDIS_URL'))
|
||||
if (redisUrl) {
|
||||
values.REDIS_URL = redisUrl
|
||||
writeEnvValues('realtime', { REDIS_URL: redisUrl })
|
||||
}
|
||||
const trigger = await promptTrigger()
|
||||
if (trigger) Object.assign(values, trigger)
|
||||
const storage = await promptStorage(simAfter.vars, false)
|
||||
if (storage) Object.assign(values, storage)
|
||||
Object.assign(values, await promptSignInProviders(simAfter.vars, APP_URL))
|
||||
Object.assign(values, await promptEmail(simAfter.vars))
|
||||
const security = await promptSecurity(simAfter.vars)
|
||||
Object.assign(values, security.sim)
|
||||
if (Object.keys(security.mirrorToRealtime).length > 0) {
|
||||
writeEnvValues('realtime', security.mirrorToRealtime)
|
||||
}
|
||||
Object.assign(values, await promptUnlocks(simAfter.vars))
|
||||
}
|
||||
if (Object.keys(values).length > 0) writeEnvValues('sim', values)
|
||||
|
||||
let script = 'dev:full'
|
||||
if (detection.specs.hostMemGb < 16) {
|
||||
script = await p.select({
|
||||
message: `Low RAM detected (${detection.specs.hostMemGb}GB) — which dev server?`,
|
||||
options: [
|
||||
{
|
||||
value: 'dev:full:minimal-registry',
|
||||
label: 'Minimal block registry (recommended)',
|
||||
hint: 'much lower memory — loads fewer integration blocks in dev',
|
||||
},
|
||||
{
|
||||
value: 'dev:full',
|
||||
label: 'Full registry',
|
||||
hint: 'every block available — can use 4-5GB+ on its own',
|
||||
},
|
||||
],
|
||||
initialValue: 'dev:full:minimal-registry',
|
||||
})
|
||||
}
|
||||
return {
|
||||
startNow: await p.confirm({
|
||||
message: `Start Sim now? (bun run ${script})`,
|
||||
initialValue: true,
|
||||
}),
|
||||
script,
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,298 @@
|
||||
import { spawnSync } from 'node:child_process'
|
||||
import { getErrorMessage } from '@sim/utils/errors'
|
||||
import type { Detection } from '../detect.ts'
|
||||
import { ensureDocker } from '../docker.ts'
|
||||
import { generateSecret, ROOT } from '../env-files.ts'
|
||||
import { SetupError } from '../errors.ts'
|
||||
import { waitFor } from '../probes.ts'
|
||||
import * as p from '../prompter.ts'
|
||||
import { glyph, theme } from '../theme.ts'
|
||||
|
||||
const RELEASE = 'sim-dev'
|
||||
const NAMESPACE = 'sim-dev'
|
||||
const LOCAL_CONTEXT_PREFIXES = ['kind-', 'docker-desktop', 'minikube', 'orbstack']
|
||||
|
||||
/**
|
||||
* `input` is piped on stdin rather than passed as arguments — argv is readable
|
||||
* by any process on the machine, so secrets must never travel that way.
|
||||
*/
|
||||
function run(command: string, args: string[], failMessage: string, input?: string): string {
|
||||
// `helm upgrade --install ./helm/sim` uses chart paths relative to the repo
|
||||
// root, so pin cwd regardless of where the wizard was invoked from.
|
||||
const result = spawnSync(command, args, { encoding: 'utf8', input, cwd: ROOT })
|
||||
if (result.status !== 0) {
|
||||
throw new Error(`${failMessage}: ${result.stderr.trim() || result.stdout.trim()}`)
|
||||
}
|
||||
return result.stdout
|
||||
}
|
||||
|
||||
function isLocalContext(context: string): boolean {
|
||||
return LOCAL_CONTEXT_PREFIXES.some((prefix) => context === prefix || context.startsWith(prefix))
|
||||
}
|
||||
|
||||
const LOCAL_SERVER_HOSTS = new Set([
|
||||
'127.0.0.1',
|
||||
'localhost',
|
||||
'0.0.0.0',
|
||||
'::1',
|
||||
'kubernetes.docker.internal',
|
||||
'host.docker.internal',
|
||||
])
|
||||
|
||||
/** True when a context's API server is a loopback/host address — i.e. a local cluster. */
|
||||
export function isLocalKubeContext(context: string): boolean {
|
||||
const server = contextServerHost(context)
|
||||
return server !== null && LOCAL_SERVER_HOSTS.has(server)
|
||||
}
|
||||
|
||||
/**
|
||||
* Liveness probe — a kubeconfig entry can outlive a stopped or deleted cluster
|
||||
* (kind clusters are Docker containers that don't restart on their own), so a
|
||||
* context looking local is no guarantee its API server answers.
|
||||
*/
|
||||
function clusterReachable(context: string): boolean {
|
||||
return (
|
||||
spawnSync('kubectl', ['cluster-info', '--context', context, '--request-timeout=5s'], {
|
||||
stdio: 'ignore',
|
||||
}).status === 0
|
||||
)
|
||||
}
|
||||
|
||||
/** The API server host a context points at, or null if kubectl can't resolve it. */
|
||||
function contextServerHost(context: string): string | null {
|
||||
const result = spawnSync(
|
||||
'kubectl',
|
||||
[
|
||||
'config',
|
||||
'view',
|
||||
'--minify',
|
||||
'--context',
|
||||
context,
|
||||
'-o',
|
||||
'jsonpath={.clusters[0].cluster.server}',
|
||||
],
|
||||
{ encoding: 'utf8' }
|
||||
)
|
||||
if (result.status !== 0) return null
|
||||
try {
|
||||
return new URL(result.stdout.trim()).hostname
|
||||
} catch {
|
||||
return null
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* POSIX-quote a value for a copyable shell hint. A kube-context passes only a
|
||||
* prefix check, so it can still contain whitespace or shell metacharacters that
|
||||
* would break the `--context` argument (or run embedded syntax) when copied.
|
||||
* Ordinary context names stay bare; only unsafe ones get single-quoted.
|
||||
*/
|
||||
function shq(value: string): string {
|
||||
if (/^[A-Za-z0-9._/-]+$/.test(value)) return value
|
||||
return `'${value.replace(/'/g, `'\\''`)}'`
|
||||
}
|
||||
|
||||
async function ensureLocalContext(detection: Detection): Promise<string> {
|
||||
if (!detection.binaries.helm || !detection.binaries.kubectl) {
|
||||
throw new SetupError('kubernetes mode needs kubectl and helm on PATH.', [
|
||||
`install them: ${theme.command('brew install kubectl helm')}`,
|
||||
])
|
||||
}
|
||||
const context = detection.kubeContext
|
||||
if (context && isLocalContext(context)) {
|
||||
// The name is only a hint — a remote cluster can be named like a local one
|
||||
// (e.g. "kind-prod"). Verify the API server is a loopback/host address before
|
||||
// defaulting to "yes", so generated secrets can't silently ship to a remote
|
||||
// cluster on a blind Enter.
|
||||
const server = contextServerHost(context)
|
||||
if (server && LOCAL_SERVER_HOSTS.has(server)) {
|
||||
if (!clusterReachable(context)) {
|
||||
// The context is local but its cluster isn't answering — stopped or
|
||||
// deleted. Don't offer it (helm would just fail); fall through to the
|
||||
// kind path, which starts a stopped "sim" cluster or creates one.
|
||||
p.log.warn(
|
||||
`Context "${context}" points at a local cluster that isn't responding — it looks stopped or deleted. The wizard will start or recreate a kind cluster instead.`
|
||||
)
|
||||
} else {
|
||||
const useIt = await p.confirm({
|
||||
message: `Use current kube context "${context}"?`,
|
||||
initialValue: true,
|
||||
})
|
||||
if (useIt) return context
|
||||
}
|
||||
} else {
|
||||
p.log.warn(
|
||||
`Context "${context}" is named like a local cluster, but its API server${server ? ` (${server})` : ''} does not look local. Continuing would deploy the generated secrets there.`
|
||||
)
|
||||
const useIt = await p.confirm({
|
||||
message: `Deploy to "${context}" anyway?`,
|
||||
initialValue: false,
|
||||
})
|
||||
if (useIt) return context
|
||||
}
|
||||
} else if (context) {
|
||||
p.log.warn(
|
||||
`Current context "${context}" does not look like a local cluster. Deploying to remote clusters is not supported by the wizard yet — switch to a kind/docker-desktop context, or drive helm directly (see helm/sim/examples/values-production.yaml).`
|
||||
)
|
||||
}
|
||||
if (!detection.binaries.kind) {
|
||||
throw new SetupError('no local cluster available.', [
|
||||
`install kind: ${theme.command('brew install kind')} — the wizard creates the cluster for you`,
|
||||
'or enable Kubernetes in Docker Desktop settings, then re-run',
|
||||
])
|
||||
}
|
||||
await ensureDocker(true)
|
||||
const clusters = run('kind', ['get', 'clusters'], 'kind get clusters failed')
|
||||
.trim()
|
||||
.split('\n')
|
||||
.filter(Boolean)
|
||||
if (clusters.includes('sim')) {
|
||||
run('kind', ['export', 'kubeconfig', '--name', 'sim'], 'kind export kubeconfig failed')
|
||||
if (clusterReachable('kind-sim')) {
|
||||
p.log.step('Reusing existing kind cluster "sim"')
|
||||
} else {
|
||||
// The cluster exists in kind but isn't answering — its node containers are
|
||||
// stopped (a Docker/machine restart). Start them and wait for the API.
|
||||
const spin = p.spinner()
|
||||
spin.start('kind cluster "sim" is stopped — starting it…')
|
||||
const nodes = run('kind', ['get', 'nodes', '--name', 'sim'], 'kind get nodes failed')
|
||||
.trim()
|
||||
.split('\n')
|
||||
.filter(Boolean)
|
||||
for (const node of nodes) spawnSync('docker', ['start', node], { stdio: 'ignore' })
|
||||
const up = await waitFor(() => Promise.resolve(clusterReachable('kind-sim')), 60_000, 2000)
|
||||
if (!up) {
|
||||
spin.stop(`${glyph.fail} kind cluster "sim" would not start`)
|
||||
throw new SetupError('the kind cluster "sim" exists but will not come up.', [
|
||||
`inspect it: ${theme.command('docker ps -a --filter name=sim-control-plane')}`,
|
||||
`recreate it: ${theme.command('kind delete cluster --name sim')}, then re-run ${theme.command('bun run setup')}`,
|
||||
])
|
||||
}
|
||||
spin.stop('kind cluster "sim" started')
|
||||
}
|
||||
} else {
|
||||
const spin = p.spinner()
|
||||
spin.start('Creating kind cluster "sim"…')
|
||||
run('kind', ['create', 'cluster', '--name', 'sim'], 'kind create cluster failed')
|
||||
spin.stop('kind cluster "sim" ready')
|
||||
}
|
||||
return 'kind-sim'
|
||||
}
|
||||
|
||||
function existingReleaseSecrets(context: string): Record<string, string> | null {
|
||||
const scope = ['--kube-context', context, '-n', NAMESPACE]
|
||||
const status = spawnSync('helm', ['status', RELEASE, ...scope], { stdio: 'ignore' })
|
||||
if (status.status !== 0) return null
|
||||
const values = JSON.parse(
|
||||
run('helm', ['get', 'values', RELEASE, ...scope, '-o', 'json'], 'helm get values failed')
|
||||
) as { app?: { env?: Record<string, string> }; postgresql?: { auth?: { password?: string } } }
|
||||
const env = values.app?.env ?? {}
|
||||
const password = values.postgresql?.auth?.password
|
||||
if (
|
||||
!env.BETTER_AUTH_SECRET ||
|
||||
!env.ENCRYPTION_KEY ||
|
||||
!env.INTERNAL_API_SECRET ||
|
||||
!env.CRON_SECRET ||
|
||||
!password
|
||||
) {
|
||||
return null
|
||||
}
|
||||
return {
|
||||
BETTER_AUTH_SECRET: env.BETTER_AUTH_SECRET,
|
||||
ENCRYPTION_KEY: env.ENCRYPTION_KEY,
|
||||
INTERNAL_API_SECRET: env.INTERNAL_API_SECRET,
|
||||
CRON_SECRET: env.CRON_SECRET,
|
||||
POSTGRES_PASSWORD: password,
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Values document piped to helm on stdin instead of `--set`. `JSON.stringify`
|
||||
* quotes and escapes each value — JSON is a subset of YAML, so a secret
|
||||
* containing `#`, `:`, or a leading `*` can neither break the document nor be
|
||||
* reinterpreted as YAML syntax.
|
||||
*/
|
||||
function secretValues(secrets: Record<string, string>): string {
|
||||
const { POSTGRES_PASSWORD, ...appEnv } = secrets
|
||||
const env = Object.entries(appEnv)
|
||||
.map(([key, value]) => ` ${key}: ${JSON.stringify(value)}`)
|
||||
.join('\n')
|
||||
return `app:\n env:\n${env}\npostgresql:\n auth:\n password: ${JSON.stringify(POSTGRES_PASSWORD)}\n`
|
||||
}
|
||||
|
||||
export async function runK8sMode(detection: Detection): Promise<void> {
|
||||
// Pin every subsequent call to the context we validated: the ambient context
|
||||
// can change between detection and deploy, which would send generated
|
||||
// credentials to an unintended cluster.
|
||||
const context = await ensureLocalContext(detection)
|
||||
|
||||
const reused = existingReleaseSecrets(context)
|
||||
const secrets = reused ?? {
|
||||
BETTER_AUTH_SECRET: generateSecret(),
|
||||
ENCRYPTION_KEY: generateSecret(),
|
||||
INTERNAL_API_SECRET: generateSecret(),
|
||||
CRON_SECRET: generateSecret(),
|
||||
POSTGRES_PASSWORD: generateSecret().slice(0, 24),
|
||||
}
|
||||
if (reused) p.log.step('Reusing secrets from the existing release')
|
||||
|
||||
const spin = p.spinner()
|
||||
spin.start('helm upgrade --install (first run pulls images — this can take several minutes)…')
|
||||
try {
|
||||
run(
|
||||
'helm',
|
||||
[
|
||||
'upgrade',
|
||||
'--install',
|
||||
RELEASE,
|
||||
'./helm/sim',
|
||||
'--kube-context',
|
||||
context,
|
||||
'--namespace',
|
||||
NAMESPACE,
|
||||
'--create-namespace',
|
||||
'--values',
|
||||
'./helm/sim/examples/values-development.yaml',
|
||||
'--values',
|
||||
'-',
|
||||
'--wait',
|
||||
'--timeout',
|
||||
'15m',
|
||||
],
|
||||
'helm upgrade --install failed',
|
||||
secretValues(secrets)
|
||||
)
|
||||
} catch (error) {
|
||||
spin.stop(`${glyph.fail} helm install failed`)
|
||||
throw new SetupError(getErrorMessage(error), [
|
||||
`pod status: ${theme.command(`kubectl --context ${shq(context)} -n ${NAMESPACE} get pods`)}`,
|
||||
`stuck pods: ${theme.command(`kubectl --context ${shq(context)} -n ${NAMESPACE} describe pod <name> | tail -20`)}`,
|
||||
'ImagePullBackOff on ghcr.io/simstudioai/* usually means the chart appVersion tag was never published — check Chart.yaml against ghcr',
|
||||
])
|
||||
}
|
||||
spin.stop('Release deployed, all pods ready')
|
||||
|
||||
const testSpin = p.spinner()
|
||||
testSpin.start('Running helm test…')
|
||||
const test = spawnSync('helm', ['test', RELEASE, '--kube-context', context, '-n', NAMESPACE], {
|
||||
encoding: 'utf8',
|
||||
cwd: ROOT,
|
||||
})
|
||||
if (test.status !== 0) {
|
||||
testSpin.stop(`${glyph.fail} helm test failed`)
|
||||
throw new SetupError(`helm test failed:\n${test.stdout}${test.stderr}`, [
|
||||
`pod status: ${theme.command(`kubectl --context ${shq(context)} -n ${NAMESPACE} get pods`)}`,
|
||||
`app logs: ${theme.command(`kubectl --context ${shq(context)} -n ${NAMESPACE} logs deploy/${RELEASE}-app --tail 50`)}`,
|
||||
])
|
||||
}
|
||||
testSpin.stop('helm test passed')
|
||||
|
||||
p.note(
|
||||
[
|
||||
`kubectl --context ${shq(context)} -n ${NAMESPACE} port-forward svc/${RELEASE}-app 3000:3000`,
|
||||
`kubectl --context ${shq(context)} -n ${NAMESPACE} get pods`,
|
||||
`helm uninstall ${RELEASE} --kube-context ${shq(context)} -n ${NAMESPACE} # tear down`,
|
||||
].join('\n'),
|
||||
'Reach your cluster'
|
||||
)
|
||||
}
|
||||
Reference in New Issue
Block a user