mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
8b199d719964d58f91073edbfeddf12e216bb80a
5677
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
8b199d7199 |
fix(files): serve rendered documents and stream workspace archives (#5995)
* fix(files): serve rendered documents instead of generator source Generated docs (docx/pptx/pdf/xlsx) store their generation source as the primary file; the rendered binary lives in a separate content-addressed artifact store. resolveServableDocBytes is the chokepoint that swaps one for the other, and the serve route, single-file download and ~50 tool routes all go through it. Archive compression and the public v1 download read raw bytes instead, so a generated document arrived as source text under a .docx name and Word reported it as corrupt. Both now resolve through the servable reader. Adds fetchServableWorkspaceFileBuffer beside the raw reader so the record to UserFile mapping lives in one place, preserving storageContext; the raw reader's doc comment now says it returns generation source, so the next call site has to choose deliberately. v1 also has to send the resolved content type rather than the record's source MIME, and returns a retryable 409 rather than a 500 when an artifact is still compiling. docNotReadyMessage centralizes the 409 copy, and isRenderableDocumentName is shared so the read path and its callers agree on which extensions can expand. * perf(files): stream workspace archives instead of buffering them The bulk download route materialized every selected file before writing a byte, so peak memory tracked the size of the selection. Ordinary files are now appended as lazy streams: each opens its storage read only when the archiver reaches it, so one entry is resident at a time rather than the whole archive. Generated documents still resolve to buffers first. They are the only entries whose bytes decide anything, and every status this route returns comes from them — once the first byte is written the status code is committed, so those decisions have to happen before the archive starts. The per-entry allowance and byte budget therefore govern documents only. archiver processes appended entries through a sequential queue; lazystream defers each storage read until that entry's turn, since handing the archiver an open stream per entry would hold more connections than the storage client pools. nodeReadableToWebStream moves out of input-validation.server.ts into a shared util rather than being written twice: Readable.toWeb throws ERR_INVALID_STATE when a consumer cancels while the source is still flowing, which is exactly what a cancelled download does. Trade: a storage read failing mid-archive truncates the response rather than returning 500, since the status is already sent. Documents cannot hit this. The table export route already behaves this way. * improvement(files): surface archive download errors in place Bulk and folder downloads navigated to the API route, so any rejection replaced the Files page with the raw JSON error body. Single-file download already fetched and saved the blob, so the two paths had diverged. That was survivable when the only failures were "too many files" and "too large"; resolving rendered documents adds a 409 for a still-compiling artifact, which is reachable in normal use. Both archive paths now fetch the zip and show the server's message as a toast — the route writes that copy for a person, so it is worth surfacing rather than discarding. Single-file download shows its error too instead of only logging it. * improvement(files): simplify the archive route and stop buffering real uploads Four review passes over the branch converged on the same points. Routing every office extension through the buffered path held an entire selection of genuinely uploaded documents in memory — for a check the resolver settles on the first few magic bytes. Only files stored under a generator-source content type need resolving; a real .docx serves what is stored and now streams like anything else, which is what the change was for. With resolution sequential, the AbortController cancelled nothing (its signal never reached the storage read), the success-path over-limit branch was unreachable, and the cancellation guard in the catch could not fire. A plain loop that returns at the point of failure replaces the outcome record, the sentinel, the fan-out helper at limit 1, and four post-hoc scans. ZIP_MATERIALIZE_CONCURRENCY was dead, and the aggregate over-limit message reported a byte count that was by construction under the limit. lazystream and its types are gone: Readable.from over an async generator defers the open the same way and propagates source errors natively, so the PassThrough relay went with them. The client uses requestRaw with the existing binary contract instead of a hand-built query string and a re-typed error extractor, and both archive call sites share one callback. The render headroom moves beside isRenderableDocumentName, the servable reader stops minting a request id per file, and the v1 download returns a view rather than a second full copy. * improvement(files): keep the lazy entry stream in byte mode Readable.from defaults to object mode; these chunks are bytes headed for an archive, so the mode is now explicit. Also makes the fire-and-forget bulk download call consistent with its sibling. * improvement(files): correct the servable reader's 409 note Batch callers build the response from docNotReadyMessage rather than through docNotReadyResponse, so the doc comment now describes the outcome instead of naming one of the two helpers. * improvement(files): correct two comments that described the wrong behavior The resolution loop's comment claimed at most one rendered document is resident. It is not: every resolved buffer is held until the archive is assembled, bounded by the request's remaining budget. The loop resolves one at a time, it does not release one at a time. RENDERED_DOCUMENT_HEADROOM_BYTES was documented as bounding the expansion beyond the declared size, but it is passed straight through as maxBytes and is an absolute ceiling on the rendered document — renamed to MAX_RENDERED_DOCUMENT_BYTES so the name matches. Download failures also carry a fallback message, so a transport error surfaces something better than a bare "Failed to fetch". * fix(files): put streamed files and rendered documents on one byte budget The budget only counted rendered documents, so streamed entries never reserved their declared bytes. A mixed selection could clear the up-front declared-size gate and still ship close to two full limits: ordinary files up to the cap, plus documents drawing a fresh cap of their own. Streamed entries ship exactly what they declared, so their share is known before anything is read and is now reserved up front; documents draw from what is left. The budget-exhausted rejection also quoted declared sizes, which for a selection of small sources reads as a tiny total exceeding the limit. That case has no knowable byte count — the documents have not been rendered — so it now says what happened without inventing a number. * fix(files): attach the download anchor and handle a finalize rejection A detached anchor's click() works in current browsers, but every other download helper in the app attaches first, and a silent no-op on this path would look exactly like a download that never started. Not worth the ambiguity for two DOM operations. archive.finalize() was fired with void, so a failure after the response had started became an unhandled rejection on top of the stream error event that already fails the response. It is caught and logged now. * fix(files): guard the archive download against concurrent clicks Navigating to the route meant a second click just re-navigated. Fetching the archive instead means each click starts another download that holds the whole zip in tab memory, and the Download button gave no sign anything was happening. The button now reflects the in-flight download alongside the existing move and delete states. The ref guard is there as well as the state because two clicks in the same tick would both pass a state check. * fix(files): resolve every generated-document source type, not four of five The archive keyed off a locally enumerated set of source content types that omitted text/x-pdflibjs — the isolated-vm PDF generator. Those files failed the check and streamed their generator source under a .pdf name, which is the exact corruption this change exists to fix. A canonical five-member set already existed in the file viewer; enumerating a fourth copy was the mistake, not the missing string. The set moves to file-utils beside the other file-type predicates, the viewer imports it, and the route asks isGeneratedDocumentSourceType rather than carrying its own list. A parameterized test pins all five, so adding a generator without extending the set fails. * test(files): validate the produced archive with a real zip reader Every existing test mocks the storage stream and reads the result back with the same JS library that wrote it, which cannot catch the failure this route exists to fix: an archive a real zip reader will not open. This drives 5 MB of non-repeating bytes from an fs stream through archiver, the lazy entry generator and the Node-to-web bridge, then checks the output with the operating system's unzip and compares the extracted bytes. Skips rather than fails where unzip is unavailable. * fix(files): bound the markdown export's embedded-asset bundling The embed list is produced by scanning the document body, so its length and the bytes behind it are whatever the author put there. Every referenced asset was downloaded at once with no concurrency bound, no per-file cap and no total cap, so one request could materialize an unbounded number of unbounded files. Its sibling bulk-download route already had a count cap and a byte cap. Metadata is now resolved first, which bounds the download from declared sizes before a byte is read and moves the authorization check off the download path. Assets are then fetched with bounded concurrency and a per-file cap; a single unreadable or oversized asset drops out of the bundle rather than failing the export, which is what the previous allSettled did. * fix(files): mount rendered documents into the sandbox, not generator source Mounting a generated document handed the sandbox its generator source under a .docx name, so a python-docx or openpyxl script failed on a file that looked fine and the agent debugged code that was correct. Both branches were wrong, not just the buffered one: with cloud storage the mount presigns record.key, which is the raw source object, so swapping the buffered read alone would have fixed only local storage. Source-backed documents now always take the servable path — they are bounded by the render ceiling, so routing them through the web process instead of presigning is affordable. The text/binary decision also keyed off record.type, which for these files is text/x-docxjs, so a resolved binary would have been decoded as UTF-8. It reads the resolved content type now. * chore(deps): regenerate the lockfile against staging's dependency rework Staging reworked the OTel split, dropped dead deps and declared emcn peers, so the lockfile is regenerated on top of that rather than carrying a stale merge. The archive dependency is the only addition; lazystream stays as archiver's transitive dep now that it is no longer declared directly. Regenerated incrementally rather than from scratch: a clean resolve fails the bunfig age gate because minimumReleaseAgeExcludes lists only the darwin and linux-gnu @next/swc binaries, not the musl and windows variants that a full re-resolve also pulls. * fix(files): budget sandbox mounts on rendered bytes, not declared source size A source-backed document declares the size of its generator, not of the document, so the per-file and aggregate mount pre-checks were reading a number unrelated to what was about to be mounted — tiny sources cleared the guard and only afterwards was the running total updated with the real length. Those pre-checks now apply only to files that serve what they declare. A document is capped by the read instead, at the smaller of the per-file limit and the budget left, and a size rejection is reported in the same terms as the other mount limits. Same invariant the archive route already had to learn: declared size is not a bound once bytes can be rendered. * test(files): cover the two surfaces added without tests The markdown export route had no test file at all, and the sandbox mount's source-backed branch was exercised by nothing — both were changed in this PR and one of them shipped a defect a reviewer caught within minutes. Export: the embed-count rejection, the declared-bytes rejection landing before any asset leaves storage, the per-asset download cap, an unreadable asset dropping out rather than failing the export, and an unauthorized asset never being read. Mount: a generated document never presigning its raw key even on cloud storage, its rendered bytes mounting as base64 rather than being utf-8 decoded off the source MIME, and the budget rejecting on rendered length. * fix(files): stop the export caps from rejecting documents that used to work The caps I picked would have failed exports that previously succeeded: a screenshot-heavy document can hold more than 100 embeds, and 100 embeds of unoptimised PNGs can exceed 100 MB. Adding a limit to bound memory is right; setting it below real usage turns a memory risk into a broken feature. The byte ceiling now matches the bulk-download route, so the two export surfaces reject at the same size rather than at two invented ones. The count is only a guard on the metadata lookups that run before the byte check, so it moves well above any hand-authored document instead of sitting where a legitimate one could reach it. |
||
|
|
6aa3e0e29b |
fix(logs): match copilot log-grep patterns with RE2 (linear time) (#5987)
* fix(logs): match copilot log-grep patterns literally, never as a regex
`grepSpans` compiled a caller-supplied pattern with a bare `new RegExp`, whose
only guard was a `catch` for syntax errors. Trace text is attacker-influenced —
a workflow can emit arbitrarily long uniform runs into its own block outputs —
and matching runs synchronously on the shared event loop, so an authenticated
caller could choose both the pattern and the input and stall every request on
the instance.
Patterns are now matched as an escaped, case-insensitive literal. Screening was
implemented first and abandoned, because each rule only excludes the shapes
someone thought to enumerate:
- `safe-regex2` (already used by the guardrails validator) documents itself as
having false negatives. It screens star height only, so it passes `(a|a)*b`,
measured >61s on V8 at 30 characters.
- Rejecting quantified groups on top of that catches `(a|a)*b`, and every
attack `safe-regex2` catches, but still passes `a*a*b` — adjacent quantifiers
over overlapping character sets — measured 213s on JSC and 132s on V8 over a
10k-character run, well inside what a block output can hold.
An escaped literal is linear on every engine, needs no dependency, and does not
depend on which runtime serves the request. `safe-regex2` stays a dependency for
its existing guardrails/PII callers; it is simply not relied on here.
`query_logs` loses regex matching. Its catalog entry describes `pattern` as
"greps" rather than promising regex, but it is generated from a contract in
another repository and cannot be updated here, so a `patternNotice` is returned
whenever a pattern contains regex syntax — the caller is told the pattern was
taken literally instead of reading zero matches as "not in the trace".
Two bounds are kept as backstops: a cumulative match-time budget charging only
time inside `test`/`exec` (never the blob-store reads the scan awaits between
matches, which would truncate slow-but-legitimate greps under load), and a
total scanned-character cap.
* improvement(logs): only flag genuine regex intent, document the truncated invariant
Review follow-ups on the literal log-grep.
The `patternNotice` fired on any regex metacharacter, which includes `.` — so
ordinary literal searches (`example.com`, `file.pdf`, `v1.2.3`,
`block_1.output`) were told their pattern "was not interpreted as a regex",
inviting the agent to retry a search that had in fact worked. Measured 10 false
notices across 19 realistic literal searches. `REGEX_INTENT` now keys on actual
regex intent — escape classes, character class, group, alternation, a
leading/trailing anchor, or a repetition quantifier — which drops that to 0/19
while still flagging all 10 regex attempts tested.
`truncated` gains a TSDoc block stating its invariant: it reports that trace was
left unread, not that a budget was reached. Every point that skips work goes
through `done()`, which sets it, so a budget exhausted by the final match
correctly reports `false`. Raised by Cursor Bugbot; behavior is right, the
contract was just implicit.
Also drops `snippetAround`'s `maxChars` parameter, which every caller passed
from the state object it already receives.
* fix(logs): match log-grep patterns with RE2 instead of dropping regex support
Restores full regex support on copilot log-grep, which the previous commits in
this PR had removed to close the ReDoS. Deleting the capability was the wrong
trade: RE2JS is a port of RE2 that matches in time linear in the input, so no
pattern can blow up regardless of the text, and callers keep their regexes.
Measured through the real compile path, against a 100k-character adversarial
input — 10x the run that took 213s on JSC / 132s on V8:
(a+)+$ 13.0ms (\w+\s?)*$ 10.7ms
(a|a)*b 6.7ms (x+x+)+y 4.0ms
a*a*b 8.1ms ^(\d+)*$ 0.0ms
Two escape hatches keep this honest:
- A pattern with no regex metacharacter takes the built-in engine. Semantics are
identical when there is nothing to interpret, and RE2JS costs ~100x more per
byte (~25ms/MB), so the common plain-text search stays native.
- Lookaround and backreferences are not in RE2. RE2JS rejects them at compile
time, so those fall back to a literal and return a `patternNotice` naming the
constructs — a narrow, reported gap rather than a silent behavior change.
The match-time and scanned-character budgets stop being formalities: RE2JS's
throughput is what they now bound, not backtracking.
Also collapses the old test-then-exec double scan — `compilePattern` returns a
single `find` returning a match index, which `recordIfMatch` passes straight to
`snippetAround`, so each field is scanned once instead of twice.
re2js@2.8.6 is MIT, pure JavaScript (no native addon, so nothing changes for the
oven/bun image), server-only, and published 2026-07-05 — clearing the 7-day
bunfig minimumReleaseAge gate without an exclusion.
* fix(security): route every caller-supplied regex through the linear-time engine
Extends the RE2 fix to the three other places that compiled a caller-supplied
pattern and ran it on the shared event loop, behind one shared helper —
`lib/core/security/linear-regex.ts` — rather than four local copies.
- `lib/copilot/vfs/operations.ts` — copilot VFS grep took a caller pattern over
caller-supplied file content with no screen at all.
- `lib/guardrails/validate_regex.ts` (`validateRegex`) — a guardrail rule's
pattern matched against LLM output, screened only by `safe-regex2`.
- `lib/chunkers/regex-chunker.ts` — the KB chunker's split pattern. This one
screened for catastrophic backtracking by running the pattern against six
probes including `'a'.repeat(10000)` and rejecting anything over 50ms — but
it measured elapsed time *after* the match returned, so the screen was the
denial of service it existed to prevent (`a*a*b` on that probe: 213s). The
probe is deleted rather than repaired.
Deliberately unchanged: `validateRegexPattern`. Those patterns are persisted for
Presidio, a separate service where a slow pattern times out instead of stalling
this loop, and Presidio's Python engine accepts constructs RE2 does not — so
narrowing it to the RE2 subset would reject patterns that work. Its TSDoc now
says plainly that its `safe-regex2` screen is a courtesy, not a ReDoS defense,
and points in-process callers at `compileLinearRegex`. `safe-regex2` therefore
stays a dependency, used only there.
Lookaround is the standard way to split while keeping the delimiter (`(?=#\s)`)
and RE2 cannot represent it, so the chunker keeps the built-in engine for those
patterns — the only remaining path that can backtrack, bounded by the existing
500-character cap. RE2JS `split` was verified identical to `String.split` across
six representative chunking patterns.
Adds 27 tests for the shared helper (every catastrophic pattern asserted linear
over a 50k-character input) plus regression tests at each site, plus the first
tests for `validateRegex`. Also removes the now-duplicated `escapeRegExp` and
literal-finder locals from `log-views.ts`.
* fix(security): drop safe-regex2, keep lookaround splits on the linear engine
Audit follow-ups, plus the Cursor round-4 finding.
Removes `safe-regex2` entirely. Its last caller was `validateRegexPattern`, the
PII custom-pattern boundary, where it was pure cost: it screens star height
alone and passes `(a|a)*b` and `a*a*b`, while rejecting patterns that work —
lookbehind and optional groups could not be saved as custom PII rules at all.
Nor could any screen be sound there: those patterns execute in Presidio's Python
engine, which backtracks on shapes RE2 accepts, so RE2-representability says
nothing about their runtime. Validation is now syntax-only and the dependency is
gone.
Cursor flagged the chunker's lookaround fallback as reintroducing backtracking,
which was correct. The fallback is removed. Verified first that no bounded probe
can replace it: at a probe size small enough to survive an exponential pattern
(24 chars), `(?=a*a*b)` completes in 0.0ms and passes, because polynomial blowup
only shows on long input — small probes miss it and large probes hang, the same
trap as the original guard.
Instead `compileLookaroundSplit` puts the two idioms that actually need
lookaround onto RE2: splitting on `(?=X)` is slicing at every match start of X,
and on `(?<=X)` at every match end, neither of which needs the assertion. Both
`(?=#\s)` and the pre-existing `(?<=</section>)` case keep working, now in
linear time; anything else RE2 cannot represent is rejected with an actionable
message rather than run on the backtracking engine.
Also from the audit:
- `literalRegex` recompiled its `RegExp` on every call to dodge `lastIndex`
state from the `g` flag. VFS grep calls it per line, so that was a compile per
line. It now keeps one non-global instance for `test`/`find` and a global one
for `split`.
- `validateRegex` and VFS grep log a warning when a pattern degrades. A
guardrail rule that used lookaround now fails closed, which reads as the
guardrail tripping on every input, and VFS grep's return shape has nowhere to
report that a pattern was taken literally — both need to be findable in logs.
Measured: RE2 costs 6-13x native on per-line matching (20k-line, 1.66MB file
greps in 7ms), against ~100x on single multi-megabyte strings.
* test(pii): drop the custom-pattern editor's backtracking-screen assertion
CI caught this; my local runs only covered `lib/**` and missed the `.tsx`
counterpart of removing `safe-regex2`.
The editor asserted that `(a+)+$` renders "potentially unsafe". That screen is
gone, so the assertion is inverted: the test now pins that `(a+)+$`, lookbehind
and optional groups all render cleanly, which is the point of the removal —
they are valid Presidio patterns that could not be saved before. Syntactically
invalid patterns still surface "Invalid regex".
* fix(chunkers): support lookaround with an affix, and keep JS whitespace semantics
Three independent reviewers with no shared context audited the RE2 migration.
Two found the same high-severity regression, and one found the divergence that
silently rewrites stored data. Both are fixed here.
**Lookaround with an affix threw.** `compileLookaroundSplit` only accepted a
pattern that was *entirely* one assertion, so `(?<=\.)\s+` — split after the
period — and `(?<=[.!?])\s+(?=[A-Z])` — the textbook sentence splitter — fell
through to the chunker's `throw`. Those are persisted KB configs, and the
compile runs in the constructor, so an existing knowledge base would stop
re-ingesting with a hard error. The TSDoc claiming the bare assertions "cover
the reason a split pattern reaches for lookaround" was simply wrong.
Fixed generally rather than by adding another special case: splitting never
needs the assertion, only the span the delimiter consumes, so `(?<=X)Y(?=Z)`
compiles to `(?:X)(Y)(?:Z)` and group 1's span is exactly what `String.split`
removes. One rule now covers every combination, and the four shapes above match
`String.prototype.split` output exactly.
**`\s` silently lost 14 codepoints.** RE2's `\s` is ASCII-only; ECMAScript's
also covers `\v`, NBSP, U+2028/9, U+202F, U+3000, U+FEFF and the rest of
`\p{Zs}`. A reviewer measured 9 of 35 realistic pattern/document pairs chunking
differently through the real `RegexChunker` — NBSP from PDF/DOCX/HTML
extraction, U+3000 in CJK, U+202F in French text. That rewrites stored chunk
boundaries and embeddings for a document that never changed, with nothing
thrown and nothing logged. `translateToRe2` now rewrites `\s`/`\S` to the
verified-equivalent RE2 class, and `\uXXXX` to `\x{XXXX}` (RE2 rejects the
former outright, turning a non-ASCII delimiter into a hard failure).
Also from the audit:
- `truncated`'s new TSDoc asserted an invariant the code violates — reaching
`maxMatches` sets it even when nothing was left unread. Reworded to what it
actually means.
- The chunker's rejection message blamed "backreferences and lookaround" for
repeat-count and `\uXXXX` rejections too.
- `escapeRegExp` was a byte-identical copy of `executor/constants.ts:472` and
exported for nobody; now module-private.
- The guardrails wand prompt told the LLM to emit "valid JavaScript regex",
which reliably produces `(?=.*\d)`-style patterns that now fail every check.
- The PII editor's TSDoc still described the removed backtracking screen.
- VFS grep's TSDoc omitted that invalid patterns now match literally.
Test gaps the audit exposed: the match-time budget's only mechanism was
unguarded (deleting the accumulation left every test green) — now covered by a
mutation-verified test; the escape test survived dropping `.` from the class;
the split-parity test dodged the one real divergence; and two test names
described the built-in engine that no longer runs.
* fix(security): correct three defects in the lookaround split decomposition
A fourth reviewer, given no context beyond "these scanners are new and
unreviewed", differential-tested them against the built-in engine over ~62k
comparisons and found three real defects in code added an hour earlier. All
were reachable through RegexChunker with plausible patterns, and all silently
produced different chunks rather than failing.
1. **Top-level alternation was reshaped into a different pattern.**
`(?<=\.)\s+|\n\n` means `((?<=\.)\s+)|(\n\n)`, but decomposing it produced
`(?:\.)(\s+|\n\n)` — demanding the period before *both* branches. It could
lose boundaries (`"Heading\n\nAlpha beta.\nGamma delta."` split in 2, not 3)
and invent them. No grouping recovers the original meaning, so patterns whose
middle alternates at the top level are now declined outright, which also
removes an asymmetry: `\n\n|(?<=\.)\s` already threw.
2. **A capturing group inside the lookbehind stole group 1.** `(?<=(a))b` cut at
the assertion instead of the delimiter, and where the group did not
participate `start(1)` was -1, so `find` and `test` contradicted each other.
The middle is now captured by name, immune to numbering.
3. **Consuming the assertion text swallowed every other boundary.** The compiled
form matched `behind + middle + ahead`, so any boundary whose lookahead
doubles as the next one's lookbehind was skipped: `(?<=\w)\s+(?=[A-Z])` over
`"A B C D"` split 2 of 3 gaps. Each iteration now resumes at the end of the
delimiter rather than the end of the match. This was documented as affecting
only self-overlapping delimiters; that was wrong and far too narrow.
The reviewer confirmed `translateToRe2` and `closingParen` clean across
exhaustive escape-state, character-class, and paren-matching probes, and that
nothing added is superlinear.
One divergence remains and is now documented rather than overstated: when
delimiters themselves overlap (a single-character middle whose matches abut, as
in `(?<=\w).(?=\w)`), a match starting behind the cursor is dropped instead of
emitting the empty segment the built-in engine produces. Splitters that consume
whitespace or punctuation between tokens do not overlap and match exactly.
* test(security): pin engine parity with a differential suite
Every defect in this module has been a silent semantic divergence rather than a
crash, and each was found by comparing engines over a corpus in a throwaway
script. That comparison now lives in the suite: ~700 pattern/document pairs plus
match/index/ignore-case parity, checked against the built-in engine on every
run, with the accepted divergences enumerated and individually pinned.
It earned its place immediately by falsifying my own documentation. The
"self-overlapping delimiter" caveat was stale — fixing the boundary-advance in
the previous commit also made `(?=#{1,6}\s)`, `(?=aa)` and `(?=##)` exact, so
the entry asserting they still diverge failed. One real divergence remains and
is now stated precisely: a lookbehind whose body self-overlaps (`(?<=aa)` over
`aaaaa`). Making that exact means restarting the scan one character past each
match start, which is quadratic on a multi-megabyte document and forfeits the
linear guarantee — so it is a deliberate trade, not an oversight.
Mutation-verified against the three bugs that actually escaped review:
disabling the `\s` translation fails 12 tests, changing the boundary-advance
fails 2, and removing the top-level-alternation guard fails 3. That last one
initially passed, because the corpus had no alternation pattern in it — the gap
is closed and the mutation now fails as it should.
|
||
|
|
a042f0bd0f |
chore(deps): fix OTel version split, drop dead deps, declare emcn peers (#5994)
* chore(deps): fix OTel version split, drop dead deps, declare emcn peers
- Pin @opentelemetry/{resources,sdk-metrics,sdk-trace-base,sdk-trace-node}
to exact 2.7.1 so they match sdk-node's pins instead of floating to 2.8.0.
The carets meant app code built spans with 2.8.0 and passed them into
NodeSDK from 2.7.1, which only worked by duck-typing.
- Declare the 13 packages @sim/emcn imports but never declared, as peers
mirrored into devDeps. @radix-ui/react-dismissable-layer had no
declaration anywhere in the repo and resolved only transitively.
- Remove ffmpeg-static: its binary downloads via postinstall, but it is not
in trustedDependencies and Docker installs with --ignore-scripts, so the
accessSync branch never succeeded and both call sites always fell through
to system ffmpeg.
- Remove critters + experimental.optimizeCss: Next only loads critters from
the Pages Router renderer, and apps/sim is App Router only.
- Make simstudio-ts-sdk zero-dependency by dropping node-fetch for native
fetch; engines >=18.
- Remove unused @vercel/og and postgres from docs, dotenv/inquirer/listr2
from the CLI, and yaml from the root.
- Move @aws-sdk/client-appconfig from the root to apps/sim, its only consumer.
- Delete the apps/sim overrides block; Bun only honors top-level overrides.
- Bump free-email-domains 1.2.25 -> 1.9.70 (4,779 -> 13,059 domains).
- Validate object/array variables with JSON.parse instead of JSON5, matching
what the executor actually parses.
- Swap the changelog GitHub icon off lucide to GithubOutlineIcon, matching
the navbar chip on the same page.
- Unify @types/node on 24.2.1 and lucide-react on ^0.511.0; bump chalk to 5
and image-size to 2.
* fix(deps): complete the OTel pin, restore SDK error detail, revert email list
Follow-ups from an independent audit of the previous commit.
- Pin @opentelemetry/sdk-node and the three otlp-http exporters to exact
0.217.0. Pinning only their four dependents was self-reversing: sdk-node
0.219.0 requires core 2.8.0 exactly, so the next update would have
silently rebuilt the split this PR removes.
- Declare @opentelemetry/core (2.7.1). It is imported by
lib/copilot/request/go/propagation.ts but resolved only by hoisting, and
it is the OTel package with the most version churn in the tree.
- Pin @radix-ui/react-dismissable-layer to exact 1.1.13 in @sim/emcn. All
five transitive parents pin it exactly; a caret would fork a second copy
on 1.1.14, which is the duplicate-context bug the declaration prevents.
- Surface error.cause in simstudio-ts-sdk. Native fetch reports network
failures as a bare "fetch failed" and puts the reason on cause, so every
DNS/TLS/refused error was reaching callers with no diagnostic content.
- Revert free-email-domains to 1.2.25. Upstream now merges the free-domain
list with two disposable-email blocklists, so 1.9.70 classifies real
organization domains as free — UK charities, some companies and
universities, and the JP/KR ISP domains APAC SMBs use for business mail.
The demo form blocks submission on that check, so a false positive costs
the booking entirely. Worth doing deliberately, not inside a deps change.
- Lower packages/cli engines to >=18. chalk 5 and commander 11 both accept
>=16 and the source uses no Node 20 API, so >=20 only produced EBADENGINE
for Node 18 users.
---------
Co-authored-by: Waleed Latif <waleed@simstudio.ai>
|
||
|
|
e4cedcb91f |
improvement(landing): lead the hero H1 with the AI workspace keywords (#5996)
The H1 opened with "Sim is the", spending its first three words on a brand token already carried by the title tag, the meta description, and the hero's sr-only summary. Lead with the terms people search instead. Relaxes the landing H1 rule in both places that asserted the brand had to appear in the heading itself; the sr-only summary still opens by naming Sim, so entity consistency for answer engines is unchanged. |
||
|
|
fc396fcbf5 |
improvement(settings): match standalone shell chrome to the workspace one (#5992)
Account, organization, and self-host settings framed both columns in a rounded border and sat the pair inside an outer p-2, so the same section looked different depending on whether it was reached from a workspace. Mirrors WorkspaceChrome instead: the sidebar column is flush and borderless against the app surface, and only the content pane carries the rounded border. |
||
|
|
75b8b6f3e5 |
feat(settings): self-host settings plane, Sim wordmark in sidebar (#5990)
* feat(settings): self-host settings plane, Sim wordmark in sidebar
Chat keys were only reachable at a standalone /account/settings/chat-keys
page that nothing linked to. They now live on a dedicated self-host plane
alongside the two other settings a self-hoster needs from the managed
service.
- new /selfhost/settings/{general,billing,chat-keys} plane, open to any
signed-in user
- chat keys move off the account plane entirely; no isHosted gate
- registry `unified` projection is now optional (mirroring `planes`), so a
section can opt out of the editor sidebar
- plane items resolve their own description and throw when one is missing,
since a plane-only section has no unified projection to inherit from
- settings sidebar shows the Sim wordmark linking to /?home instead of a
Back chip; drops the now-dead backHref prop
- `bun run setup` runs `bun install` first so a fresh clone is one command
* docs(readme): point Chat keys at the self-host settings plane
The account-plane URL stopped resolving when chat keys moved to
/selfhost/settings.
* improvement(settings): make the sidebar wordmark a plane attribute
Replacing the Back chip everywhere was too broad — account and
organization are reached from inside the app, so Back is right there.
Self-host is reached from outside it (the CLI wizard, the README), so it
leads with the brand mark instead.
SETTINGS_PLANE_CHROME declares that per plane, keyed on
StandaloneSettingsPlane so adding a plane forces the decision rather than
defaulting silently. It also absorbs the shell's parallel plane-label map.
* fix(settings): hide self-host Chat keys on non-hosted deployments
Chat keys are issued by the managed service and useCopilotKeys is
`enabled: isHosted` for that reason. Moving the section onto the self-host
plane dropped its hosted gate, so a self-hosted deployment rendered a
Chat keys nav item whose list could never populate.
Restores the gate on the plane that now owns the URL. sim.ai is unaffected
— the section stays visible there to every signed-in user, which is the
surface self-hosters are pointed at.
|
||
|
|
23bc210bf8 |
fix(seo): noindex non-production sim.ai hosts, redirect indexed 404s (#5988)
* fix(seo): noindex non-production sim.ai hosts, redirect indexed 404s
Non-production deployments (dev.sim.ai, staging.sim.ai) serve the same
build as www.sim.ai and were fully crawlable. Send X-Robots-Tag noindex
for those hosts, and 301 seven marketing paths that an external SEO
audit found returning 404.
* fix(seo): parse only the first entry of a comma-joined forwarded host
A multi-value x-forwarded-host survives the port split intact, so
endsWith('.sim.ai') matched on the trailing entry and could apply
noindex to the canonical site. Normalize with the same
split(',')[0].trim() the rest of the repo uses for forwarded headers.
* fix(seo): drop the /security redirect, complete the urls test mock
security.txt advertises /security as its RFC 9116 Policy URI, so a
permanent redirect to marketing would mislead that link and shadow a
real policy page added later.
urlsMock is installed globally and documents itself as carrying every
real export, but was missing CANONICAL_SITE_HOST and
isNonCanonicalSimHost — any test loading proxy.ts would have called
undefined in track().
|
||
|
|
4edc169db3 |
fix(files): preserve nested folder structure when archiving workspace files (#5985)
The file_compress tool flattened every entry to its leaf name, so asking Sim to package a folder produced a zip with all files at the root and collision suffixes instead of the folder layout. Archive entry paths now mirror the workspace folder structure, with the ancestor chain the whole selection shares dropped so a single folder is not nested under its parents. The bulk download route reuses the same builder instead of its own copy of the path sanitizing and dedup logic. |
||
|
|
334c7c81bc |
fix(security): suppress the CodeQL password-hash alert on sha256Hex (#5982)
CodeQL fails #5963 (the v0.7.46 release PR) with 2 high-severity js/insufficient-password-hash alerts on packages/security/src/hash.ts. Both are false positives that surfaced now because
|
||
|
|
9c1094352d |
fix(content): stop wide markdown tables breaking the mobile layout (#5984)
* fix(content): stop wide markdown tables breaking the mobile layout Library and blog posts render GFM tables straight into the prose container with no width bound. Comparison tables run up to nine columns, so on a phone the table set the article's content width and body copy was clipped at both edges with no way to scroll to it — the table simply ran off the canvas. 19 of 20 library posts contain a table. Wraps tables in an `overflow-x-auto` container so that scrolling is confined to the table's own axis, and gives the table a `min-w` so columns do not crush to one word per line inside it. Both surfaces share `mdxComponents`, so this covers blog posts too. Also adds `lib/**` to Tailwind's content globs. `lib/content/mdx.tsx` emits classes like any component, but only `app`, `components`, and `pages` were scanned, so a class unique to that file was silently never generated — `min-w-[520px]` here, and already `list-outside`, `text-[19px]`, and `text-[0.9em]` on the existing paragraph and list styles. Verified against a real browser: without the glob the rule is absent and computed `min-width` stays `0px`. Generated CSS grows 538 bytes minified (+0.26%), all additions, nothing removed. * chore(content): type the table renderer instead of any Uses `ComponentPropsWithoutRef<'table'>` so the destructured and forwarded props are checked. The surrounding renderers in this file still take `any`; converting them is a separate cleanup, not something to fold into a mobile layout fix. |
||
|
|
66ac015c4c |
fix(library): generate every post cover from one template (#5980)
* fix(library): generate every post cover from one template Three posts shipped an `ogImage` pointing at a file that was never committed, so the library index rendered broken images and their `og:image`, JSON-LD, and sitemap entries all 404'd. Several others were authored without the brand font loaded or with the title clipping off the bottom edge. Covers were hand-made per post with no generator, which is why they drifted. Adds `bun run library:covers`, rendering each cover from the post's frontmatter title using the reference template already encoded in the docs OG route, and regenerates all 20 so the grid is uniform. Line widths come from the font's real advance metrics rather than an average-glyph-width estimate: the template joins words with non-breaking spaces to dodge a Satori space-measurement bug, which leaves hyphens as the only fallback break points, so an under-measured line breaks mid-compound. Also drops six orphaned `cover.png` sources left over from the JPEG compression pass in #5528. * fix(library): re-render covers every run and add a sync check Covers are derived artifacts, so skipping outputs that already exist left an image showing the old title after a post's frontmatter `title` changed. Every run now re-renders from scratch; rendering is deterministic, so an unchanged title re-encodes to identical bytes and a full run stays a no-op in git. Replaces `--force` (now the default) with `--check`, which renders in memory and compares against the committed bytes without writing, so CI can catch both a stale cover and the missing-cover case that caused the original breakage. * fix(library): compare decoded pixels in the cover sync check Byte-equality on the mozjpeg output assumed portable encoder bytes. libvips/mozjpeg does not guarantee that across OS and CPU, so identical input can encode differently on a contributor's machine or a Linux CI runner and fail the check for no real reason — exactly where the check was meant to run. Decodes both images to greyscale and compares mean absolute difference instead, which discards encoder variance while still testing what the check is about. Measured on this cover set: re-encoding an identical render with a deliberately different encoder moves it ~0.26, a one-word title change moves it ~12; the threshold of 2 sits between them with ~8x margin. * fix(library): count redrawn pixels in the cover sync check Averaging the difference diluted a local edit across all 810,000 pixels. Changing a title's "2026" to "2027" moved the mean by 0.42 — under the tolerance that absorbed encoder noise — so the check passed a cover still showing the old year. Counts pixels that moved more than 48 greyscale levels instead. Measured on this cover set, that one-character edit redraws 2,559 pixels while three deliberately different encodes of an identical render (quality 60/70 without mozjpeg, quality 95 with) redraw none, so the count separates real drift from encoder variance in both directions. * fix(library): parse frontmatter with gray-matter and split oversized tokens Two issues in the cover generator, neither reachable from a current title. The hand-rolled frontmatter regex could disagree with `gray-matter`, which is what renders the page and its `og:title`. On a double-quoted escape or a block scalar the cover would have rendered a title the page never shows, with `--check` calling it in sync. Uses `gray-matter` directly so there is one parser. `wrapTitleLines` only breaks between space-separated words, so a token wider than the title box on its own stayed on an overflowing line, and the non-breaking spaces left Satori no recourse but to break it at a hyphen — the mid-compound break this layout exists to prevent. Oversized tokens now split here, at hyphens first and per-character only for something like a URL, and a font size is accepted only if every line measures within the box. All 20 covers re-render byte-identically, so neither change alters current output. |
||
|
|
6d48444525 |
fix(docs): render native block icons instead of the two-letter fallback (#5981)
* fix(docs): render native block icons instead of the two-letter fallback
The Table and Logs pages (and every other native resource block) showed a
two-letter text fallback because the generated icon map never contained them.
Four separate causes in scripts/generate-docs.ts:
- The icon-map allowlist had drifted behind NATIVE_RESOURCE_BLOCK_TYPES, the
set the docs writer uses. Key the exception off that set so the map cannot
fall behind the pages that consume it.
- extractIconNameFromContent only matched identifiers ending in `Icon`, so
Logs (`icon: Library`) resolved to nothing. Match any identifier, excluding
bare JS literals.
- The map imported everything from `@/components/icons`, so an icon sourced
from `@sim/emcn/icons` could not resolve. Imports are now grouped by the
module each icon is actually imported from.
- Trigger-only pages (slack_app, twilio) and hand-written pages (a2a) had no
entry at all. Seed provider icons from the trigger definitions.
Also fixes three regen bugs found while verifying the output:
- A comment reading "this becomes `hideFromToolbar: true`" in slack.ts was
matched as the property itself, so a clean regen dropped Slack from the
integrations catalog and reduced slack.mdx to a 29-line stub. Property
probes now run against comment-stripped source.
- 16 hand-written *-service-account guides were unregistered, so the stale-doc
cleanup deleted them on every regen. Registered them, and cleanup now refuses
to delete any page holding MANUAL-CONTENT (this also restores the intros on
file.mdx and twilio.mdx).
- Trigger outputs referenced as a constant (`outputs: SLACK_TRIGGER_OUTPUTS`)
resolved to nothing, dropping whole Output tables. Constants and sibling
modules now resolve, which also restores 319 lines on clickup.mdx.
Removes the language selector from the docs navbar.
Regenerated docs are included; remaining content deltas are tool-definition
drift since the last regen.
* improvement(docs): drop the preview-gated slack_app page, document managed_agent
- slack_oauth is reachable only through the preview-gated slack_v2 block, so
documenting it published an unreleased surface under its own slack_app page.
Triggers whose every hosting block sets `preview: true` are now excluded from
the docs and the icon map. Triggers no block claims are untouched, so
standalone webhook providers keep their pages.
- Adds the MANUAL-CONTENT intro to managed_agent.mdx, matching the other
integration pages. Verified it survives a regen.
* fix(docs): stop truncating quoted descriptions, tighten the cleanup guard
Review findings from round 1.
- parseSubBlockObject read string properties with a single `['"]…[^'"]+…['"]`
character class, which ends the match at the first quote of either kind. Any
description holding an apostrophe inside a double-quoted string was cut
mid-word ("Your app", "Found in your Zoom app"). Matches the opening quote to
its own closing quote now, reusing the alternation the tool-description
extractor already used. Restores full text across calendly, gmail,
google_sheets, hubspot, intercom, whatsapp, and zoom.
- The stale-doc cleanup guard tested for a bare `MANUAL-CONTENT-START`
substring, so a stray or unterminated marker would pin a stale page that has
nothing recoverable. It now gates on what extractManualContent actually
returns.
|
||
|
|
54c3e1cba5 |
improvement(chat): defer resource-menu list queries until the menu opens (#5973)
* improvement(chat): defer resource-menu list queries until the menu opens The `+` resource menu and the `@`-mention menu both hydrate from `useAvailableResources`, which fired nine workspace-wide list queries and rebuilt ten item groups on mount — from two always-mounted call sites, for menus that were closed. Gate it on each menu's own open state. - add an `enabled` option to `useAvailableResources`; pass `open` from `AddResourceDropdown` and `PlusMenuDropdown` - add `options.enabled` to the six query hooks that lacked it, matching the convention already used by `useWorkspaceFiles` / `useKnowledgeBasesQuery` / `useLogsList` - skip the tab-name lookup's five list queries when there are no tabs to label - drop the `= []` destructuring defaults: a literal default allocates a fresh array every render while `data` is undefined, busting the group memo in exactly the disabled state - move the hook call into `PlusMenuDropdown`, which owns the open state, removing the `availableResources` pass-through from `usePromptEditor` - delete the `existingKeys`/`isOpen` plumbing — one consumer never read it; resolve it at the single call site instead * fix(chat): warm resource lists on editor focus so fast mentions resolve Deferring the lists to menu-open left a window where `selectActive()` saw an empty candidate list, and its `false` return falls through to submitting the message with the mention unresolved. Start hydrating on first focus — the earliest reliable signal a mention may be coming — which closes the window while keeping the lists off the page-load path. * fix(chat): never submit a mention while its candidate lists are loading Warming on focus shrank the race but gave no readiness guarantee: an empty candidate list still meant "no match" to the keydown handler, which falls through and submits the message with the mention as raw text. Make the distinction explicit. `useAvailableResources` now reports `isHydrating`, and `PlusMenuHandle.selectActive` returns `selected | empty | hydrating` instead of a boolean. The editor swallows Tab/Enter while hydrating and only lets them through once the lists have settled, where an empty result is a genuine no-match. `isHydrating` keys off `isPending`, not `data === undefined`, so a failed list settles instead of blocking the key forever. |
||
|
|
67ffdac13f | improvement(mship): mcp persistence in chat | ||
|
|
ae49070bea |
fix(slack): stop requesting the unapproved scopes, gate them behind an opt-in flag (#5977)
* fix(slack): stop requesting the unapproved mention/assistant/DM scopes * feat(slack): gate the approval-pending scopes behind an opt-in env flag |
||
|
|
bd61603701 |
feat(tiktok): unhide integration (#5978)
* feat: unhide TikTok integration * test: remove TikTok visibility assertion --------- Co-authored-by: Bill Leoutsakos <billleoutsakos@Bills-MacBook-Pro.local> |
||
|
|
49adf91c20 | improvement(pii): clarify custom regex replacement placeholder (#5975) | ||
|
|
104a6fcb16 |
fix(tables): make the atomic per-cell merge the only row-write path (#5970)
Table rows store all cells in one jsonb column; row updates did a full-object read-modify-write (row-level last-write-wins). updateRow now always merges changed cells via a shared jsonbMergePatch helper (data = data || {changed}::jsonb) under the row lock, and batchUpdateRows does the same per-row, so concurrent edits to different cells of the same row no longer clobber each other. Removed the never-beneficial replace mode; fix lives entirely in the service layer.
|
||
|
|
1a4bfe4c59 |
fix(setup,compose): bundle Redis, fix socket reconnect, and harden the setup wizard (#5964)
* feat(compose,setup): bundle Redis, always configure it, fix lifecycle detection
Compose shipped no redis service at all — REDIS_URL was ${REDIS_URL:-} in both app and realtime, so every self-hosted stack ran without it. Storage silently falls back to PostgreSQL, but the pub/sub channels (live Chat task-status, table events) have no fallback, so live updates never arrived.
- compose (prod + local): add a redis:7-alpine service with a healthcheck, default REDIS_URL to redis://redis:6379, and make app/realtime depend on it being healthy. Not published to the host — only the containers need it, and binding 6379 would collide with a local Redis. An external REDIS_URL in root .env still overrides. Deliberately not written into root .env: doctor pings REDIS_URL from the host, and a compose-internal hostname would fail that probe the same way DATABASE_URL would.
- dev mode: configure Redis in quick too. Quick uses a new non-interactive ensureRedis (adopt whatever answers, else start the managed container, warn only if Docker is unavailable); custom keeps the ladder, with corrected copy — the old prompt claimed Redis was only for multi-replica.
- lifecycle: detect compose stacks via 'docker compose ls' instead of probing '-f <file> ps' in the working directory. Compose derives the project name from the directory it was started in, so the old probe found a stack only when run from the checkout that launched it (a globally linked sim never could) and listed the same stack once per candidate file. compose ls reports the real project and its config file, so one stack yields one install from anywhere; non-Sim projects are filtered by compose filename. Every compose op now runs in that stack's directory.
- lifecycle: distinguish 'Docker unreachable' from 'nothing installed'. With the daemon down, status reported containers as 'absent' and suggested re-running setup; it now says Docker is down and marks state unknown.
* fix(setup): don't start managed Postgres with a password the volume will ignore
POSTGRES_PASSWORD only applies when initdb runs on an empty data directory. The sim-postgres-data volume outlives its container (sim down keeps it, docker rm keeps it, and the wizard's own recreate path keeps it), and inspectManagedContainer recovers the password from the *container*, not the volume — so once the container is gone the password is unrecoverable.
Setup then generated a fresh password and ran against the initialized volume. Postgres kept its original password and rejected every connection with 'password authentication failed for user postgres', which surfaced as a misleading 'container did not become healthy'.
Detect an already-bootstrapped volume (PG_VERSION present) before choosing a password, and ask: supply the existing password, or delete the volume and start fresh (double-confirmed, since that destroys data). Refusing both fails with the exact docker volume rm command instead of looping.
* improvement(setup): default to Docker Compose and sharpen the run-mode copy
Compose was listed first but only preselected when Docker happened to be running — with Docker stopped the cursor sat on 'Local dev', steering people toward a source checkout when they wanted to run Sim. Compose mode calls ensureDocker(true), which offers to start Docker Desktop, so a stopped daemon is no reason to change the default.
Also tightens the hints to say what each mode is for: run bundled Sim (fastest way to start), work on Sim itself, test a production-style k8s deploy.
* fix(compose): point the browser socket at :3002 so it stops reconnecting
The stack publishes the app on 3000 and realtime on 3002 with no reverse proxy between them, but NEXT_PUBLIC_SOCKET_URL defaulted to empty — which tells the browser client to use the page origin. :3000/socket.io answers 308 (a Next redirect), not a Socket.IO handshake, so the client failed and retried forever. Default it to http://localhost:3002; a proxied deployment overrides it (or sets it empty to use the page origin).
Also give COPILOT_API_KEY and SIM_AGENT_API_URL empty defaults so every compose command stops printing 'variable is not set' warnings. The app already falls back to the prod copilot backend when SIM_AGENT_API_URL is blank.
* feat(setup): pass SIM_AGENT_API_URL through, and warn on a half-set mothership
Sim devs testing against a non-prod mothership export SIM_CLI_AUTH_ORIGIN so the Chat key is minted there, but nothing carried the matching backend URL into the install — the app kept defaulting to prod copilot, which rejects a staging key with 'Invalid API key'.
Persist SIM_AGENT_API_URL when it is exported, so later docker compose up / dev runs stay on that backend instead of reverting to prod once the shell is gone:
SIM_CLI_AUTH_ORIGIN=https://www.staging.sim.ai \
SIM_AGENT_API_URL=https://www.staging.copilot.sim.ai \
bun run setup
Setting only the auth origin is the trap, so that combination warns. Neither set is the self-hoster default and stays silent — no prompts, no flags.
* fix(setup): survive a vanished port owner, and stop flagging our own containers
Two failures from one compose re-run:
- 'Kill it for me' crashed setup with 'kill() failed: ESRCH: No such process'. The owner list is an lsof snapshot, so the process can exit before the signal lands — which is the outcome we wanted, not an error. ESRCH now counts as freed, EPERM warns that it must be stopped by hand, and anything else warns; the loop re-probes either way instead of aborting a setup that had already written .env.
- Compose mode demanded 3000/3002 be free even when this stack was the one holding them, so re-running setup against a running install reported its own realtime container as a blocker and offered to kill Docker's listener. 'docker compose up -d' reconciles its own containers, so skip the check when the project already has some. A foreign process is still caught, and a foreign container still surfaces as a bind error from compose.
* fix(csp,setup): permit the socket origin the client actually uses; encode DSN passwords
Review round on #5964:
- The socket reconnect was a CSP bug, not a URL bug. getSocketUrl() already falls back to localhost:3002 for a localhost page, but generateRuntimeCSP gated that same fallback on isDev — and compose runs NODE_ENV=production, so connect-src omitted ws://localhost:3002 and the browser blocked the handshake. Key the fallback on the app URL being localhost instead, mirroring getSocketUrl. Revert the compose NEXT_PUBLIC_SOCKET_URL default: an explicit value suppresses the page-origin fallback that reverse-proxied self-hosts depend on, and ':-' treats empty as unset so the documented escape hatch could not work either. LOCALHOST_HOSTNAMES is duplicated locally because csp.ts is loaded by next.config.ts before @/ aliases resolve.
- Percent-encode the password when building the Postgres DSN. A user-supplied password containing @ : / # does not merely re-parse to the wrong host — it fails to parse as a URL at all, so a correct password surfaced as a connection failure.
- Tell 'Postgres rejected this password' apart from 'Postgres never started'. On the keep-the-volume path a wrong password left a healthy server and the old generic 'container did not become healthy' error, which is the confusion this change set exists to remove.
Adds a CSP regression test for the unset-socket-URL production case; verified it fails against the previous condition.
* improvement(setup): make k8s mode end somewhere usable, and show install progress
Two things made k8s mode the least satisfying path.
The services are ClusterIP, so a successful install left nothing on :3000 — 'Sim is ready' was true about the cluster and useless to the user, who had to notice and run a port-forward by hand. Compose opens a browser and dev offers to start the server; k8s now offers the forward the same way and runs it in the foreground so Ctrl-C ends it. Realtime gets its own forward (kubectl takes one resource per invocation) or the editor socket fails; it is a child in the same process group, so the terminal's Ctrl-C reaches it, and it is killed explicitly when the app forward exits.
'helm --wait' then blocked for minutes with a single static spinner, so a slow image pull looked identical to a wedged install. Run helm asynchronously and poll the cluster, so the spinner reports '3/3 pods ready · 1 starting'. CronJob-owned pods are excluded: the chart schedules a lot of them (36 on a running cluster here) and they finish as Completed, which would swamp the count and make readiness jitter for reasons unrelated to the install. Restarting pods are surfaced too — a cold cluster restarts realtime while Postgres comes up, and a silent spinner made that look like nothing was happening.
* fix(setup): identify Sim compose projects by content, not filename
Cursor (High): composeInstalls treated any project whose config basename was docker-compose.prod.yml or docker-compose.local.yml as a Sim install. Those names are common, and sim reset runs 'compose down -v' — so a stranger's stack could have had its volumes destroyed.
I introduced that reach. The previous ROOT-scoped '-f' probe was implicitly safe because it could only ever see the project in this checkout; switching to a global 'compose ls' to find stacks started elsewhere means projects must be identified by content instead. Read the config file Docker recorded and require a Sim marker (the published app image, or the app Dockerfile this repo builds), so both the prod and local variants match while an unrelated file with the same name does not. An unreadable or since-deleted file is left unmanaged rather than assumed ours.
Verified against a decoy nginx compose file using our exact filename: ignored, while both real Sim compose files still match.
* fix(setup): scope the compose port skip to published ports; print both k8s forwards
Review round on #5964:
- ensureComposePortsFree skipped conflict handling whenever the project had any container running, so leftover db/redis (which publish neither app port) waved through a foreign process on :3000 — it then surfaced as a raw compose bind error instead of the prompt. Read the host ports the project actually publishes and skip only those; the remaining ports still get the full check. Reading from the containers rather than the file matters because what counts is what is bound right now.
- The post-install note and the skip path documented only the app forward, while offerPortForward runs two. Skipping the prompt or copying the printed command left the editor's socket dead — the exact failure this change set exists to fix. Both commands now come from one forwardCommands() helper, so what is printed and what is run cannot drift.
* fix(setup): one source for the k8s forwards, and surface a dead realtime forward
Third round on the same theme, so fix it at the root rather than at another call site.
- lifecycle's k8sReachHints (used by sim start/restart) still restated an app-only forward, recreating the dead editor socket the setup path had just been fixed for. forwardCommands is now exported and consumed there, so every place that tells a user how to reach a ClusterIP release derives it from one definition.
- The realtime forward was spawned with stdio ignored and never checked, so a busy :3002 or a missing service killed it silently while the app forward kept running — indistinguishable from success until the editor won't connect. Keep its stderr, warn on an exit we did not ask for, and stay quiet on the intentional kill.
* fix(setup): pin the compose project on every lifecycle op
composeInstalls records the real project name from 'compose ls' and status and the destructive confirms print it, but every op ran 'compose -f <file>' with only cwd set — so Compose re-derived the project from that directory. The derived name is frequently not the recorded one: a directory is lowercased and stripped of dots (Sim.Demo_Test derives simdemo_test), and an explicit -p or COMPOSE_PROJECT_NAME at creation diverges outright. stop/down/reset could therefore act on a different project than the one named in the confirm, and reset runs 'down -v'.
Route every op through composeArgs(), which pins '-p <recorded project>'. cwd stays, since the file's own relative paths still resolve against it. Verified with a stack started as -p pinned-name from a directory deriving simdemo_test: the old form found 0 of its containers, the pinned form finds them.
* fix(setup): warn on both halves of a mothership mismatch
mothershipOverride warned only when SIM_CLI_AUTH_ORIGIN was set without SIM_AGENT_API_URL, while its own copy said to set both or neither. The reverse is the same failure mirrored: with only SIM_AGENT_API_URL set, the Chat key is still minted against the default prod auth origin and then validated against the override, which rejects it — silently, which is exactly what this helper exists to prevent.
Warn on either asymmetry, and read the default origin from one constant shared with the handoff so the message can't claim an origin the code no longer uses.
* fix(setup): warn about a half-set mothership before minting the key
mothershipOverride ran two steps after promptCopilotKey, so a half-set override minted a key against one environment, stored it, and only then warned that the other environment would reject it. Worse on a re-run: promptCopilotKey offers to keep an existing COPILOT_API_KEY and defaults to yes, so the bad key survives.
Move the override ahead of the key prompt in both compose and dev, so the warning arrives while it can still change the outcome — the user can abort and set the missing half before anything is minted. Nothing in the override depends on the key, so the order is free.
|
||
|
|
7f4cc38fb7 |
improvement(tables): show the lock chip only when a table is actually locked (#5967)
* improvement(tables): show the lock chip only when a table is actually locked - The header chip rendered whenever an admin had the flag on, so an unlocked table permanently carried a "Lock settings" entry. Header space is for state: the chip now appears only once something is locked and names the mode. The admin route to the panel on an unlocked table is the breadcrumb dropdown, which already had it - Keep the Append-only name when the schema is locked too. Append-only describes the row semantics and a schema lock doesn't change them; the detail line calls out the locked columns instead * improvement(tables): resolve the lock flag server-side and record lock changes fully - Drop NEXT_PUBLIC_TABLE_LOCKS. A feature flag's gating lives in AppConfig, which has no client counterpart, so mirroring it into a public env var meant AppConfig couldn't control the UI at all and org/user clauses could never reach the client — only global on/off. The page now resolves the flag with session context and passes it down, per the add-feature-flag skill. Embedded renders default to false, failing closed; enforcement of stored locks is unaffected either way - Record the previous locks alongside the new ones and name the transitions in the audit description, so the log answers who locked what without expanding metadata. Forward the request for IP / user-agent capture * fix(tables): resolve the lock flag with the same context on both gates The page passed userId/orgId while the PATCH gate resolved the flag with no context, so an org- or user-targeted rollout would show the settings panel and then 403 on save — the rollout path the previous commit exists to enable. Both now key on the workspace's host organization rather than the viewer's active one, matching the convention getWorkspaceHostContextForViewer documents: active-org describes the account, not the workspace host. The route's lookup runs only when a lock is actually being turned on. |
||
|
|
0dcbc56ef6 |
feat(tables): per-table mutation locks (schema/insert/update/delete) (#5960)
* feat(tables): per-table mutation locks (schema/insert/update/delete) Adds four independent, admin-only locks to a table so a workspace can make it append-only, read-only, or schema-frozen. Enforced at the lib/table service layer rather than the routes, so the API, workflow blocks, and Mothership (which calls the services directly) are all covered by one assert. Violations return 423; changing a lock requires workspace admin and returns 403. * fix(tables): address review — CI snapshot, lock-aware jobs, UI gating - Add the drizzle meta snapshot the hand-written migration was missing, which made CI regenerate the lock columns as a phantom 0271 - Stop a queued or retried delete/update job that starts after its lock is enabled; a run that has already committed pages still finishes, since pages are never rolled back and aborting would leave an uncancelled half-done state - Map TableLockedError to 423 on the internal single-row DELETE (was a 500) - Gate the lock settings UI behind NEXT_PUBLIC_TABLE_LOCKS so it can't open onto a Save that 403s against the server-side flag - Respect the delete lock on column drops in the grid, matching assertColumnDestructive, instead of failing only on click * fix(tables): close lock gaps in async import and undo/redo - Assert insert (and delete for replace) at the start of the import worker. Replace deleted every row before the first insert assert, so an insert-locked table was wiped and then failed, leaving it empty. - Run the delete/update worker start checks through assertRowDelete / assertRowUpdate instead of reading the lock flags directly, so the worker and its enqueue site apply identical rules — including the workflow-column exemption a bulk update was being cancelled despite. - Make undo/redo verb mapping direction-aware for column ops: dropping or retyping a column needs the delete lock clear too, matching assertColumnDestructive, so those steps are skipped instead of stranded. * fix(tables): stop locks over-blocking workflow runs, duplicate, and imports - Run / Re-run / Stop no longer inherit the cell-edit gate: they write only workflow-output columns (which the update lock exempts) and Stop is a cancel - Duplicate needs only the insert lock — it inserts a full copied row, so it stays available on an append-only table - updateWorkflowGroup asserts the destructive rule only when a patch actually drops or remaps output columns; rename / autoRun / mapping edits need just the schema lock - Gate the import-async enqueue on insert (and delete for replace) so a locked table 423s instead of claiming the write-job slot and failing in the worker - Disable Import CSV under an insert lock, and wire the expanded cell editor and the previously-unused stable add-row handler to the lock-aware flags * fix(tables): keep column menu usable and let locks always be cleared - Pass the blocked-delete handler instead of undefined so ColumnOptionsMenu still mounts on a delete-locked table; withholding it hid the entire menu, including Insert column, which the delete lock must not affect - Restore ColumnHeaderMenu readOnly to permission-only: it swaps the header for a static label, which was also disabling pin, column-select and open-config - Assert the schema lock at the import-async enqueue when createColumns is set - Allow a locks PATCH that only clears locks even when the feature flag is off, and keep the settings entry reachable on a locked table, so flipping the kill switch can't strand a table with locks nobody can remove * fix(tables): scope the update-lock exemption and re-check locks per page - Make the workflow-output carve-out opt-in via `computedWrite`, set only by the cell-write path. It was caller-agnostic, so an ordinary API caller could PATCH a workflow-output column on an update-locked table - Re-read the lock before every page in the delete and update runners. The single worker-start check went stale immediately, so enabling a lock could not stop a job already deleting or overwriting rows. Committed pages still stay committed, exactly as with an explicit cancel. Import keeps its one-shot check: a half-imported table needs the deletes the lock now forbids - Route Insert column to the locked-action modal under a schema lock instead of leaving it live (regular header) or hiding the whole menu (group header) * fix(tables): revalidate locks in-transaction and unbreak backfill + partial unlock - Re-assert the lock inside each page-write transaction, under the same advisory lock updateTableLocks takes, so check-then-write is atomic against a lock change. The per-page check alone still left the page-selection await between the assert and the write - Pass computedWrite from the backfill runner: batchUpdateRows is the workflow-output backfill path, so scoping the exemption had made rebuilding outputs 423 on an update-locked table while live cell writes still worked - Gate the feature flag on a lock actually going off->on rather than on an all-false payload. The settings UI always submits all four flags, so with the flag off an admin could not clear one lock while another stayed on - Hoist the lock-kind -> flag map to lib/table/types as TABLE_LOCK_FLAGS * fix(tables): guard import batches in-transaction and split paste by verb - Re-assert the insert lock inside each import batch's insert transaction, under the same advisory lock updateTableLocks takes. Reversing the earlier "let the file finish" call: an admin can lift the lock to clean up a partial import, so honouring it beats letting the rest of the file land - Gate paste per verb. Overwriting existing rows is an update and extending past the last row is a full-row insert, so a paste-append still works on an append-only table while an overwrite explains itself. Refuses the whole paste rather than applying half of it - Space opens the same row editor as double-click, so it now follows the update lock instead of filling in a form that 423s on save * fix(tables): guard the remaining import writes and explain blocked Shift+Enter - Route the replace-mode wipe, the inferred-schema write and createColumns through the same guardBatch transaction as the batch inserts, so a delete or schema lock committed while the file is downloading or being sampled is seen instead of the job-start snapshot - Shift+Enter takes the manual-add path, so it now opens the lock modal like the Add row button instead of returning silently * improvement(tables): surface blocked lock actions as a toast, not a modal - Replace TableLockedModal with a warning toast carrying a "Lock settings" action button for admins. Being told you can't edit shouldn't cost a dismiss click, and the button still routes admins straight to the panel - Dedupe by id so repeated attempts on a locked cell replace one notice instead of stacking a column of them - Move the copy into lock-copy.ts alongside the rest of the lock vocabulary and tighten it for toast length * fix(tables): don't show the update-lock notice to users without write access Double-click checked the update lock before any permission check, so a read-only member on a locked table got "Editing rows is locked" — misleading (the lock isn't why they can't edit) and it swallowed the expanded-cell viewer, which is a legitimate read-only affordance. * improvement(tables): trim the update-lock toast copy * fix(tables): map append-import locks to 423 and stop conflating locks with permissions - The sync append branch returns instead of rethrowing, so the outer catch's mapper never saw a TableLockedError and every lock violation became a 500. Map it in that catch, with a regression test (replace mode already rethrows) - Stop mounting the workflow-group column menu for users without edit access: passing the blocked handlers unconditionally made it appear for read-only members and report a lock even on an unlocked table - Enter/F2 now raises the same lock notice as double-click and Space instead of silently doing nothing - guardBatch returns the freshly-read definition so addTableColumnsWithTx asserts live state; a schema lock cleared mid-import no longer fails the createColumns step. The snapshot pre-asserts now only run as the fallback for callers that pass no revalidator - Append-only no longer labels a table whose schema is also locked, which claimed columns were mutable when they weren't - Import dialog withholds Replace on a delete-locked table and create-column on a schema-locked one, instead of offering a configuration that only 423s * fix(tables): revalidate sync imports under the advisory lock, explain every blocked key - The sync append/replace paths asserted only the request-start snapshot, so a lock committed while the CSV was parsed still let the write through. Both now re-read under the schema advisory lock at the top of their own transaction, taken before acquireRowOrderLock so the order stays advisory -> rows_pos -> definitions - Delete/Backspace, Cmd+D, typeahead and cut raised no notice on an update-locked table; they now explain the lock, and stay silent for users without write access - The multipart import fetch threw a plain Error, dropping the 423 status, so the lock self-heal never ran and the stale detail cache survived. It throws ApiClientError now and both import-into-table hooks call handleTableLockRejection - Force append at submit when the delete lock landed while the dialog was open with Replace already selected * fix(tables): restore the workspace ownership check on copilot row deletes Switching deleteRow/deleteRowsByIds to take a TableDefinition dropped the workspaceId argument that previously scoped the query, and the two branches I added loaded the table without the ownership comparison every other operation in the tool performs — letting a caller delete rows from a table in another workspace. Both now reject a foreign table as not found. |
||
|
|
8329dac4c5 |
feat(tables): add select & multi-select column types (#5873)
* feat(tables): add select/multiselect column types (backend) Adds two enum-style column types where the column declares a fixed set of options (stable id + name + palette color) and every cell is constrained to them. - COLUMN_TYPES gains `select` / `multiselect`; SELECT_COLORS is a fixed, theme-aware palette mapping 1:1 to Badge color variants (no raw hex) - ColumnDefinition.options carries the option set. Cells store option *ids* (a string for select, string[] for multiselect) so renaming or recoloring an option never rewrites row data - Row validation enforces membership; coercion tolerantly maps an option *name* to its id for tool/import writes and drops unmatched entries - validateColumnDefinition enforces non-empty, unique-id, unique-name and valid-color option sets, and rejects options on non-select columns - Contract gains selectOptionSchema plus a cross-field refine requiring options exactly on select types; routes thread options through type changes and a new options-only updateColumnOptions path No migration: column config already lives in the user_table_definitions schema JSONB blob. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014e4MvFV1szLhNLap1CCNUd * feat(tables): select/multiselect column UI Surfaces the new enum column types in the tables grid. - New `select-field/` module: SelectPill (colored option via the shared Badge palette), SelectValueEditor (one ChipDropdown-backed picker reused by every edit surface), SelectOptionsEditor (add/rename/recolor/remove) - Option colors are picked from inline squircle swatches — no labels, no nested dropdown. Each swatch's fill is a Badge in that variant, so the palette stays single-sourced and theme-aware - Cells render option pills; an empty select cell shows a muted "None" so it reads as a dropdown. The single-select menu always offers "None" to clear - Wired into all three edit surfaces (inline cell, expanded popover, row modal) plus the type picker and column-type icons - ChipDropdown gains `defaultOpen`/`onOpenChange` so the inline cell editor can open on mount and commit when the menu closes; open state is now controlled in both single and multi modes Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014e4MvFV1szLhNLap1CCNUd * fix(tables): auto-fit select columns on option labels, not ids Column auto-resize measured `String(val)` for every non-json/date column, which for a select cell is the opaque option id (and for multiselect the comma-joined id array). Selecting auto-fit on a select column therefore sized it to ids the user never sees. Measure the resolved option names instead, matching what the pills actually render. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014e4MvFV1szLhNLap1CCNUd * feat(tables): default select options to grey, drop color picker for now Removes the per-option swatch picker and defaults every option to the neutral gray pill. The `color` field stays in the data model and the SELECT_COLORS palette/contract are untouched, so a picker can be re-added later as a pure UI change with no migration. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014e4MvFV1szLhNLap1CCNUd * fix(tables): guard select type conversion, escape, and required clear Addresses review findings on the select/multiselect column types: - Column type conversion now checks each existing value against the target option set (resolve by id or name); a `select`/`multiselect` change is blocked when values don't fit, instead of accepting shapes that later coercion would strand or silently drop - Inline select editor discards its draft on Escape (matching the text/date editors) rather than committing on menu close - A required single-select no longer offers "None" — clearing to null could never be committed Exports resolveSelectOptionId from validation for the conversion gate. * fix(tables): block emptying required multiselect, skip no-op cell writes Round 2 review fixes: - A required multiselect can no longer be emptied — the toggle that would remove the last option is ignored, since an empty selection can't be committed (server rejects it). Mirrors the required single-select "None" guard. - Inline cell save now compares old vs new value structurally, so a no-op edit (e.g. opening a multiselect and closing it unchanged, producing a new array reference) no longer writes a row update or pushes an undo entry. Uses JSON compare for arrays/objects, matching the existing optionsEqual convention; primitives keep the === fast path. * fix(tables): expanded multiselect string value + same-type options drop Round 3 review fixes: - Expanded-cell popover normalizes a multiselect's initial value with toSelectedIds, so a single option-id string (the normal shape right after a select->multiselect conversion, before the row is rewritten) is no longer collapsed to an empty array and cleared on save. - Column PATCH now only routes to updateColumnType on a real type change. updateColumnType early-returns on an unchanged type, so a payload repeating the current type together with new options previously applied neither; an unchanged type with options now routes to the options-only update. Fixed in both the internal and v1 column routes. * improvement(tables): inline select edit + idiomatic options editor - Double-clicking a select/multiselect cell now opens the inline option dropdown (like date/number) instead of the fixed-height text popover, which rendered just a small dropdown floating in a large empty container. Removes the now-unreachable ExpandedSelectEditor from the expanded-cell popover. - Options editor's add/remove controls match the sibling Filter UI: a ghost, muted "Add option" button with a small plus, and an X remove icon. * improvement(tables): inline select cell uses bare DropdownMenu The inline select editor used a ChipDropdown pill, which reads as a foreign form control inside a grid cell. It now renders the canonical DropdownMenu anchored to the cell — an invisible full-cell trigger (no pill, no label text), options as their colored pills with a check on the selected ones, and "None" as a proper menu item. The cell keeps showing its pills while the menu is open instead of going blank. SelectValueEditor (the ChipDropdown pill) is now used only by the row modal, where a pill form control is idiomatic. This lets the defaultOpen/onOpenChange additions to the shared ChipDropdown be reverted — it's back to its original behavior with no consumers depending on the change. * improvement(tables): unify select into one type with a multiple flag - Collapse `select` + `multiselect` into a single `select` column type with a `multiple` boolean. A cell holds one option id when single, an array when multiple. Validation/coercion, the contract, routes, and all cell/editor code now branch on `column.multiple` instead of a separate type. - The column-config sidebar exposes an "Allow multiple" toggle (create and edit) and no longer offers "Unique" on select columns. - Give the closed select cell a dropdown affordance: it renders emcn's `comboboxVariants` chrome (border + chevron, compacted to the row height), with the pills inside and the chevron flipping while the menu is open. - Single↔multiple toggles reconcile lazily on the next row write (single tolerates an array by taking its first resolved option). * improvement(tables): revert select cell to chip-only view Drop the combobox chrome (border + chevron) from the select cell — the plain option pills read better. Reverts the comboboxVariants barrel export too, since nothing else uses it. * fix(tables): block multiple→single select switch when cells have >1 option Switching a select column from multiple to single would silently keep only the first option of any multi-valued cell. updateColumnOptions now scans the rows on that transition and errors ("N row(s) have multiple options selected") instead of dropping data, mirroring the type-change compatibility guard. * improvement(tables): rename select "Allow multiple" toggle to "Multiselect" * fix(tables): drop removed select options from cells instead of stale pills When an option is deleted, cells that referenced it previously rendered a gray fallback pill labeled with the raw internal id. They now drop the orphaned id and fall back to empty ("None"). Editors seed their selection from the still- valid ids too, so editing/saving such a cell writes the cleaned value. * improvement(tables): use dashed add-row button for select options Match the sub-block list editors (filter/sort builders): a full-width, dashed-border ghost "Add option" button instead of a plain text button. * perf(tables): don't refetch rows on metadata-only column saves; fix copy log - Add/update column are metadata-only server-side (row data never changes), so they now invalidate just the schema, not the rows. Saving a column (e.g. editing select options) no longer triggers a full rows refetch/flash; cells re-render from the refetched schema. Delete-column and workflow-group ops keep the rows invalidation since they do change row data. - "Failed to copy rows" logged an empty {} for DOMExceptions; log the extracted message so the real reason (e.g. lost transient activation) is visible. * fix(tables): migrate legacy multiselect columns; readable contract error - Tables created before the select/multiselect merge stored `type: 'multiselect'`, which the removed type no longer accepts — every response carrying such a column failed contract validation and column edits threw. getTableById now maps `multiselect` → `select` + `multiple: true` on read (covers reads, column ops via withLockedTable, and their responses); the migrated shape persists on the table's next schema write. - requestJson's "Response failed contract validation" now appends a short "field: reason" summary of the Zod issues instead of a bare message, so the failing field is visible. * revert(tables): drop legacy multiselect read-migration Only local dev tables ever held the removed `multiselect` type (the feature never shipped), so the read-time backward-compat mapping isn't warranted. Keeps the readable contract-validation error from the same change. * improvement(tables): add select options by typing into a trailing row Replace the "Add option" button with a trailing empty row: the first keystroke materializes the option and focus jumps into it (cursor at end) so typing flows straight through. Enter in an option jumps back to the trailing row to add the next one. Removing a row is unchanged. * fix(tables): export select columns as option names, not ids CSV/JSON export serialized the raw stored value for select cells — an option id (or an array of ids for multi) — so downloads showed opaque `opt_…` ids. Export now resolves ids to option names: CSV comma-joins multi names; JSON resolves values before the id→name key translation. Ids with no matching option (deleted) are dropped. * fix(tables): resolve select option ids to names across all read/consume surfaces Stored select cells hold opaque option ids; only the cell renderer and write coercion handled id<->name. Every other read/serialize/search boundary leaked the id. Centralize translation in lib/table/select-values and apply it at each boundary: - reads return names: sync export (CSV+JSON), function/snapshot mounts, clipboard copy/cut, workflow-tool + mothership reads, v1 API rows - filter/search: tolerant name->id on filter operands (row-wire INTERNAL_JWT, mothership, v1), select-aware SQL predicates (operator whitelist, multiselect empty), grid Cmd-F matches resolved names, UI filter option picker sends ids - sort: select columns order alphabetically by option name (id->name CASE) - multiselect writes split a comma-delimited string; mothership add/update_column forward options/multiple * refactor(tables): remove vestigial per-option color from select columns Colors were never exposed (no picker; every option hardcoded to gray), only kept in the model for a future picker. Drop the field entirely: SelectOption is now { id, name }, pills render a fixed neutral badge, and validation/contract no longer carry a color. Existing stored options keep a harmless color key that is stripped on the next read/write. * feat(tables): generate select option ids in the copilot handler from agent-supplied names The mothership agent authors select options by name only; the copilot tool handler now mints the stable option id (preserving any id it's re-sent, so existing cell data survives an options edit) before calling the table service — the model never authors the cell key. Applied to create (per select column in the schema), add_column, and update_column. * refactor(tables): collapse row read boundaries onto one outbound seam Turning a stored row into its outward form needs two translations — keys (column id to name) and values (select option id to name) — and every boundary composed them by hand. That shape leaked twice: table-change triggers and workflow-column/enrichment inputs both built name-keyed rows without resolving select values, so workflows saw opt_a1b2 instead of "Open". Neither had any test coverage; neither is live (select is unreleased on staging). lib/table/cell-format.ts now fuses both into a single pass: - namedRowMapper(columns) — keys + values together, so they cannot drift - fillMissingColumns(named, columns) — widens to match a headers array - mapInputValues(data, columns, mappings) — for id-keyed input mappings Migrated the 12 hand-composed pairings and the inlined copy in export-runner, then deleted rowDataIdToName, resolveRowSelectValues and selectColumnsOf so the keys-without-values shape is unrepresentable. Also formats dates in the read-only expanded popover, which rendered raw storage to viewer-role users. CSV is untouched (export-format.ts has no diff), so export bytes are unchanged. * fix(tables): stream the persisted cell value, not the raw workflow output A workflow column writing into a select column persists correctly — updateRow coerces the option name to its stored id — but it coerces its own merged copy, so the live SSE snapshot still carried the raw value. The grid resolved that name as an option id, found nothing, and rendered the cell EMPTY from the moment the workflow completed until the next refetch: the write looked like it had failed at exactly the moment the user was watching it land. Coerce a copy of the event outputs through the same inbound switch the persist path uses. A copy because the patch object is identity-compared for the progress writer's retry bookkeeping. This aligns every type, not just select — a date now streams canonically, and a value the DB would reject streams as null instead of briefly showing a value that was never stored. * fix(tables): keep option ids stable when the agent edits a select column The agent authors select options as bare names, so normalizeSelectOptionsInput minted a fresh id for every one. On update_column that replaced the whole option list with new ids — orphaning every cell that referenced the old ones and silently clearing the column's data on what looks like an additive edit (adding "Medium" to ["Low", "High"] wiped both existing values). Match incoming names against the column's current options (case-insensitively) and reuse those ids; only genuinely new options get a fresh one. Found while writing the catalog description that promised this behavior. Also regenerates the tool catalog now that the copilot registry advertises select columns, options and multiple. * fix(tables): filter multi-select columns by membership, not equality A multi-select cell stores an array of option ids, and Postgres containment does not match a scalar against an array — `{"t":["a"]} @> {"t":"a"}` is false. So every multi-select filter compiled to a predicate that could never be true and silently returned nothing. Split the select operator whitelist by cardinality: single-select keeps eq/ne/in/nin, multi-select takes contains/ncontains, both keep isEmpty. The membership clause wraps the operand (`data @> '{"t":["a"]}'`), which still uses the same GIN index. The equality shorthand (`{ tags: 'opt_a' }`) compiles to membership too — it bypasses the operator whitelist, so it had to be handled or it would keep failing silently. Also resolves $contains/$ncontains operands name → id, so an agent or API caller filtering by option name reaches the right rows, and narrows the filter UI's operator list per cardinality. * chore(tables): drop a stale doc comment and trim two restating ones The expanded-cell-popover TSDoc claimed workflow and boolean cells are read-only there, but the editability rule has no workflow check and the inline comment beside it says the opposite — a comment that contradicts the code is worse than none. The accurate rule already sits next to what enforces it. * fix(tables): close review findings on select column conversion and editing Bugbot round on the select work — eight findings, all real: - Converting away from select left opaque option ids in every cell. Migrate ids to option names in the same transaction (multi joins comma-separated), and run the compatibility check against the name, not the id. - Multiselect to text silently dropped arrays; text targets now reject structured values instead of nulling them on the next write. - Converting to single-select accepted multi-valued cells, which the next coerce would quietly truncate to the first id. Same guard updateColumnOptions already had. - Empty strings had become compatible with every target type, so a text column with blank cells could convert to number. Narrow that back to select. - Copilot update_column rejected a multiple-only payload with 'options is required', contradicting the catalog. Fall back to the column's options. - Empty multiselect rendered ChipDropdown's 'All' label, reading as if every option were selected. - Opening and dismissing an empty multiselect wrote a row update, since null and [] compared unequal. - Converting a unique column to select stranded the constraint with no way to clear it. Adds the first tests for the conversion rules. * fix(tables): migrate select cells in both directions and refetch rows after Follow-up round on the conversion fix — the migration was one-directional. - Converting TO select left cells holding option names, which every reader resolves by id, so populated cells rendered as None and matched no filter. - Toggling single→multi left scalar ids while multi filters compile to array containment, which a scalar never matches — pre-toggle rows silently dropped out of their own column's filters. - useUpdateColumn settled with the schema-only invalidation, whose premise ('stored row data never changes') these migrations broke, so the grid kept serving pre-migration cells. Both directions now share set-based helpers keyed off a jsonb map, and the to-select map keys ids as well as names so re-running is a no-op. The client refetches rows when the payload carries a type or multiple change. * fix(tables): align select cell migration with the compatibility check Two spots where the check and the migration disagreed, each leaving a cell the check had already accounted for: - resolveSelectOptionId matches option names case-insensitively, so a cell like 'open' against option 'Open' passed the convert-to-select check, but the migration map was case-sensitive and left it as the raw name. The map now keys the folded name too (duplicate names are already rejected case-insensitively, so it is unambiguous) and lookups try the exact form first. - Converting away from select, the check treats an orphaned option id as null while the migration passed it through, so a column with deleted-option cells could land an opaque opt_ id in a number/date/boolean cell. Orphans now null, matching selectValueForConversion; a multi cell drops them and nulls when nothing survives. This reverses the pass-through choice from the earlier round — consistency with the check is what matters, and an orphan is a value the UI never rendered. * fix(tables): treat an emptied multiselect as empty everywhere `[]` is the canonical empty multiselect but it is not falsy, so every emptiness check that only looked for null/undefined/'' read a cleared selection as filled: - dependent workflow groups became eligible off an empty cell, and enrichments with a required input ran on rows they should have skipped - marking a column required did not count `[]` rows, even though write-path validation rejects an empty required multiselect One `isEmptyCellValue` predicate now backs the dep, output-filled and enrichment checks, and the required-constraint guard counts '[]' too. Also normalizes '' on conversion to select — compatibility admits it, so it has to land as null (single) or [] (multi) rather than a bare string the column's own validation would then reject. * fix(tables): route an unchanged type with options to the options update updateColumnType early-returns when the type is unchanged, so an agent that restated newType: 'select' alongside a new option set had those options silently dropped while the tool still reported success. The HTTP columns route already guards this; the copilot path now mirrors it — only a genuine type change goes to updateColumnType, anything else with options or multiple goes to updateColumnOptions. A payload that only restates the current type is a no-op and now reports the live schema rather than an undefined one. * fix(tables): restore select options on column-delete undo; map v1 option errors to 400 - The delete-column undo snapshot never captured options or multiple, so re-creating a select column was rejected outright (it is invalid with no option set) and the saved cell data — which is option ids — had nothing to attach to. Undo of a select column deletion simply could not succeed. - The v1 add-column route mapped only 'already exists' and 'maximum column' to 400, while the internal route also maps invalid-column and option errors. A bad select option set surfaced as a 500 on the public API. * fix(tables): clear unique in the service when converting a column to select The clearing lived only in the column-config sidebar, so a conversion through the v1 API or the copilot tool left unique: true stranded on a select column — where the constraint compares the stored option id, capping each option at one row for the whole table, and the sidebar hides the toggle so it could never be cleared again. Moving it into updateColumnType gives all three callers the same behavior; the sidebar's payload is now belt-and-braces rather than the only enforcement. * fix(tables): clear removed select options, reject unique selects, resolve pasted names - updateColumnOptions now drops ids for options that no longer exist, so a removed option can't leave orphans that block edits on required columns - validateColumnDefinition and updateColumnConstraints reject unique on select - cleanCellValue resolves pasted option names to ids client-side so the optimistic cache holds ids and the pill renders immediately * fix(tables): refetch rows when a select option is removed Removing an option rewrites cells server-side, so the schema-only invalidation left the cache holding orphaned ids — hidden by the grid but still visible to emptiness checks, filters, and dependent-group eligibility. * fix(tables): let a multiselect round-trip through text Converting multiselect to text flattens cells to `Alpha, Beta`, but converting back read that as one unknown option and rejected the change. Both the compatibility check and the cell migration now split a comma string the way the write path already did, through one shared helper. * fix(tables): scope the empty-selection guard to multiselect; order the option ref map - cellValuesEqual treated null and [] as equal for every type, so clearing a stored [] on a json cell (or writing one into a null cell) never saved - migrateCellsToSelectIds built one flat id/name map, letting an option whose name equals another option's id overwrite that id and repoint its cells * fix(tables): prune stale select filters; reject unique+select before converting An applied filter outlives the schema it was built against, so converting a column to select (or toggling multiple) could strand an operator the server rejects, failing every subsequent rows query until the filter was cleared by hand. Prune those conditions in useTable, above every consumer of the rows query key, and share one operator whitelist with the filter picker. A payload setting both newType: "select" and unique: true ran two separate locked transactions, committing the conversion and then failing the constraint. Both routes now reject that pair before either runs. * fix(tables): scope bulk ops to the pruned filter; guard unique+select in copilot useTable pruned the filter for the rows query but kept it private, so select-all run/stop/delete still sent the raw one — targeting a predicate the grid wasn't displaying and that the server rejects. Return it and use it for every server-bound scope; the filter button and editor keep showing what the user configured so the stale rule can still be repaired. The copilot update_column path ran the same two-transaction sequence the HTTP routes now reject, so a convert-to-select plus unique payload committed the conversion and then failed. Same guard applied. Also adds the filter to selectedRunScope's deps — it was read but not listed, so a filter change while select-all was active carried a stale scope. * fix(tables): reject the multiple flag on non-select columns A create/add payload could persist multiple: true on a string column, where it sits inert until updateColumnType inherits it via `data.multiple ?? column.multiple` — silently turning an intended single-select into a multiselect and rewriting every cell as an array. Rejected now in both the service validator and the shared contract refine, which the internal and v1 column routes both consume. * fix(tables): block required blank conversions; sort multiselect by option name `required` only rejects null/undefined on a write, so a required string column legitimately holds ''. Converting it to a required select stored null (or [] for a multi), and every later update of that row then failed its own required check. Reject a blank source value when the target select is required. Multiselect sort compared the raw id-array text, since `["opt_b","opt_a"]` matches no single-id CASE branch. It now resolves elements to names and joins them in stored order — the same text the grid renders and an export writes. * fix(tables): validate a conversion against the required flag the request sets A single PATCH converting an optional column to select while also setting required: true checked blanks against the column's CURRENT required flag, so the conversion committed and the constraint write then failed — an error response with the type change already persisted. updateColumnType now takes the requested `required` and validates against the constraint the column ends up with. Empty rows are also counted when the target is required (they were skipped outright), so the null case fails up-front too rather than at the constraint write. * fix(tables): one emptiness predicate, guard required option removal, order aggregates The conversion pre-flight and the constraint write each had their own idea of "empty", and rows missing the column key are filtered out of the conversion's row set entirely — so a combined type + required PATCH passed the pre-flight, committed the conversion, then failed the constraint. Both now call one countEmptyCells helper, which is the actual fix: they can no longer disagree. Removing the last option a required cell holds nulled it, producing exactly the state updateColumnConstraints refuses to create. Blocked with a count of the rows that would be stranded. The multiselect rewrite aggregates carried no ORDER BY while sort and export preserve stored order via WITH ORDINALITY; they now do too. * fix(tables): restore in/nin on select filters and assert the whitelists agree The client operator set was a hand transcription of the server's and dropped $in/$nin, so the picker never offered them and — worse — pruneFilterForColumns silently discarded an existing $in filter the server would have accepted. Both server sets are now exported and a test maps the UI operators onto them through UI_TO_WIRE_OPERATOR, so the two can't drift again by hand. * fix(tables): stop Find crashing on a null multiselect cell jsonb_array_elements_text throws "cannot extract elements from a scalar" on a JSON null, which is exactly what a multiselect cell holds once it is cleared, cut, or has its last option removed — so search failed for the whole table. Gate the array arm on jsonb_typeof, falling back to the single mapping so a scalar left over from a single->multi toggle stays searchable. Swept the other array expansions: the rest are inside UPDATEs whose WHERE already filters to arrays, or are CASE-guarded. This was the only unguarded one. * fix(tables): clear removed options before the cardinality guard and migration The multi->single guard counted options the same request was dropping, so trimming a cell to one option and turning multiselect off in one save was rejected. Reordering fixes that and a latent corruption: the migration keeps a multi cell's FIRST element, which could be a removed id sitting ahead of a kept one, so the surviving option would be discarded and the dead one kept. Removal now runs first, against the pre-toggle cell shape, then the guard counts what actually remains, then the shape migration runs. Find also joined multiselect names unordered while sort and export preserve stored order, so a search for the displayed label could miss. * fix(tables): validate option removal against the pending required flag; show the applied filter The stranded-options guard read the column's current `required`, so a removal paired with `required: true` cleared cells and then failed the constraint write — options change committed behind an error — while a removal paired with `required: false` was blocked despite the column no longer being required. It now takes the requested value, and also pre-flights already-empty rows when required is newly imposed. The filter chip and editor keyed off the raw filter while rows, runs, stops and deletes used the pruned one, so the UI could claim an active filter the grid wasn't reflecting. Both now read the applied filter. * fix(tables): gate the unique guard on the resulting type, not on a type change Pairing unique: true with an options or multiple update on a column that is ALREADY select skipped the guard, so updateColumnOptions committed and the separate constraint write then failed — the options change landing behind a 400. Gating on `updates.type ?? currentColumn.type` covers the conversion and the options-only case with one condition, on both column routes and the copilot update_column path. * fix(tables): stop coercing select option ids in filter values Option ids are caller-supplied strings, so parseScalar turned an id of "1" or "true" into a number or boolean; JSONB containment then compared the wrong type and matched nothing while the picker still showed a valid option. filterRulesToFilter takes the columns and keeps a select value as text, for $in lists too. pruneFilterForColumns forwards them as well — its round-trip through rules would otherwise re-coerce the ids it just preserved. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
892750b4be |
fix(autolayout): keep branches in stable rows across layers (#5877)
* fix(autolayout): keep branches in stable rows across layers * fix(autolayout): container overlap on insert and error-branch alignment Lay out contained groups before their containers so a container carries the height it renders at before siblings are placed beside it. Targeted layout sized containers from a hypothetical child layout it never applied, so a container whose children stay frozen rendered taller than the space reserved for it and overlapped its neighbour. Also mark containers resized by a child edit so following blocks shift clear, and derive the error handle offset from the layout's own block height instead of the raw height field. |
||
|
|
19c3b6f47d |
feat(setup): setup wizard with browser-based Chat key handoff (#5911)
* feat(setup): setup wizard with browser-based Chat key handoff
Adds `bun run setup` and `bun run doctor` for local installs, and replaces
the wizard's paste-your-Chat-key step with a browser handoff that never puts
the key in a URL.
* improvement(setup): drop the paste-a-key fallback, simplify consent copy
The browser handoff is now the only path — the wizard waits on a spinner
instead of racing a paste prompt. Consent card leads with "Connect your
terminal" and moves the match-the-code disclaimer into the description.
* fix(setup): pin kube context, keep secrets out of argv, validate reused keys
Review findings from #5911:
- helm/kubectl now run against the validated context instead of the ambient one
- helm values are piped on stdin rather than passed as --set arguments
- ENCRYPTION_KEY/API_ENCRYPTION_KEY are checked for the 64-hex format the app
requires, not just length, so an unusable key is replaced rather than kept
- the managed Redis container's published port is read back instead of assumed
* refactor(copilot): one module for Chat API key operations
list/generate/delete each repeated the same /api/validate-key envelope in
their route. They now share callValidateKey in lib/copilot/server/api-keys.ts,
which also keeps the display masking server-side so the full key can only ever
leave at creation.
* improvement(setup): reuse shared helpers, parallelize probes, drop dead code
- PKCE verifier/state/pairing code now use generateSecureToken, generateRandomHex
and generateShortId instead of hand-rolled randomBytes; the pairing loop's
modulo was unbiased only because 256 % 32 == 0
- new sha256Base64Url in @sim/security/hash so both sides of the PKCE exchange
derive the challenge from one implementation
- isUsableSecret moved beside SECRET_KEYS so setup and doctor apply the same
rule; doctor previously passed a key setup would replace
- isTruthy narrowed to true/1, matching the app it claims to mirror — it accepted
yes/on, so a flag could read on in doctor and off in the app
- checkLive runs its five probes concurrently (~17s serial worst case)
- detection overlaps the banner animation instead of queueing behind it
- glyph.fail/glyph.warn at 13 sites that bypassed the constant; removed unused
prompter exports, a dead ENV_PATHS re-export, and an unused export keyword
* fix(setup): make doctor understand the compose env layout
Compose writes a single root .env (what docker-compose reads via env_file) but
the checks required the three per-app files, so a successful compose install was
followed by doctor printing three failures and exiting 1 — and the whole
coherence catalog was skipped because it keyed off apps/sim/.env existing.
Layout is now derived from what's on disk and every check consults it: file and
schema checks iterate the layout's targets, consistency reports skip when
there's only one file to mirror, and coherence/live read the layout's primary
file. The wizard's existing-config detection counts root for the same reason —
a compose install used to read as unconfigured and re-run from scratch.
* feat(cli-auth): device-authorization poll flow, drop the loopback listener
The CLI no longer binds a local port. It generates a request id + poll secret,
opens /cli/auth, and polls /api/cli/auth/poll over TLS while the user approves
in the browser — so the flow works over SSH and inside containers, where the
browser and terminal don't share a machine.
- approve stores the approval keyed by request id (session-authed, userId from
the session only); poll verifies the secret before an atomic claim, so an
observer of the semi-public request id can neither mint nor cancel it
- pairing code stays as the anti-phishing compare; no key ever crosses the
browser; done page just confirms
- removes the loopback listener, /token exchange, buildCliHandoffUrl, and
validateCliCallbackUrl (+ its tests) — nothing hands a key to a URL anymore
* fix(setup): reuse an existing managed Postgres container instead of colliding
A running sim-postgres fell through to `docker run --name sim-postgres` and died
on the name conflict; a stopped one failed with "no DATABASE_URL to reach it"
because the generated password only lived in the env files a fresh clone lacks.
Both facts are recoverable from Docker: the ladder now reads the published port
and password back via `docker inspect` and reuses the container (starting it if
stopped). A container that won't answer prompts before recreating, and never
drops the data volume silently.
* improvement(setup): audience-first run-mode hints
Each run mode now names who it's for — compose for self-hosting/evaluating, dev
for contributing to Sim, k8s for rehearsing a production deploy — with the live
detection state (Docker/kube/VM) appended.
* fix(cli-auth): retry a failed mint, port container/port fixes to Redis + k8s
Review findings from #5911:
- poll now reserves the mint with an atomic NX lock instead of deleting the
approval up front, so a failed mint (e.g. mothership blip) is retried by the
next poll instead of forcing a fresh browser approval; the lock still prevents
a double-mint and its TTL frees the slot if the caller dies
- setup reuses/recreates an unhealthy managed sim-redis instead of colliding on
the name (Redis has no data volume, so it removes and recreates without a prompt)
- k8s failure-path hints carry --context, matching the success-path hints, so a
changed ambient context can't send diagnostics to the wrong cluster
- compose port-free waits for a killed port to actually release before
re-checking; SIGKILL is async, so the immediate re-check re-saw the port
* fix(setup): harden mint cleanup, Windows browser, container detection, helm cwd
Review findings from #5911:
- a post-mint completeApproval failure no longer routes into releaseMint — the
mint lock now outlives the approval (shared TTL), so a cleanup blip can't leave
a re-mintable window and orphan a key; cleanup is best-effort after the key ships
- compose doctor --fix writes the feature-flag twin to the layout's primary env
(root .env on a compose install), not always apps/sim/.env
- Windows opens the browser via `cmd /c start "" <url>` — `start` is a shell
builtin, so spawning it directly ENOENT'd and the handoff never opened
- managed-container detection filters loosely and pins the exact name in code;
Docker's `name=^x$` anchor matches the internal `/x` form and often missed,
skipping the reuse branch
- the shared helm/kind run helper pins cwd to the repo root, matching helm test,
so `helm upgrade --install ./helm/sim` works from any working directory
* feat(chat-keys): standalone manage page, drop from settings nav, refresh README
- Add /account/settings/chat-keys — a linkable page to view, create, and revoke Chat API keys
- Remove Chat keys from the settings sidebar (account + unified nav) and its render branches
- README: replace Docker Compose + Manual Setup with the bun run setup wizard; drop the manual COPILOT_API_KEY step, point to the manage page
* fix(setup): per-key reason in the secret-replacement warning
Cursor: the warn hardcoded '64-character hex key', but only ENCRYPTION_KEY/API_ENCRYPTION_KEY require that — BETTER_AUTH_SECRET/INTERNAL_API_SECRET only need length >= 32. Use the existing secretRequirement(key) helper so each replaced key reports its actual requirement.
* fix(setup): compose doctor schema, cross-platform binary detection, quoted context hints
- Doctor: for the compose (root) env layout, require only the secrets compose has no interpolation default for (BETTER_AUTH_SECRET/ENCRYPTION_KEY/INTERNAL_API_SECRET). DATABASE_URL/BETTER_AUTH_URL/NEXT_PUBLIC_APP_URL come from docker-compose ${VAR:-default}, so a healthy compose install no longer fails doctor.
- Binary detection: use Bun.which instead of which (which is absent on Windows), so kubectl/helm/kind/docker resolve cross-platform.
- k8s diagnostic hints: POSIX-quote the kube-context so a context with whitespace/metacharacters can't break or inject into a copied command.
* fix(setup): quote kube-context in the helm uninstall tear-down hint too
The tear-down hint used --kube-context ${context} raw while the sibling kubectl hints already used shq(); a context with whitespace/metacharacters could break or inject into the copied command. All copyable k8s hints now go through shq(context).
* feat(setup): sim lifecycle CLI — start/stop/status/logs/down/reset
Turn the setup entry into a 'sim' command umbrella so there's one place to run everything, not scattered docker/bun commands. Adds a global bin (bun link) + a bun run sim fallback.
- Detects how you're running (compose file / managed dev containers / helm release) from disk + docker/helm state — no persisted mode. Ambiguous installs prompt.
- start/stop/restart/logs work per mode; down removes containers (volumes kept); reset archives .env + wipes managed data; both destructive verbs confirm first.
- status shows detected mode, container states, and app/realtime health.
- Wizard outro + README now point at the sim commands and the one-time bun link.
* feat(setup): 'bun run sim' is the primary entry; bare invocation prints help
- Lead usage/wizard-outro/README with 'bun run sim <cmd>' (works with zero PATH setup); global bare 'sim' via bun link is an optional upgrade, with the ~/.bun/bin PATH caveat spelled out (Homebrew's bun omits it).
- Bare 'sim' now prints help instead of launching the wizard; the wizard is 'sim setup'. The 'setup' npm script passes the keyword so 'bun run setup' is unchanged.
* fix(setup): quote the auth URL for cmd /c start on Windows
Cursor (High): cmd re-parses the command line and treats & in the query string as a command separator, so cmd /c start opened a URL truncated at the first &, breaking the key flow on win32 (the handoff URL always has request/challenge/pairing). Quote the URL and pass args verbatim so & stays literal.
* fix(setup): verify kube-context is really local; lengthen CLI handoff wait
- k8s: a context named like a local cluster (kind-*, docker-desktop) can actually point at a remote API server. Verify the server host is loopback/docker-internal before defaulting the 'use this context?' confirm to yes; otherwise warn and default to no, so generated secrets can't ship to a remote cluster on a blind Enter.
- cli-auth: bump the device-flow wait from 3 to 15 minutes so first-time users have time to sign up, wait for the email OTP, and approve before the terminal stops polling. The server-side approval record keeps its own short TTL, so a longer client wait only costs cheap rate-limited polls.
* fix(setup): only manage k8s lifecycle on a verified-local context
Greptile: sim down/reset used the ambient kube-context, so switching context after setup could uninstall a same-named sim-dev release from the wrong cluster. Gate k8sInstall on the same locality check the wizard uses (API server is loopback/docker-internal) via a shared isLocalKubeContext helper — the wizard only ever deploys locally, so a remote current-context is never treated as a Sim install.
* fix(setup): doctor skips placeholder secrets when seeding; reset names its target
- checks: the missing-file autofix copied shared keys from apps/sim/.env whenever truthy, including .env.example placeholders — doctor --fix could seed unusable secrets into realtime/db env files. Skip placeholders, matching autofixForMissing.
- lifecycle: reset now names the exact install (k8s context / compose file / dev containers) in its confirm, so a destructive reset can't silently hit the wrong same-named install after a context switch (down already names the context).
* fix(cli-auth): size the poll rate limit to the poll cadence; honor Retry-After
The poll route used the default public-IP bucket (10 burst, 5/min) but the CLI polls every 2s (30/min), so it 429'd within ~20s — worse behind a slow dev cold-compile. Give the endpoint a bucket matched to its cadence (60 burst, 60/min); it's not a brute-force surface (unknown request id returns pending, minting needs the 256-bit verifier). Also make the CLI honor Retry-After and back off on 429 so a shared-NAT per-IP limit degrades gracefully instead of hammering.
* fix(setup): check ports before starting the dev server, not just compose
Local dev auto-start spawned bun run dev:full with no port check, so it silently started a server that couldn't bind when 3000/3002 were already taken (e.g. another worktree's dev server). Extract compose's port-conflict resolver into a shared ensurePortsFree(ports) and run it before the dev start too — kill/recheck/leave, same as compose. Leaving the ports skips the auto-start with guidance instead of failing; compose still treats it as fatal.
* fix(setup): verify the kube cluster is reachable, not just local
A kubeconfig context can outlive its cluster — a kind cluster gets deleted or its Docker container stops (Docker/machine restart), but the context entry remains, pointing at a dead API-server port. The wizard checked the context looked local and handed it to helm, which failed with 'cluster unreachable'.
Add a clusterReachable() liveness probe: only offer the current context when it actually answers; if a local context is dead, fall through to the kind path. There, if kind still knows 'sim' but it's stopped, start its node containers and wait for the API; if it's gone, create fresh. Either way the user gets a working cluster instead of a cryptic helm failure.
* fix(helm): point appVersion at published image tags (v-prefixed, current)
The chart's appVersion was "0.6.73", but CI publishes GHCR tags with a v prefix (its release-commit regex captures v0.7.45). Since sim.image defaults every image tag to Chart.AppVersion, a default helm install requested ghcr.io/simstudioai/{simstudio,realtime,migrations}:0.6.73 — a tag that has never existed — so app and realtime sat in ImagePullBackOff and helm --wait failed with 'progress deadline exceeded'. Any self-hoster installing with default values hit this, not just the setup wizard.
Set appVersion to v0.7.45 (latest release on main; all three images verified present on ghcr) and bump the chart version to 1.1.1. Verified with helm lint, helm template (all images render as v0.7.45), and a live helm upgrade on a kind cluster where the new pods pull successfully while the old 0.6.73 pods remain in ImagePullBackOff.
* Revert "fix(helm): point appVersion at published image tags (v-prefixed, current)"
This reverts commit
|
||
|
|
ca77908e20 |
fix(ci): skip lifecycle scripts on CI installs (#5959)
* fix(ci): set up Node 22 for every job that runs bun install isolated-vm's install script now runs on install (#5935), and upstream only publishes prebuilds for Node 22 (ABI 127) and Node 24 (ABI 137). Jobs without an explicit setup-node inherit the runner default, Node 20, where prebuild-install finds nothing and falls back to node-gyp — which crashes on Node 20 with "webidl.util.markAsUncloneable is not a function", failing bun install --frozen-lockfile outright. This broke Create GitHub Release on main. - add setup-node 22 to ci.yml create-release and deploy-trigger-dev - add setup-node 22 to migrations.yml migrate - bump publish-cli.yml from the EOL Node 18 to 22, matching publish-ts-sdk.yml All eight jobs that run bun install now pin Node 22, matching the repo's declared engines.node >= 22.19.0. * fix(ci): skip lifecycle scripts on CI installs No CI job needs a compiled native module. isolated-vm landed in December 2025 and .npmrc blocked all lifecycle scripts until #5935, so CI ran green for ~7 months with it never built: next.config.ts and trigger.config.ts both external it, every test mocks it, and the only real require lives in isolated-vm-worker.cjs, which no CI job spawns. All three Dockerfiles already install with --ignore-scripts and rebuild it by hand. Building it in CI therefore buys nothing and couples every job to prebuild availability for the pinned Node. Upstream ships prebuilds for two ABIs only (Node 22/24), so the next setup-node bump would resurface the same opaque node-gyp failure that broke Create GitHub Release on main. - pass --ignore-scripts to all 8 CI bun install invocations - retarget the Setup Node comments at engines.node >= 22.19.0, which is the standalone reason for the pin now that scripts no longer run * fix(ci): drop redundant setup-node from bun-only jobs Once lifecycle scripts are skipped, nothing in create-release, migrate, or deploy-trigger-dev invokes node: they run bun run scripts/create-single-release.ts, bun run db:push plus bun run scripts/migrate.ts, and bunx trigger.dev deploy. All three were green without setup-node for months — the isolated-vm install script was the only thing that ever needed it, and --ignore-scripts covers that. setup-node stays where it is load-bearing: publish-cli and publish-ts-sdk need it for npm publish and its registry-url auth, and test-build/docs-embeddings already had it. publish-cli keeps the Node 18 -> 22 bump: that setup-node is required, and 18 has been EOL since April 2025, so it now matches publish-ts-sdk. |
||
|
|
cafc9aa140 |
test(data-drains): cover the migrated detail and create surfaces (#5958)
Executes the two new views rather than reasoning about them, and locks the
behaviours this migration could silently drop:
- the detail carries every column the table used to show (source,
destination, cadence, last run) and never renders credentials
- run sizes report sub-kilobyte writes instead of flooring to '0 Bytes', and
a multi-gigabyte run stays in GB
- Run now stays disabled while a drain is disabled, as the row menu did
- delete leaves only after the request resolves, and stays put on failure
- create gates on name plus a complete destination, sends the right
destination branch, and toasts on failure
- every labelled field in all seven destination forms resolves to a real
control id, and every select carries an accessible name
The first run caught a defect: dropping toLowerCase from humanizeConfigKey
left keys rendering Title Case ('Force Path Style') while the TSDoc still
promised sentence case. Now sentence case with initialisms preserved —
'Access key ID', 'Service account JSON'.
|
||
|
|
8a2ae25d78 | improvement(agent-streaming): add way to opt in for workflow executions (#5956) | ||
|
|
49bf3f2c68 |
fix(settings): a11y labels, URL-backed member search, and design-system cleanup (#5955)
* fix(settings): a11y labels, URL-backed member search, and design-system cleanup * fix(secrets): use useId for autofill salt so it survives hydration |
||
|
|
e39045d94e |
improvement(data-drains): move drains to the fullscreen list/detail pattern (#5954)
* improvement(data-drains): move drains to the fullscreen list/detail pattern Data drains was the last settings surface still on a Table plus a create modal. It now matches Skills and Custom Tools: - the list is SettingsResourceRow rows in a clickable button, with the source, destination, cadence, and last run on the row and a Disabled tag when a drain is paused - clicking a row opens a detail sub-view: the drain's actions (Run now / Test connection / Delete), its enabled toggle, resolved destination config, and recent run history — replacing the expanding table row - creating a drain is a fullscreen view instead of a modal, reusing the destination form registry so the two never fork - a drain's detail is deep-linkable via `data-drain-id`, pushed on open and replaced on close, matching custom-tool-id and custom-block-id The detail is read-only apart from the enabled toggle: destination credentials are never returned by the API, so changing one means recreating the drain. Alignment work the migration surfaced: - the destination registry rendered its 36 fields with ChipModalField, whose label is byte-identical to a SettingsSection header — on a page that made section titles and field labels indistinguishable. All of them, and the create view's own fields, now use SettingRow like every other page-level detail view - the shared source/destination/cadence label maps moved to labels.ts, now that three files need them - run status uses Badge with a dot rather than hand-coloured text, and byte counts use formatFileSize — the old /1024 math showed a 5 GB export as "5242880.0 KB" - copy follows the sibling surfaces: Create drain, Disabled, Run queued, Connection test passed, and "No drains found matching …" - seed the list cache on create so landing on the new drain's detail doesn't depend on the invalidation refetch succeeding * fix(data-drains): wire destination field labels to their controls The ChipModalField -> SettingRow migration dropped the label/control association that ChipModalField generated for free, so screen readers announced the destination fields unnamed and clicking a label focused nothing. All 35 id-capable controls now carry an id with a matching htmlFor on their row; the one ChipSelect has no id prop, matching how the existing SettingRow + combobox rows render. Also pass includeBytes to formatFileSize — without it any run writing under 1 KB reported '0 Bytes'. * fix(data-drains): name the select fields for assistive tech ChipSelect takes an aria-label that lands on its trigger button, so the four select fields no longer announce as unnamed buttons — the Datadog site plus the create view's source, cadence, and destination type. The visible SettingRow label stays; this only gives the trigger an accessible name, since ChipSelect exposes no id for htmlFor to point at. |
||
|
|
2272d4c0f0 |
fix(skills): show the new skill after creating it (#5949)
* fix(skills): show the new skill after creating it The create page navigated to the new skill's detail route, but the unsaved-changes guard immediately undid it: `isDirty` was derived from `createSkill.isSuccess`, so on success the guard's effect fired `history.back()` to pop its sentinel entry and cancelled the navigation that had just run. The header stayed on "New skill" with the fields intact, and nothing confirmed the save. The guard now exposes `release()` to retire itself for the rest of the mount, and the create page calls it before navigating with `replace` so the sentinel entry is consumed rather than stacked. Also from a cleanup pass over the same surfaces: - `useCreateSkill` resolves the created row, so the page reads `created.id` instead of re-deriving it from the response list - seed the list cache with the upsert's authoritative list verbatim; the id-merge kept the previous ordering and appended the new skill last - treat placeholder data as loading in the detail page, which could otherwise flash "Skill not found." on a workspace switch - seed credential-detail drafts on id change rather than in a value-keyed effect, so a background refetch can't clobber an in-progress edit - toast on save and on delete failure, which were both silent * improvement(skills): align the custom tools row and harden the guard lifecycle Follow-ups from an audit of the skills and custom tools surfaces. Custom tools row, now matching the skills row exactly: - the trailing arrow could be squeezed by a long description; the shared row owns `flex-shrink-0` for its trailing slot, so the five callers that hand-rolled it (and the two that forgot) all get it - drop a redundant `cursor-pointer` — Tailwind v3's preflight already sets it on `button` - the tile chrome was defined twice, in ResourceTile and again in SettingsResourceRow, despite ResourceTile's doc claiming to be the single source. Both now share RESOURCE_TILE_BASE/RESOURCE_TILE_FILL, and ResourceTile's redundant wrapper div is gone - skills' row uses the text-sm/text-caption tokens instead of literal pixels; the added gap-[1px] offsets the 21px->20px line-height so the row stays 39px Guard and cache correctness: - the delete path released the guard after the await, but the delete is optimistic — the row leaves the cache first, so the form went clean mid-flight and popped the sentinel before the release landed. Release up front and rearm if the request fails - hold the loading frame across the whole optimistic delete instead of flashing "Skill not found." on the way out - custom tools had the same placeholder-data bug just fixed in skill detail, and worse: a deep-linked id could resolve against the workspace just left - keep cached rows the create response omits, so a concurrent create isn't dropped until the refetch lands * fix(skills): retire the in-app back guard on release too release() suppressed the unload warning and the browser Back trap, but the in-app back link still keyed only on isDirty — so the Skills chip could open the unsaved-changes modal after a successful create, while the drafts were still populated and the navigation was already in flight. * fix(skills): track a sentinel consumed while the guard is released release() intentionally leaves the seeded history entry in place, but it also drops the popstate listener — so Back during the released window (an optimistic delete's round-trip) consumed that entry with nothing to record it. hasSentinelRef stayed true, and a failed delete's rearm() then skipped re-seeding, leaving the surface with no Back confirm despite unsaved edits. The released branch now keeps a bookkeeping-only popstate listener that clears the ref, so rearm() seeds a fresh entry when the old one is gone. |
||
|
|
919a98d00f |
refactor(settings): fold verified domains into SSO, move group-detail state to nuqs, design-system cleanup (#5950)
* refactor(settings): fold verified domains into SSO and use shared primitives Verified domains only gates SSO, so managing it on a separate page meant discovering the requirement after filling out the whole IdP form and then navigating away mid-setup. Move it into the SSO page as a section above the provider config and drop the standalone page, its nav entry, and both route branches. Align the surfaces with the shared settings primitives rather than bespoke chrome, matching whitelabeling/custom-blocks/access-control: - SSO's local FormField (muted labels) is replaced by the shared SettingRow, so its fields read like every other settings page. SettingRow gains optional `optional` and `error` props to absorb what FormField did — additive, so existing consumers are untouched. - The domains section is built from SettingsSection, SettingRow, SettingsResourceRow, and SettingsEmptyState instead of hand-rolled cards. Also drop the redundant Upload/Change buttons in whitelabeling: the logo and wordmark thumbnails were already clickable, so the button was a second control for the same action. Remove still appears once an image is set. * fix(settings): move group-detail view state to nuqs and clean up design-system drift Access control's group detail kept its tab, three search boxes, and three status filters in useState, so a `?group-id=` link always landed on General and a filter was lost on reload — the parent already puts the group id in the URL. The three tabs never render together, so search and status share one param each rather than carrying three mutually-exclusive keys, and switching tabs resets both. Closing the detail clears all three alongside group-id in one batched write, so nothing lingers on the list URL. Design-system fixes from a cleanup pass over the surfaces this branch touched: - Restore accessible names lost when the whitelabeling Upload buttons were removed. The thumbnail is now the only click target, and it contained just an icon, so it announced as an unlabeled button; the icon-only Remove had the same problem. Both now carry aria-labels reflecting their state. - Use Chip, not the legacy Button, for the domain actions — Button is ~26px against the 30px ChipInput beside it, so "Add domain" sat visibly short. - SettingRow now uses the emcn Info component; a bare svg as Tooltip.Trigger was neither focusable nor nameable. It also stops re-specifying Label's own default styling. - Use the new SettingRow error prop for the group name instead of a hand-rolled error paragraph, which is what the prop was added for. - Hoist the block-category lookup out of a sort comparator, size-* over h/w, name the staleTime constants the rules require, and import RowActionsMenu from its barrel. * fix(settings): alias the old /settings/domains path to SSO Folding verified domains into the SSO page dropped /settings/domains, so bookmarks and shared links 404'd instead of landing where domains now live. Both alias maps already exist for exactly this (organization/'members', subscription/'billing'); add domains -> sso to each. * chore(settings): adopt ChipCopyInput, named staleTime constants, and a11y labels * fix(settings): reset group detail params on open and drop issuer mono styling |
||
|
|
f43b52c569 |
fix(realtime): evict revoked collaborators from live workflow rooms (#5917)
* fix(realtime): evict revoked collaborators from live workflow rooms via periodic read-access re-validation * fix * fix(realtime): close join/eviction race and make sweep cleanups independent Re-authorize immediately before socket.join so an in-flight join cannot reverse a sweep eviction, and run the sweep's best-effort cleanups independently so a room-state failure cannot skip the presence broadcast. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(realtime): single-flight role resolution and retry failed eviction cleanup Coalesce concurrent role resolutions per (user, workflow) so a slow stale read can never overwrite a recorded revocation, and defer failed eviction room-state cleanups into a per-sweep retry queue so collaborators are not left with a stale presence entry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(realtime): scope room removal to the target workflow and detect swallowed cleanup failures Honor the workflowIdHint as the target room in both room managers so removing a stale room cannot clobber the mapping of a room the socket has since moved to, and confirm eviction cleanup via the returned workflowId plus an unswallowed mapping read so Redis failures actually defer into the retry queue. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(realtime): treat unconfirmed removals as failures and refresh role cache on fresh verify Treat any null removal result as a failed cleanup (the sweep always passes the target room, so null only means failure — including with expired mapping keys), move the same-room rejoin guard to a synchronous check immediately before the removal, and record verifyWorkflowAccess's fresh decision into the role cache so a re-granted user is not blocked by a stale cached revocation. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(realtime): prefer a mid-flight recorded role decision over the in-flight query result If a fresh authoritative read (join-time verify) records a decision while a single-flighted resolution's query is in flight, keep the recorded decision instead of overwriting it with the potentially stale result. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(realtime): isolate the security scan from the Redis cleanup lane Run the revocation scan (local sockets + DB only) and the best-effort room-state cleanup as independently-guarded lanes so a hanging Redis command can stall only presence cleanup, never revocation enforcement. Evictions now enqueue cleanup instead of awaiting it, and the scan no longer reads presence for a fallback role. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(realtime): bound authorization waits in the revocation scan Race each socket's authorization check against a per-socket timeout and cap the whole pass with a budget below the sweep interval, so a hanging DB query skips that socket for the pass (never evicting on uncertainty) instead of wedging the scan lane and starving subsequent ticks. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(realtime): round-robin the revocation scan so hung checks cannot starve later sockets Resume each scan pass after the last target the previous pass processed, so a fixed prefix of hanging authorization checks can never repeatedly consume the pass budget and leave sockets behind it unexamined. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9666533d00 |
fix(chat): sweep the text shimmer left-to-right again (#5926)
#5671 rewrote the shimmer-sweep keyframes to run 100% -> -100%, which is already left-to-right, but left the `reverse` from #5650 on the animation shorthand. The two flips cancel into a right-to-left sweep on every ShimmerText consumer: subagent labels, tool-call rows, and the new agent-stream thinking chrome. Drop `reverse` so direction lives in exactly one place — the keyframes. |
||
|
|
fe184d3695 |
improvement(whatsapp): validate + improve integration skill for file inputs/outputs (#5942)
* improvement(whatsapp): validate + improve integration skill for file inputs/outputs * fix lint * add whatsapp subblock migration |
||
|
|
d64739cf4f |
fix(ci): unblock @next/swc, lockfile-keyed node_modules, per-image runner sizing (#5945)
* fix(ci): key node_modules sticky disk on the lockfile hash * improvement(ci): per-image Blacksmith runner sizing + cold-build memory preflight * fix(deps): exclude @next/swc binaries from the release-age gate * fix(deps): pin @next/swc binaries so frozen installs get a compiler * docs(ci): explain ARM runner sizing rationale |
||
|
|
17d77795b4 |
feat(providers): prompt caching capability + usage-based cache pricing (#5922)
* improvement(providers): validation pass, and stream tool loop improvements * remove deploy options correctly * fix * feat(providers): prompt caching capability and usage-based cache pricing Replace the arbitrary cached-rate heuristic with a single cache-aware pricing function, and add prompt caching as an opt-in capability for Anthropic. Pricing: priceModelUsage in cost-policy.ts is now the only place cache arithmetic happens. Provider adapters normalize their wire shape into ModelUsage (input always excludes cache buckets); the pricing function never branches on provider. This removes five divergent behaviors, including the !!request.context heuristic that gave Router and Evaluator an unearned 10x input discount, and the overwrite that silently billed Anthropic cache reads and writes at zero. Also parses OpenAI cache_write_tokens, previously ignored. Caching: Anthropic gets a capability-gated advanced switch that places cache_control on the last tool and last system block; system is now always a TextBlockParam array. OpenAI gets a stable per-block prompt_cache_key with no UI, since its caching is automatic. * fix(providers): route OpenAI and Gemini block cost through cache-aware pricing Cache-aware pricing only reached trace segments. The billable block cost still called calculateCost on the cache-inclusive prompt total, so OpenAI cache hits and Gemini implicit-cache hits were charged at the full input rate and GPT-5.6+ cache writes went unbilled. Both providers now accumulate cache buckets and price through priceModelUsage, matching the Anthropic token convention where input excludes cache reads and writes. Cached counts are clamped to the prompt total so an over-reporting payload cannot bill more input than the request contained. * fix(streaming): redact tool payloads on selected outputs in public chat Redaction only ran on the empty-selection branch, but a deployment almost always selects outputs, so it was dead in the case it exists for. Selecting toolCalls streamed the raw arguments and results to a public chat client in a chunk frame, and providerTiming carried thinking content the same way. Both paths now extract from the sanitized block output rather than the raw log: the streamed selected output, which is the reachable vector, and the final envelope. Sanitizing the source rather than per selected path means a newly selectable field cannot reopen the hole. * refactor(providers): drop unreachable billing fallbacks Every provider pricing helper took a policy parameter no caller passed. Worse than dead: passing one would have double-applied the margin the central layer already applies. Removed, so providers can only price at list. Also removed guards that cannot fire. The central fallback normalized cache buckets no provider can reach it with (all three that report cache usage price themselves) and did so at a 1x write multiplier no vendor charges. priceModelUsage re-validated token counts the adapter had already clamped, and applyModelCostPolicy defaulted a required total field. Validation now happens once, in the adapter that parses the vendor payload and is the only layer that knows cache buckets are a subset of the prompt total. |
||
|
|
290e52c3f3 | fix(ci): build app on 16vcpu runner to stop cold-cache OOM kills (#5944) | ||
|
|
6ec2385928 | chore(ci): run companion-pr-check on Blacksmith (#5943) | ||
|
|
90f6708270 |
improvement(helm): hygiene pass — CI gating, strict values schema, ESO v1 default (#5939)
* improvement(helm): hygiene pass — CI gating, strict values schema, ESO v1 default, ci values Closes the gaps from a best-practices audit of the chart (template-level conformance was already clean: full label set, 82 unit tests, kubeconform- valid renders): - new Helm Chart workflow gates every chart change: helm lint, the 82 helm-unittest cases, kubeconform validation of default + all-components renders (k8s 1.29 strict, CRD catalog), and a render of all 10 example values files - helm/sim/ci/ values files (chart-testing convention) so the chart lints and templates cleanly out of the box with dummy secrets - values.schema.json declares all 30 top-level keys (16 were invisible) and sets root additionalProperties: false, so top-level typos fail fast - externalSecrets.apiVersion defaults to v1: current ESO releases removed the v1beta1 compatibility path in 2026, so the old default produced rejected manifests on new installs; NOTES/values comments updated - wait-for-postgres init container gets requests/limits (the only container in the chart without them; broke ResourceQuota'd namespaces) - drop the telemetry Prometheus scrape config for app/realtime — neither exposes /metrics, so it was dead config that also rendered a realtime target with realtime disabled - chart 1.2.0 with README upgrade notes * fix(helm): version-agnostic schema-error assertions in the secret-length suite Helm v4 phrases schema rejections as 'minLength: got N, want 32' while v3 says 'String length must be greater than or equal to 32'. The suite grepped for 'minLength' only, so it passed on local helm v4 and failed on CI's v3.16.4 — the enforcement itself works on both. Patterns now assert the key name plus either wording. Caught by the new Helm Chart workflow on its very first run. * improvement(helm): kind install test + chart version-bump gate in CI Benchmarked against the flagship OSS charts (ingress-nginx, argo-cd, kube-prometheus-stack, grafana, bitnami, cert-manager): sim already exceeds most of them on validation rigor (strict schema — 4 of 6 ship none; 82 unit tests vs argo-cd's zero; kubeconform manifest validation none of them run), but every top community chart repo actually installs the chart on a kind cluster in chart CI — the one majority practice we lacked. Adds: - install job: kind cluster, helm install with ci/default + a new small-footprint ci/kind-values.yaml overlay (default app requests of 4Gi can't schedule on a CI node), --wait, then the chart's helm test hook, with pod/event/log diagnostics on failure - version-bump job (PR-only): fails when helm/sim/** changes without a Chart.yaml version increment — argo/kps/grafana all enforce this * fix(helm): declare naming overrides in the strict schema; SHA-pin CI actions - nameOverride/fullnameOverride are consumed by sim.name/sim.fullname but were never in values.yaml, so root additionalProperties: false rejected Helm's standard naming overrides — both now declared (a helper-wide sweep confirmed they were the only template-read keys missing), with a smoke regression test that installs under the strict schema and asserts the override lands in resource names - CI supply-chain hardening: checkout/setup-helm/kind-action pinned to full commit SHAs (tag comments retained) and the kubeconform archive verified against its published sha256 * fix(helm): immutable unittest runner and fail-closed version gate in CI - helm-unittest runs via the project's official docker image pinned by immutable sha256 digest instead of a plugin install from a mutable git tag - the version-bump gate fetches the full base ref (a --depth=1 fetch could leave no merge base), computes the merge-base and diff outside the if so any git failure fails the job instead of falling into the skip branch * fix(ci): run the helm-unittest container as the runner UID The digest-pinned image runs as a non-root user that cannot write into the runner-owned bind mount (it creates tests/__snapshot__, absent from the checkout since empty dirs aren't tracked). Standard bind-mount pattern: --user "$(id -u):$(id -g)" with HOME=/tmp for helm's cache. * fix(helm): appVersion points at a real GHCR tag (v0.7.44) The kind install job caught this on its first full run: the default image tag (Chart.AppVersion 0.6.73) returns 404 on GHCR for all three images — the registry's tags are v-prefixed — so an unpinned default install could never pull. Updated appVersion to v0.7.44 (verified 200 for simstudio, realtime, and migrations manifests), stale values comment refreshed, and an upgrade note added. The CI kind values deliberately stay tag-free so the job keeps exercising the true default path. |
||
|
|
f0f98708fc | refactor(blocks): rename Sim block to 'Sim Chat' for clarity (#5933) | ||
|
|
f224d02e27 |
improvement(library): align enterprise copy with canonical source (#5936)
Reconcile compliance and traction wording across three library guides against the canonical data in lib/compare/data/sim.ts, and tighten a couple of related phrasings for consistency. |
||
|
|
91a9c20acc |
fix(deps): build isolated-vm on install via root trustedDependencies (#5935)
.npmrc set ignore-scripts=true globally, which overrode Bun's trustedDependencies allowlist and blocked every lifecycle script — including the repo's own root scripts. isolated-vm therefore shipped unbuilt and each contributor had to npm rebuild it by hand. The .npmrc line landed in the May 2025 Bun migration, seven months before isolated-vm was introduced. Two later commits added isolated-vm to trustedDependencies (root, then apps/sim) and both were silently inert. - delete .npmrc; lifecycle policy now lives solely in root trustedDependencies - add isolated-vm to root trustedDependencies - drop the workspace-level trustedDependencies from apps/sim (Bun only reads the root list, so it never had any effect) - drop the .npmrc entry from CODEOWNERS Naming any package in trustedDependencies replaces Bun's curated default-trust list, so only isolated-vm and sharp may run install scripts. Docker is unaffected — all three Dockerfiles pass --ignore-scripts explicitly and rebuild isolated-vm against the Node ABI by hand. |
||
|
|
5a8d21904d |
fix(sso): surface DNS verification failures and the provider auto-append gotcha (#5931)
* fix(sso): surface DNS verification failures and the provider auto-append gotcha Post-merge audit follow-ups for verified domains (#5909): - The host field handed admins the FQDN `_sim-challenge.acme.com`. GoDaddy, Namecheap, Hover and most cPanel panels append the zone to whatever is typed, yielding `_sim-challenge.acme.com.acme.com` — the record looks right in their panel but never verifies, and our 422 tells them to wait 48 hours. Add a hint under the field (and a docs callout) telling those admins to enter just the label. - DNS failures were logged at debug, but production log level is ERROR, so a blocked-egress or SERVFAIL condition was invisible to us and misreported to the admin as "record not found yet". Log infrastructure-class failures at warn with the DNS error code; keep the genuinely-absent codes at debug. - Trim the joined TXT value before comparing: several DNS panels pad the stored string, which otherwise fails an exact match forever. - Lower the resolver to 2s/1 try. c-ares multiplies timeout across servers and retries by ~7x, so the previous 5s/2-try config could block a verify request for ~35s when resolvers are unreachable. - Cover `checkDomainTxtRecord` — the function that decides whether the gate opens had no tests. Adds exact match, chunk-joined value, match among unrelated records, padded value, near-miss, another org's token, absent record, infrastructure failure, and empty-response cases. * fix(sso): log DNS infrastructure failures at error so prod actually surfaces them Production's default minimum log level is ERROR, so the warn introduced in the previous commit was still filtered out — the fault stayed invisible exactly as before. A resolver failure that is not 'record absent' is a genuine infrastructure error, so ERROR is both the visible and the honest severity. * fix(sso): state the zone-removal rule instead of a wrong subdomain hint The hint computed the bare label as the first segment of the challenge host, so for a subdomain like eng.acme.com it advised entering `_sim-challenge` when the host relative to the acme.com zone is `_sim-challenge.eng` — following it would publish the record on the wrong name and verification would never succeed, the exact failure the hint exists to prevent. Deriving the real zone needs the Public Suffix List (acme.co.uk defeats naive label-stripping), so state the rule instead: enter the host with the trailing zone removed. Docs show both the apex and the subdomain form. |
||
|
|
6105376593 |
improvement(ship): pre-flight regenerate artifacts + parallel audit suite (#5927)
* improvement(ship): pre-flight regenerate artifacts + parallel audit suite Add a two-phase pre-ship step: (A) regenerate every committed artifact (agent-stream-docs, skills, contract syncs) in parallel so generated files never drift into a CI failure, then (B) run lint plus the full audit suite CI's Lint and Test job enforces, in parallel, aborting ship on any failure. Regenerate the ship command projections. * fix(ship): scope Phase A to in-repo generators + propagate failures Cursor review: (1) mship:generate is an umbrella over the 9 contract generators and biome-formats the shared generated dir, so running it alongside its constituents races/corrupts; it also reads an external copilot-contract source that ENOENTs in a bare worktree. (2) bare wait swallowed generator exit codes. Narrow Phase A to the always-in-repo generators (agent-stream-docs, skills) and collect each job's exit status in both phases; document domain generators as on-demand only. * fix(ship): make pre-flight status checks exit non-zero on failure Cursor/Greptile: grep && echo ❌ || echo ✅ always exits 0, so a failed generator/audit didn't actually gate ship (regressed the sequential bun-run checks whose non-zero exit agents rely on). Replace both Phase A and Phase B aggregations with an if-grep that exits 1 on any failure. * fix(ship): gate lint exit in Phase B pre-flight Cursor: bare bun run lint before the audit grep meant a non-zero lint was ignored — if audits then passed, the block exited 0 and ship continued. Gate lint with || { echo …; exit 1; } so an unfixable lint error aborts before the audits run. |
||
|
|
e83d8b84b4 |
content(library): add observability, procurement, and MCP security guides (#5928)
* content(library): add observability, procurement, and MCP security guides Three AEO guides, each dated into a recent empty day in the posting calendar (2026-07-19, 07-21, 07-22) so publication is spread across days rather than dumped on one. - ai-agent-observability: what observability is, why APM falls short for non-deterministic agents, what to instrument per lifecycle stage, and the signals to track - ai-agents-in-procurement: what procurement agents do, buy-vs-build, where they add value, and how to start narrow - mcp-security: tool poisoning, confused-deputy/OAuth flaws, and supply-chain risk, plus how to build and govern MCP servers securely Every third-party factual claim carries an outbound citation to a verified primary source (OpenTelemetry semconv, Fiddler, LangChain and PwC surveys, Icertis/ProcureCon, Ironclad, GEP, the MCPoison/CurXecute CVEs on NVD, Anthropic's MCP announcement, RFC 8707/9700, OAuth 2.1, the MCP auth spec), and each post carries 3 internal links to related library posts. One claim from the source copy — a "November 2025" date on the WhatsApp MCP exfil case — could not be verified and was dropped; the case itself is cited to Docker's writeup. Covers come from the autogeneration pipeline via the standard /library/<slug>/cover.jpg path. * fix(library): correct two unverifiable claims in the new guides Accuracy pass on the three new posts turned up two claims that could not be substantiated: - HIPAA compliance: the procurement post claimed Sim has "SOC2 and HIPAA compliance," but the canonical compliance data (lib/compare/data/sim.ts) states SOC2 only, and explicitly that Sim offers self-hosting "beyond SOC2, rather than additional certifications." Removed HIPAA; kept SOC2 plus self-hosting for data residency. (The same claim exists in ~5 pre-existing library posts and should be corrected separately.) - "more than 18,000 servers were listed on MCP Market": no source substantiates this figure. Replaced with "thousands of community-built servers," which the ecosystem supports, keeping the Anthropic and MCP Market links. Every other third-party claim was verified against a primary source (both CVEs on NVD, RFC 8707 title, RFC 9700 as January 2025, and all six survey statistics). |
||
|
|
70313cd2ad |
feat(providers): add Claude Opus 5 model (#5925)
* feat(providers): add Claude Opus 5 model * fix(providers): route Opus 5 through adaptive thinking Opus 5 supports only adaptive thinking (no extended thinking / budget_tokens), same as Opus 4.8/4.7. Add opus-5 to supportsAdaptiveThinking so thinking requests use thinking.type: adaptive instead of falling through to budget_tokens extended thinking, which Opus 4.7+ rejects with a 400. * docs(agent): regenerate agent-stream docs for Opus 5 |
||
|
|
513292f17b |
feat(sso): DNS domain verification gating org SSO registration (#5909)
* feat(sso): DNS domain verification gating org SSO registration
Add org-scoped domain ownership verification (DNS TXT challenge) as the
security precondition for configuring SSO. Closes the first-come domain-claim
vuln where any org could wire another company's domain to its own IdP.
- New sso_domain table + migration 0266; existing org SSO domains are
grandfathered as verified so live tenants are unaffected
- Verified-domains settings UI (enterprise-gated) with add/verify/remove
- Register route now requires a verified domain for org-scoped registration;
personal SSO and already-grandfathered domains are unaffected
- Self-host register script writes the verified sso_domain row directly, so
script-driven registration stays backwards compatible
* fix(sso): harden domain verification against concurrency + fix CI lint
Addresses review findings on state invariants under concurrent/failed writes:
- Add unique index on (organization_id, domain) so concurrent claims can't
create duplicate pending rows; POST re-reads and stays idempotent on conflict
- Verify flips the row only if it's still the exact pending challenge checked
(guards deletion/token-rotation mid-DNS-lookup) and maps the partial unique
index violation to 409 instead of an unhandled 500
- Wrap the self-host script's provider write + verified-domain upsert in a
transaction so a failed ownership write can't leave a provider committed
- Format 0266 snapshot/journal with biome (fixes @sim/db lint:check)
* fix(sso): re-check domain verification before provider write (TOCTOU)
The register gate checked the verified sso_domain row only at handler entry,
then ran OIDC discovery before writing the provider. A verified row removed
during that window could still complete registration. Extract the check into a
closure and call it both as an entry fast-fail and authoritatively right before
registerSSOProvider, alongside the existing domain-conflict re-check.
* fix(sso): stop rotating verification token on idempotent re-add
Re-adding a pending domain rotated its verification token, which invalidated a
TXT record the admin may have already published and — under two concurrent
re-adds — could return a token the racing write had already superseded, so the
admin's DNS record would never verify. Return the existing row unchanged
instead; the pending token is always shown in the UI, so it is never lost.
* fix(sso): close register TOCTOU with compensating delete + harden edges
Audit-driven hardening:
- Close the residual register TOCTOU: registerSSOProvider is create-only (throws
if the providerId exists), so a compensating delete after the write is provably
safe — it can only remove the just-created row. If verification was revoked
during the write, roll the provider back and 403.
- Verify is now idempotent under concurrency: a same-org row already flipped to
verified by a racing request returns 200, not a confusing 409.
- Grandfather backfill + self-host script now match normalizeSSODomain's dominant
transforms (lower + trim + strip leading wildcard) so a non-canonical legacy
domain can't miss the runtime gate's lookup. Prod backfill result is unchanged.
- Cleanup: drop dead default export, align card radius to sibling convention.
* fix(sso): redact domain tokens from non-admins + fix script stale-update
Round-5 review findings:
- GET /domains redacted the pending TXT verification token (a management
secret) to any org member. Now only owner/admins read it; members see the
list and status without it. Non-Enterprise orgs get an empty list (entitlement
flag only), never the domains/tokens.
- Self-host script decided update-vs-insert from a read taken OUTSIDE the
transaction; a provider deleted mid-flight made the UPDATE match zero rows
silently while the verified-domain upsert still committed (orphaned domain).
The decision now happens inside the transaction from the UPDATE's row count.
* docs(sso): drop unshipped enforce-SSO / auto-join copy from verified domains
Verified domains currently only gate SSO configuration. Remove the
forward-looking references to enforcing SSO and auto-joining members (deferred
to a later release) from the docs, settings copy, nav description, and schema
comment so we don't promise unshipped features.
* fix(sso): guard rollback to new providers only + Enterprise-gate domain removal
Round-6 review findings:
- The compensating provider rollback now only fires when the provider did not
exist before this request (providerExistedBefore). registerSSOProvider is
create-only today so reaching the rollback already implies a fresh create, but
this makes the safety local and future-proof: if Better Auth ever allowed
updating an existing provider, a revoked-verification rollback must not delete
that pre-existing row.
- DELETE /domains now requires an Enterprise plan like add/list/verify, so all
domain mutations share one entitlement (the UI already hides removal from
non-Enterprise orgs). Adds a delete-route test.
* fix(sso): roll back the SSO provider by row id, not logical keys
The compensating rollback deleted by (providerId, orgId). providerId is unique,
so if this request's row were deleted and recreated by a concurrent registration
in the narrow window before the rollback, the logical-key delete would remove
that other request's provider. Delete by the primary-key id registerSSOProvider
returns instead, so only the exact row this request created is ever removed.
* chore(sso): final-review polish — trim script read, unify copy, doc migration edge
Cosmetic cleanup from a final 4-track adversarial review (no bugs found in the
new logic):
- Self-host script: narrow the pre-transaction existence read to select({ id })
instead of SELECT * (it only feeds a log line now).
- Unify invalid-domain copy ("for example acme.com") and the verified-elsewhere
409 wording ("is already verified by another organization") across routes.
- p-3 shorthand on the domain row card.
- Document the migration's rare two-orgs-share-a-domain grandfather behavior
(login unaffected; validated no such duplicates in prod).
* fix(sso): apply attribute mapping + make SSO edit work; drop dead guard
Two pre-existing SSO bugs the final review surfaced (prod has one SSO org, RVW,
script-registered with the default mapping, so neither change affects it):
- Attribute mapping was passed at the top level of the register payload, which
Better Auth ignores — it reads oidcConfig.mapping / samlConfig.mapping. Nest it
so custom mappings actually apply. (Default mapping is unchanged, so existing
logins are unaffected.)
- Editing an SSO provider was broken: registerSSOProvider is create-only and
threw on the existing providerId → generic 500. Route now detects a provider
the caller already owns and updates it via Better Auth's updateSSOProvider, and
surfaces Better Auth's own error status/message instead of a blanket 500.
Also drops the now-unnecessary providerExistedBefore guard (the rollback deletes
by the created row's primary-key id and register is create-only) and the earlier
final-review polish (script read, unified copy, migration edge note).
Smoke-test SSO login + edit on staging before merge (auth-path change).
* fix(sso): require null org on personal-mode provider lookups (gate bypass)
The personal branch of both provider-ownership lookups keyed on
(providerId, userId) without requiring organizationId IS NULL. Because org
providers store userId = their creator and providerId is globally unique, an org
admin could send a personal-mode request (no orgId) — which skips the membership
check and the domain-verification gate — yet still match, and then via the new
update path move, their org's provider to an unverified domain. Add
isNull(organizationId) to the personal branch of both clauses so it can only
match a genuinely personal provider, matching the route's own isOwnedByCaller.
Found by an adversarial review of the update path added in 394bda9f7.
* fix(sso): script updates the observed provider by id, not providerId
Inside the registration transaction the script updated WHERE providerId — the
logical key. If the observed provider was deregistered and a replacement created
with the same providerId before the transaction ran, that update would clobber
the replacement's config and ownership. Update the specific observed row by its
primary-key id instead; if it's gone we insert, which fails cleanly on the
providerId unique constraint rather than overwriting the replacement.
* fix(sso): script upserts provider via delete-then-insert (no unique constraint)
sso_provider.provider_id is a plain (non-unique) index and prod holds legitimate
duplicates, so the previous "update by id, else insert" could create a duplicate
provider when the observed row was deregistered and replaced before the
transaction — the fallback insert would succeed. Delete every row for the
providerId then insert exactly one, inside the transaction, so the providerId
ends up as exactly this config atomically. Linked accounts key on the providerId
string (not the row id), so existing logins are unaffected.
* fix(sso): guard compensating-delete row id so rollback can't silently no-op
* chore(sso): regenerate migration as 0268 after merging staging
Staging landed migrations 0266/0267, colliding with our 0266. Removed our
migration, merged staging, and regenerated cleanly with drizzle-kit as
0268_sso_domain_verification (identical sso_domain table + indexes), then
re-appended the grandfather backfill. api-validation baseline reconciled to 973
(staging 970 + our 3 domain routes). Also make the register-route test's
registerSSOProvider mock return an id so the guarded compensating delete runs.
* refactor(sso): share normalizeSSODomain via @sim/utils so script matches gate
The self-host script canonicalized SSO domains with a minimal inline transform
(lower+trim+wildcard) that diverged from the app's full normalizeSSODomain
(protocol, port, path, trailing dot, email local part) — equivalent spellings
could store a different ownership key than the runtime gate looks up. Move
normalizeSSODomain into @sim/utils/sso-domain (a pure function) so the register
route, the domain-claim route, and the script all use the identical canonicalizer.
The script now skips the verified-domain record when SSO_DOMAIN isn't a valid
registrable domain instead of storing a malformed key.
|
||
|
|
444c415a0b |
improvement(data-retention): docs for overrides + PII redaction, fix wedged saves (#5905)
* fix(data-retention): clamp sub-day retention values so saves aren't wedged A stored value under 12 hours rounded to '0' days on load and was re-sent as 0, which the contract rejects (min 24) — blocking every save on the page, including unrelated fields. Clamp hours->days into the contract's range on read, and throw instead of emitting 0/NaN on write. * improvement(docs): rewrite data retention for workspace overrides + PII redaction - Document PII redaction (Logs / Workflow input / Block outputs stages, entity types, languages, custom regex patterns) - Replace the stale 'no per-workspace overrides' section with the retention-policies list and override inheritance - Correct log retention (also covers background job logs) and soft deletion (adds Chat conversations, KB documents) - Add PII + override screenshots, refresh the main one |
||
|
|
e6eef4a4bf |
feat(library): What Is an MCP Server? (#5919)
Co-authored-by: Sim Pi Agent <pi@sim.ai> |
||
|
|
7120cdc901 |
fix(providers): final regenerated stream must not re-call tools (empty chat answers) (#5915)
* fix(providers): stop the final regenerated stream from re-calling tools and clobbering the answer
After the silent tool loop settles, OpenAI Responses and Gemini re-issue a
streaming request purely to stream the answer as prose — but with tools still
attached and auto tool choice, a reasoning model can re-decide to call a tool
there. Streamed calls are never executed on this path, so the run ends with a
dead function call, an empty streamed answer, and the stream callback
clobbering the tool loop's settled text with '' (deployed chat rendered
{"content": ""}).
Force tool_choice 'none' / functionCallingConfig NONE on the regeneration and
keep the tool loop's settled answer whenever the stream ends without text.
* feat(openai): settled tool chips on the regenerated answer stream
The silent Responses tool loop has no live stream while tools run, so opted-in
consumers saw no tool chips at all for OpenAI. The loop now records each
executed call and prepends settled tool_call_start/end pairs (name + status
only) to the agent-events stream ahead of the regenerated answer. Runs without
a sink never see these events, so legacy output is unchanged.
* fix(providers): apply the regeneration guard fleet-wide
Audit of every silent tool loop for the same race fixed for OpenAI/Gemini
(final regenerated stream re-calls a tool that is never executed, ending with
an empty answer that clobbers the settled text):
- anthropic (both implementations): tool_choice {type:'none'} on the
regeneration (tools must stay — history carries tool_use blocks) + keep the
settled answer when the stream ends without text
- groq: was re-applying the ORIGINAL tool_choice, so forced-tool runs
re-forced the tool on the regeneration — guaranteed dead call; now 'none'
+ fallback
- deepseek, mistral, cerebras, azure-openai (legacy chat path), openrouter,
xai: 'auto' -> 'none' + fallback
- bedrock: fallback only — Bedrock's ToolChoice has no 'none' and toolConfig
is required when history carries toolUse blocks
Already guarded (no change): meta, sakana, nvidia, vllm, litellm, baseten,
together, fireworks, kimi, zai, ollama.
|