Files
sim/scripts/check-api-validation-contracts.ts
T
Vikhyath MondretiandClaude 7798e83489 feat(function): custom sandboxes (#6071)
* feat(sandboxes): workspace dependency sets for Function blocks

Named package sets a Function block can import from. The server
canonicalizes and hashes the list; E2B prebuilds a content-addressed
template per set, Daytona installs per execution. Create/edit is gated to
Max or Enterprise via the shared workspace entitlement check; execution is
deliberately ungated, so a downgraded workspace keeps running what it
already built.

Also on this branch:

- Extract the duplicated dropdown/combobox option-fetch lifecycle into
  use-fetched-options. Only combobox had the dependency-change reset, so
  every dropdown with dependsOn + fetchOptions cleared its list and never
  repopulated until reopened.
- Collapse the repeated Max-tier entitlement check onto one
  hasMaxTierWorkspaceAccess, shared by inbox, live sync, and sandboxes.
- Resolve a personal payer's block state through getEffectiveBillingStatus
  in getBillingEntityBlockStatus, so the client-side Max gates agree with
  the server-side ones when blockOrgMembers' fan-out is stale.
- Carve the Daytona dependency install out of the caller's execution
  budget instead of stacking on top of it.

Co-Authored-By: Claude <noreply@anthropic.com>

* chore(db): regenerate the sandboxes migration as 0273

Staging claimed 0271 and 0272 while this branch was out, so the hand-authored
0271_workspace_sandboxes was dropped before the merge and regenerated on top
of the merged schema. Same DDL; drizzle emits plain CREATE TABLE/INDEX rather
than the hand-added IF NOT EXISTS, which matches the repo default — that
idempotent form is only needed for files with CONCURRENTLY ops below an
embedded COMMIT. Regenerating also restores the meta snapshot the
hand-authored migration never had.

Co-Authored-By: Claude <noreply@anthropic.com>

* chore(db): drop the sandboxes migration ahead of the staging merge

Staging independently claims idx 0273, so remove ours before merging to
avoid an add/add conflict on the drizzle migration index. Regenerated at
the next free index once the merge lands.

Co-Authored-By: Claude <noreply@anthropic.com>

* chore(db): regenerate the sandboxes migration as 0275

Staging took 0273 and 0274, so the sandboxes DDL lands at the next free
index. The emitted SQL is byte-identical to the dropped 0273.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(billing): consolidate the Max-tier entitlement onto one predicate

The Max tier was spelled five ways. The odd one out — `isMax`, defined as
`isPro(plan) && credits >= 25000` — excluded both `team_25000` and
`enterprise`, and it was the sole input to the personal-workspace cap. A
delinquent Max-for-Teams org admin got 1 personal workspace while a
delinquent Max individual got 10. Only free/pro_6000/pro_25000 were tested,
so the two broken tiers were unpinned.

Separately, the server gate and the client `hasUsableMaxAccess` were
independent copies of the same rule. The settings sidebar renders Sandboxes
and Sim Mailer from the client one while the API answers 403 from the
server one, so any drift renders a feature unlocked that the API refuses.

- `MAX_TIER_CREDITS` is derived from the `CREDIT_TIERS` table; `isMaxTier`
  in plan-helpers is now the single definition, shared by the server gates,
  the client derivation, `getPlanTypeForLimits`, `plan-view`, and the cap
- `hasWorkspaceTierAccess(id, predicate, { intent, onMissingWorkspace })`
  becomes the one org-vs-personal payer fork. `intent: 'active-use'` means
  active and not billing-blocked; `'retention'` means active/past_due with
  block state ignored, so the inbox teardown guard keeps its fail-open
  semantics instead of implying them through a duplicated fork
- `isWorkspaceOnEnterprisePlan`'s personal branch now applies the status and
  block checks its own org branch always had, and its TSDoc names its real
  consumer (copilot BYOK, not Access Control)
- the client live-sync gate gained the server's `isHosted` branch, so a
  self-hosted deploy with billing on no longer locks an interval the API
  accepts. It reads both flags directly rather than taking one as a
  parameter the callers sourced from the same module
- `sqlIsPro`/`sqlIsTeam` escape the `_` LIKE wildcard, matching the already
  correct hand-rolled filter in seat-drift
- deletes the `TERMINAL_SUBSCRIPTION_STATUSES` and `ENTITLED_STATUSES`
  shadow constants, and corrects three test mocks that asserted `trialing`
  was entitled or usable

`max-tier-parity.test.ts` asserts the client and server answers match for
every plan name. Both new guards were checked against the old code: the
parity test fails 3 assertions with the previous predicate, and the
self-hosted test fails without the `isHosted` branch.

Co-Authored-By: Claude <noreply@anthropic.com>

* chore(db): drop the sandboxes migration ahead of the staging merge

Staging has claimed 0275 (table_views) and 0276 (drop_legacy_folder_tables)
since the last merge, so our 0275_workspace_sandboxes collides on the index.

Dropping ours first — the .sql, meta/0275_snapshot.json, and the journal
entry — leaves packages/db/migrations byte-identical to the merge-base, so
the merge sees no add/add conflict at all. Regenerated on the far side.

Ours is the droppable side: plain additive DDL with no hand edits, which
drizzle reproduces exactly. Staging's migrations are hand-written and must
survive.

Co-Authored-By: Claude <noreply@anthropic.com>

* chore(db): regenerate the sandboxes migration as 0277

Staging claimed 0275 (table_views) and 0276 (drop_legacy_folder_tables), so
the sandboxes migration dropped before the merge comes back on top as 0277.

The emitted SQL is byte-identical to what was dropped — the original had no
hand edits, so there is nothing to reapply. It is purely additive: two enums,
sandbox_image and workspace_sandbox, their two FKs and six indexes. That it
regenerated unchanged also confirms the schema.ts auto-merge was correct —
had it lost staging's legacy-folder-table drops, drizzle would have emitted
CREATE TABLE for them here.

Snapshot chain is continuous (0273 -> 0277, each prevId matching the previous
id) and the table counts track the DDL: 100 -> 101 (table_views) -> 99
(legacy folder tables dropped) -> 101 (the two sandbox tables).

Co-Authored-By: Claude <noreply@anthropic.com>

* feat(sandboxes): gate on the enterprise feature flags, drop the rollout switch

Sandboxes shipped behind `custom-sandboxes`, an AppConfig rollout flag falling
back to a `CUSTOM_SANDBOXES` secret. That made it the only Max-gated surface
with no self-hosted path: `INBOX_ENABLED` can force Sim Mailer on for an
operator running their own billing, and `ENTERPRISE_ENABLED` turns on the
other nine features at once, but neither reached sandboxes. A self-hoster had
to find a separately-named variable that was not part of that family, and one
running with billing enabled could not enable it at all.

Sandboxes now joins the enterprise feature set and the rollout flag is gone:

- `sandboxes` is an `EnterpriseFeature` with `SANDBOXES_ENABLED` and its
  `NEXT_PUBLIC_` twin, so the master switch and the per-feature override both
  reach it like every sibling
- `hasWorkspaceSandboxAccess` takes the inbox's shape exactly — the override
  wins, then a deployment without billing is unrestricted, then the workspace
  payer needs usable Max or Enterprise
- the settings nav gains `selfHostedOverride`, so the section resolves through
  the same path as Sim Mailer instead of a second entitlement AND-ed in
- `custom-sandboxes`, the `CUSTOM_SANDBOXES` secret, the now-unreachable
  `SANDBOXES_UNAVAILABLE` 403 copy, and the route's kill-switch branch are
  deleted

Its legacy default is `true`, matching `inbox`: the gate already returns true
whenever billing is off, so `false` would leave the nav override disagreeing
with the gate that answers the request. Self-hosted builds run on the
operator's own E2B/Daytona credentials, so there is no Sim-side cost to
withhold — the docs now say so, since enabling the feature without a provider
configured is the obvious trap.

The new gate tests run with billing enabled on purpose; the `!isBillingEnabled`
bail would otherwise answer every case and hide whether the override is wired.
Verified by deleting the override line — exactly the one assertion fails.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): let the language menu match its trigger width

`matchTriggerWidth={false}` exists for the opposite case — a narrow trigger
whose option labels would truncate, letting the menu grow past it. The language
field is a full-width form control with two short labels, so the override
shrank the menu to "JavaScript" and pinned it to the right edge instead.

The default (`true`) is correct here. Every other consumer passing `false` is a
genuinely narrow trigger — a role picker in a member row, a table filter chip.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): re-queue a build when resolution finds the image unusable

`ensureSandboxImage` only ran when a sandbox was saved, so resolution treated
an unusable image as terminal and told the user to go fix a definition that was
never wrong. Three states stuck permanently until someone re-saved in Settings:

- a build that failed
- a build whose worker died mid-flight, stranding the row in `building`
- every sandbox created while the deployment ran a `runtime` provider, after a
  switch to a `prebuilt` one — `runtime` writes no image rows at all, so the
  whole fleet resolved to "no completed build" with nothing to repair it

Resolution now re-queues through the registry's existing idempotent entry point
before failing, and says a build is on its way instead of pointing at Settings.
The conflict guard already claims only a `failed` row or a stale `pending`/
`building` one, so executions arriving during a healthy build enqueue nothing —
no thundering herd from a hot workflow.

The registry is imported dynamically for the same reason `sandboxDb` is: it
pulls `@sim/db` into the static graph, which this module keeps out of the
executor bundle. That also avoids a cycle, since the registry imports
`invalidateSandboxResolution` from here. A repair that itself fails is logged
and swallowed — it must never replace the build error naming the sandbox.

Verified by deleting the repair call: exactly the three new assertions fail.

Co-Authored-By: Claude <noreply@anthropic.com>

* improvement(sandboxes): let the picker show just the sandbox name

The label read "Test · Python · 1 package". The block's own list is already
scoped to the language its sibling `language` subblock selects, so the language
repeated on every row said nothing, and the package count is decoration next to
the name that identifies the sandbox.

The language stays for the one caller that cannot filter — agent tool-input
renders this field under a synthetic id where the sibling `language` value is
unreachable, so its list spans both languages and the name alone is ambiguous.
That is the same missing value which disables filtering, so `showLanguage` is
derived from it directly rather than passed independently and left to drift.

A failed build is still marked: that suffix is the difference between a
selection that runs and one that does not.

Passing the flag also means dropping `.map(toSandboxOption)` for an explicit
arrow — `Array.map` hands the index to the second parameter.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): show the sandbox name on the block card, not its uuid

The card printed "443f4934-26ab-44ab-8...". `resolveDropdownLabel` only reads a
subblock's static `options` array, and the sandbox picker is a `combobox` whose
options load asynchronously, so its array is empty and the raw stored id fell
through to the label.

Resolved the same way skills and tools already are: a `resolveSandboxLabel` in
the display layer, fed from the shared sandbox list query — the same cache entry
the picker reads, so this adds no request.

Two deliberate scopings:

- the query is subscribed only for the sandbox row. `SubBlockRow` is memoized
  per subblock, and the list query polls while a build is in flight, so an
  unconditional hook would re-render every row on the canvas on each poll tick
- the resolver matches the field id, not just the type. There is no dedicated
  subblock type for it, and matching `combobox` alone would relabel unrelated
  pickers

An id with no matching sandbox resolves to null rather than a guess, so a
deleted sandbox falls through to the caller's placeholder. The template preview
surface is left alone: it is explicitly hook-free and passes empty lists for
tools and skills too.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): hide the Sandboxes section with no provider configured

Entitlement decides whether a workspace may author sandboxes; nothing decided
whether anything could run one. A self-hosted deployment with SANDBOXES_ENABLED
but no E2B or Daytona credentials got a fully functional tab whose output no
Function block could select — the picker is gated on the provider vars, the tab
was not.

Both navigation planes now drop the section when neither
NEXT_PUBLIC_SANDBOX_ENABLED nor the pre-Daytona NEXT_PUBLIC_E2B_ENABLED is set —
the same pair the picker's `showWhenEnvSet` reads, so the two cannot disagree.
Dropped rather than locked: an upgrade does not conjure a provider.

The unified plane drops it in `buildUnifiedSettingsNavigation` rather than in the
sidebar's filter, because the sidebar's `selfHostedOverride` short-circuit runs
before its `requiresMax` check and would have revealed the tab anyway. It reads
the browser twins, not the server's `isRemoteSandboxEnabled`, since this module
renders on both sides.

The predicate is a function, not a module constant, because the constant form was
untestable and ambient: the env mock falls through to `process.env`, and
`apps/sim/.env` (gitignored, so absent on CI) sets NEXT_PUBLIC_E2B_ENABLED=true.
The nav tests passed locally and failed 6 assertions with the flag cleared. They
now pin both flags, so the suite is identical with and without a local env file —
verified by running it both ways.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(sandboxes): correct three claims the code no longer makes

The Sandboxes section described behavior two commits on this branch changed, and
led with an internal detail no reader needs.

- entitlement is no longer Max/Enterprise only: self-hosted deployments unlock
  sandboxes with SANDBOXES_ENABLED, and the section is hidden outright when a
  deployment has no sandbox provider, which is the state a self-hoster is most
  likely to hit and least likely to diagnose
- a build that is not Ready is no longer terminal. It is queued again on the next
  run, so the advice is to wait and re-run, not to go edit a package list that
  was never wrong
- deleting a sandbox frees its build once nothing else references it. Builds are
  shared by content, so this is the one place a reader could reasonably assume
  deletion is immediate

Dropped the `ModuleNotFoundError` aside: what the old code did instead is not
something a reader needs to know to use the feature.

The page is hand-written — `function` has category 'blocks' and is absent from
`NATIVE_RESOURCE_BLOCK_TYPES`, so generate-docs skips it and these edits will not
be overwritten.

Co-Authored-By: Claude <noreply@anthropic.com>

* feat(sandboxes): release the provider image when nothing references it

Deleting a sandbox only removed its row, leaving the built template in E2B until
the 30-day retention sweep — up to a month of paying to store an image nothing
could select. Editing a package list had the same effect on the old content
address, which is the more common case since every edit re-points the sandbox.

`releaseSandboxImage(specHash)` now deletes the provider image and its row from
both paths. It reuses the sweep's provider call and its ordering: image first,
row second, so a refused delete leaves the row for the sweep to retry rather than
orphaning a remote template nothing points at.

Two guards make eager deletion safe:

- builds are keyed by content, not by workspace, so two workspaces declaring the
  same package list share one image. The release no-ops while any sandbox still
  references the hash — otherwise one workspace's delete would break the other's
- an in-flight build is left alone rather than raced; the sweep collects it once
  it settles

Called detached from both routes. The row is already committed by then, so the
user's action has succeeded whatever the provider says, and awaiting would hold a
UI delete open on a remote call the sweep would retry anyway. Every failure inside
is logged and swallowed for the same reason.

E2B's delete verified against their API reference: DELETE /templates/{templateID}
with X-API-Key, 204 on success. The existing implementation already matched, so
this commit only adds the call sites and the guards.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): rate-limit the automatic rebuild, drop the one-off status dot

Two follow-ups to the resolution repair.

The repair had no rate limit. `ensureSandboxImage` re-claims a `failed` row on
sight, and a bad package name fails in seconds, so the in-flight guard never
closed the window: a workflow on a one-minute schedule would enqueue a build a
minute against a package list that will never resolve, each one real provider
build compute. Before the repair existed resolution simply threw, so this was
introduced with it.

The two callers want different things, so the cooldown is opt-in. A save is a
person explicitly asking for another attempt and still retries immediately;
resolution passes `FAILED_BUILD_RETRY_COOLDOWN_MS` and gets at most one attempt
per window no matter how often the workflow runs. Ten minutes: long enough that
per-minute runs cannot drive per-minute builds, short enough that a transient
registry outage clears within the hour.

The status line loses its colour dot. `size-[6px] rounded-full` appeared in
exactly one file in the repo, so it was a new primitive rather than a pattern,
and it duplicated state the text colour already carries — the label now turns
`--text-error` on a failed build, which is what every other status row in
settings does. `ChipTag` was the wrong home for this: its variants are
`mono`/`invite`, with no semantic tone, so a status version would have meant
overriding its chrome from the consumer.

Also corrects the docs line this changes: a failed build is retried periodically,
and saving is the way to retry now, so "wait a moment and run again" no longer
describes it.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): claim the image row and its reference check in one statement

Greptile P1. Reading references in one statement and deleting in another left a
window — a wide one, since a provider delete is a network call — where a second
workspace could declare the same package list, inherit the `ready` row, and have
its next run fail against a template already on its way out. Content addressing
is what makes that reachable: the image is shared, so one workspace's delete can
strand another's sandbox.

The reference check now lives in the conditional DELETE itself, so winning the
delete is the proof that nothing referenced the hash. A workspace that adopts the
hash first makes the delete match nothing and the release becomes a no-op.

Claiming the row before the provider call would otherwise strand a template
nothing points at if the provider then refused, so that path puts the row back
and the retention sweep inherits the retry — the same property the previous
ordering had.

The sweep is deliberately left as it is: its equivalent window needs a hash
unreferenced AND unused for 30 days, and its provider-first ordering encodes the
documented retry-on-refusal behaviour this path now reproduces explicitly.

No transaction is opened. The provider call sits between discrete statements
rather than inside one, so no pooled connection is held across it — which is why
this uses a conditional delete instead of the repo's `pg_advisory_xact_lock`
pattern, whose lock only releases at commit.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): route the retention sweep through the same image claim

Cursor and Greptile both flagged the sweep as still carrying the interleaving
just fixed in releaseSandboxImage, and they are right — the reason given for
leaving it alone last round does not survive scrutiny.

That reason was that provider-first ordering encodes retry-on-refusal, so making
the claim atomic would trade a race for an orphaned template. The release path
already answers that: claim the row, and put it back if the provider refuses. The
sweep can have both properties too.

The rarity argument was also weaker than stated. The sweep nominates up to 200
candidates and then works through them eight network deletes at a time, so its
check-to-delete gap is seconds to minutes — wider than the window that was just
closed, not narrower.

Both callers now share `claimAndDeleteImage`, which owns the whole contract: the
unreferenced check lives inside the DELETE, the provider call runs only after the
claim succeeds, and a refusal restores the row. Having written that ordering twice
is what let the two paths drift, so it exists once now.

The sweep's query becomes a nomination step only. Its retention cutoff is passed
into the claim rather than trusted from the earlier read, so a candidate that
stops qualifying mid-sweep fails its claim and is skipped instead of losing its
image.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): rebuild a hash adopted while its image was being deleted

Greptile's third pass on this path, and a case the previous two did not cover:
the adopter starting a *fresh build* rather than inheriting a ready row.

Claiming removes the registry row, so between that and the provider delete
finishing, a workspace can declare the same package list, get a new row, and start
a build under the same content-derived imageRef — which the in-flight delete then
removes.

The window itself is inherent. The registry row and the provider template are two
systems with no shared transaction, so it can be narrowed but not closed. A Redis
lock would not close it either: acquireLock returns true when Redis is absent, so
it cannot be a correctness guarantee for self-hosted. Holding a Postgres advisory
lock would, but only by pinning a pooled connection for the length of a provider
call, which is a worse trade.

What was avoidable is the adopter finding out the slow way. Its row is new and
healthy-looking, so nothing noticed: resolution only repairs a row that is missing
or failed, and a failed one waits out the retry cooldown first. The release path
now re-checks after the delete and re-enqueues, so the rebuild starts immediately
instead of one failed run plus a cooldown later. A build already in flight is left
to the conflict guard, since it may still outlive the delete.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): reclaim a ready row whose image was deleted underneath it

Greptile found the hole the previous commit left, and it is the case that made the
claim in that commit's message wrong: this one is permanent, not transient.

If a re-adopted hash reaches `ready` before the in-flight provider delete lands —
plausible, since E2B layer caching can rebuild an identical spec in seconds — the
row looks healthy while its imageRef points at nothing. Resolution repairs a row
that is missing or failed, never one claiming to be ready, so nothing recovers it.
The sandbox stays broken until someone re-saves it by hand.

`rebuildIfReadopted` called `ensureSandboxImage` with no options, whose conflict
guard reclaims only a failed or stale in-flight row, so it silently did nothing in
exactly that case.

The release path now passes `imageKnownGone`, which widens the re-claim to any
settled row rather than only a failed one. It is the one caller that knows the
image is gone regardless of what the row says. An in-flight build is still left
alone: it either recreates the template it was building or fails into the normal
repair path, and resetting it would only add a duplicate build.

The three ways a settled row may be re-claimed now sit in one `settledRebuildBranch`
helper — any settled row when the image is known gone, a failed one after the
cooldown for an automatic caller, a failed one immediately for a person — because
inlining the third case is what hid the gap.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): let a same-spec save retry a failed build

Cursor Bugbot. `scheduleSandboxBuild` sat inside the changed-hash branch, so a save
that did not alter the package list never reached the registry. The comment above
it described the opposite — that an unchanged spec finds a ready row and enqueues
nothing — which is what `ensureSandboxImage` does, but only if it is called.

That made the docs wrong too. They tell a reader to save the sandbox again to retry
a failed build immediately, and this branch is exactly why that did nothing: the
only way to retry was to edit the package list into a different hash, which is not
what someone recovering from a transient registry failure wants to do.

The call is now unconditional and the registry decides what a save costs, which is
what its conflict guard is for: a ready or in-flight row is left alone, a failed one
is re-claimed at once. Releasing the previous image stays behind the hash check,
since only a changed hash orphans one. Cache invalidation is unchanged —
`scheduleSandboxBuild` already does it, which is why the else branch existed.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(sandboxes): correct the image cache's staleness invariant

Cursor Bugbot found that a released image can still be served from another
replica's cache. The finding is real, and the reason it went unnoticed is that the
cache documented an invariant which eager release quietly broke.

It claimed a `ready` row is terminal for its spec hash, so a cached hit could not
go stale in a way that matters. That held while the only ways a row changed were an
edit (new hash) or a delete (caught by the `workspace_sandbox` read). Releasing an
image eagerly made a `ready` row disappear with the hash unchanged, so the premise
no longer holds and the comment was actively misleading to the next reader.

No behaviour change here — the exposure is bounded at IMAGE_TTL_MS on replicas
other than the one that ran the release, and it self-heals once the entry expires
and the row read finds nothing. Closing it properly needs cross-replica
invalidation or a provider-error path that invalidates on "template not found",
both of which are larger than a review fix; the comment now says so instead of
implying the problem cannot exist.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(sandboxes): note that a JavaScript sandbox needs an import to apply

Cursor Bugbot pointed out that `useRemoteSandbox` keys on detected static
import/require and never on the selected sandbox, so JavaScript without one runs
locally and the selection has no effect.

Keeping the behaviour: honouring the selection would force those blocks remote,
and the large-value-ref guard immediately below would then reject code that runs
fine today. Documenting it instead, next to the picker, since a selection that
silently does nothing is only surprising if nothing says so.

Python is unaffected — it always runs remotely, so its sandbox always applies.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): stop create mode surviving a return to an open sandbox

Cursor Bugbot. Create mode and having a sandbox open are mutually exclusive, but
nothing enforced it, so both could be set at once — and the screen then lied about
which sandbox its Delete pointed at.

With `isCreating` true and `selectedId` restored, `baseline` is null, so the editor
renders an empty "New sandbox" form, while the Delete action is built from
`selected` and still targets the restored sandbox. An admin looking at a blank
create form could delete a sandbox it never named.

Two ways in, both closed:

- Browser Forward after starting a new sandbox restores `selectedId` without going
  through `closeEditor`. The render-time sync that already drops a stale draft now
  also leaves create mode, which is the same class of correction and the reason
  that block exists.
- "New sandbox" set `isCreating` without clearing `selectedId`, so the same
  contradiction was reachable without touching history at all. It now clears the
  selection, with `history: 'replace'` because switching mode is not a destination.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(ci): pin the sandbox flag in the second nav catalog test, bump the chart

Two CI failures, both mine.

`app/workspace/[workspaceId]/settings/navigation.test.ts` asserts the unified
catalog and was left on ambient env. Dropping the Sandboxes section without a
sandbox provider made it 26 items instead of 27 on CI, which has no
`apps/sim/.env` — the same trap already fixed in the sibling
`components/settings/navigation.test.ts`, in the one file that was missed.

Fixing it needs `vi.hoisted` rather than the sibling's `beforeEach`, because this
file reads `allNavigationItems`, built once at module load; a hook would run after
the value it is trying to influence already exists.

The chart gate is separate: this branch adds sandbox settings to
`helm/sim/values.yaml`, and the workflow requires a Chart.yaml bump whenever
`helm/sim/**` changes. Additive config, so 1.3.0 -> 1.4.0 by SemVer.

Verified by running the whole suite with the flags forced off, not just the two
navigation files — no other test depends on a local env file.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): keep the row restore to a refused delete only

Cursor and Greptile, independently, on the same code. `deleteImage` and
`rebuildIfReadopted` shared one try/catch, so a rebuild failure after a *successful*
provider delete was handled as if the provider had refused: the catch put the
claimed row back, `ready` status and all, pointing at a template that no longer
exists.

That is the one state resolution cannot repair — it fixes a row that is missing or
failed, never one claiming to be ready — so it reintroduced the permanent breakage
an earlier commit had just closed, through the error path rather than the happy one.

Restoring now belongs strictly to a refused delete. Once the template is gone the
row stays gone, and the rebuild runs past that catch. The rebuild also swallows its
own failures: it follows a delete that already succeeded, so it must not be reported
as a failed release, and inside the sweep it must not reject the rest of its chunk.
The adopter's next run still reaches the normal repair path.

The regression test drives a rebuild failure and asserts no row is restored. It
fails against the original shape — rebuild inside the shared try, no inner catch —
which is what the two reviewers were describing.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): drop the dead row when a re-adopt rebuild cannot be scheduled

Greptile, one layer under the previous fix. Making the post-delete rebuild swallow
its own failures kept it from being reported as a failed release, but left the
adopter's row claiming a `ready` image whose template is already deleted — the one
state resolution cannot repair, since it rebuilds a row that is missing or failed
and never one that says ready.

So the row is now dropped when the rebuild does not take. That turns the adopter
into the missing-row case, which the next execution repairs on its own, instead of
a sandbox that stays broken until someone re-saves it by hand. A failure to drop it
is logged at error, because at that point two writes in a row have failed and there
is nothing further this path can do.

Also gives the release tests a default "nothing re-adopted" select. Without it the
rebuild threw on an unstubbed mock and the cleanup delete overwrote the predicate
the claim assertions read, so two of them were passing on the wrong statement.

Co-Authored-By: Claude <noreply@anthropic.com>

* feat(sandboxes): repair a missing image at create, where the truth is observable

Six review rounds narrowed the window between deleting a shared template and
another workspace adopting its content hash, and each fix exposed the next facet.
They all share a cause: the registry row and the provider template are two systems
with no shared transaction, so any scheme that keeps them in step is guessing.

Create is the one step that does not have to guess. It either gets a sandbox or it
does not, so a `ready` row pointing at a deleted template now corrects itself the
first time it is used, rather than needing someone to re-save the sandbox.

- `SandboxImageBuilder.isMissingImage` asks the provider to classify its own
  failure. Prebuilt-only, because a runtime provider has no image to miss
- E2B answers it off `NotFoundError`, which the SDK maps from a 404. The only
  resource a create names is the template, and the two subclasses that describe
  other calls — a missing file, an exited sandbox — are excluded. The classifier
  stays deliberately narrow: treating auth or rate-limit failures as a missing
  image would turn a provider outage into a build storm
- `repairMissingSandboxImage` invalidates the cache, rebuilds with
  `imageKnownGone` (no cooldown, since this observed the image is gone rather than
  inferring it), and returns copy telling the author to run again
- `ResolvedSandbox` carries `specHash` so the failing execution can name what to
  rebuild

This subsumes the open facets rather than adding another guard beside them: the
stale per-replica cache, an adopter left `ready` against a deleted ref, and a
rebuild that never took all end at the same place — the next run repairs itself.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(sandboxes): key the build trigger by attempt, not by spec

Cursor Bugbot. The Trigger.dev idempotency key was the content address alone, so a
second attempt at the same spec was deduped against the first: the SDK returns the
finished run instead of starting one, and the row that `ensureSandboxImage` just
flipped to `pending` sits there with no worker. Nothing can re-claim a `pending`
row until it goes stale, so a retry inside the 5-minute TTL did nothing for the
next half hour.

That silently disabled every repair path — save-to-retry, which the docs name
explicitly, and both the resolution and create-time rebuilds.

The key's own comment already said it exists "to collapse concurrent saves of the
same spec into one build, not to suppress a retry after one failed". The conditional
update above it is what actually collapses concurrent saves: only one caller gets a
row back, so only one ever reaches the trigger. Keying by the claim's `updatedAt`
keeps that property and makes each genuine attempt distinct, while a duplicate
delivery of one attempt still collapses.

Co-Authored-By: Claude <noreply@anthropic.com>

* feat(sandboxes): create a sandbox from the picker, and fix three UI papercuts

The Function block's sandbox field now pins a "Create Sandbox" row above its
options, matching the "Create Skill" / "Create Tool" rows it sits beside, so
authoring a package list no longer means leaving the workflow for Settings. The
row is declared by the field (`createAction`) rather than hardcoded by id;
block configs are read by the serializer and executor, so the name maps to a
modal in the picker rather than carrying a component.

Two things the modal has to get right. It seeds the new sandbox's language from
the sibling the list is scoped by, or a sandbox created off a JavaScript block
would land in the Python list and vanish. And the created option is held locally
until a real fetch carries it, or the field would sit on a raw uuid until
hydration answered.

Also:
- The Sandboxes icon was the Logs block's icon (`blocks/blocks/logs.ts`), in
  both the settings nav and the list rows. It is the Function block's now.
- "Default image (no extra packages)" claimed something untrue: E2B and Daytona
  base images both ship with packages installed.
- A new sandbox opened in Python while the Function block defaults to
  JavaScript. The test pins the two together rather than the literal.

Draft shape and helpers moved out of the editor component into `utils.ts` —
three consumers now, and it makes the defaults testable without a DOM.

* feat(settings): one Max-plan wall, and give the create modal the same one

The create-sandbox modal answered a non-Max workspace with a red line under a
form it could never submit, and no way to act on it. It now renders the same
wall the Settings > Sandboxes tab does — heading, one sentence on what the plan
unlocks, and an Upgrade to Max chip — instead of the fields.

That wall existed twice already (sandboxes and Sim Mailer), so this extracts it
rather than adding a third copy. `SettingsUpgradeNotice` owns the copy rhythm
and the route, and `compact` trades the page's full-height centering for a
modal's. Both settings consumers now compose it; neither keeps its own markup.

The action lands on billing, which `resolveSettingsHref` already redirects to
the plan-comparison page for a member who cannot manage billing — so it is a
route to explore plans, never a dead end. The chip stays hidden for non-admins,
exactly as the settings pages had it.

A non-admin on an entitled workspace gets the muted reason rather than the
upgrade wall: buying a plan is not what is in their way.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-30 16:33:35 -07:00

1385 lines
50 KiB
TypeScript

#!/usr/bin/env bun
import { readdir, readFile } from 'node:fs/promises'
import path from 'node:path'
const ROOT = path.resolve(import.meta.dir, '..')
const API_DIR = path.join(ROOT, 'apps/sim/app/api')
const CONTRACTS_DIR = path.join(ROOT, 'apps/sim/lib/api/contracts')
const QUERY_HOOKS_DIR = path.join(ROOT, 'apps/sim/hooks/queries')
const SELECTOR_HOOKS_DIR = path.join(ROOT, 'apps/sim/hooks/selectors')
const BASELINE = {
totalRoutes: 1000,
zodRoutes: 1000,
nonZodRoutes: 0,
} as const
const BOUNDARY_POLICY_BASELINE = {
routeZodImports: 0,
routeLocalSchemaRoutes: 0,
routeLocalSchemaConstructors: 0,
routeZodErrorReferences: 0,
clientHookZodImports: 0,
clientHookLocalSchemaFiles: 0,
clientHookLocalSchemaConstructors: 0,
clientHookRawFetches: 0,
clientSameOriginApiFetches: 0,
doubleCasts: 8,
rawJsonReads: 6,
untypedResponses: 0,
annotationsMissingReason: 0,
} as const
const INDIRECT_ZOD_ROUTES = new Set([
'apps/sim/app/api/demo-requests/route.ts',
// Input-less session-bound GET: nothing to validate; response is
// contract-typed via `satisfies InvitationDetails` in the route.
// Public updater feed: input-less GET, session-less, returns YAML (not JSON),
// so it can't be JSON-contract-bound. Wrapped in withRouteHandler.
'apps/sim/app/api/desktop/update/latest-mac.yml/route.ts',
'apps/sim/app/api/invitations/route.ts',
'apps/sim/app/api/logs/export/route.ts',
'apps/sim/app/api/tools/docusign/route.ts',
// Better Auth handles its own validation for the catch-all route below.
'apps/sim/app/api/auth/[...all]/route.ts',
// Better Auth handles validation for the Stripe webhook handler.
'apps/sim/app/api/auth/webhook/stripe/route.ts',
// Routes with no client-supplied input that previously had no-op
// `z.object({}).strict().parse({})` guards. The boundary contract for
// these routes is "no input", and they consume validated data only via
// session/headers handled by `getSession()` / Better Auth.
'apps/sim/app/api/auth/oauth/connections/route.ts',
'apps/sim/app/api/auth/providers/route.ts',
'apps/sim/app/api/auth/socket-token/route.ts',
'apps/sim/app/api/desktop/auth/handoff/route.ts',
'apps/sim/app/api/workspaces/invitations/route.ts',
// Internal cron entry point that authenticates via `Authorization: Bearer
// CRON_SECRET` and ignores query/body. The boundary contract is "no
// client-supplied input"; query params from external callers are not
// consumed.
'apps/sim/app/api/schedules/execute/route.ts',
// Routes with no client-supplied input. Auth is handled via session/cron/internal
// tokens and there are no params, query, or body to validate. Previously had
// no-op `validateSchema(noInputSchema, {})` guards.
'apps/sim/app/api/health/route.ts',
'apps/sim/app/api/settings/allowed-providers/route.ts',
'apps/sim/app/api/settings/allowed-integrations/route.ts',
'apps/sim/app/api/settings/allowed-mcp-domains/route.ts',
'apps/sim/app/api/cron/cleanup-tasks/route.ts',
'apps/sim/app/api/cron/cleanup-soft-deletes/route.ts',
'apps/sim/app/api/cron/cleanup-stale-executions/route.ts',
'apps/sim/app/api/cron/cleanup-sandbox-images/route.ts',
'apps/sim/app/api/cron/renew-subscriptions/route.ts',
'apps/sim/app/api/cron/reconcile-billing-seats/route.ts',
'apps/sim/app/api/cron/reconcile-inbox-entitlement/route.ts',
'apps/sim/app/api/cron/run-data-drains/route.ts',
'apps/sim/app/api/logs/cleanup/route.ts',
'apps/sim/app/api/knowledge/connectors/sync/route.ts',
'apps/sim/app/api/webhooks/outbox/process/route.ts',
'apps/sim/app/api/webhooks/cleanup/idempotency/route.ts',
// Shared Slack app event ingest. The body is an opaque, HMAC-verified Slack
// event envelope (varies per event type) read via parseWebhookBody; there is
// no client contract to bind — authenticity is enforced by signature.
'apps/sim/app/api/webhooks/slack/route.ts',
'apps/sim/app/api/webhooks/slack/custom/[credentialId]/route.ts',
'apps/sim/app/api/resume/poll/route.ts',
// MCP routes that take only auth context (no client-supplied params/query/body).
'apps/sim/app/api/mcp/discover/route.ts',
'apps/sim/app/api/mcp/tools/stored/route.ts',
// MCP OAuth callback is the provider redirect target — the response is HTML
// that closes the popup, so the JSON-mode contract framework doesn't fit.
// Validation is enforced via state lookup + session-vs-row userId match.
'apps/sim/app/api/mcp/oauth/callback/route.ts',
// Deprecated Copilot MCP surface: these routes are gated to always return
// 410 Gone and consume no client-supplied input.
'apps/sim/app/api/mcp/copilot/route.ts',
'apps/sim/app/api/mcp/copilot/.well-known/oauth-authorization-server/route.ts',
'apps/sim/app/api/mcp/copilot/.well-known/oauth-protected-resource/route.ts',
// Deprecated v1 headless copilot chat API: gated to always return 410 Gone
// and consumes no client-supplied input.
'apps/sim/app/api/v1/copilot/chat/route.ts',
])
/**
* Routes baseline-allowed to use `await request.json()` / `await req.json()`
* directly (without an inline `// boundary-raw-json:` annotation).
*
* These are legitimately partial: tolerant body parses (`.catch(() => ({}))`),
* JSON-RPC envelopes that need their own dispatch, multi-stage MCP routes that
* read pre-parsed bodies, and routes whose Zod-backed migration is queued
* behind a separate contract / schema authoring step. New routes must NOT
* introduce raw `await request.json()` reads — annotate the call with
* `// boundary-raw-json: <reason>` instead.
*/
const RAW_JSON_BASELINE_ROUTES = new Set([
'apps/sim/app/api/billing/portal/route.ts',
'apps/sim/app/api/copilot/api-keys/generate/route.ts',
'apps/sim/app/api/copilot/api-keys/validate/route.ts',
'apps/sim/app/api/copilot/chat/abort/route.ts',
'apps/sim/app/api/folders/[id]/restore/route.ts',
'apps/sim/app/api/invitations/[id]/accept/route.ts',
'apps/sim/app/api/invitations/[id]/reject/route.ts',
'apps/sim/app/api/invitations/[id]/route.ts',
'apps/sim/app/api/knowledge/[id]/documents/route.ts',
'apps/sim/app/api/knowledge/[id]/documents/[documentId]/chunks/route.ts',
'apps/sim/app/api/mcp/serve/[serverId]/route.ts',
'apps/sim/app/api/mcp/servers/route.ts',
'apps/sim/app/api/mcp/servers/[id]/route.ts',
'apps/sim/app/api/mcp/servers/test-connection/route.ts',
'apps/sim/app/api/mcp/tools/discover/route.ts',
'apps/sim/app/api/mcp/tools/execute/route.ts',
'apps/sim/app/api/mcp/workflow-servers/route.ts',
'apps/sim/app/api/mcp/workflow-servers/[id]/route.ts',
'apps/sim/app/api/mcp/workflow-servers/[id]/tools/route.ts',
'apps/sim/app/api/mcp/workflow-servers/[id]/tools/[toolId]/route.ts',
'apps/sim/app/api/organizations/route.ts',
'apps/sim/app/api/organizations/[id]/invitations/route.ts',
'apps/sim/app/api/organizations/[id]/members/route.ts',
'apps/sim/app/api/organizations/[id]/transfer-ownership/route.ts',
'apps/sim/app/api/resume/[workflowId]/[executionId]/[contextId]/route.ts',
'apps/sim/app/api/speech/token/route.ts',
'apps/sim/app/api/table/[tableId]/rows/route.ts',
'apps/sim/app/api/tools/file/manage/route.ts',
'apps/sim/app/api/workspaces/invitations/batch/route.ts',
'apps/sim/app/api/workspaces/[id]/route.ts',
'apps/sim/app/api/workspaces/[id]/files/[fileId]/route.ts',
'apps/sim/app/api/workspaces/[id]/files/[fileId]/content/route.ts',
])
const CONTRACT_IMPORT_PATTERN = /\bfrom\s+['"]@\/lib\/api\/contracts(?:\/[^'"]*)?['"]/
const SERVER_VALIDATION_IMPORT_PATTERN = /\bfrom\s+['"]@\/lib\/api\/server(?:\/validation)?['"]/
const SCHEMA_PARSE_PATTERN = /\b\w+Schema\.(?:safeParse|parse)\(/
const CONTRACT_SERVER_HELPER_PATTERN = /\bparseToolRequest\(/
const CANONICAL_HELPER_USAGE_PATTERN =
/\b(?:isZodError|validationErrorResponse|validationErrorResponseFromError|getValidationErrorMessage)\s*\(/
const CONTRACT_MAP_PARSE_PATTERN =
/\b\w+ContractsByPath[\s\S]{0,600}\.(?:body|query|params)!?\.(?:safeParse|parse)\(/
/**
* Matches `from 'zod'` and any zod subpath import like `from 'zod/v4'` or
* `from 'zod/mini'`. The capturing-group-free alternation keeps this safe to
* use with `.test(...)` and `.replace(...)` callers.
*/
const ZOD_IMPORT_PATTERN = /\bfrom\s+['"]zod(?:\/[^'"]+)?['"]/
const ZOD_REQUIRE_PATTERN = /\brequire\(['"]zod(?:\/[^'"]+)?['"]\)/
const ZOD_SCHEMA_CONSTRUCTOR_PATTERN =
/\bz\.(?:object|string|number|boolean|array|enum|nativeEnum|union|discriminatedUnion|record|literal|tuple|preprocess|coerce|date|unknown|any|instanceof|custom|lazy)\s*\(/g
const ZOD_ERROR_PATTERN = /\bZodError\b|\bz\.ZodError\b/
const SKIP_DIRS = new Set(['node_modules', '.next', '.turbo', 'coverage'])
const WIRE_TYPE_DECLARATION_PATTERN =
/(?:^|\n)\s*(?:export\s+)?(interface|type)\s+([A-Z]\w*(?:Response|Result))\b(?=\s*(?:=|extends|\{))/g
const CONTRACT_DERIVED_WIRE_TYPE_PATTERN =
/\b(?:ContractJsonResponse|ContractJsonErrorResponse|z\.(?:input|output|infer))\b/
const RAW_FETCH_PATTERN = /\bfetch\(/g
const RAW_FETCH_HELPER_GUARD_PATTERN = /(?:requestJson|requestRaw|prefetchJson|preFetchJson)$/
/**
* Matches `fetch(` (with optional whitespace, including newlines) followed by
* a string literal — single quote, double quote, or template literal —
* whose first character is `/api/`. This catches same-origin internal API
* fetches in any non-test source file under `apps/sim/**` that aren't an
* `app/api/**\/route.ts` server handler. Template literals with leading
* interpolations (e.g. `${base}/api/foo`) are intentionally NOT matched
* because they're rare and could trigger false positives on non-`/api/` URLs.
*/
const SAME_ORIGIN_API_FETCH_PATTERN = /\bfetch\(\s*[`'"]\/api\//g
const DOUBLE_CAST_PATTERN = /\bas unknown as\b/g
/**
* Matches `await request.json()` / `await req.json()` and the multi-line
* `await request.clone().json()` clone-then-read variant. Both forms read
* the request body without going through `parseRequest(...)` / a contract
* and count toward the `rawJsonReads` ratchet.
*
* `\s` is multi-line (handles common Prettier/Biome formatting where the
* `.clone()` and `.json()` calls land on separate lines).
*/
const RAW_JSON_READ_PATTERN =
/\bawait\s+(?:request|req)\s*(?:\.\s*clone\s*\(\s*\))?\s*\.\s*json\s*\(\s*\)/g
/**
* Matches `schema:` followed directly by a "validates nothing" zod construct.
* Three forms are treated equivalently:
* 1. `schema: z.unknown()` — no validation at all.
* 2. `schema: z.object({}).passthrough()` — validates only that the value
* is an object; allows any keys/values.
* 3. `schema: z.record(z.string(), z.unknown())` — validates only that the
* value is a string-keyed object; values are arbitrary.
*
* Anchored on the literal `schema:` token so that nested `z.unknown()` /
* `z.object({}).passthrough()` uses inside an otherwise-typed object schema
* are NOT flagged — only the top-level response declaration in
* `defineRouteContract({ ..., response: { mode: 'json', schema: ... } })`.
*/
const UNTYPED_RESPONSE_PATTERN =
/\bschema\s*:\s*(?:z\.unknown\s*\(\s*\)|z\.object\s*\(\s*\{\s*\}\s*\)\s*\.passthrough\s*\(\s*\)|z\.record\s*\(\s*z\.string\s*\(\s*\)\s*,\s*z\.unknown\s*\(\s*\)\s*\))/g
const RAW_FETCH_ANNOTATION_PREFIX = '// boundary-raw-fetch:'
const DOUBLE_CAST_ANNOTATION_PREFIX = '// double-cast-allowed:'
const RAW_JSON_ANNOTATION_PREFIX = '// boundary-raw-json:'
const UNTYPED_RESPONSE_ANNOTATION_PREFIX = '// untyped-response:'
const SOURCE_FILE_EXTENSIONS = /\.(?:ts|tsx)$/
const TEST_FILE_PATTERN = /(?:\.test|\.spec)\.(?:ts|tsx)$/
const TEST_HELPER_FILE_PATTERN = /(?:^|\/)test-[^/]+\.ts$/
const TEST_DIR_SEGMENT_PATTERN = /(?:^|\/)(?:__tests__|testing)(?:\/|$)/
/**
* Skips user-uploaded content stored under `apps/sim/uploads/...` (workspace
* file uploads, etc.). Does NOT match `apps/sim/lib/uploads/...`, which is
* source code for the uploads subsystem.
*/
const USER_UPLOADS_DIR_PATTERN = /(?:^|\/)apps\/sim\/uploads(?:\/|$)/
const SOURCE_SKIP_DIRS = new Set([
'node_modules',
'.next',
'.turbo',
'coverage',
'dist',
'__tests__',
'testing',
])
type AnnotationKind = 'raw-fetch' | 'double-cast' | 'raw-json' | 'untyped-response'
interface AnnotationResult {
allowed: boolean
missingReason: boolean
}
interface RawFetchFinding {
path: string
line: number
preview: string
}
interface SameOriginApiFetchFinding {
path: string
line: number
preview: string
}
interface DoubleCastFinding {
path: string
line: number
preview: string
}
interface RawJsonFinding {
path: string
line: number
preview: string
}
interface UntypedResponseFinding {
path: string
line: number
preview: string
}
interface AnnotationMissingReasonFinding {
path: string
line: number
kind: AnnotationKind
}
interface RouteAudit {
path: string
usesZod: boolean
hasZodImport: boolean
schemaConstructorCount: number
hasZodErrorReference: boolean
hasBodyRead: boolean
hasQueryRead: boolean
hasFormDataRead: boolean
hasParamsContext: boolean
}
interface WireTypeFinding {
path: string
name: string
line: number
}
interface QueryHookAudit {
path: string
hasZodImport: boolean
schemaConstructorCount: number
adHocWireTypes: WireTypeFinding[]
}
interface FamilyStats {
total: number
zod: number
nonZod: number
}
type BoundaryPolicyKey = keyof typeof BOUNDARY_POLICY_BASELINE
interface BoundaryPolicyMetric {
key: BoundaryPolicyKey
label: string
current: number
}
interface PrintOnlyBoundaryPolicyMetric {
label: string
current: number
}
async function walk(
dir: string,
shouldIncludeFile: (fileName: string) => boolean,
results: string[] = []
): Promise<string[]> {
const entries = await readdir(dir, { withFileTypes: true })
for (const entry of entries) {
if (SKIP_DIRS.has(entry.name)) continue
const fullPath = path.join(dir, entry.name)
if (entry.isDirectory()) {
await walk(fullPath, shouldIncludeFile, results)
} else if (shouldIncludeFile(entry.name)) {
results.push(fullPath)
}
}
return results
}
function lineNumberForIndex(content: string, index: number): number {
let line = 1
for (let i = 0; i < index; i++) {
if (content.charCodeAt(i) === 10) line += 1
}
return line
}
/**
* Inspects up to three consecutive non-empty preceding lines for an
* opt-out annotation matching the given kind. The annotation is allowed
* when the matching prefix is followed by a non-empty reason. When the
* prefix is present but the reason is empty, `missingReason` is set so
* the audit can flag and fail on dangling annotations.
*/
function extractAnnotation(
content: string,
lineIndex: number,
kind: AnnotationKind
): AnnotationResult {
const prefix =
kind === 'raw-fetch'
? RAW_FETCH_ANNOTATION_PREFIX
: kind === 'double-cast'
? DOUBLE_CAST_ANNOTATION_PREFIX
: kind === 'raw-json'
? RAW_JSON_ANNOTATION_PREFIX
: UNTYPED_RESPONSE_ANNOTATION_PREFIX
const lines = content.split('\n')
let inspected = 0
for (let i = lineIndex - 1; i >= 0 && inspected < 3; i -= 1) {
const trimmed = lines[i]?.trim() ?? ''
if (trimmed.length === 0) continue
inspected += 1
if (!trimmed.startsWith('//')) {
return { allowed: false, missingReason: false }
}
const prefixIndex = trimmed.indexOf(prefix)
if (prefixIndex === -1) continue
const reason = trimmed.slice(prefixIndex + prefix.length).trim()
if (reason.length === 0) {
return { allowed: false, missingReason: true }
}
return { allowed: true, missingReason: false }
}
return { allowed: false, missingReason: false }
}
/**
* Walks `apps/sim/**` and optionally `packages/**` for `.ts` / `.tsx`
* source files, excluding tests, build artifacts, and coverage output.
* Kept separate from `walk(API_DIR)` because the source-wide audit has
* different exclusion rules than the route audit.
*/
async function walkAllSourceFiles(root: string, includePackages: boolean): Promise<string[]> {
const roots = includePackages
? [path.join(root, 'apps/sim'), path.join(root, 'packages')]
: [path.join(root, 'apps/sim')]
const results: string[] = []
for (const start of roots) {
await walkSourceTree(start, results)
}
return results
}
async function walkSourceTree(dir: string, results: string[]): Promise<void> {
let entries
try {
entries = await readdir(dir, { withFileTypes: true })
} catch {
return
}
for (const entry of entries) {
if (SOURCE_SKIP_DIRS.has(entry.name)) continue
const fullPath = path.join(dir, entry.name)
if (entry.isDirectory()) {
await walkSourceTree(fullPath, results)
continue
}
if (!SOURCE_FILE_EXTENSIONS.test(entry.name)) continue
if (TEST_FILE_PATTERN.test(entry.name)) continue
if (TEST_HELPER_FILE_PATTERN.test(entry.name)) continue
if (TEST_DIR_SEGMENT_PATTERN.test(fullPath)) continue
if (USER_UPLOADS_DIR_PATTERN.test(fullPath)) continue
results.push(fullPath)
}
}
function isContractFetchHelperCall(line: string, matchIndex: number): boolean {
const before = line.slice(0, matchIndex)
return RAW_FETCH_HELPER_GUARD_PATTERN.test(before)
}
function buildPreview(line: string): string {
return line.trim().slice(0, 160)
}
function findRawFetchFindings(
filePath: string,
content: string
): {
findings: RawFetchFinding[]
exemptions: number
missingReasons: AnnotationMissingReasonFinding[]
} {
const relativePath = path.relative(ROOT, filePath)
const lines = content.split('\n')
const findings: RawFetchFinding[] = []
const missingReasons: AnnotationMissingReasonFinding[] = []
let exemptions = 0
for (let i = 0; i < lines.length; i += 1) {
const line = lines[i] ?? ''
RAW_FETCH_PATTERN.lastIndex = 0
let match: RegExpExecArray | null
while ((match = RAW_FETCH_PATTERN.exec(line)) !== null) {
if (isContractFetchHelperCall(line, match.index)) continue
const annotation = extractAnnotation(content, i, 'raw-fetch')
if (annotation.missingReason) {
missingReasons.push({ path: relativePath, line: i + 1, kind: 'raw-fetch' })
findings.push({ path: relativePath, line: i + 1, preview: buildPreview(line) })
continue
}
if (annotation.allowed) {
exemptions += 1
continue
}
findings.push({ path: relativePath, line: i + 1, preview: buildPreview(line) })
}
}
return { findings, exemptions, missingReasons }
}
function findDoubleCastFindings(
filePath: string,
content: string
): {
findings: DoubleCastFinding[]
exemptions: number
missingReasons: AnnotationMissingReasonFinding[]
} {
const relativePath = path.relative(ROOT, filePath)
const lines = content.split('\n')
const findings: DoubleCastFinding[] = []
const missingReasons: AnnotationMissingReasonFinding[] = []
let exemptions = 0
for (let i = 0; i < lines.length; i += 1) {
const line = lines[i] ?? ''
DOUBLE_CAST_PATTERN.lastIndex = 0
while (DOUBLE_CAST_PATTERN.exec(line) !== null) {
const annotation = extractAnnotation(content, i, 'double-cast')
if (annotation.missingReason) {
missingReasons.push({ path: relativePath, line: i + 1, kind: 'double-cast' })
findings.push({ path: relativePath, line: i + 1, preview: buildPreview(line) })
continue
}
if (annotation.allowed) {
exemptions += 1
continue
}
findings.push({ path: relativePath, line: i + 1, preview: buildPreview(line) })
}
}
return { findings, exemptions, missingReasons }
}
/**
* Inspect a route file for `await request.json()` / `await req.json()` reads.
*
* Returns one finding per unannotated read. Routes in
* `RAW_JSON_BASELINE_ROUTES` are baseline-allowed: their reads still appear
* in `findings` so the `rawJsonReads` ratcheted metric counts them, but they
* are NOT required to carry per-line `// boundary-raw-json: <reason>`
* annotations. The ratchet's enforcement is by file count
* (see `buildBoundaryPolicyMetrics`): adding a raw read in a route outside
* the baseline pushes the unique-file count above `BOUNDARY_POLICY_BASELINE.rawJsonReads`.
*
* Annotated reads (`// boundary-raw-json: <reason>` on one of the three
* preceding non-empty lines) are treated as exemptions and excluded from
* `findings`. An annotation with the prefix but an empty reason is flagged
* via `missingReasons` and still counts as a finding.
*/
function findRawJsonFindings(
filePath: string,
content: string
): {
findings: RawJsonFinding[]
exemptions: number
missingReasons: AnnotationMissingReasonFinding[]
} {
const relativePath = path.relative(ROOT, filePath)
const lines = content.split('\n')
const findings: RawJsonFinding[] = []
const missingReasons: AnnotationMissingReasonFinding[] = []
let exemptions = 0
for (let i = 0; i < lines.length; i += 1) {
const line = lines[i] ?? ''
RAW_JSON_READ_PATTERN.lastIndex = 0
while (RAW_JSON_READ_PATTERN.exec(line) !== null) {
const annotation = extractAnnotation(content, i, 'raw-json')
if (annotation.missingReason) {
missingReasons.push({ path: relativePath, line: i + 1, kind: 'raw-json' })
findings.push({ path: relativePath, line: i + 1, preview: buildPreview(line) })
continue
}
if (annotation.allowed) {
exemptions += 1
continue
}
findings.push({ path: relativePath, line: i + 1, preview: buildPreview(line) })
}
}
return { findings, exemptions, missingReasons }
}
/**
* Inspect a contracts file for "validates nothing" response schema
* declarations. Three forms are treated equivalently and all count toward
* the `untypedResponses` ratchet:
* - `schema: z.unknown()`
* - `schema: z.object({}).passthrough()`
* - `schema: z.record(z.string(), z.unknown())`
*
* Anchored on `schema:` so nested uses inside an otherwise-typed object
* (e.g. `output: z.unknown()` inside `z.object({ ... })`) are NOT flagged.
* Each callsite must carry a `// untyped-response: <reason>` annotation on
* one of the three preceding non-empty lines. Annotated callsites become
* exemptions; un-annotated callsites become findings that count toward the
* `untypedResponses` ratchet. An annotation with the prefix but an empty
* reason is flagged via `missingReasons` and still counts as a finding.
*/
function findUntypedResponseFindings(
filePath: string,
content: string
): {
findings: UntypedResponseFinding[]
exemptions: number
missingReasons: AnnotationMissingReasonFinding[]
} {
const relativePath = path.relative(ROOT, filePath)
const lines = content.split('\n')
const findings: UntypedResponseFinding[] = []
const missingReasons: AnnotationMissingReasonFinding[] = []
let exemptions = 0
for (let i = 0; i < lines.length; i += 1) {
const line = lines[i] ?? ''
UNTYPED_RESPONSE_PATTERN.lastIndex = 0
while (UNTYPED_RESPONSE_PATTERN.exec(line) !== null) {
const annotation = extractAnnotation(content, i, 'untyped-response')
if (annotation.missingReason) {
missingReasons.push({ path: relativePath, line: i + 1, kind: 'untyped-response' })
findings.push({ path: relativePath, line: i + 1, preview: buildPreview(line) })
continue
}
if (annotation.allowed) {
exemptions += 1
continue
}
findings.push({ path: relativePath, line: i + 1, preview: buildPreview(line) })
}
}
return { findings, exemptions, missingReasons }
}
function isClientHookFile(filePath: string): boolean {
const normalized = filePath.replace(/\\/g, '/')
return (
normalized.startsWith(`${QUERY_HOOKS_DIR}/`) || normalized.startsWith(`${SELECTOR_HOOKS_DIR}/`)
)
}
/**
* Identifies `apps/sim/app/api/**\/route.ts` API route handlers. Same-origin
* `/api/` fetch scanning skips these — server-side fetches from inside a
* route handler are a different concern and are not what this ratchet is
* trying to catch.
*/
function isApiRouteHandler(filePath: string): boolean {
const normalized = filePath.replace(/\\/g, '/')
if (!normalized.startsWith(`${API_DIR}/`)) return false
return normalized.endsWith('/route.ts')
}
/**
* Inspect a non-API-route source file under `apps/sim/**` for raw
* `fetch('/api/...')`, `fetch("/api/...")`, or ``fetch(`/api/...`)`` calls.
*
* Each callsite is either an exemption (annotated with
* `// boundary-raw-fetch: <reason>` on one of the three preceding non-empty
* lines) or a finding. The ratchet's enforcement is by unique-file count:
* introducing a same-origin `/api/` fetch without an annotation pushes the
* unique-file count above `BOUNDARY_POLICY_BASELINE.clientSameOriginApiFetches`
* and fails the audit.
*
* Scanning runs over the entire file content (not line-by-line) so that
* multi-line constructs like `fetch(\n \`/api/...\`)` are still caught.
* An annotation with the prefix but an empty reason is flagged via
* `missingReasons` and still counts as a finding.
*/
function findSameOriginApiFetchFindings(
filePath: string,
content: string
): {
findings: SameOriginApiFetchFinding[]
exemptions: number
missingReasons: AnnotationMissingReasonFinding[]
} {
const relativePath = path.relative(ROOT, filePath)
const lines = content.split('\n')
const findings: SameOriginApiFetchFinding[] = []
const missingReasons: AnnotationMissingReasonFinding[] = []
let exemptions = 0
SAME_ORIGIN_API_FETCH_PATTERN.lastIndex = 0
let match: RegExpExecArray | null
while ((match = SAME_ORIGIN_API_FETCH_PATTERN.exec(content)) !== null) {
const lineNumber = lineNumberForIndex(content, match.index)
const lineIndex = lineNumber - 1
const line = lines[lineIndex] ?? ''
const annotation = extractAnnotation(content, lineIndex, 'raw-fetch')
if (annotation.missingReason) {
missingReasons.push({ path: relativePath, line: lineNumber, kind: 'raw-fetch' })
findings.push({ path: relativePath, line: lineNumber, preview: buildPreview(line) })
continue
}
if (annotation.allowed) {
exemptions += 1
continue
}
findings.push({ path: relativePath, line: lineNumber, preview: buildPreview(line) })
}
return { findings, exemptions, missingReasons }
}
function routeFamily(routePath: string): string {
const relative = routePath.replace(/^apps\/sim\/app\/api\//, '')
const [first, second] = relative.split('/')
if (first === 'tools') return `tools/${second ?? 'unknown'}`
if (first === 'v1') return second ? `v1/${second}` : 'v1'
return first ?? 'unknown'
}
function hasZodUsage(relativePath: string, content: string): boolean {
if (ZOD_IMPORT_PATTERN.test(content) || ZOD_REQUIRE_PATTERN.test(content)) {
return true
}
if (
/\bparseRequest\(/.test(content) &&
/\bfrom\s+['"]@\/lib\/api\/server['"]/.test(content) &&
CONTRACT_IMPORT_PATTERN.test(content)
) {
return true
}
if (
CONTRACT_IMPORT_PATTERN.test(content) &&
(SCHEMA_PARSE_PATTERN.test(content) || CONTRACT_MAP_PARSE_PATTERN.test(content))
) {
return true
}
if (CONTRACT_IMPORT_PATTERN.test(content) && CONTRACT_SERVER_HELPER_PATTERN.test(content)) {
return true
}
if (
CONTRACT_IMPORT_PATTERN.test(content) &&
SERVER_VALIDATION_IMPORT_PATTERN.test(content) &&
CANONICAL_HELPER_USAGE_PATTERN.test(content)
) {
return true
}
if (
SERVER_VALIDATION_IMPORT_PATTERN.test(content) &&
/\b(?:isZodError|validationErrorResponseFromError)\b/.test(content) &&
SCHEMA_PARSE_PATTERN.test(content)
) {
return true
}
return INDIRECT_ZOD_ROUTES.has(relativePath)
}
function auditRoute(filePath: string, content: string): RouteAudit {
const relativePath = path.relative(ROOT, filePath)
const schemaConstructorCount = [...content.matchAll(ZOD_SCHEMA_CONSTRUCTOR_PATTERN)].length
return {
path: relativePath,
usesZod: hasZodUsage(relativePath, content),
hasZodImport: ZOD_IMPORT_PATTERN.test(content) || ZOD_REQUIRE_PATTERN.test(content),
schemaConstructorCount,
hasZodErrorReference: ZOD_ERROR_PATTERN.test(content),
hasBodyRead: /\brequest\.json\(\)|\breq\.json\(\)/.test(content),
hasQueryRead: /\.searchParams\b|new URL\([^)]*\)\.searchParams/.test(content),
hasFormDataRead: /\.formData\(\)/.test(content),
hasParamsContext: /\bparams\b/.test(content) && /\bPromise<\{|\bRouteContext\b/.test(content),
}
}
function findAdHocWireTypes(filePath: string, content: string): WireTypeFinding[] {
const findings: WireTypeFinding[] = []
const relativePath = path.relative(ROOT, filePath)
for (const match of content.matchAll(WIRE_TYPE_DECLARATION_PATTERN)) {
const kind = match[1]
const name = match[2]
const declarationStart = match.index ?? 0
const declarationPreview = content.slice(declarationStart, declarationStart + 800)
if (name.endsWith('QueryResult')) continue
if (CONTRACT_DERIVED_WIRE_TYPE_PATTERN.test(declarationPreview)) continue
if (kind === 'type') {
const equalsIndex = declarationPreview.indexOf('=')
const typeBody = declarationPreview.slice(equalsIndex + 1).trimStart()
if (!typeBody.startsWith('{') && !/^[A-Z]\w*\s*&\s*\{/.test(typeBody)) {
continue
}
}
findings.push({
path: relativePath,
name,
line: lineNumberForIndex(content, declarationStart),
})
}
return findings
}
function auditQueryHook(filePath: string, content: string): QueryHookAudit {
const relativePath = path.relative(ROOT, filePath)
const schemaConstructorCount = [...content.matchAll(ZOD_SCHEMA_CONSTRUCTOR_PATTERN)].length
return {
path: relativePath,
hasZodImport: ZOD_IMPORT_PATTERN.test(content) || ZOD_REQUIRE_PATTERN.test(content),
schemaConstructorCount,
adHocWireTypes: findAdHocWireTypes(filePath, content),
}
}
function printFamilyStats(audits: RouteAudit[]) {
const families = new Map<string, FamilyStats>()
for (const audit of audits) {
const family = routeFamily(audit.path)
const stats = families.get(family) ?? { total: 0, zod: 0, nonZod: 0 }
stats.total += 1
if (audit.usesZod) {
stats.zod += 1
} else {
stats.nonZod += 1
}
families.set(family, stats)
}
console.log('\nFamily breakdown:')
for (const [family, stats] of [...families.entries()].sort((a, b) => a[0].localeCompare(b[0]))) {
console.log(
` ${family.padEnd(32)} total=${String(stats.total).padStart(3)} zod=${String(stats.zod).padStart(3)} nonZod=${String(stats.nonZod).padStart(3)}`
)
}
}
function printRiskyNonZodRoutes(audits: RouteAudit[]) {
const risky = audits.filter(
(audit) =>
!audit.usesZod &&
(audit.hasBodyRead || audit.hasQueryRead || audit.hasFormDataRead || audit.hasParamsContext)
)
console.log(`\nNon-Zod routes with parsed inputs: ${risky.length}`)
for (const audit of risky.slice(0, 50)) {
const markers = [
audit.hasBodyRead ? 'json' : null,
audit.hasFormDataRead ? 'formData' : null,
audit.hasQueryRead ? 'query' : null,
audit.hasParamsContext ? 'params' : null,
].filter(Boolean)
console.log(` ${audit.path} (${markers.join(', ')})`)
}
if (risky.length > 50) {
console.log(` ... ${risky.length - 50} more`)
}
}
function printAllNonZodRoutes(audits: RouteAudit[]) {
if (!process.argv.includes('--list-non-zod')) return
const nonZod = audits.filter((audit) => !audit.usesZod)
console.log(`\nAll non-Zod routes: ${nonZod.length}`)
for (const audit of nonZod) {
const markers = [
audit.hasBodyRead ? 'json' : null,
audit.hasFormDataRead ? 'formData' : null,
audit.hasQueryRead ? 'query' : null,
audit.hasParamsContext ? 'params' : null,
].filter(Boolean)
console.log(` ${audit.path}${markers.length > 0 ? ` (${markers.join(', ')})` : ''}`)
}
}
function buildBoundaryPolicyMetrics(
routeAudits: RouteAudit[],
queryHookAudits: QueryHookAudit[],
rawFetchSummary: { findings: RawFetchFinding[]; exemptions: number },
sameOriginApiFetchSummary: {
findings: SameOriginApiFetchFinding[]
exemptions: number
},
doubleCastSummary: { findings: DoubleCastFinding[]; exemptions: number },
rawJsonSummary: { findings: RawJsonFinding[]; exemptions: number },
untypedResponseSummary: { findings: UntypedResponseFinding[]; exemptions: number },
annotationsMissingReason: AnnotationMissingReasonFinding[]
): {
ratchetedMetrics: BoundaryPolicyMetric[]
printOnlyMetrics: PrintOnlyBoundaryPolicyMetric[]
} {
const routeZodImports = routeAudits.filter((audit) => audit.hasZodImport)
const routeLocalSchemaRoutes = routeAudits.filter((audit) => audit.schemaConstructorCount > 0)
const routeLocalSchemaConstructors = routeLocalSchemaRoutes.reduce(
(total, audit) => total + audit.schemaConstructorCount,
0
)
const routeZodErrorReferences = routeAudits.filter((audit) => audit.hasZodErrorReference)
const queryHookZodImports = queryHookAudits.filter((audit) => audit.hasZodImport)
const queryHookLocalSchemaFiles = queryHookAudits.filter(
(audit) => audit.schemaConstructorCount > 0
)
const queryHookLocalSchemaConstructors = queryHookLocalSchemaFiles.reduce(
(total, audit) => total + audit.schemaConstructorCount,
0
)
const queryHookAdHocWireTypes = queryHookAudits.flatMap((audit) => audit.adHocWireTypes)
return {
ratchetedMetrics: [
{
key: 'routeZodImports',
label: 'route files importing zod',
current: routeZodImports.length,
},
{
key: 'routeLocalSchemaRoutes',
label: 'route files with local schema constructors',
current: routeLocalSchemaRoutes.length,
},
{
key: 'routeLocalSchemaConstructors',
label: 'route local schema constructor calls',
current: routeLocalSchemaConstructors,
},
{
key: 'routeZodErrorReferences',
label: 'route files referencing ZodError',
current: routeZodErrorReferences.length,
},
{
key: 'clientHookZodImports',
label: 'client hook files importing zod',
current: queryHookZodImports.length,
},
{
key: 'clientHookLocalSchemaFiles',
label: 'client hook files with local schema constructors',
current: queryHookLocalSchemaFiles.length,
},
{
key: 'clientHookLocalSchemaConstructors',
label: 'client hook local schema constructor calls',
current: queryHookLocalSchemaConstructors,
},
{
key: 'clientHookRawFetches',
label: 'client hook raw fetch() calls',
current: rawFetchSummary.findings.length,
},
{
key: 'clientSameOriginApiFetches',
label: 'apps/sim files with raw same-origin /api/ fetch() calls',
current: new Set(sameOriginApiFetchSummary.findings.map((finding) => finding.path)).size,
},
{
key: 'doubleCasts',
label: 'as unknown as double-casts (non-test)',
current: doubleCastSummary.findings.length,
},
{
key: 'rawJsonReads',
label: 'route files with raw await request.json() reads',
current: new Set(rawJsonSummary.findings.map((finding) => finding.path)).size,
},
{
key: 'untypedResponses',
label:
'contract untyped response schemas (z.unknown / z.object({}).passthrough / z.record)',
current: untypedResponseSummary.findings.length,
},
{
key: 'annotationsMissingReason',
label: 'audit annotations missing reason',
current: annotationsMissingReason.length,
},
],
printOnlyMetrics: [
{
label: 'client hook ad-hoc wire Response/Result types',
current: queryHookAdHocWireTypes.length,
},
{
label: 'client hook raw fetch() exemptions (annotated)',
current: rawFetchSummary.exemptions,
},
{
label: 'apps/sim raw same-origin /api/ fetch() callsites',
current: sameOriginApiFetchSummary.findings.length,
},
{
label: 'apps/sim raw same-origin /api/ fetch() exemptions (annotated)',
current: sameOriginApiFetchSummary.exemptions,
},
{
label: 'as unknown as double-cast exemptions (annotated)',
current: doubleCastSummary.exemptions,
},
{
label: 'route raw await request.json() annotated exemptions',
current: rawJsonSummary.exemptions,
},
{
label: 'contract untyped response annotated exemptions',
current: untypedResponseSummary.exemptions,
},
],
}
}
function printBoundaryPolicyMetric(metric: BoundaryPolicyMetric) {
const baseline = BOUNDARY_POLICY_BASELINE[metric.key]
const delta = metric.current - baseline
const deltaText = delta === 0 ? 'at baseline' : `${delta > 0 ? '+' : ''}${delta} vs baseline`
console.log(` ${metric.label}: ${metric.current} (${deltaText})`)
}
function printRawFetchAndDoubleCastMetrics(
rawFetchFindings: RawFetchFinding[],
sameOriginApiFetchFindings: SameOriginApiFetchFinding[],
doubleCastFindings: DoubleCastFinding[],
rawJsonFindings: RawJsonFinding[],
untypedResponseFindings: UntypedResponseFinding[],
annotationsMissingReason: AnnotationMissingReasonFinding[],
rawFetchExemptions: number,
sameOriginApiFetchExemptions: number,
doubleCastExemptions: number,
rawJsonExemptions: number,
untypedResponseExemptions: number
) {
console.log('\nRaw fetch and double-cast metrics:')
console.log(` client hook raw fetch() calls: ${rawFetchFindings.length}`)
console.log(` client hook raw fetch() exemptions (annotated): ${rawFetchExemptions}`)
const sameOriginFiles = new Set(sameOriginApiFetchFindings.map((finding) => finding.path)).size
console.log(
` apps/sim files with raw same-origin /api/ fetch() calls: ${sameOriginFiles} (baseline ${BOUNDARY_POLICY_BASELINE.clientSameOriginApiFetches})`
)
console.log(
` apps/sim raw same-origin /api/ fetch() callsites: ${sameOriginApiFetchFindings.length}`
)
console.log(
` apps/sim raw same-origin /api/ fetch() exemptions (annotated): ${sameOriginApiFetchExemptions}`
)
console.log(` as unknown as double-casts (non-test): ${doubleCastFindings.length}`)
console.log(` as unknown as double-cast exemptions (annotated): ${doubleCastExemptions}`)
const rawJsonFiles = new Set(rawJsonFindings.map((finding) => finding.path)).size
console.log(
` route files with raw await request.json() reads: ${rawJsonFiles} (baseline ${BOUNDARY_POLICY_BASELINE.rawJsonReads})`
)
console.log(` route raw await request.json() reads (callsites): ${rawJsonFindings.length}`)
console.log(` route raw await request.json() annotated exemptions: ${rawJsonExemptions}`)
console.log(
` contract untyped response schemas (z.unknown / z.object({}).passthrough / z.record): ${untypedResponseFindings.length}`
)
console.log(` contract untyped response annotated exemptions: ${untypedResponseExemptions}`)
console.log(` audit annotations missing reason: ${annotationsMissingReason.length}`)
console.log(' raw fetch examples:')
for (const finding of rawFetchFindings.slice(0, 25)) {
console.log(` ${finding.path}:${finding.line} ${finding.preview}`)
}
if (rawFetchFindings.length > 25) {
console.log(` ... ${rawFetchFindings.length - 25} more`)
}
console.log(' same-origin /api/ fetch examples:')
for (const finding of sameOriginApiFetchFindings.slice(0, 25)) {
console.log(` ${finding.path}:${finding.line} ${finding.preview}`)
}
if (sameOriginApiFetchFindings.length > 25) {
console.log(` ... ${sameOriginApiFetchFindings.length - 25} more`)
}
console.log(' double-cast examples:')
for (const finding of doubleCastFindings.slice(0, 25)) {
console.log(` ${finding.path}:${finding.line} ${finding.preview}`)
}
if (doubleCastFindings.length > 25) {
console.log(` ... ${doubleCastFindings.length - 25} more`)
}
console.log(' raw await request.json() examples:')
for (const finding of rawJsonFindings.slice(0, 25)) {
console.log(` ${finding.path}:${finding.line} ${finding.preview}`)
}
if (rawJsonFindings.length > 25) {
console.log(` ... ${rawJsonFindings.length - 25} more`)
}
console.log(' untyped response schema examples:')
for (const finding of untypedResponseFindings.slice(0, 25)) {
console.log(` ${finding.path}:${finding.line} ${finding.preview}`)
}
if (untypedResponseFindings.length > 25) {
console.log(` ... ${untypedResponseFindings.length - 25} more`)
}
console.log(' annotations missing reason (must be 0 for --enforce-boundary-baseline):')
for (const finding of annotationsMissingReason.slice(0, 25)) {
console.log(` ${finding.path}:${finding.line} (${finding.kind})`)
}
if (annotationsMissingReason.length > 25) {
console.log(` ... ${annotationsMissingReason.length - 25} more`)
}
console.log(
' annotation forms: `// boundary-raw-fetch: <reason>` (raw fetch in client hook OR same-origin /api/ fetch outside an API route handler), `// double-cast-allowed: <reason>` (double-cast), `// boundary-raw-json: <reason>` (raw request.json read), `// untyped-response: <reason>` (z.unknown() / z.object({}).passthrough() / z.record(z.string(), z.unknown()) response schema)'
)
}
function printBoundaryContractDrift(
routeAudits: RouteAudit[],
queryHookAudits: QueryHookAudit[],
sameOriginApiFetchFindings: SameOriginApiFetchFinding[],
untypedResponseFindings: UntypedResponseFinding[],
ratchetedMetrics: BoundaryPolicyMetric[],
printOnlyMetrics: PrintOnlyBoundaryPolicyMetric[]
) {
const zodImportRoutes = routeAudits.filter((audit) => audit.hasZodImport)
const localSchemaRoutes = routeAudits.filter((audit) => audit.schemaConstructorCount > 0)
const zodErrorRoutes = routeAudits.filter((audit) => audit.hasZodErrorReference)
const zodImportQueryHooks = queryHookAudits.filter((audit) => audit.hasZodImport)
const localSchemaQueryHooks = queryHookAudits.filter((audit) => audit.schemaConstructorCount > 0)
const adHocWireTypes = queryHookAudits.flatMap((audit) => audit.adHocWireTypes)
const sameOriginApiFetchFiles = [
...new Set(sameOriginApiFetchFindings.map((finding) => finding.path)),
].sort()
console.log('\nBoundary policy drift:')
console.log(' ratcheted metrics:')
for (const metric of ratchetedMetrics) {
printBoundaryPolicyMetric(metric)
}
console.log(' print-only heuristics:')
for (const metric of printOnlyMetrics) {
console.log(` ${metric.label}: ${metric.current}`)
}
console.log(
' ratchet enforcement: pass --enforce-boundary-baseline to fail on ratcheted metric increases'
)
console.log(' ratchet update: lower BOUNDARY_POLICY_BASELINE after reducing a ratcheted count')
console.log('\nBoundary policy examples:')
console.log(' route zod import examples:')
for (const audit of zodImportRoutes.slice(0, 25)) {
console.log(` ${audit.path}`)
}
if (zodImportRoutes.length > 25) {
console.log(` ... ${zodImportRoutes.length - 25} more`)
}
console.log(' route local schema constructor examples:')
for (const audit of localSchemaRoutes.slice(0, 25)) {
console.log(` ${audit.path} (${audit.schemaConstructorCount})`)
}
if (localSchemaRoutes.length > 25) {
console.log(` ... ${localSchemaRoutes.length - 25} more`)
}
console.log(' route ZodError reference examples:')
for (const audit of zodErrorRoutes.slice(0, 25)) {
console.log(` ${audit.path}`)
}
if (zodErrorRoutes.length > 25) {
console.log(` ... ${zodErrorRoutes.length - 25} more`)
}
console.log(' client hook zod import examples:')
for (const audit of zodImportQueryHooks.slice(0, 25)) {
console.log(` ${audit.path}`)
}
if (zodImportQueryHooks.length > 25) {
console.log(` ... ${zodImportQueryHooks.length - 25} more`)
}
console.log(' query-hook local schema constructor examples:')
for (const audit of localSchemaQueryHooks.slice(0, 25)) {
console.log(` ${audit.path} (${audit.schemaConstructorCount})`)
}
if (localSchemaQueryHooks.length > 25) {
console.log(` ... ${localSchemaQueryHooks.length - 25} more`)
}
console.log(' query-hook ad-hoc wire type examples:')
for (const finding of adHocWireTypes.slice(0, 25)) {
console.log(` ${finding.path}:${finding.line} ${finding.name}`)
}
if (adHocWireTypes.length > 25) {
console.log(` ... ${adHocWireTypes.length - 25} more`)
}
console.log(' apps/sim same-origin /api/ fetch file examples:')
for (const filePath of sameOriginApiFetchFiles.slice(0, 25)) {
console.log(` ${filePath}`)
}
if (sameOriginApiFetchFiles.length > 25) {
console.log(` ... ${sameOriginApiFetchFiles.length - 25} more`)
}
console.log(' contract untyped response schema examples:')
for (const finding of untypedResponseFindings.slice(0, 25)) {
console.log(` ${finding.path}:${finding.line} ${finding.preview}`)
}
if (untypedResponseFindings.length > 25) {
console.log(` ... ${untypedResponseFindings.length - 25} more`)
}
}
function boundaryPolicyFailures(metrics: BoundaryPolicyMetric[]): string[] {
return metrics
.filter((metric) => metric.current > BOUNDARY_POLICY_BASELINE[metric.key])
.map(
(metric) =>
`${metric.label} increased from ${BOUNDARY_POLICY_BASELINE[metric.key]} to ${metric.current}`
)
}
async function auditQueryHooks(): Promise<QueryHookAudit[]> {
const queryHookFiles = await walk(
QUERY_HOOKS_DIR,
(fileName) => /\.(ts|tsx)$/.test(fileName) && !/\.test\.(ts|tsx)$/.test(fileName)
)
const audits: QueryHookAudit[] = []
for (const filePath of queryHookFiles) {
const content = await readFile(filePath, 'utf8')
audits.push(auditQueryHook(filePath, content))
}
return audits
}
async function main() {
const checkOnly = process.argv.includes('--check')
const enforceBoundaryBaseline = process.argv.includes('--enforce-boundary-baseline')
const routeFiles = await walk(API_DIR, (fileName) => fileName === 'route.ts')
const audits: RouteAudit[] = []
const rawJsonFindings: RawJsonFinding[] = []
const annotationsMissingReason: AnnotationMissingReasonFinding[] = []
let rawJsonExemptions = 0
for (const filePath of routeFiles) {
const content = await readFile(filePath, 'utf8')
audits.push(auditRoute(filePath, content))
const rawJson = findRawJsonFindings(filePath, content)
rawJsonFindings.push(...rawJson.findings)
rawJsonExemptions += rawJson.exemptions
annotationsMissingReason.push(...rawJson.missingReasons)
}
const queryHookAudits = await auditQueryHooks()
const sourceFiles = await walkAllSourceFiles(ROOT, true)
const rawFetchFindings: RawFetchFinding[] = []
const sameOriginApiFetchFindings: SameOriginApiFetchFinding[] = []
const doubleCastFindings: DoubleCastFinding[] = []
let rawFetchExemptions = 0
let sameOriginApiFetchExemptions = 0
let doubleCastExemptions = 0
const appsSimRoot = path.join(ROOT, 'apps/sim')
for (const filePath of sourceFiles) {
const content = await readFile(filePath, 'utf8')
const normalized = filePath.replace(/\\/g, '/')
if (isClientHookFile(filePath)) {
const rawFetch = findRawFetchFindings(filePath, content)
rawFetchFindings.push(...rawFetch.findings)
rawFetchExemptions += rawFetch.exemptions
annotationsMissingReason.push(...rawFetch.missingReasons)
}
if (
normalized.startsWith(`${appsSimRoot}/`) &&
!isApiRouteHandler(filePath) &&
filePath !== path.join(ROOT, 'scripts', 'check-api-validation-contracts.ts')
) {
const sameOrigin = findSameOriginApiFetchFindings(filePath, content)
sameOriginApiFetchFindings.push(...sameOrigin.findings)
sameOriginApiFetchExemptions += sameOrigin.exemptions
annotationsMissingReason.push(...sameOrigin.missingReasons)
}
const doubleCast = findDoubleCastFindings(filePath, content)
doubleCastFindings.push(...doubleCast.findings)
doubleCastExemptions += doubleCast.exemptions
annotationsMissingReason.push(...doubleCast.missingReasons)
}
const contractFiles = await walk(CONTRACTS_DIR, (fileName) => /\.ts$/.test(fileName))
const untypedResponseFindings: UntypedResponseFinding[] = []
let untypedResponseExemptions = 0
for (const filePath of contractFiles) {
const content = await readFile(filePath, 'utf8')
const untyped = findUntypedResponseFindings(filePath, content)
untypedResponseFindings.push(...untyped.findings)
untypedResponseExemptions += untyped.exemptions
annotationsMissingReason.push(...untyped.missingReasons)
}
const { ratchetedMetrics, printOnlyMetrics } = buildBoundaryPolicyMetrics(
audits,
queryHookAudits,
{ findings: rawFetchFindings, exemptions: rawFetchExemptions },
{
findings: sameOriginApiFetchFindings,
exemptions: sameOriginApiFetchExemptions,
},
{ findings: doubleCastFindings, exemptions: doubleCastExemptions },
{ findings: rawJsonFindings, exemptions: rawJsonExemptions },
{ findings: untypedResponseFindings, exemptions: untypedResponseExemptions },
annotationsMissingReason
)
const totalRoutes = audits.length
const zodRoutes = audits.filter((audit) => audit.usesZod).length
const nonZodRoutes = totalRoutes - zodRoutes
console.log('API validation route audit')
console.log(` total routes: ${totalRoutes}`)
console.log(` Zod-backed routes: ${zodRoutes}`)
console.log(` non-Zod routes: ${nonZodRoutes}`)
console.log(
` baseline: total=${BASELINE.totalRoutes} zod=${BASELINE.zodRoutes} nonZod=${BASELINE.nonZodRoutes}`
)
printFamilyStats(audits)
printRiskyNonZodRoutes(audits)
printAllNonZodRoutes(audits)
printBoundaryContractDrift(
audits,
queryHookAudits,
sameOriginApiFetchFindings,
untypedResponseFindings,
ratchetedMetrics,
printOnlyMetrics
)
printRawFetchAndDoubleCastMetrics(
rawFetchFindings,
sameOriginApiFetchFindings,
doubleCastFindings,
rawJsonFindings,
untypedResponseFindings,
annotationsMissingReason,
rawFetchExemptions,
sameOriginApiFetchExemptions,
doubleCastExemptions,
rawJsonExemptions,
untypedResponseExemptions
)
if (!checkOnly) return
const failures: string[] = []
if (totalRoutes > BASELINE.totalRoutes) {
failures.push(`route count increased from ${BASELINE.totalRoutes} to ${totalRoutes}`)
}
if (nonZodRoutes > BASELINE.nonZodRoutes) {
failures.push(
`non-Zod routes increased from ${BASELINE.nonZodRoutes} to ${nonZodRoutes} (${zodRoutes} Zod-backed routes)`
)
}
if (enforceBoundaryBaseline) {
failures.push(...boundaryPolicyFailures(ratchetedMetrics))
}
if (failures.length > 0) {
console.error('\nAPI validation audit failed:')
for (const failure of failures) {
console.error(` - ${failure}`)
}
process.exit(1)
}
console.log('\nAPI validation audit passed.')
}
void main().catch((error) => {
console.error('API validation audit failed:', error)
process.exit(1)
})