* fix(workflow-edges): enforce edge/block validation server-side, not just client-side
Dragging a connection that creates a cycle correctly refused to render
client-side, but the cyclic edge was still queued for realtime persistence
and written to the DB, so it reappeared after refresh. Root cause: cycle
detection (and several other edge/block rules) only lived in the client
Zustand store and was never enforced by the realtime persistence layer,
which is the actual source of truth on reload.
- Move wouldCreateCycle, edge scope-boundary, annotation-only-block,
duplicate-edge, and block-name-conflict checks into @sim/workflow-types
so the client store, the collaborative queueing layer, and
apps/realtime's DB write path all share one implementation
- Wire these into apps/realtime/database/operations.ts's edge-add and
block-rename handlers as the authoritative gate
- Client-side behavior is unchanged (same call sites, same error messages,
same rule ordering) — verified via existing + new test coverage
* fix(workflow-edges): address review findings on realtime edge validation
- Select triggerMode when fetching blocks for edge-add validation —
isKnownWorkflowTriggerBlock checked block.triggerMode but the column
was never fetched from the DB, so trigger-mode blocks could still
receive an incoming edge (Cursor Bugbot)
- Make filterUniqueWorkflowEdges incremental, so two duplicate edges
within the same BATCH_ADD_EDGES payload are also deduped instead of
both surviving (Greptile)
* fix(workflow-edges): normalize empty-string handles in duplicate-edge check
filterUniqueWorkflowEdges compared handles with ??, so a `sourceHandle: ''`
edge wasn't recognized as a duplicate of an existing null-handle edge —
even though both get persisted as the same null value at insert time
(edge.sourceHandle || null). Falsy-coalesce in the comparison so '' and
null/undefined are treated as the same "no handle" state everywhere.
(Greptile)
* improvement(workflow-edges): dedup realtime edge-add validation, reuse block-name-conflict helper
/simplify pass on the already-merged-quality PR before sign-off:
- Extract filterEdgesForPersist in apps/realtime/src/database/operations.ts:
the single-edge ADD and batch BATCH_ADD_EDGES handlers hand-inlined the
same six-step validation pipeline (missing block, protected target,
annotation-only, trigger-target, scope boundary, duplicate, cycle) and
each independently re-fetched blocksById/existingEdgesForCycleCheck. Two
copies of one rule in the same file was exactly the drift risk this PR
otherwise closes across client/server. One shared helper now backs both.
Net -63 lines despite the new shared function.
- Fix a real bug this surfaced: droppedCounts was keyed by the free-text,
block-id-bearing scope-boundary message, so it could never aggregate
across edges/runs. Now keyed by a stable 'scope boundary' reason.
- use-collaborative-workflow.ts's collaborativeUpdateBlockName still
hand-rolled the empty/reserved/duplicate block-name-conflict checks this
PR centralized as getWorkflowBlockNameConflict (already adopted by
store.ts). Switched it to the shared helper, which also fixes a latent
check-order mismatch between the two (this pre-check ran reserved before
duplicate; the store's real gate — after this PR's own store.ts change —
runs duplicate before reserved, so they could disagree on which toast a
name that was both reserved and duplicate would surface).
* docs(workflow-edges): document the per-workflow write-serialization invariant
No behavior change. Address Greptile P1 on filterEdgesForPersist ('concurrent
duplicate writes can persist without a per-workflow write guard') by
documenting, at the actual mechanism, why the concern doesn't apply here:
persistWorkflowOperation's leading 'UPDATE workflow SET updatedAt ... WHERE
id = workflowId' already takes a row lock that serializes every operation
(including edge adds) for a given workflowId — a second concurrent call
blocks on that UPDATE until the first transaction commits or rolls back, so
the validate-then-insert sequence in filterEdgesForPersist can never
interleave across two writers on the same workflow.
Verified empirically, not just by reading: ran two concurrent transactions
against a throwaway local Postgres against the exact statement shape used
here (UPDATE the parent row, sleep to simulate the read/validate window,
insert, commit). The second transaction's UPDATE blocked for the full
duration of the first's transaction and only proceeded once the first
committed — confirming the row lock, not any application-level guard,
already provides the serialization Greptile flagged as missing.
Added a comment at the lock site (not a second, redundant advisory lock)
so a future change can't silently break this invariant by making the
UPDATE conditional/skippable as a perceived no-op optimization.
* fix(workflow-edges): validate edges in BATCH_ADD_BLOCKS before persisting
Real gap Cursor's PR summary flagged ('BATCH_ADD_BLOCKS edge inserts... not
fully covered by the new server pipeline'), confirmed by reading the code:
this handler bulk-inserted the edges from a block-paste/duplicate/import
payload directly into workflowEdges with zero validation — no missing-block,
protected-target, annotation-only, trigger-target, scope-boundary,
duplicate, or cycle check. A client sending edges through this operation
instead of BATCH_ADD_EDGES could bypass every rule this PR otherwise
enforces server-side, exactly the class of gap the PR exists to close.
Routes it through the same filterEdgesForPersist used by the other two
edge-add handlers. Runs after the block insert in this same handler, so the
shared helper's blocksById lookup also sees the blocks this same batch just
inserted (a transaction observes its own prior writes).
Two tool entries of the same type inside an Agent block's tool-input array
(e.g. two Table tools) shared a single canonical-mode override keyed by
${toolType}:${canonicalId}, so switching basic/advanced mode on one field
silently switched it on every other instance of the same tool type -
including at execution time, where the wrong basic/advanced value could be
resolved for the second tool.
Rescope the override key to the tool's position in its tool-input array
(${toolIndex}:${canonicalId}) instead of its type, and thread that index
through every consumer: the editor (read + write), execution
(agent-handler/providers), search-index, and fork/promote remapping.
* feat(db): resolve DATABASE_URL per role (DATABASE_URL_<ROLE> with fallback)
* fix(db): pin realtime process to SIM_DB_ROLE=realtime so both pools share the role
Without it, the realtime process left SIM_DB_ROLE unset: the shared @sim/db
client defaulted role to 'web' (web pool profile + DATABASE_URL_WEB) while
socketDb used 'realtime', so the two pools diverged after cutover. Set it at the
process level (bootstrap + dev/start scripts), mirroring DB_APP_NAME, so the
shared client and socketDb both resolve the realtime profile and URL.
* perf(db): drive Postgres pool size + application_name from per-role profiles
Replace ad-hoc DB_APP_NAME sizing with a per-role profile map keyed by
SIM_DB_ROLE (web/trigger/realtime), defaulting to web. Trigger machines
open a small pool instead of 15 to avoid PgBouncer connection exhaustion.
Also size realtime's separate socketDb pool down to 10.
* fix(db): throw on invalid SIM_DB_ROLE instead of silently using web pools
* fix(db): use Object.hasOwn for SIM_DB_ROLE validation to avoid prototype keys
* feat(db): attribute Postgres connections by runtime via application_name
* improvement(db): label migration-runner connection sim-migrate; trim DB_APP_NAME comment
* fix(db): label realtime's shared @sim/db connections sim-realtime too
The realtime process uses both its own socketDb pool and the shared @sim/db
client (handlers, preflight, permissions). Only socketDb was labeled, so the
shared client defaulted to sim-app, mislabeling much of realtime's DB traffic.
Set DB_APP_NAME=sim-realtime at the process level (bootstrap before the dynamic
@/index import for prod; dev/start scripts for local) so both clients report it.
* improvement(governance): org-ws-credential roles clarity
* revert isHosted
* improvement(credentials): code cleanup
* address comments
* make kb cascade delete on user hard delete
* revert env flags
* chore(db): drop local 0242 migration to regenerate after merging staging
Our 0242 collides with staging's 0242. Remove it (and its snapshot +
journal entry) so the KB-cascade migration can be regenerated with the
correct number on top of the merged staging migrations.
* chore(db): regenerate kb→workspace cascade migration as 0243
Regenerated via drizzle-kit generate on top of the merged staging
migrations (staging took 0242). Re-applied the safety edits: NOT VALID
+ separate VALIDATE on the FK re-add, and the -- migration-safe note on
the DROP. check:migrations passes.
* improve copy
* update docs
* feat(realtime): preflight schema-compatibility check on startup
The socket service authorizes every connection with a full-row query against
the workflow table. When a deploy ships a realtime image whose compiled schema
is ahead of/behind the live DB (e.g. a column dropped by a migration the image
predates), that query fails on every request and silently breaks persistence —
yet the process stays up and the shallow /health probe keeps returning 200, so
the deploy looks healthy while serving nothing.
Run one representative workflow query before listen(): a schema mismatch throws,
propagates to the entrypoint, and the task exits non-zero and never goes healthy,
so CodeDeploy auto-rolls-back instead of shifting traffic onto broken tasks.
Schema-class errors (undefined column/table/function) fail fast; connection-class
errors retry with backoff so a cold DB at boot does not flap. Runs once at
startup, never on the per-probe LB health check, to avoid a DB blip mass-
terminating the fleet (cascading failure).
* fix(realtime): unwrap cause for schema codes, drop sleep after final attempt
- isSchemaMismatch now walks the error.cause chain — drizzle wraps the driver
error, so the SQLSTATE often lives on the inner cause, not the outer throw.
Without this a wrapped 42703/42P01 was retried 5x and mis-reported as
"database unreachable" instead of failing fast.
- No longer sleeps after the final failed attempt (~6-10s of dead wait that
undermined the fail-fast contract); sleep now only happens between attempts.
- Tests: assert sleep is called exactly 4 times on exhaustion, and add a
wrapped-cause fail-fast case.
* chore(utils): migrate to shared random/ID utilities and add enforcement linting
- Replace all Math.random(), crypto.randomUUID(), crypto.randomBytes(), nanoid, and uuid usages with shared @sim/utils/random and @sim/utils/id helpers across 72 files
- Add new @sim/utils exports: deepClone, omit, filterUndefined (object), truncate (string), backoffWithJitter, parseRetryAfter (retry), getErrorMessage (errors)
- Sweep all getErrorMessage, sleep, deepClone callsites across 500+ files to use shared utilities
- Add Biome noRestrictedImports rule to catch nanoid, uuid, and crypto named imports at lint time
- Add scripts/check-utils-enforcement.ts to catch Math.random and crypto.* global property access
- Add check:utils script to package.json
* chore(utils): replace deepClone wrapper with structuredClone built-in
deepClone() was a one-line wrapper around structuredClone(), which is
universally available in Node 17+ and all modern browsers. Removing the
abstraction reduces indirection and means contributors don't need to
learn a project-specific name for a well-known built-in.
- Remove deepClone from packages/utils/src/object.ts and index.ts
- Replace all 17 call sites with structuredClone() directly
- Update check:utils script suggestion text
- Update CLAUDE.md and global.md docs
* fix(utils): add missing biome noRestrictedImports rule and correct truncate docs
- Add noRestrictedImports to biome.json under style — bans nanoid and uuid
package imports at lint time (crypto.randomUUID/randomBytes are caught by
the check:utils grep script which handles global property access)
- Correct truncate() TSDoc and parameter name: sliceLength makes it clear
that total output length is sliceLength + suffix.length, matching the
behavior all callers were already written to expect
* fix(utils): add missing getErrorMessage imports at 4 call sites
The sweep agents added getErrorMessage calls without the corresponding
import in 4 files, causing test failures. Added the missing imports.
* fix(utils): fix build errors from getErrorMessage sweep and retry.ts Turbopack issue
- Fix retry.ts cross-file import: Turbopack cannot resolve './random.js' for
internal package imports; inline the jitter crypto call directly
- Add missing getErrorMessage imports to 32 files where the sweep added calls
without the corresponding import (caught by type-check and test runs)
- Remove accidental getErrorMessage import from crowdstrike/query/route.ts
which has its own domain-specific getErrorMessage for parsing CrowdStrike's
JSON error format
- Fix use-sub-block-value.ts type error from structuredClone narrowing:
add 'as T' cast at emitValue callsite (safe — valueCopy is always a
structural copy of newValue)
* fix(tools): use toError in crowdstrike catch block instead of local getErrorMessage
The catch block was calling the local getErrorMessage function which
parses CrowdStrike API JSON responses, not JavaScript Error objects.
Use toError(error).message to correctly extract the message from a
caught value in this context.
* improvement(repo): restructuring to make realtime image narrower scoped
* improvements
* chore(repo): rebase fixes and quality improvements for realtime split
Addresses merge-time issues and gaps from the realtime app split:
- Retarget stale vi.mock paths to @sim/workflow-persistence/subblocks
- Restore README branding, fix AGENTS.md script reference
- Restore TSDoc on workflow-persistence subblocks helpers
- Use toError() from @sim/utils/errors in save.ts
- Add vitest config + local mocks so @sim/audit tests run standalone
- Move socket.io-client to devDependencies in apps/realtime
- Add missing package COPY steps to docker/app.Dockerfile
- Add check:boundaries/check:realtime-prune scripts and wire into CI
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* refactor(security): consolidate crypto primitives into @sim/security
Move general-purpose crypto primitives out of apps/sim into the
@sim/security package so both apps/sim and apps/realtime can share them.
@sim/security exports (all pure, dependency-free):
./compare safeCompare (constant-time HMAC-wrapped equality)
./encryption encrypt/decrypt (AES-256-GCM, iv:cipher:tag format)
./hash sha256Hex
./tokens generateSecureToken (base64url)
Migrate apps/sim call sites to use these + @sim/utils helpers:
crypto.randomUUID() -> generateId() from @sim/utils/id
createHash('sha256').digest -> sha256Hex
timingSafeEqual on hashed hex -> safeCompare
new Promise(setTimeout) -> sleep from @sim/utils/helpers
No behavior change: encryption format, digest output, and token
length are preserved exactly.
* refactor(copilot): use toError in remaining otel/finalize sites
Replace the last two `error instanceof Error ? error : new Error(String(error))`
patterns with toError from @sim/utils/errors. Completes the sweep of clean
candidates — no behavior change.
* refactor(security): consolidate HMAC-SHA256 primitives into @sim/security
Adds hmacSha256Hex and hmacSha256Base64 to @sim/security/hmac and migrates
15 webhook providers plus 5 other hot paths (deployment token signing,
outbound webhook requests, workspace notification delivery, notification
test route, Shopify OAuth callback) off bare `createHmac` calls. Secret
parameter accepts `string | Buffer` to cover base64-decoded Svix-style
secrets (Resend) and MS Teams' HMAC scheme. AWS SigV4 signing in S3 and
Textract tools intentionally retains direct `createHmac` usage — its
multi-step key derivation chain doesn't fit a generic helper.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* chore(packages): post-audit test + packaging polish
- Add safeCompare unit tests (identity, length mismatch, hex-nibble diff).
- Add Buffer-secret cases to hmac tests to lock in Svix/MS-Teams contract.
- Declare `reactflow` as a peerDependency on @sim/workflow-types — only used for type imports.
- Add a barrel export to @sim/workflow-persistence for consumers that prefer package-level imports; subpath exports retained.
- Document the data-field invariant in load.ts for loop/parallel subflow patching.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* chore(realtime): address PR review feedback
- Remove redundant SOCKET_PORT=3002 env from Dockerfile runner stage
(env.PORT already defaults to 3002 via zod schema).
- Reorder PORT fallback so an explicitly-set SOCKET_PORT wins over
the schema default for PORT; keeps SOCKET_PORT functional as an
override instead of dead code.
- Add dedicated type-check CI step for @sim/realtime so TS errors
surface pre-deploy (the Dockerfile runs source TS via Bun and has
no implicit build-time type check).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* chore(realtime): remove unused SOCKET_PORT env var
SOCKET_PORT has lived in the socket server since the June 2025 refactor
but was never actually set in any deploy config — docker-compose.prod,
helm values/templates, .env.example, and docs all use PORT or the 3002
default exclusively. No self-hoster was ever pointed at SOCKET_PORT, so
removing it is safe.
Simplifies realtime port resolution to `env.PORT` (zod-validated with a
3002 default) and drops the orphaned sim-side schema entry.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
---------
Co-authored-by: Waleed Latif <walif6@gmail.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>