Commit Graph
10 Commits
Author SHA1 Message Date
Theodore Li e369997eac fix(db): encode Date binds in raw sql, document the fetch_types contract (#6040)
* fix(db): encode Date binds in raw sql, document the fetch_types contract

* fix(testing): drop template placeholders from mock error message for biome
2026-07-29 13:06:34 -04:00
Theodore Li 12fb4a9db1 feat(db): auto-apply tracked script data migrations in db:migrate (#5497)
* feat(db): auto-apply tracked script data migrations in db:migrate

* fix(db): reset session lock_timeout before script migrations, guard journal insert

* improvement(db): re-verify advisory-lock session before script migrations
2026-07-07 20:55:42 -04:00
Waleed 2fa3dd65bc fix(db): retry the migration connection on transient slot exhaustion (#5226)
* ci(migrations): skip db:migrate on merges that change no migration files

Every push to main/staging ran db:migrate against the production/staging
database even when the merge changed no schema, so a no-op migration would dial
the DB and fail whenever it was at its connection limit (53300, slots reserved
for SUPERUSER) — red-X'ing UI-only merges.

Add a detect-migrations job (dorny/paths-filter on packages/db/migrations/**)
and pass the result into the reusable migrations workflow, which now skips the
apply step when no migration files changed. The migrate job still runs so
downstream build/deploy jobs that need it are never skipped, and the flag
defaults to 'true' so manual dispatch and any unknown value always apply
migrations — the gate only ever skips a provably-empty change.

* fix(db): retry the migration connection on transient slot exhaustion

The migration opens its session on the first query (the advisory-lock
acquire). When the deploy database briefly exhausts every non-superuser
connection slot at peak, that connect fails with 53300 ("remaining connection
slots are reserved for roles with the SUPERUSER attribute") and the whole
deploy's migrate step errors out — even when the spike clears within seconds.

Add a bounded connectWithRetry() before acquiring the lock that retries 53300,
the 08xxx connection_exception class, and the driver's transport errors with
backoff (10 attempts, ~90s ceiling). Non-transient errors (auth, bad config)
still fail fast. The migration is a single short-lived session, so waiting out
a transient spike is far safer than failing the deploy.

* ci: drop the migration paths-filter gate (out of scope)

Revert the detect-migrations gate carried over from the closed CI PR; we are
fixing the connection failure at its source (migrate.ts connection retry)
rather than gating db:migrate, which the reviewers correctly noted could leave
a previously-merged migration unapplied after a failed deploy.
2026-06-26 12:21:11 -07:00
Theodore Li a68d38ae75 feat(db): attribute Postgres connections by runtime via application_name (#5211)
* feat(db): attribute Postgres connections by runtime via application_name

* improvement(db): label migration-runner connection sim-migrate; trim DB_APP_NAME comment

* fix(db): label realtime's shared @sim/db connections sim-realtime too

The realtime process uses both its own socketDb pool and the shared @sim/db
client (handlers, preflight, permissions). Only socketDb was labeled, so the
shared client defaulted to sim-app, mislabeling much of realtime's DB traffic.
Set DB_APP_NAME=sim-realtime at the process level (bootstrap before the dynamic
@/index import for prod; dev/start scripts for local) so both clients report it.
2026-06-25 17:11:23 -04:00
Waleed 284edf0b85 improvement(db): opt-in read-replica client + migration runner hardening (#4955)
* improvement(db): add opt-in read-replica client and harden migration runner

* fix(db): detect wrapped lock-timeout errors, jittered retries, direct migration DSN support

* fix(db): audit fixes — unparameterized SET, primary reads for authz scoping, shared error helper

* fix(db): pin migration session — disable max_lifetime recycling, guard backend pid across retries

* fix(db): gate pid session guard to direct migration connections (pooler pids legitimately vary)

* improvement(billing): route display-only usage aggregations to the read replica via executor threading

* chore(db): drop call-site justification comments

* chore(db): tighten doc comments

* ci(migrations): map optional direct-connection DSN secrets

* fix(data-drains): add stability window to time cursors so late-visible rows are never skipped

* fix(billing): resolve pagination cursor on primary; replica for member ledger display

* fix(billing): usage-limits returns cost and limit from one computation
2026-06-10 18:33:58 -07:00
Waleed b2a485e164 fix(db): serialize concurrent migrations with a Postgres advisory lock (#4939)
* fix(db): serialize concurrent migrations with a Postgres advisory lock

Deployments start N app replicas at once, each with a migration sidecar.
drizzle migrate() has no cross-process lock, so all N read
__drizzle_migrations, all see the same migration pending, and all apply it
concurrently — one wins, the losers run the same DDL against already-mutated
state and exit 1 (e.g. DROP TABLE "form" -> table does not exist /
TaskFailedToStart). Wrap migrate() in a session-level pg_advisory_lock so
runners serialize: the winner migrates, the losers block, then re-read and
find nothing pending. Session locks auto-release on disconnect, so a crashed
runner never wedges the lock.

* fix(db): guard pg_advisory_unlock so it cannot mask a successful migration

If the explicit unlock throws (e.g. connection drops in the window after
migrate() commits), the exception bubbled to the outer catch and exited 1 —
falsely reporting a failed migration to the deploy orchestrator. The session
lock auto-releases on disconnect anyway, so swallow and log instead.

* refactor(db): move unlock-guard rationale to TSDoc helper
2026-06-09 21:54:27 -07:00
Vikhyath MondretiandCursor 49673055be improvement(logs): object storage backed tracespans (#4787)
* improvement(logs): obj storage backed tracespans

* fix storage write context

* fix tests

* address comments

* address comments

* chore(db): remove migration 0219 to regenerate after staging merge

Drops the 0219_robust_shard SQL, its snapshot, and the journal entry so the
trace-spans/cost schema migration can be regenerated on top of the latest
staging migration chain (avoids a number collision with staging's migrations).

Co-authored-by: Cursor <cursoragent@cursor.com>

* improvement(billing): accurate per-member usage via shared ledger helper

Per-member/per-user usage in the org-member routes now adds the usage_log
ledger to the currentPeriodCost baseline (which is no longer incremented),
via a shared getOrgMemberLedgerByUser helper to avoid repeating the
subscription→period→ledger lookup across the admin and member-facing routes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* regen migrations

* update migration

* address comments

* more code cleanup

* incorrect type cast

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-29 13:47:32 -07:00
Theodore Li 786c6f0607 fix(db): disable statement_timeout for migrations (#4714)
* fix(db): disable statement_timeout for migrations

* fix(ci): route migration workflow through guarded migrate.ts
2026-05-22 14:23:06 -04:00
Vikhyath Mondreti 41a1b50ace improvement(migrations): log better errors (#4260) 2026-04-21 22:06:05 -07:00
abhinavDhulipala 7971a64e63 fix(setup): db migrate hard fail and correct ini env (#3946) 2026-04-04 16:22:19 -07:00