perf(dev): re-enable the Turbopack dev filesystem cache (5.4x faster restarts) (#6151)

* perf(dev): re-enable the Turbopack dev filesystem cache (5.4x faster restarts)

`turbopackFileSystemCacheForDev` has been `false` since #5408 — a landing-page
homepage redesign whose description covers hero cards, feature-card aspect
ratios, eyebrow chips and a voice-input button color, and never mentions
Turbopack, caching, or dev performance. It was collateral, not a decision, and
it overrode the Next default (true since v16.1).

It is not the flag #6078/#6080 measured. That A/B was `...ForBuild` and its
conclusion stands — the build cache is a 3.2x regression and stays off. The two
flags look alike and are opposite decisions; both are now commented as such.

Measured on `/workspace/[workspaceId]/w`, n=3 per arm, SIGINT between runs:

  cache OFF   31.4s / 30.1s / 31.9s   RSS 9.0-9.8 GB
  cache ON     5.6s /  5.6s /  5.5s   RSS 4.4-5.1 GB

5.4x faster restarts, ~2x less resident memory. Cold compile against an empty
cache is unchanged (~32s either way) — the cache only pays back on restart,
which is the loop that actually hurts.

The cache is unbounded on disk: the abandoned one on this machine had reached
78 GB across 1,848 SST files, and a stale cache is slower to read back, so left
alone it erodes the win it exists to provide. `prune-turbopack-cache.ts` runs on
`predev` and drops it past a cap (default 20 GB, `SIM_TURBOPACK_CACHE_MAX_GB` to
override); `bun run dev:cache:prune` forces it. It never blocks `next dev` on a
maintenance failure.

Adds a `dev-performance` skill recording the cost model, the reference numbers,
and the benchmarking method — including that stopping the server with `kill -9`
mid-cache-write discards the cache and makes this exact win read as no win.

* improvement(dev): chain the cache prune into dev scripts instead of a predev hook

Review read the root `bun run dev` path as bypassing the `predev` hook and so
never capping the newly-enabled cache. Turbo does fire `pre*` hooks — verified
live, the run prints the prune before `next dev` — but the concern is fair in
that the guarantee rested on package-manager lifecycle semantics that are
invisible at the call site.

Chaining it explicitly removes the question entirely: every `dev` variant now
runs `bun run dev:cache:cap && …`, which holds on any invocation path, is
visible in the command itself, and drops the three duplicated `predev:*` entries
for one shared script.

Verified on both paths — direct `bun run dev` and root `turbo run dev`, the
latter printing:

  sim:dev: $ bun run dev:cache:cap && next dev --port 3000
  sim:dev: $ bun run ../../scripts/prune-turbopack-cache.ts

* docs(dev): document cache-corruption recovery, the cost of enabling the cache

Stress-tested the failure mode rather than assuming it: deliberately corrupting
an SST block makes Turbopack abort with a FATAL panic — it does not self-heal.

  FATAL: An unexpected Turbopack error occurred.
  Cache corruption detected: checksum mismatch in block 4 of 00000221.sst

`bun run dev:cache:prune` and restart fixes it; verified the canvas serves 200
again afterwards. Documented in the skill and in the script's header, since the
symptom is a hard crash and the remedy is not guessable.

This is the honest cost of turning the cache on. It is worth paying — a 5.4x
faster restart against a rare, loud, single-command failure — but it should be
written down rather than discovered.

Worth distinguishing from the adjacent case: an ordinary hard kill does *not*
corrupt the cache. Turbopack discards a partially-written cache and rebuilds it
silently, which is exactly why a `kill -9`-based benchmark reads as "no cache
win" (noted in the benchmarking section).

* refactor(dev): drop the dev-performance skill, keep its findings at the code

A whole skill was too much for what this is. The parts that are load-bearing —
why the two lookalike cache flags are opposite decisions, the measured numbers,
the corruption remedy, and the benchmarking trap — now live in the config and
script they describe, where someone changing the flag actually reads them.

The trap is the piece worth keeping: `next dev` compiles on demand so startup
time is meaningless, and stopping the server with `kill -9` makes Turbopack
discard a partially-written cache and rebuild silently — which reads as 'the
cache does nothing' and is how this flag stayed wrong for a month.

Dropped rather than relocated: generic advice that was not specific to this repo
(antivirus, Docker-on-macOS, orphaned processes) and a measured no-op
(`optimizePackageImports` for lucide-react changed nothing, 31.6s vs 31.7s).

* docs(dev): record the measured cost and concurrency behaviour of cache pruning

Stress-tested the maintenance path rather than assuming it is free.

Cost: the size walk is ~30ms on a real cache and ~85ms at 2,000 files — under 2%
of a 4.2s warm restart, and invisible against a cold one. It runs before every
dev start, so it needed to be cheap; it is.

Concurrency: pruning while a dev server is live (which happens when a second
server is started from the same checkout) does not crash it. The running server
keeps its in-memory state and kept serving HTTP 200 with zero panics. It does
stop persisting for the rest of that session, so its next start is cold once —
verified recovering at 23.4s then 4.5s. Worth writing down because the directory
silently never reappears mid-session, which looks like a bug if you go looking.

The cap is a backstop, not routine: a normal session sits at 1-2 GB against a
20 GB default.

* fix(dev): cap every app's Turbopack cache, not just apps/sim

`apps/docs` is a Next app too (`next dev --port 3001`) and overrides nothing, so
it uses the Next default where the dev filesystem cache is on. It already had an
uncapped 1.1 GB cache here, and the root `bun run dev` (`turbo run dev`) starts
it — so a teammate using the documented command was accumulating a cache nothing
would ever prune.

The script now resolves its target from the working directory instead of
hardcoding `apps/sim`, and each app chains its own cap. Per-app rather than one
sweep on purpose: a single pass would let one app's dev start delete a cache
another app is holding open, which costs that session its persistence.

Verified both: `apps/sim` and `apps/docs` each report and cap their own 1.1 GB
cache, and both dev servers start clean (`Ready in 299ms` / `229ms`, docs serving).

* refactor(dev): drop dev:cache:prune in favour of the existing dev:clean

`dev:cache:prune` duplicated `dev:clean`, which already existed in `apps/sim` and
does strictly more (`rm -rf .next/dev/cache` covers the Turbopack cache plus the
fetch and image caches). Two commands for one job is worse than one, and the
docs pointed at the newer, narrower of the two.

Removes it from both apps and gives `apps/docs` the `dev:clean` that `apps/sim`
already had, so the recovery command is the same everywhere. `dev:cache:cap`
stays — it is the chained step, used by more than one dev variant, and naming it
keeps the relative script path out of each command.

Verified `dev:clean` is a real remedy: corrupt a cache block, run it, restart —
canvas serves 200 with no panic.

Also corrects an overstatement. A damaged cache does not *always* abort
Turbopack; whether it panics depends on whether the damaged region is read, so
it is not reliably reproducible. Both notes now say "can abort" and give the same
remedy either way.
This commit is contained in:
Waleed
2026-08-01 11:27:35 -07:00
committed by GitHub
parent 273ca06ddb
commit 28abbc94c6
4 changed files with 155 additions and 5 deletions
+114
View File
@@ -0,0 +1,114 @@
#!/usr/bin/env bun
/**
* Caps the size of Turbopack's dev filesystem cache.
*
* `experimental.turbopackFileSystemCacheForDev` is ON (see `apps/sim/next.config.ts`)
* because it makes dev-server restarts 5.4x faster and cuts RSS ~1.9x. The tradeoff
* is that the cache is an append-heavy LSM store with no upstream GC: it grows with
* every route compiled and every code change, and it is never pruned. An abandoned
* cache in this repo reached 78 GB across 1,848 SST files.
*
* A stale cache is also slower to read back, so an unbounded one eventually erodes
* the win it exists to provide. This script drops the whole cache once it crosses a
* threshold; Turbopack rebuilds it on the next run (one cold compile, ~32s).
*
* Also the remedy when Turbopack aborts with `Cache corruption detected:
* checksum mismatch` — it does not self-heal from a corrupted cache, so prune
* and restart. (An ordinary hard kill does not corrupt it; Turbopack discards a
* partial write and rebuilds silently.)
*
* Runs automatically before every `dev` script. Also invokable directly:
* bun run scripts/prune-turbopack-cache.ts # prune if over the cap
* bun run scripts/prune-turbopack-cache.ts --force # always prune
* bun run scripts/prune-turbopack-cache.ts --dry-run # report only
*
* Override the cap with `SIM_TURBOPACK_CACHE_MAX_GB` (default 20).
*
* The cap is a backstop, not routine maintenance. A normal session's cache sits
* around 1-2 GB; the 78 GB case was a month of accumulation while the feature was
* effectively abandoned. Measured cost of the size walk: ~30ms on a real cache,
* ~85ms at 2,000 files — under 2% of a warm restart.
*
* Safe to run while a dev server is live, which happens if a second server is
* started from the same checkout: the running server keeps its in-memory state
* and continues serving (verified — no crash, no corruption). It does stop
* persisting for the rest of that session, so its *next* start is cold once, then
* back to normal. Nothing to clean up.
*
* Note for anyone benchmarking the cache: stop the dev server with SIGINT, not
* `kill -9`. Turbopack discards a partially-written cache and rebuilds silently,
* so a hard kill makes a real cache win read as no win at all.
*/
import { existsSync, statSync } from 'node:fs'
import { readdir, rm, stat } from 'node:fs/promises'
import path from 'node:path'
const DEFAULT_MAX_GB = 20
const BYTES_PER_GB = 1024 ** 3
/**
* Resolved from the working directory, not hardcoded to one app: every Next app
* in the monorepo has its own dev cache (`apps/docs` runs `next dev` too and, with
* no override, uses the Next default where the cache is on). Each app caps its
* own, so one app's dev start never prunes a cache another app is holding open.
*/
const CACHE_DIR = path.resolve(process.cwd(), '.next', 'dev', 'cache', 'turbopack')
/** Sums the size of every regular file under `dir`, skipping symlinks. */
async function directorySize(dir: string): Promise<number> {
let total = 0
const entries = await readdir(dir, { withFileTypes: true, recursive: true })
for (const entry of entries) {
if (!entry.isFile()) continue
try {
total += (await stat(path.join(entry.parentPath, entry.name))).size
} catch {
// Turbopack compacts while we walk; a file vanishing mid-scan is expected.
}
}
return total
}
function parseMaxGb(): number {
const raw = process.env.SIM_TURBOPACK_CACHE_MAX_GB
if (!raw) return DEFAULT_MAX_GB
const parsed = Number(raw)
if (!Number.isFinite(parsed) || parsed <= 0) {
console.warn(
`[turbopack-cache] ignoring invalid SIM_TURBOPACK_CACHE_MAX_GB=${raw}, using ${DEFAULT_MAX_GB}`
)
return DEFAULT_MAX_GB
}
return parsed
}
async function main() {
const force = process.argv.includes('--force')
const dryRun = process.argv.includes('--dry-run')
if (!existsSync(CACHE_DIR) || !statSync(CACHE_DIR).isDirectory()) return
const maxGb = parseMaxGb()
const bytes = await directorySize(CACHE_DIR)
const gb = bytes / BYTES_PER_GB
const sizeLabel = `${gb.toFixed(1)} GB`
if (!force && gb <= maxGb) {
if (dryRun) console.log(`[turbopack-cache] ${sizeLabel} (cap ${maxGb} GB) — keeping`)
return
}
const reason = force ? 'forced' : `over the ${maxGb} GB cap`
if (dryRun) {
console.log(`[turbopack-cache] ${sizeLabel} is ${reason} — would prune (dry run)`)
return
}
console.log(`[turbopack-cache] ${sizeLabel} is ${reason} — pruning; next start recompiles cold`)
await rm(CACHE_DIR, { recursive: true, force: true })
}
main().catch((error) => {
// Never block `next dev` on a cache-maintenance failure.
console.warn('[turbopack-cache] prune skipped:', error)
})