mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
perf(tools): generate serializable tool metadata artifacts (#6153)
* perf(tools): generate serializable tool metadata artifacts
Adds `scripts/sync-tool-metadata.ts`, which projects the executable tool
registry down to the data half nobody needs a closure for, plus typed accessors
over the result. No consumer is rewired yet — that is the next PR.
`@/tools/registry` is a ~9,000-line barrel over 4,366 tools. Each `ToolConfig`
mixes plain data (`params`, `outputs`, `name`) with closures (`request.headers`,
`transformResponse`, `directExecution`, `postProcess`), and those closures reach
every integration's SDK client and parser — which is why reaching the barrel
costs ~4,700 modules. Every client-reachable caller was audited: none of them
need a closure. They need `outputs`, `params`, or an existence check.
Two artifacts, not one. `outputs` is ~4 MB of the ~8 MB and has a single
consumer, so it is emitted separately and exposed from its own module; callers
needing only params never load it.
The data is a JSON string parsed at runtime rather than an imported `.json` or
an object literal. That is not stylistic — with `resolveJsonModule` (enabled
repo-wide) a `.json` import makes TypeScript infer a literal type for all 4,366
entries:
tsc --noEmit, baseline 12.6s
tsc --noEmit, with `.json` imports 8m07s (38x)
tsc --noEmit, with string literals 12.0s
An ambient `declare module` does not short-circuit it (measured: 8m18s), and an
object literal is the same inference work. A single string literal is one cheap
token for the compiler and the bundler, and `JSON.parse` beats evaluating the
equivalent literal at runtime.
The generator refuses to emit any function value, so shipping executable config
to the client fails loudly instead of silently. `hosting` and `schemaEnrichment`
are excluded on those grounds — both hold functions and are server-only.
Also strips empty param entries: the registry has one (`stt_deepgram_v2`, an
`undefined`) which crashes callers that read `param.type` while iterating.
`JSON.stringify` drops `undefined` on its own, so the guard is there for an
explicit `null` — which serializes faithfully and would reach consumers — and to
warn either way.
Wires `tool-metadata:check` into CI alongside the other generated-contract
gates, and ignores the generated directory in biome (it exceeds the 1 MB limit
and was being skipped with a notice on every commit).
Adds a `tool-registry-boundary` skill covering which module to import, the three
non-obvious properties of the artifacts, and how to verify an edge is actually
cut — the canvas route reaches the registry through four redundant paths, so
cutting one alone moves the module count by ~1.
* fix(tools): harden the metadata accessors against inherited keys
Review found two real defects in the generated-metadata layer.
`JSON.parse` returns an object with the normal prototype, so a bare bracket
lookup resolved inherited members: `getToolMetadata('constructor')` returned a
*function* typed as `ToolMetadata`, and `getToolOutputsMetadata('toString')`
likewise — silently violating the accessors' documented "undefined if unknown"
contract. Guarded with `Object.hasOwn`, with a parameterised regression test
over `constructor`, `toString`, `valueOf`, `hasOwnProperty` and `__proto__`.
The generator's no-functions scan also gave up past ten levels of nesting. Param
and output schemas nest arbitrarily, so a deeper closure would have been dropped
silently by `JSON.stringify` while generation reported success — shipping an
incomplete schema and defeating the guarantee the scan exists to provide. The
depth cap is gone; a `WeakSet` handles the cycles that exposes.
* docs(tools): tell tool authors to regenerate the metadata artifacts
A new tool now has a second registration step. Client code reads `params` and
`outputs` from the generated artifacts rather than from the registry, so a tool
added without regenerating them is registered but invisible to the UI — and CI
fails on the stale artifacts.
`add-tools` and `add-integration` are where someone actually adds a tool, so the
step goes in both, next to the registry edit and in each checklist.
* docs(blocks): note when a block change needs tool-metadata regeneration
Adding a block alone needs no regeneration — it references existing tool IDs and
changes no tool's shape. But a change that touches a tool alongside the block
does, and this is where that is easy to miss: a block's `outputs` are authored
to match its tools' outputs, and the UI now reads those from the generated
metadata, so a stale artifact makes the block's declared outputs disagree with
what the panel renders (and fails CI).
Completes the tool-authoring surface alongside add-tools and add-integration.
* docs(tools): cover tool removal in the regeneration guidance
The three tool-authoring skills said to regenerate after adding or changing a
tool, but not after removing one. Removal is equally breaking and equally
guarded: deleting a tool from `tools/registry.ts` without regenerating fails
`tool-metadata:check` (verified — exit 1), so a contributor following the skill
literally would have hit a CI failure the skill never warned about.
This commit is contained in:
@@ -0,0 +1,232 @@
|
||||
#!/usr/bin/env bun
|
||||
/**
|
||||
* Projects the executable tool registry down to its serializable metadata.
|
||||
*
|
||||
* `apps/sim/tools/registry.ts` is a ~9,000-line barrel importing all 4,300+
|
||||
* tools. Each `ToolConfig` mixes plain data (`params`, `outputs`, `name`) with
|
||||
* closures (`request.headers`, `transformResponse`, `directExecution`,
|
||||
* `postProcess`), and it is those closures — and the SDK clients and API
|
||||
* helpers they reach — that make the barrel cost ~4,700 modules to compile.
|
||||
*
|
||||
* No client-reachable caller needs a closure. They need `outputs` (block output
|
||||
* inference), `params` (serialization and the tool-input panel), or merely
|
||||
* whether an id exists. So this script emits that data on its own, letting those
|
||||
* callers read tool metadata without pulling the registry.
|
||||
*
|
||||
* Two artifacts rather than one, because `outputs` is ~4 MB of the ~6 MB and has
|
||||
* a single consumer — keeping it separate means callers that only need `params`
|
||||
* or an id check don't pay for it:
|
||||
*
|
||||
* tools/generated/tool-metadata.ts id -> { name, description, version, params, oauth }
|
||||
* tools/generated/tool-outputs.ts id -> outputs
|
||||
*
|
||||
* Each artifact holds its data as one JSON string parsed at runtime — see
|
||||
* `serialize()` for why an imported `.json` or an object literal is not viable
|
||||
* at this size.
|
||||
*
|
||||
* Usage:
|
||||
* bun run scripts/sync-tool-metadata.ts # write artifacts
|
||||
* bun run scripts/sync-tool-metadata.ts --check # fail (exit 1) if stale
|
||||
*/
|
||||
import { mkdir, readFile, writeFile } from 'node:fs/promises'
|
||||
import { dirname, resolve } from 'node:path'
|
||||
import { fileURLToPath } from 'node:url'
|
||||
import { tools } from '../apps/sim/tools/registry'
|
||||
import type { ToolConfig } from '../apps/sim/tools/types'
|
||||
|
||||
const SCRIPT_DIR = dirname(fileURLToPath(import.meta.url))
|
||||
const ROOT = resolve(SCRIPT_DIR, '..')
|
||||
const GENERATED_DIR = resolve(ROOT, 'apps/sim/tools/generated')
|
||||
const METADATA_PATH = resolve(GENERATED_DIR, 'tool-metadata.ts')
|
||||
const OUTPUTS_PATH = resolve(GENERATED_DIR, 'tool-outputs.ts')
|
||||
|
||||
/**
|
||||
* Fields copied into `tool-metadata.ts`. Every one must be plain data.
|
||||
*
|
||||
* Deliberately excluded: `request`, `transformResponse`, `directExecution`,
|
||||
* `postProcess` (closures, and the whole reason the registry is expensive);
|
||||
* `hosting` and `schemaEnrichment` (contain predicates/`enrichSchema`, and are
|
||||
* only consumed server-side); `outputs` (emitted separately).
|
||||
*/
|
||||
const METADATA_FIELDS = ['name', 'description', 'version', 'params', 'oauth'] as const
|
||||
|
||||
type ToolRecord = Record<string, ToolConfig>
|
||||
|
||||
/**
|
||||
* Recursively locates any function value, which must never reach the artifacts.
|
||||
*
|
||||
* Unbounded in depth on purpose: param and output schemas nest arbitrarily, and
|
||||
* a depth cap would let a deeply-nested closure through — `JSON.stringify` drops
|
||||
* it silently, so the artifact would ship an incomplete schema while generation
|
||||
* reported success. `seen` guards the cycles that removing the cap exposes.
|
||||
*/
|
||||
function findFunctionPaths(
|
||||
value: unknown,
|
||||
path: string,
|
||||
found: string[],
|
||||
seen = new WeakSet<object>()
|
||||
): void {
|
||||
if (found.length >= 10 || value == null) return
|
||||
if (typeof value === 'function') {
|
||||
found.push(path)
|
||||
return
|
||||
}
|
||||
if (typeof value !== 'object') return
|
||||
if (seen.has(value as object)) return
|
||||
seen.add(value as object)
|
||||
|
||||
if (Array.isArray(value)) {
|
||||
value.forEach((item, i) => findFunctionPaths(item, `${path}[${i}]`, found, seen))
|
||||
return
|
||||
}
|
||||
for (const [key, item] of Object.entries(value as object)) {
|
||||
findFunctionPaths(item, `${path}.${key}`, found, seen)
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Drops empty param entries so consumers may iterate params unguarded.
|
||||
*
|
||||
* The registry contains one (`stt_deepgram_v2`), which crashes any caller that
|
||||
* reads `param.type` while iterating. That entry is `undefined`, which
|
||||
* `JSON.stringify` would drop anyway; this guard additionally covers an explicit
|
||||
* `null` — which serializes faithfully and would reach consumers — and surfaces
|
||||
* either case as a warning rather than silently.
|
||||
*/
|
||||
function normalizeParams(toolId: string, params: ToolConfig['params'] | undefined) {
|
||||
const normalized: Record<string, unknown> = {}
|
||||
let dropped = 0
|
||||
for (const [paramId, config] of Object.entries(params ?? {})) {
|
||||
if (config == null) {
|
||||
dropped++
|
||||
continue
|
||||
}
|
||||
normalized[paramId] = config
|
||||
}
|
||||
if (dropped > 0) {
|
||||
console.warn(`[tool-metadata] ${toolId}: dropped ${dropped} empty param entr(y/ies)`)
|
||||
}
|
||||
return normalized
|
||||
}
|
||||
|
||||
function build(registry: ToolRecord) {
|
||||
const metadata: Record<string, unknown> = {}
|
||||
const outputs: Record<string, unknown> = {}
|
||||
|
||||
// Sorted so the artifacts are stable across runs regardless of registry order.
|
||||
for (const toolId of Object.keys(registry).sort()) {
|
||||
const tool = registry[toolId] as ToolConfig & Record<string, unknown>
|
||||
if (!tool) continue
|
||||
|
||||
const entry: Record<string, unknown> = { id: tool.id ?? toolId }
|
||||
for (const field of METADATA_FIELDS) {
|
||||
if (field === 'params') {
|
||||
entry.params = normalizeParams(toolId, tool.params)
|
||||
} else if (tool[field] !== undefined) {
|
||||
entry[field] = tool[field]
|
||||
}
|
||||
}
|
||||
metadata[toolId] = entry
|
||||
if (tool.outputs !== undefined) outputs[toolId] = tool.outputs
|
||||
}
|
||||
|
||||
const offenders: string[] = []
|
||||
findFunctionPaths(metadata, 'metadata', offenders)
|
||||
findFunctionPaths(outputs, 'outputs', offenders)
|
||||
if (offenders.length > 0) {
|
||||
throw new Error(
|
||||
`Refusing to emit tool metadata: found non-serializable values at:\n ${offenders.join('\n ')}\n` +
|
||||
`Add the offending field to the exclusion list in ${'scripts/sync-tool-metadata.ts'}.`
|
||||
)
|
||||
}
|
||||
|
||||
return {
|
||||
metadata: serialize(
|
||||
metadata,
|
||||
'toolMetadata',
|
||||
'/** Serializable metadata for every built-in tool, keyed by tool id. */'
|
||||
),
|
||||
outputs: serialize(
|
||||
outputs,
|
||||
'toolOutputs',
|
||||
'/** Declared output shapes for every built-in tool, keyed by tool id. */'
|
||||
),
|
||||
toolCount: Object.keys(metadata).length,
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Escapes a JSON document into a single-quoted JavaScript string literal.
|
||||
*
|
||||
* Single quotes rather than double: JSON is dense with `"`, which would need
|
||||
* escaping inside a double-quoted literal and inflates the file by ~90%.
|
||||
*/
|
||||
function toJsStringLiteral(json: string): string {
|
||||
const escaped = json
|
||||
.replace(/\\/g, '\\\\')
|
||||
.replace(/'/g, "\\'")
|
||||
.replace(/\n/g, '\\n')
|
||||
.replace(/\r/g, '\\r')
|
||||
// Valid raw inside a JSON string, but line terminators in a JS literal.
|
||||
.replace(/\u2028/g, '\\u2028')
|
||||
.replace(/\u2029/g, '\\u2029')
|
||||
return `'${escaped}'`
|
||||
}
|
||||
|
||||
/**
|
||||
* Emits the data as a string parsed at runtime, rather than as an object
|
||||
* literal or an imported `.json`.
|
||||
*
|
||||
* Both of the obvious alternatives are unusable at this size. A `.json` import
|
||||
* (with `resolveJsonModule`, which this repo enables) makes TypeScript infer a
|
||||
* literal type for all 4,300+ entries: it took `tsc --noEmit` from **12.6s to
|
||||
* 8m07s**, a 38x regression, and an ambient `declare module` does not
|
||||
* short-circuit it. A generated object literal costs the same, since it is the
|
||||
* same inference work.
|
||||
*
|
||||
* A single string literal is one cheap token for the compiler and the bundler,
|
||||
* and `JSON.parse` on a large payload is faster at runtime than evaluating the
|
||||
* equivalent object literal.
|
||||
*
|
||||
* The trade-off is that these files diff as one line. That is acceptable for a
|
||||
* generated artifact nothing reads by eye and CI verifies wholesale.
|
||||
*/
|
||||
function serialize(entries: Record<string, unknown>, exportName: string, doc: string): string {
|
||||
const literal = toJsStringLiteral(JSON.stringify(entries))
|
||||
return `// Generated by scripts/sync-tool-metadata.ts — do not edit.
|
||||
// Regenerate with: bun run tool-metadata:generate
|
||||
|
||||
${doc}
|
||||
const ${exportName}: Record<string, unknown> = JSON.parse(
|
||||
${literal}
|
||||
)
|
||||
|
||||
export default ${exportName}
|
||||
`
|
||||
}
|
||||
|
||||
async function main() {
|
||||
const checkOnly = process.argv.includes('--check')
|
||||
const { metadata, outputs, toolCount } = build(tools as ToolRecord)
|
||||
|
||||
if (checkOnly) {
|
||||
const [existingMetadata, existingOutputs] = await Promise.all([
|
||||
readFile(METADATA_PATH, 'utf8').catch(() => null),
|
||||
readFile(OUTPUTS_PATH, 'utf8').catch(() => null),
|
||||
])
|
||||
if (existingMetadata !== metadata || existingOutputs !== outputs) {
|
||||
throw new Error('Generated tool metadata is stale. Run: bun run tool-metadata:generate')
|
||||
}
|
||||
console.log(`✓ tool metadata in sync (${toolCount} tools)`)
|
||||
return
|
||||
}
|
||||
|
||||
await mkdir(GENERATED_DIR, { recursive: true })
|
||||
await Promise.all([
|
||||
writeFile(METADATA_PATH, metadata, 'utf8'),
|
||||
writeFile(OUTPUTS_PATH, outputs, 'utf8'),
|
||||
])
|
||||
console.log(`✓ wrote tool metadata for ${toolCount} tools`)
|
||||
}
|
||||
|
||||
await main()
|
||||
Reference in New Issue
Block a user