mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
* feat(embeddings): multi-provider Embeddings block on a shared core
The Embeddings block was OpenAI-only with a bare fetch: no batching, no
retry, no metering, and no hosted-key support. Meanwhile the knowledge-base
indexing path already had a real multi-provider engine. Nothing bridged the
two, so the block could not reach Gemini and the KB engine could not be
reached from a workflow.
Extract the shared core into lib/embeddings/ first, then build breadth on
top of it, so both the KB path and the block resolve models and providers
from one catalog and one set of adapters instead of a third parallel
implementation.
- lib/embeddings/: catalog, client, key resolution, batching, L2
normalization, and adapters for OpenAI, Azure OpenAI, Gemini, Cohere,
and Mistral
- lib/knowledge/embeddings.ts becomes a thin KB wrapper with its exported
signatures unchanged; the 1536-dimension vector invariant does not move
- one tool per provider from a shared factory, behind a single
/api/tools/embeddings route and contract
- new `embeddings` block type; the `openai` block is left functionally
untouched and only leaves the discovery surfaces via hideFromToolbar
plus sunset.replacedBy, so placed instances keep working unmigrated
- openai_embeddings is now an alias of embeddings_openai, so legacy
instances pick up batching, retry, and metering with no visible change
* fix(embeddings): report an unsupported dimension as a client error
The route validated the model and the provider match up front but left
`dimensions` to be checked inside embed(), where resolveDimensions throws
and the generic catch maps it to 502. A typo in the block's dimension
field, or a reference expression resolving to an out-of-range value, was
reported as an upstream gateway failure rather than bad input.
Resolve dimensions in the route alongside the other boundary checks and
return 400. The throw stays the single source of the message, so the two
call sites cannot drift.
Adds route tests covering auth, the response shape, each boundary
rejection, input normalization, and the 502 path for genuine provider
failures.
* fix(embeddings): only send a dimension when the caller asked to reduce
resolveDimensions() returns the model's native size when no reduction is
requested, and that resolved value was handed straight to the adapter. The
adapters guard on `dimensions !== undefined`, so the field was always
populated and always sent.
Models that support Matryoshka reduction accept their own native size, so
this was invisible for text-embedding-3-*, gemini-embedding-001,
embed-v4.0, and codestral-embed. Models that do not support the parameter
at all reject it outright: every unreduced request to text-embedding-ada-002
and mistral-embed failed with a 400, which is both of the models whose
catalog entry has no supportedDimensions.
Track the caller's explicit reduction separately from the resolved
dimensionality. The resolved value still drives reporting and billing; only
the requested one reaches the wire.
Found by driving the live provider matrix against all four providers.
* test(knowledge): de-flake the sync-engine suite
Every test dynamically imported the module under test, so the first one to
run paid the whole cold-load cost inside its own 10s timeout and failed
intermittently under load.
The dynamic imports were working around a hoisting problem: mockMapTags is
a top-level const read by a vi.mock factory, and vi.mock is hoisted above
it, so a static import of the module under test crashes with a
use-before-initialization error. Declaring the mock through vi.hoisted()
removes that constraint, which is the pattern the testing guidelines
already call for.
One static import replaces 42 dynamic ones. The file drops from ~15s to
~2s and passed 5 consecutive runs.
* fix(embeddings): drop a capability the selected model no longer offers
The per-model Dimensions and Task Type dropdowns each share one subblock
id, and nothing clears a stored subblock value when its dependsOn fields
change — dependsOn only feeds rendering. A choice made for one model
therefore outlives a switch to another.
Picking 3072 on text-embedding-3-large and switching to -3-small left 3072
stored while the dropdown offered at most 1536, and the block forwarded it.
Same for a task type: 'similarity' chosen on Gemini survived a switch to
Cohere, which has no equivalent input type.
The guards only checked that the model declared the capability at all, not
that the value was one it lists. Check membership so a stale value falls
back to the model's native size, or is omitted, instead of being sent and
rejected. The user cannot have deliberately chosen an option the dropdown
stopped presenting.
* feat(embeddings): use the latent-constellation mark for the block icon
Replaces the scatter-plot-on-axes placeholder with a centre node, four
neighbours, and the rays between them — a point and its nearest neighbours
in embedding space, which is what the block actually produces. The axes
mark read as a generic chart and said nothing specific to embeddings.
Nodes are filled so they hold their shape at small sizes. The rays carry
less weight than the nodes to keep the hierarchy, but at 1.6/0.9 rather
than the 1.4/0.75 they were drawn at, so they do not thin out to loose
dots in the 14px block-search row.
Kept byte-identical between the app and docs icon sets.
* fix(embeddings): declare the outputs the legacy openai block returns
openai_embeddings became an alias of embeddings_openai, so the legacy
block's runtime payload gained `provider` and `dimensions`. Its declared
outputs still listed only embeddings/model/usage, so the tag picker never
offered two fields every run demonstrably returns, and downstream blocks
could not reference them.
Declaring them is additive and does not touch execution. Asserts the
legacy block's output keys match the replacement's, since both run the
same tool and neither should expose fields the other lacks.
* fix(copilot): resolve same-id subblock variants before validating
A block may declare one field id several times, each variant conditioned
on another field — the embeddings block declares model, dimensions, and
taskType once per provider, and the image and video generators do the
same. Validation keyed a map by id alone, so whichever variant was
declared last silently became the validator for every write to that
field.
Programmatic edits to an embeddings block were therefore checked against
Mistral's option lists whatever the saved provider: `text-embedding-3-small`
was rejected as not one of mistral-embed/codestral-embed, and dimensions
valid only elsewhere (3072, 768) could not be set at all. Values that
happened to overlap the last variant passed, so automation saw partial
success rather than a clean failure.
Keep every candidate per id and pick the one whose condition holds,
evaluating against the mutation's inputs merged over the block's saved
values so a partial write still resolves. When no condition matches, fall
back to the union of all variants' options rather than guessing.
Conditions still never gate whether a field may be written — that was a
deliberate choice and a hidden field stays writable. They only select
which definition describes the field, and an unresolved condition widens
the accepted set instead of narrowing it.
* fix(copilot): prefer a conditioned variant over an unconditioned catch-all
An unconditioned same-id variant matches every set of values, so it would
shadow a genuinely selected variant purely by being declared first. Prefer
a variant that actually asserted something about the current values.
No block in the registry currently declares a catch-all ahead of a
conditioned variant on a field where it would change validation, so this
is a guard against the pattern rather than a fix for a live case.
* chore(embeddings): scope this branch to the multi-provider block
Two changes made while building the Embeddings block are not part of it and
ship separately, so their files are restored to staging here:
- copilot edit-workflow validation resolving same-id conditional subblock
variants. The embeddings block surfaced it, but it is a platform fix
affecting ~20 blocks that declare a field id more than once, and it
narrows what programmatic edits accept — that deserves its own review.
- the sync-engine test de-flake, which is unrelated test hygiene.
Both are preserved in full on feat/embeddings-full-snapshot.
Note this restores the reported bug where a programmatic edit to an
embeddings block validates model/dimensions against the last-declared
provider variant. The block is unaffected in the editor and at runtime.
* fix(embeddings): honor per-model token limits and bound the JSON input path
Review round 1.
Batching used one 8,000-token constant for every model, inherited from the
knowledge-base engine this branch extracted. `batchByTokenLimit` truncates
any single text above the limit it is given, so that constant both sent
oversized input to models with a lower ceiling and silently dropped content
models with a higher one accept:
- Gemini declares 2,048, so a 3,000-token text passed through whole and the
provider rejected it, surfacing as a 502. This also affected knowledge-base
indexing on staging, which uses the same constant.
- Cohere declares 128,000, so anything past 8,000 was truncated for no reason.
Batch against the selected model's own `maxInputTokens` instead. Using the
per-input ceiling as the per-batch budget also keeps every individual text
within it.
The contract bounds the array arm of `input`, but a JSON-encoded array
arrives as a plain string and `normalizeInput` only expands it after
validation — so neither the 1,000-input cap nor the non-empty checks applied
to the reference-expression path the route was written to accept. `"[]"`
also reported success with no vectors. Re-check the normalized list so the
bounds hold for both shapes.
* chore(embeddings): regenerate tool metadata for the new embedding tools
CI's tool-metadata:check gate failed: registering embeddings_openai,
embeddings_gemini, embeddings_cohere, and embeddings_mistral left the
generated tool-ids/metadata/outputs artifacts stale.
* fix(embeddings): project before batching, and keep the sunset block's docs icon
Review round 2.
Projection ran inside callEmbeddingAPI, after batchByTokenLimit had already
measured and truncated the original text. The projector rewrites resolved
secrets to placeholders, which changes length, so batching sized against a
string that was never sent: a lengthening projection then pushed input past
the model's ceiling and the provider rejected it, and a shortening one
discarded document content that would have fit.
Project once up front, then batch the projected text, so truncation measures
what actually goes to the provider. This also keeps projection to exactly one
call per embed(), so no retry can re-project.
Separately, marking the legacy openai block hideFromToolbar dropped it from
the generated docs icon map, which only retains hidden blocks when they are
versioned. integrations/openai.mdx is deliberately kept — docsLink is baked
into every placed instance — so BlockInfoCard lost its icon and fell back to
a text tile. A sunset block keeps its docs page for the same reason a hidden
versioned block does, so the generator now treats it the same way.
The sim-side integrations map still omits it, which is intended: that feeds
the discovery page a sunset block should not appear on, and placed blocks
render from the registry's own icon reference.
* fix(embeddings): override stale block params instead of omitting them
Review round 3.
The generic handler merges the params() result over the original inputs
(`{ ...inputs, ...transformedParams }`), so omitting a key leaves the stale
value in place. The previous round dropped an unsupported taskType or
dimensions by omission, which was therefore a no-op through the executor
path: a reduction or task type chosen for one model still reached the tool
after a model switch.
Rewrite each stale field to an explicit `undefined`, which does override in a
spread.
Same class of bug for `model` itself, which was forwarded whenever present
without checking it belongs to the selected provider. Every provider's model
dropdown shares the `model` id, so switching provider kept the previous
provider's model and failed at the route as a mismatch. It now falls back to
the provider's default unless the saved model actually belongs to it.
Tests assert the merged result rather than the returned object, since the
return shape alone cannot distinguish an omitted key from an overridden one —
which is exactly why the previous fix looked correct and was not.
* fix(embeddings): discount the batch ceiling when the tokenizer is foreign
Review round 4.
Batching measures with tiktoken, which only has encodings for OpenAI models —
every other id falls back to cl100k_base. Gemini's 2048, Cohere's 128k, and
Mistral's 8192 were therefore enforced in OpenAI token units, so an input near
one of those ceilings could still be rejected upstream or trimmed more than
needed.
A true fix needs per-provider tokenizers, which the repo does not have:
estimateTokenCount is a chars-per-token heuristic, and truncation needs a real
encode/decode pair to slice on a token boundary. So the ceiling is discounted
for foreign tokenizers rather than trusted exactly.
The discount is one-sided on purpose. Overshooting means the provider rejects
the whole request; undershooting only trims a text that was already at the
limit, so the margin errs toward the second.
resolveBatchTokenCeiling is a pure function tested directly, rather than
inferred from truncation behavior, so the guarantee holds per model as the
catalog grows.
* fix(embeddings): keep the batch ceiling exact and warn before truncating
Review round 5. Reverts the safety margin from round 4.
The two review findings were in direct tension: round 4 flagged that a
foreign model's ceiling is measured in tiktoken units, and the margin added
to absorb that error reintroduced the round 3 harm — valid content truncated
below the provider's declared limit.
The margin was the wrong trade. It swapped a loud failure for a silent one:
an undercount surfaces as a provider rejection the caller can see and act on,
while shortening an embedding's input produces a degraded vector that is
indistinguishable from a good one at every layer above it. Silent quality
loss in a retrieval index is the worse outcome, and it is also the harder one
to ever notice.
So the declared ceiling is applied exactly, and truncation is no longer
silent: an input above the limit now logs a warning naming the model, the
limit, and whether the count was approximate. hasApproximateTokenCount
records which models are counted with a foreign tokenizer without being used
to shrink anything.
The tokenizer imprecision itself remains, and cannot be fixed without
per-provider BPE the repo does not have — estimateTokenCount is a
chars-per-token heuristic, and truncation needs a real encode/decode pair to
slice on a token boundary.
* refactor(embeddings): drop dead surface and enforce OpenAI's item cap
Audit follow-ups on the multi-provider embeddings work:
- Enforce OpenAI's documented 2048-entry `input` array cap in the OpenAI and
Azure adapters. Nothing bounded item count on the OpenAI path — batching
bounds tokens per request, so a batch of many short inputs could exceed it.
- Make the provider item cap single-source. It was declared both on the catalog
entry and on the adapter, read through a `??`; the adapter is the wire-protocol
owner, so the catalog copy is gone.
- Have the knowledge-base view call `getKbEligibleModels()` instead of
re-deriving the same `kbEligible` filter inline.
- Remove dead surface: the unused `EMBEDDING_TASK_TYPES` constant,
`EmbeddingToolDefinition`, `HOSTED_KEY_PROVIDERS`, and the five request-body
fields (`workspaceId`, `workflowId`, `executionId`, `userId`,
`useHostedCostTracking`) the route never reads.
- Trim `@/lib/embeddings` to what callers outside the module use.
- Drop the route's manual request-id plumbing; `withRouteHandler` supplies it.
- Fix two comments that had drifted onto the wrong declaration.
* fix(embeddings): normalize reduced Cohere output; correct OpenAI token ceiling
Second validation pass against provider documentation.
- Cohere: normalize locally when `output_dimension` reduces below native.
Cohere documents the parameter as Matryoshka truncation but never states that
it renormalizes, and an unnormalized vector silently skews cosine similarity.
`l2Normalize` is idempotent, so this is a no-op if Cohere already returns unit
vectors and a correctness fix if it does not. Covered by a test that fails
without it.
- OpenAI: raise the per-input ceiling from 8191 to the 8192 the API reference
documents, so a maximal input is no longer truncated by one token.
- Share the OpenAI response type with the Azure adapter instead of declaring an
identical copy, mirroring how the mail providers share `_nodemailer`.
- Rewrite the Gemini item-cap comment to say the 100-item limit is observed
rather than documented, which is what Google's reference actually supports.
Docs: add a manual intro to the Embeddings page covering providers, models,
inputs, outputs, and comparability rules. The generated Input tables are empty
because `createEmbeddingTool` builds params programmatically and the docs
generator only reads literals, so the manual section carries that reference.
* fix(embeddings): split per-input and per-request token limits; close provider gaps
Four gaps found in the validation pass.
Gemini token counts were estimated, not measured. `BatchEmbedContentsResponse`
carries `usageMetadata.promptTokenCount`; without reading it the client fell back
to tiktoken, which has no Gemini encoding and silently used `cl100k_base` — the
wrong tokenizer on a count knowledge-base runs bill against.
`maxInputTokens` was doing two jobs: the per-input ceiling that decides
truncation, and the per-request budget that decides how many inputs share a
batch. These are different provider limits, and conflating them meant Cohere
packed batches against its 128k per-document ceiling while OpenAI's documented
300,000-token request cap went unenforced. They are now separate fields.
Truncation moves out of `batchByTokenLimit` and into `embed`, so it happens once,
against the per-input ceiling, and always logs. The request budget is floored at
that ceiling — a budget below it would truncate inputs the provider accepts.
Batch sizes are unchanged everywhere except Gemini, which rises from 2048 to the
8192 the other providers already used.
codestral-embed now offers its documented 3072 maximum. Its API default is 1536,
so the offered sizes straddle the default; the catalog invariant relaxes from
"native size first" to "native size present", which is what the block relies on.
The Mistral API-key field no longer differs from the other three. Sim stocks
`MISTRAL_API_KEY` — `mistral_parse` already hides its key field on hosted — so
one field with `hideWhenHosted` replaces the conditional pair.
Docs: correct the API-key row, which described the old Mistral-only behavior.
* refactor(embeddings): derive block options from the catalog; use shared helpers
Findings from a four-angle quality review.
Reuse: `splitByItemLimit` and `processWithConcurrency` were reimplementations of
`chunkArray` (`@sim/utils`) and `mapWithConcurrency`
(`@/lib/core/utils/concurrency`), so `lib/embeddings/batching.ts` is gone. That
helper's doc forbade a throwing mapper; embedding legitimately wants a failed
batch to fail the call, since a partial vector set is not a usable result, so the
contract is reworded to cover both intents rather than forked.
The block no longer hand-copies the catalog. Its model, task-type, and dimension
dropdowns are derived from `EMBEDDING_MODELS`, which deletes roughly 150 lines of
literals that had to be kept in step by a drift test. The comment claiming this
was impossible was wrong: `generate-docs.ts` only reads `subBlocks` looking for
an `id: 'operation'` entry, which this block does not have. Verified by
regenerating — `embeddings.mdx` and `integrations.json` come out byte-identical.
Single-sourced two maps that were stated twice: BYOK provider ids (which encode
the non-obvious gemini -> google mapping) and the per-provider default model.
The route previously took its default from `getModelsForProvider(provider)[0]`,
which silently depended on catalog key order.
Azure's `endpoint` and `apiVersion` are required on their own context type
instead of optional on the shared one, so the adapter can no longer be built
without them and emit an `undefined/...` URL.
Also: contract enums now `satisfies` the catalog unions so they cannot drift,
the barrel exports only what callers outside the module use, the redundant
`requestedDimensions` field is a parameter, the bare `getEmbeddingModelInfo()`
call is a named `assertKbEmbeddingModel`, and the route checks payload size
before scanning entries rather than copying the body first.
* docs(embeddings): correct comments that drifted from the code
A comment pass over the feature found four that no longer matched what they sat
on, all introduced by earlier rounds of this work.
The contract's `satisfies` note promised that adding a catalog provider could
not leave the wire enum stale. It cannot deliver that: `satisfies` proves every
listed member is valid, not that the list is exhaustive, so an addition stays
silently absent. Reworded to say what it does and does not catch.
The client cited Gemini as a provider that omits usage, which the Gemini adapter
now contradicts — it reads `usageMetadata.promptTokenCount`. Every adapter
defines `parseTokens`, so the fallback is about a response lacking a usage block,
not about a particular provider.
`l2Normalize` documented only Gemini, though Cohere now calls it for a different
and stronger reason, and "normalizes in place" read as mutation when the function
returns a copy.
The route's new size-guard comment claimed it avoids copying the payload; nothing
there copies. The real reason is that summing lengths gates before the per-entry
character scan.
Also: split the derived-sub-block TSDoc so both constants carry hover text, gave
the payload cap its own doc, dropped one comment that restated a signature, and
tightened two long blocks without losing a fact.
* fix(docs): generate tool inputs for factory-built tools
The four embeddings tools rendered header-only Input tables. `extractToolInfo`
finds a tool's `params` by regex over the tool's own file, and these files hold
nothing but a `createEmbeddingTool({...})` call — the params live in the
factory's module. There was already a fallback for a same-file `...spread` base,
so this adds the cross-module equivalent: follow the factory's import and read
`params` from there.
Two things surfaced once the tables populated.
`hosting` was not in the set of keys that terminate the `params` capture, so the
non-greedy match ran past it to `request:` and swallowed the whole hosting block.
Every tool with a `hosting:` section between `params:` and `request:` was
publishing `pricing` and `rateLimit` as if they were user-facing inputs — this
drops those rows from eight unrelated integration pages as well.
The shared apiKey description was a template literal, which the regex emitted
verbatim as `${name} API key`. It is now a static string, matching how every
other tool in the repo declares one.
Docs: the Embeddings page keeps a prose intro in its MANUAL-CONTENT block like
other integrations, with the hand-written input/output tables removed now that
the generated ones are correct. The sunset `openai` page loses its
`encodingFormat` row — page generation skips hidden blocks, so that page is
frozen and would otherwise keep advertising a parameter the aliased tool no
longer accepts.
---------
Co-authored-by: Waleed Latif <walif6@gmail.com>
4294 lines
152 KiB
TypeScript
Executable File
4294 lines
152 KiB
TypeScript
Executable File
#!/usr/bin/env ts-node
|
|
import fs from 'fs'
|
|
import path from 'path'
|
|
import { fileURLToPath, pathToFileURL } from 'url'
|
|
import { isVersionedType, stripVersionSuffix } from '@sim/utils/string'
|
|
import { glob } from 'glob'
|
|
import type { BlockCategory } from '../apps/sim/blocks/types'
|
|
import { IntegrationType } from '../apps/sim/blocks/types'
|
|
|
|
console.log('Starting documentation generator...')
|
|
|
|
/**
|
|
* Cache for resolved const definitions from types files.
|
|
* Key: "toolPrefix:constName" (e.g., "calcom:SCHEDULE_DATA_OUTPUT_PROPERTIES")
|
|
* Value: The resolved properties object
|
|
*/
|
|
const constResolutionCache = new Map<string, Record<string, any>>()
|
|
|
|
const __filename = fileURLToPath(import.meta.url)
|
|
const __dirname = path.dirname(__filename)
|
|
const rootDir = path.resolve(__dirname, '..')
|
|
|
|
const BLOCKS_PATH = path.join(rootDir, 'apps/sim/blocks/blocks')
|
|
const DOCS_OUTPUT_PATH = path.join(rootDir, 'apps/docs/content/docs/en/integrations')
|
|
const ICONS_PATH = path.join(rootDir, 'apps/sim/components/icons.tsx')
|
|
const DOCS_ICONS_PATH = path.join(rootDir, 'apps/docs/components/icons.tsx')
|
|
const INTEGRATIONS_DATA_PATH = path.join(rootDir, 'apps/sim/lib/integrations')
|
|
const LANDING_INTEGRATIONS_DATA_PATH = path.join(
|
|
rootDir,
|
|
'apps/sim/app/(landing)/integrations/data'
|
|
)
|
|
const TRIGGERS_PATH = path.join(rootDir, 'apps/sim/triggers')
|
|
// Integration triggers are merged into the same per-service page as the service's
|
|
// actions (one block per integration: actions + an optional Trigger).
|
|
const TRIGGER_DOCS_OUTPUT_PATH = DOCS_OUTPUT_PATH
|
|
|
|
/**
|
|
* Hand-written integration pages in DOCS_OUTPUT_PATH that the generator must
|
|
* never clobber. Every hand-authored `*-service-account` credential guide has
|
|
* to be listed here — these pages carry no `MANUAL-CONTENT` markers and no
|
|
* backing block, so the stale-doc cleanup deletes any that go unregistered.
|
|
*/
|
|
const HANDWRITTEN_INTEGRATION_DOCS = new Set([
|
|
'index',
|
|
'a2a',
|
|
'airtable-service-account',
|
|
'asana-service-account',
|
|
'atlassian-service-account',
|
|
'attio-service-account',
|
|
'box-service-account',
|
|
'calcom-service-account',
|
|
'clickup-service-account',
|
|
'google-service-account',
|
|
'hubspot-service-account',
|
|
'hubspot-setup',
|
|
'linear-service-account',
|
|
'monday-service-account',
|
|
'notion-service-account',
|
|
'pipedrive-service-account',
|
|
'salesforce-service-account',
|
|
'shopify-service-account',
|
|
'trello-service-account',
|
|
'wealthbox-service-account',
|
|
'webflow-service-account',
|
|
'zoho-desk-service-account',
|
|
'zoom-service-account',
|
|
])
|
|
|
|
/**
|
|
* Native Sim resource blocks (category 'blocks') that still get a generated
|
|
* integration page. The writer's filter, the stale-doc cleanup, and the icon
|
|
* map must all honor this set: cleanup would otherwise delete what the writer
|
|
* emits (losing manual content), and an icon map that omits these types leaves
|
|
* their pages rendering the two-letter text fallback instead of the icon.
|
|
*/
|
|
const NATIVE_RESOURCE_BLOCK_TYPES = new Set([
|
|
'memory',
|
|
'knowledge',
|
|
'table',
|
|
'enrichment',
|
|
'logs',
|
|
'deployments',
|
|
])
|
|
|
|
/** Trigger doc pages that are hand-written and must never be overwritten. */
|
|
const HANDWRITTEN_TRIGGER_DOCS = new Set([
|
|
'index',
|
|
'start',
|
|
'schedule',
|
|
'webhook',
|
|
'rss',
|
|
'table',
|
|
'sim',
|
|
])
|
|
|
|
/** Providers whose docs are already covered by hand-written pages. */
|
|
const SKIP_TRIGGER_PROVIDERS = new Set(['generic', 'rss', 'table', 'sim'])
|
|
|
|
/**
|
|
* Maps trigger provider names (from TriggerConfig.provider) to their
|
|
* corresponding block type when the two differ. Used to resolve icon
|
|
* colours from the block registry.
|
|
*/
|
|
const PROVIDER_TO_BLOCK_TYPE: Record<string, string> = {
|
|
'microsoft-teams': 'microsoft_teams',
|
|
'google-calendar': 'google_calendar',
|
|
'google-drive': 'google_drive',
|
|
'google-sheets': 'google_sheets',
|
|
jsm: 'jira_service_management',
|
|
}
|
|
|
|
/** Human-readable display names for trigger providers. */
|
|
const TRIGGER_PROVIDER_DISPLAY_NAMES: Record<string, string> = {
|
|
airtable: 'Airtable',
|
|
ashby: 'Ashby',
|
|
attio: 'Attio',
|
|
calcom: 'Cal.com',
|
|
calendly: 'Calendly',
|
|
circleback: 'Circleback',
|
|
confluence: 'Confluence',
|
|
fathom: 'Fathom',
|
|
fireflies: 'Fireflies',
|
|
github: 'GitHub',
|
|
gmail: 'Gmail',
|
|
gong: 'Gong',
|
|
'google-calendar': 'Google Calendar',
|
|
'google-drive': 'Google Drive',
|
|
'google-sheets': 'Google Sheets',
|
|
google_forms: 'Google Forms',
|
|
grain: 'Grain',
|
|
greenhouse: 'Greenhouse',
|
|
hubspot: 'HubSpot',
|
|
imap: 'IMAP',
|
|
intercom: 'Intercom',
|
|
jira: 'Jira',
|
|
lemlist: 'Lemlist',
|
|
linear: 'Linear',
|
|
'microsoft-teams': 'Microsoft Teams',
|
|
notion: 'Notion',
|
|
outlook: 'Outlook',
|
|
resend: 'Resend',
|
|
salesforce: 'Salesforce',
|
|
servicenow: 'ServiceNow',
|
|
slack: 'Slack',
|
|
stripe: 'Stripe',
|
|
telegram: 'Telegram',
|
|
tiktok: 'TikTok',
|
|
twilio_voice: 'Twilio Voice',
|
|
typeform: 'Typeform',
|
|
vercel: 'Vercel',
|
|
webflow: 'Webflow',
|
|
whatsapp: 'WhatsApp',
|
|
zoom: 'Zoom',
|
|
}
|
|
|
|
if (!fs.existsSync(DOCS_OUTPUT_PATH)) {
|
|
fs.mkdirSync(DOCS_OUTPUT_PATH, { recursive: true })
|
|
}
|
|
|
|
// Ensure docs components directory exists
|
|
const docsComponentsDir = path.dirname(DOCS_ICONS_PATH)
|
|
if (!fs.existsSync(docsComponentsDir)) {
|
|
fs.mkdirSync(docsComponentsDir, { recursive: true })
|
|
}
|
|
|
|
/** Runtime set of valid `IntegrationType` values, derived from the canonical enum. */
|
|
const INTEGRATION_CATEGORY_VALUES: ReadonlySet<IntegrationType> = new Set(
|
|
Object.values(IntegrationType)
|
|
)
|
|
|
|
/**
|
|
* Defensive shape for blocks parsed out of source files. Fields stay loose
|
|
* (`string`) so the AST-style extractor can populate them progressively; the
|
|
* canonical taxonomy is enforced at the JSON-write boundary inside
|
|
* `writeIntegrationsJson`.
|
|
*/
|
|
interface BlockConfig {
|
|
type: string
|
|
name: string
|
|
description: string
|
|
longDescription?: string
|
|
category: string
|
|
integrationType?: string
|
|
bgColor?: string
|
|
outputs?: Record<string, any>
|
|
tools?: {
|
|
access?: string[]
|
|
}
|
|
operations?: OperationInfo[]
|
|
docsLink?: string
|
|
[key: string]: any
|
|
}
|
|
|
|
/**
|
|
* True when a block's source text marks it as an unreleased `preview: true`
|
|
* block. THE single preview gate for this script — every surface it emits
|
|
* (docs .mdx, integrations.json, icon mapping) must consult this, because a
|
|
* missed gate publishes an unreleased block to docs.sim.ai, the catalog, the
|
|
* sitemap, and OG images. Mirrors the `hideFromToolbar` source-text checks.
|
|
*/
|
|
function isPreviewSource(blockContent: string): boolean {
|
|
return /preview\s*:\s*true/.test(blockContent)
|
|
}
|
|
|
|
/**
|
|
* Blank out `//` and block comments so source-text property probes match real
|
|
* code only. Without this, prose that quotes a property — e.g. slack.ts's
|
|
* "At v2 GA this becomes `hideFromToolbar: true`" — reads as the property
|
|
* itself and silently drops the block from every generated surface.
|
|
*
|
|
* Comment bodies are replaced with spaces rather than removed so byte offsets
|
|
* stay aligned with the original content. Deliberately not applied to
|
|
* {@link isPreviewSource}: that gate is fail-closed on purpose, and a
|
|
* false positive there only over-hides an unreleased block.
|
|
*/
|
|
function stripSourceComments(content: string): string {
|
|
return content
|
|
.replace(/\/\*[\s\S]*?\*\//g, (m) => m.replace(/[^\n]/g, ' '))
|
|
.replace(/(^|[^:])\/\/[^\n]*/g, (m, prefix) => prefix + ' '.repeat(m.length - prefix.length))
|
|
}
|
|
|
|
/**
|
|
* Find the position after the matching close delimiter for an opening delimiter.
|
|
* Assumes `content[openPos]` is the opening char (e.g. `{` or `[`).
|
|
* Returns the index one past the matching close char, or -1 if unbalanced.
|
|
*/
|
|
function findMatchingClose(
|
|
content: string,
|
|
openPos: number,
|
|
openChar = '{',
|
|
closeChar = '}'
|
|
): number {
|
|
let count = 1
|
|
let pos = openPos + 1
|
|
while (pos < content.length && count > 0) {
|
|
if (content[pos] === openChar) count++
|
|
else if (content[pos] === closeChar) count--
|
|
pos++
|
|
}
|
|
return count === 0 ? pos : -1
|
|
}
|
|
|
|
interface TriggerInfo {
|
|
id: string
|
|
name: string
|
|
description: string
|
|
}
|
|
|
|
interface TriggerConfigField {
|
|
id: string
|
|
title: string
|
|
type: string
|
|
required: boolean
|
|
description?: string
|
|
placeholder?: string
|
|
}
|
|
|
|
interface TriggerFullInfo {
|
|
id: string
|
|
name: string
|
|
description: string
|
|
provider: string
|
|
polling: boolean
|
|
outputs: Record<string, any>
|
|
configFields: TriggerConfigField[]
|
|
}
|
|
|
|
interface OperationInfo {
|
|
name: string
|
|
description: string
|
|
}
|
|
|
|
interface IntegrationEntry {
|
|
type: string
|
|
slug: string
|
|
name: string
|
|
description: string
|
|
longDescription: string
|
|
bgColor: string
|
|
iconName: string
|
|
docsUrl: string
|
|
operations: OperationInfo[]
|
|
operationCount: number
|
|
triggers: TriggerInfo[]
|
|
triggerCount: number
|
|
authType: 'oauth' | 'api-key' | 'none'
|
|
oauthServiceId?: string
|
|
category: BlockCategory
|
|
integrationType: IntegrationType
|
|
tags?: string[]
|
|
landingContent?: Record<string, unknown>
|
|
}
|
|
|
|
/** A block icon component together with the module it must be imported from. */
|
|
interface IconRef {
|
|
name: string
|
|
source: string
|
|
}
|
|
|
|
/**
|
|
* Copy the icons.tsx file from the main sim app to the docs app
|
|
* This ensures icons are rendered consistently across both apps
|
|
*/
|
|
function copyIconsFile(): void {
|
|
try {
|
|
console.log('Copying icons from sim app to docs app...')
|
|
|
|
if (!fs.existsSync(ICONS_PATH)) {
|
|
console.error(`Source icons file not found: ${ICONS_PATH}`)
|
|
return
|
|
}
|
|
|
|
const iconsContent = fs.readFileSync(ICONS_PATH, 'utf-8')
|
|
fs.writeFileSync(DOCS_ICONS_PATH, iconsContent)
|
|
|
|
console.log('✓ Icons successfully copied to docs app')
|
|
} catch (error) {
|
|
console.error('Error copying icons file:', error)
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Some trigger providers have no block of their own (`slack_app`, `twilio`) yet
|
|
* still get a generated page keyed by the provider id. Seed those provider ids
|
|
* from the trigger definitions' own `icon` so their pages render the brand mark
|
|
* instead of the two-letter fallback. Never overwrites a block-derived entry —
|
|
* the block is the canonical icon source when one exists.
|
|
*/
|
|
async function addTriggerProviderIcons(iconMapping: Record<string, IconRef>): Promise<void> {
|
|
const triggerFiles = (await glob(`${TRIGGERS_PATH}/**/*.ts`)).filter((f) => !f.includes('.test.'))
|
|
const previewOnly = await collectPreviewOnlyTriggerIds()
|
|
|
|
for (const file of triggerFiles) {
|
|
const fileContent = fs.readFileSync(file, 'utf-8')
|
|
const source = stripSourceComments(fileContent)
|
|
|
|
// Pair each trigger's `id` with the `provider` that follows it in the same
|
|
// config, so files holding several trigger configs attribute each provider
|
|
// (and its icon) to the right trigger.
|
|
const configRegex =
|
|
/\bid\s*:\s*['"]([^'"]+)['"][\s\S]{0,600}?\bprovider\s*:\s*['"]([^'"]+)['"]/g
|
|
|
|
for (const match of source.matchAll(configRegex)) {
|
|
const [, triggerId, provider] = match
|
|
if (iconMapping[provider]) continue
|
|
|
|
// Preview-only triggers get no page, so they need no provider icon.
|
|
if (previewOnly.has(triggerId)) continue
|
|
|
|
const iconName = extractIconNameFromContent(source.slice(match.index))
|
|
if (!iconName) continue
|
|
|
|
iconMapping[provider] = { name: iconName, source: resolveIconSource(fileContent, iconName) }
|
|
}
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Generate icon mapping from block definitions.
|
|
* Docs need hidden historical version keys so old BlockInfoCard references and
|
|
* versioned docs links still render icons, while landing only needs visible blocks.
|
|
*/
|
|
async function generateIconMapping(options: {
|
|
includeHidden: boolean
|
|
}): Promise<Record<string, IconRef>> {
|
|
try {
|
|
console.log('Generating icon mapping from block definitions...')
|
|
|
|
const iconMapping: Record<string, IconRef> = {}
|
|
const blockFiles = (await glob(`${BLOCKS_PATH}/*.ts`)).sort()
|
|
|
|
for (const blockFile of blockFiles) {
|
|
const fileContent = fs.readFileSync(blockFile, 'utf-8')
|
|
|
|
// For icon mapping, we need ALL blocks including hidden ones
|
|
// because V2 blocks inherit icons from legacy blocks via spread
|
|
// First, extract the primary icon from the file (usually the legacy block's icon)
|
|
const primaryIcon = extractIconNameFromContent(fileContent)
|
|
|
|
// Find all block exports and their types
|
|
const exportRegex = /export\s+const\s+(\w+)Block\s*:\s*BlockConfig[^=]*=\s*\{/g
|
|
let match
|
|
|
|
while ((match = exportRegex.exec(fileContent)) !== null) {
|
|
const blockName = match[1]
|
|
const startIndex = match.index + match[0].length - 1
|
|
|
|
// Extract the block content
|
|
const endIndex = findMatchingClose(fileContent, startIndex)
|
|
|
|
if (endIndex !== -1) {
|
|
const blockContent = fileContent.substring(startIndex, endIndex)
|
|
|
|
// Check hideFromToolbar - skip hidden blocks for docs but NOT for icon mapping
|
|
const hideFromToolbar = /hideFromToolbar\s*:\s*true/.test(
|
|
stripSourceComments(blockContent)
|
|
)
|
|
|
|
// Unreleased preview blocks never reach any public surface, icon map included.
|
|
if (isPreviewSource(blockContent)) {
|
|
continue
|
|
}
|
|
|
|
// Get block type
|
|
const blockType =
|
|
extractStringPropertyFromContent(blockContent, 'type') || blockName.toLowerCase()
|
|
|
|
// Get icon - either from this block or inherited from primary
|
|
const iconName = extractIconNameFromContent(blockContent) || primaryIcon
|
|
|
|
if (!blockType || !iconName) {
|
|
continue
|
|
}
|
|
|
|
// Skip trigger/webhook/rss blocks
|
|
if (
|
|
blockType.includes('_trigger') ||
|
|
blockType.includes('_webhook') ||
|
|
blockType.includes('rss')
|
|
) {
|
|
continue
|
|
}
|
|
|
|
// Get category for additional filtering
|
|
const category = extractStringPropertyFromContent(blockContent, 'category') || 'misc'
|
|
|
|
// Exclude first-party `blocks`-category primitives (except the native
|
|
// resource blocks that still get a generated docs page) and
|
|
// core/plumbing types. Keying the exception off
|
|
// `NATIVE_RESOURCE_BLOCK_TYPES` — the same set the docs writer uses —
|
|
// keeps the icon map from drifting behind the pages that consume it.
|
|
const baseType = stripVersionSuffix(blockType)
|
|
if (
|
|
(category === 'blocks' &&
|
|
!NATIVE_RESOURCE_BLOCK_TYPES.has(baseType) &&
|
|
!HANDWRITTEN_INTEGRATION_DOCS.has(baseType)) ||
|
|
ICON_MAP_EXCLUDED_TYPES.has(blockType)
|
|
) {
|
|
continue
|
|
}
|
|
|
|
const isVersionedBlockType = isVersionedType(blockType)
|
|
/**
|
|
* A sunset block keeps its docs page — `docsLink` is baked into every
|
|
* placed instance — so it still needs an icon there, exactly like a
|
|
* hidden versioned block. Without this it renders as a text tile.
|
|
*/
|
|
const isSunsetBlockType = /sunset\s*:\s*\{/.test(stripSourceComments(blockContent))
|
|
if (
|
|
!hideFromToolbar ||
|
|
(options.includeHidden && (isVersionedBlockType || isSunsetBlockType))
|
|
) {
|
|
iconMapping[blockType] = {
|
|
name: iconName,
|
|
source: resolveIconSource(fileContent, iconName),
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
await addTriggerProviderIcons(iconMapping)
|
|
|
|
console.log(`✓ Generated icon mapping for ${Object.keys(iconMapping).length} blocks`)
|
|
return iconMapping
|
|
} catch (error) {
|
|
console.error('Error generating icon mapping:', error)
|
|
return {}
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Write the icon mapping to the docs app
|
|
* This file is imported by BlockInfoCard to resolve icons automatically
|
|
*/
|
|
/**
|
|
* Sort strings to match Biome's organizeImports order:
|
|
* case-insensitive character-by-character, uppercase before lowercase as tiebreaker.
|
|
*/
|
|
function biomeSortCompare(a: string, b: string): number {
|
|
const minLen = Math.min(a.length, b.length)
|
|
for (let i = 0; i < minLen; i++) {
|
|
const al = a[i].toLowerCase()
|
|
const bl = b[i].toLowerCase()
|
|
if (al !== bl) return al < bl ? -1 : 1
|
|
if (a[i] !== b[i]) return a[i] < b[i] ? -1 : 1
|
|
}
|
|
return a.length - b.length
|
|
}
|
|
|
|
function writeIconMapping(iconMapping: Record<string, IconRef>): void {
|
|
try {
|
|
const iconMappingPath = path.join(rootDir, 'apps/docs/components/ui/icon-mapping.ts')
|
|
|
|
// Add bare-name aliases for versioned block types so trigger provider names resolve correctly.
|
|
// e.g. github_v2 → github, fireflies_v2 → fireflies, gmail_v2 → gmail
|
|
const withAliases: Record<string, IconRef> = { ...iconMapping }
|
|
for (const [blockType, iconRef] of Object.entries(iconMapping)) {
|
|
const baseType = stripVersionSuffix(blockType)
|
|
if (baseType !== blockType && !withAliases[baseType]) {
|
|
withAliases[baseType] = iconRef
|
|
}
|
|
}
|
|
|
|
const imports = renderIconImports(Object.values(withAliases))
|
|
|
|
// Generate mapping with direct references (no dynamic access for tree shaking)
|
|
const mappingEntries = Object.entries(withAliases)
|
|
.sort(([a], [b]) => a.localeCompare(b))
|
|
.map(([blockType, iconRef]) => ` ${formatIconMapKey(blockType)}: ${iconRef.name},`)
|
|
.join('\n')
|
|
|
|
const content = `// Auto-generated file - do not edit manually
|
|
// Generated by scripts/generate-docs.ts
|
|
// Maps block types to their icon component references
|
|
|
|
import type { ComponentType, SVGProps } from 'react'
|
|
${imports}
|
|
|
|
type IconComponent = ComponentType<SVGProps<SVGSVGElement>>
|
|
|
|
export const blockTypeToIconMap: Record<string, IconComponent> = {
|
|
${mappingEntries}
|
|
}
|
|
`
|
|
|
|
fs.writeFileSync(iconMappingPath, content)
|
|
console.log('✓ Icon mapping file written to docs app')
|
|
} catch (error) {
|
|
console.error('Error writing icon mapping:', error)
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Extract operation options from the subBlock with id: 'operation' (if present).
|
|
* Returns { label, id } pairs — label is the display name, id is the option's id field
|
|
* (used to construct the tool ID as `{blockType}_{id}`).
|
|
* Parses the subBlocks array using brace/bracket counting to safely traverse
|
|
* the nested structure without eval or a full AST parser.
|
|
*/
|
|
function extractOperationsFromContent(blockContent: string): { label: string; id: string }[] {
|
|
const subBlocksMatch = /subBlocks\s*:\s*\[/.exec(blockContent)
|
|
if (!subBlocksMatch) return []
|
|
|
|
// Locate the opening '[' of the subBlocks array
|
|
const arrayStart = subBlocksMatch.index + subBlocksMatch[0].length - 1
|
|
const arrayEnd = findMatchingClose(blockContent, arrayStart, '[', ']')
|
|
if (arrayEnd === -1) return []
|
|
const subBlocksContent = blockContent.substring(arrayStart + 1, arrayEnd - 1)
|
|
|
|
// Iterate over top-level objects in the subBlocks array, looking for id: 'operation'
|
|
let i = 0
|
|
while (i < subBlocksContent.length) {
|
|
if (subBlocksContent[i] === '{') {
|
|
const j = findMatchingClose(subBlocksContent, i)
|
|
if (j === -1) break
|
|
const objContent = subBlocksContent.substring(i, j)
|
|
|
|
if (/\bid\s*:\s*['"]operation['"]/.test(objContent)) {
|
|
const optionsMatch = /options\s*:\s*\[/.exec(objContent)
|
|
if (!optionsMatch) return []
|
|
|
|
const optArrayStart = optionsMatch.index + optionsMatch[0].length - 1
|
|
const optArrayEnd = findMatchingClose(objContent, optArrayStart, '[', ']')
|
|
if (optArrayEnd === -1) return []
|
|
const optionsContent = objContent.substring(optArrayStart + 1, optArrayEnd - 1)
|
|
|
|
// Extract { label, id } pairs from each option object
|
|
const pairs: { label: string; id: string }[] = []
|
|
const optionObjectRegex = /\{[^{}]*\}/g
|
|
let m
|
|
while ((m = optionObjectRegex.exec(optionsContent)) !== null) {
|
|
const optObj = m[0]
|
|
const labelMatch = /label\s*:\s*['"]([^'"]+)['"]/.exec(optObj)
|
|
const idMatch = /\bid\s*:\s*['"]([^'"]+)['"]/.exec(optObj)
|
|
if (labelMatch) {
|
|
pairs.push({ label: labelMatch[1], id: idMatch ? idMatch[1] : '' })
|
|
}
|
|
}
|
|
return pairs
|
|
}
|
|
i = j
|
|
} else {
|
|
i++
|
|
}
|
|
}
|
|
return []
|
|
}
|
|
|
|
/**
|
|
* Extract a mapping from operation id → tool id by scanning switch/case/return
|
|
* patterns in a block file. Handles both simple returns and ternary returns
|
|
* (for ternaries, takes the last quoted tool-like string, which is typically
|
|
* the default/list variant). Also picks up named helper functions referenced
|
|
* from tools.config.tool (e.g. selectGmailToolId).
|
|
*/
|
|
function extractSwitchCaseToolMapping(fileContent: string): Map<string, string> {
|
|
const mapping = new Map<string, string>()
|
|
const caseRegex = /\bcase\s+['"]([^'"]+)['"]\s*:/g
|
|
let caseMatch: RegExpExecArray | null
|
|
|
|
while ((caseMatch = caseRegex.exec(fileContent)) !== null) {
|
|
const opId = caseMatch[1]
|
|
if (mapping.has(opId)) continue
|
|
|
|
const searchStart = caseMatch.index + caseMatch[0].length
|
|
const searchEnd = Math.min(searchStart + 300, fileContent.length)
|
|
const segment = fileContent.substring(searchStart, searchEnd)
|
|
|
|
const returnIdx = segment.search(/\breturn\b/)
|
|
if (returnIdx === -1) continue
|
|
|
|
const afterReturn = segment.substring(returnIdx + 'return'.length)
|
|
// Limit scope to before the next case/default to avoid capturing sibling cases
|
|
const nextCaseIdx = afterReturn.search(/\bcase\b|\bdefault\b/)
|
|
const returnScope = nextCaseIdx > 0 ? afterReturn.substring(0, nextCaseIdx) : afterReturn
|
|
|
|
const toolMatches = [...returnScope.matchAll(/['"]([a-z][a-z0-9_]+)['"]/g)]
|
|
// Take the last tool-like string (underscore = tool ID pattern); for ternaries this
|
|
// is the fallback/list variant
|
|
const toolId = toolMatches
|
|
.map((m) => m[1])
|
|
.filter((id) => id.includes('_'))
|
|
.pop()
|
|
if (toolId) {
|
|
mapping.set(opId, toolId)
|
|
}
|
|
}
|
|
|
|
return mapping
|
|
}
|
|
|
|
/**
|
|
* Scan all tool files under apps/sim/tools/ and build a map from tool ID to description.
|
|
* Used to enrich operation entries with descriptions.
|
|
*/
|
|
interface ToolMaps {
|
|
desc: Map<string, string>
|
|
name: Map<string, string>
|
|
}
|
|
|
|
async function buildToolDescriptionMap(): Promise<ToolMaps> {
|
|
const toolsDir = path.join(rootDir, 'apps/sim/tools')
|
|
const desc = new Map<string, string>()
|
|
const name = new Map<string, string>()
|
|
try {
|
|
const toolFiles = await glob(`${toolsDir}/**/*.ts`)
|
|
for (const file of toolFiles) {
|
|
const basename = path.basename(file)
|
|
if (basename === 'index.ts' || basename === 'types.ts') continue
|
|
const content = fs.readFileSync(file, 'utf-8')
|
|
|
|
// Find every `id: 'tool_id'` occurrence in the file. For each, search
|
|
// the next ~600 characters for `name:` and `description:` fields, cutting
|
|
// off at the first `params:` block within that window. This handles both
|
|
// the simple inline pattern (id → description → params in one object) and
|
|
// the two-step pattern (base object holds params, ToolConfig export holds
|
|
// id + description after the base object).
|
|
const idRegex = /\bid\s*:\s*['"]([^'"]+)['"]/g
|
|
let idMatch: RegExpExecArray | null
|
|
while ((idMatch = idRegex.exec(content)) !== null) {
|
|
const toolId = idMatch[1]
|
|
if (desc.has(toolId)) continue
|
|
const windowStart = idMatch.index
|
|
const windowEnd = Math.min(windowStart + 600, content.length)
|
|
const window = content.substring(windowStart, windowEnd)
|
|
// Stop before any params block so we don't pick up param-level values
|
|
const paramsOffset = window.search(/\bparams\s*:\s*\{/)
|
|
const searchWindow = paramsOffset > 0 ? window.substring(0, paramsOffset) : window
|
|
// Match against the actual opening quote so apostrophes inside a
|
|
// double-quoted description (e.g. "Find someone's email") are preserved
|
|
// rather than being treated as the closing quote and truncating the value.
|
|
const descMatch = searchWindow.match(
|
|
/\bdescription\s*:\s*(?:'([^']{5,})'|"([^"]{5,})"|`([^`]{5,})`)/
|
|
)
|
|
const nameMatch = searchWindow.match(/\bname\s*:\s*(?:'([^']+)'|"([^"]+)"|`([^`]+)`)/)
|
|
if (descMatch) desc.set(toolId, descMatch[1] ?? descMatch[2] ?? descMatch[3] ?? '')
|
|
if (nameMatch) name.set(toolId, nameMatch[1] ?? nameMatch[2] ?? nameMatch[3] ?? '')
|
|
}
|
|
}
|
|
} catch {
|
|
// Non-fatal: descriptions will be empty strings
|
|
}
|
|
return { desc, name }
|
|
}
|
|
|
|
/**
|
|
* Detect the authentication type from block content.
|
|
* Returns 'oauth' if the block uses oauth-input credentials,
|
|
* 'api-key' if it uses a plain API key field, or 'none' otherwise.
|
|
*/
|
|
function extractAuthType(blockContent: string): 'oauth' | 'api-key' | 'none' {
|
|
// Prefer the authoritative `authMode` declaration when present.
|
|
if (/authMode\s*:\s*AuthMode\.OAuth\b/.test(blockContent)) return 'oauth'
|
|
if (/authMode\s*:\s*AuthMode\.(?:ApiKey|BotToken)\b/.test(blockContent)) return 'api-key'
|
|
// Fall back to credential subBlock heuristics for blocks without authMode.
|
|
if (/type\s*:\s*['"]oauth-input['"]/.test(blockContent)) return 'oauth'
|
|
if (/\bid\s*:\s*['"](?:apiKey|api_key|accessToken)['"]/.test(blockContent)) return 'api-key'
|
|
return 'none'
|
|
}
|
|
|
|
/**
|
|
* Length-preserving copy of `content` with string-literal and comment
|
|
* interiors blanked out, so delimiter scans cannot be tripped by braces or
|
|
* quotes inside them. Indices into the result line up with indices into
|
|
* `content`.
|
|
*/
|
|
function blankStringsAndComments(content: string): string {
|
|
return content.replace(
|
|
/(['"`])(?:\\[\s\S]|(?!\1)[^\\])*\1|\/\/[^\n]*|\/\*[\s\S]*?\*\//g,
|
|
(match) => match[0] + match.slice(1, -1).replace(/[^\n]/g, ' ') + match[match.length - 1]
|
|
)
|
|
}
|
|
|
|
/**
|
|
* Extract the OAuth service id from the block's `oauth-input` credential
|
|
* subBlock. Scoped to that subBlock's object literal so `serviceId` fields on
|
|
* other subBlocks (e.g. file selectors) are never picked up. Brace matching
|
|
* runs on a blanked copy of the content so string literals and comments
|
|
* containing braces cannot skew it.
|
|
*/
|
|
function extractOAuthServiceId(blockContent: string): string | undefined {
|
|
const typeMatch = /type\s*:\s*['"]oauth-input['"]/.exec(blockContent)
|
|
if (!typeMatch) return undefined
|
|
|
|
const scannable = blankStringsAndComments(blockContent)
|
|
let depth = 0
|
|
let objectStart = -1
|
|
for (let i = typeMatch.index; i >= 0; i--) {
|
|
const char = scannable[i]
|
|
if (char === '}') depth++
|
|
else if (char === '{') {
|
|
if (depth === 0) {
|
|
objectStart = i
|
|
break
|
|
}
|
|
depth--
|
|
}
|
|
}
|
|
if (objectStart === -1) return undefined
|
|
|
|
const objectEnd = findMatchingClose(scannable, objectStart)
|
|
if (objectEnd === -1) return undefined
|
|
const subBlockContent = blockContent.substring(objectStart, objectEnd)
|
|
return /serviceId\s*:\s*['"]([^'"]+)['"]/.exec(subBlockContent)?.[1]
|
|
}
|
|
|
|
/**
|
|
* Extract the list of trigger IDs from the block's `triggers.available` array.
|
|
* Handles blocks that declare `triggers: { enabled: true, available: [...] }`.
|
|
*/
|
|
function extractTriggersAvailable(blockContent: string, fileContent?: string): string[] {
|
|
const triggersMatch = /\btriggers\s*:\s*\{/.exec(blockContent)
|
|
if (!triggersMatch) return []
|
|
|
|
const start = triggersMatch.index + triggersMatch[0].length - 1
|
|
const trigEnd = findMatchingClose(blockContent, start)
|
|
if (trigEnd === -1) return []
|
|
const triggersContent = blockContent.substring(start, trigEnd)
|
|
|
|
if (!/enabled\s*:\s*true/.test(triggersContent)) return []
|
|
|
|
const availableMatch = /available\s*:\s*\[/.exec(triggersContent)
|
|
if (!availableMatch) return []
|
|
|
|
const arrayStart = availableMatch.index + availableMatch[0].length - 1
|
|
const arrayEnd = findMatchingClose(triggersContent, arrayStart, '[', ']')
|
|
if (arrayEnd === -1) return []
|
|
const arrayContent = triggersContent.substring(arrayStart + 1, arrayEnd - 1)
|
|
|
|
// Blocks like emailbison declare `available: [...LOCAL_TRIGGER_IDS]`;
|
|
// resolve same-file const spreads to their literal entries so those
|
|
// triggers are not silently dropped from the generated data.
|
|
let resolvedContent = arrayContent
|
|
const constSource = fileContent ?? blockContent
|
|
const spreadRegex = /\.\.\.(\w+)/g
|
|
let spreadMatch: RegExpExecArray | null
|
|
while ((spreadMatch = spreadRegex.exec(arrayContent)) !== null) {
|
|
const constMatch = new RegExp(`const\\s+${spreadMatch[1]}\\s*=\\s*\\[`).exec(constSource)
|
|
if (!constMatch) continue
|
|
const constStart = constMatch.index + constMatch[0].length - 1
|
|
const constEnd = findMatchingClose(constSource, constStart, '[', ']')
|
|
if (constEnd === -1) continue
|
|
resolvedContent += constSource.substring(constStart + 1, constEnd - 1)
|
|
}
|
|
|
|
const ids: string[] = []
|
|
const idRegex = /['"]([^'"]+)['"]/g
|
|
let m
|
|
while ((m = idRegex.exec(resolvedContent)) !== null) {
|
|
ids.push(m[1])
|
|
}
|
|
return ids
|
|
}
|
|
|
|
/**
|
|
* Scan all trigger definition files and build a registry mapping trigger IDs
|
|
* to their human-readable name and description.
|
|
*/
|
|
async function buildTriggerRegistry(): Promise<Map<string, TriggerInfo>> {
|
|
const registry = new Map<string, TriggerInfo>()
|
|
const SKIP = new Set(['index.ts', 'registry.ts', 'types.ts', 'constants.ts', 'utils.ts'])
|
|
|
|
const triggerFiles = (await glob(`${TRIGGERS_PATH}/**/*.ts`)).filter(
|
|
(f) => !SKIP.has(path.basename(f)) && !f.includes('.test.')
|
|
)
|
|
|
|
for (const file of triggerFiles) {
|
|
try {
|
|
const content = fs.readFileSync(file, 'utf-8')
|
|
|
|
// A file may export multiple TriggerConfig objects (e.g. v1 + v2 in
|
|
// the same file). Extract all exported configs by splitting on the
|
|
// export boundaries and parsing each one independently.
|
|
const exportRegex = /export\s+const\s+\w+\s*:\s*TriggerConfig\s*=\s*\{/g
|
|
let exportMatch
|
|
const exportStarts: number[] = []
|
|
|
|
while ((exportMatch = exportRegex.exec(content)) !== null) {
|
|
exportStarts.push(exportMatch.index)
|
|
}
|
|
|
|
// If no typed exports found, fall back to simple regex on whole file
|
|
const segments =
|
|
exportStarts.length > 0
|
|
? exportStarts.map((start, i) => content.substring(start, exportStarts[i + 1]))
|
|
: [content]
|
|
|
|
for (const segment of segments) {
|
|
const idMatch = /\bid\s*:\s*['"]([^'"]+)['"]/.exec(segment)
|
|
const nameMatch = /\bname\s*:\s*['"]([^'"]+)['"]/.exec(segment)
|
|
const descMatch = /\bdescription\s*:\s*['"]([^'"]+)['"]/.exec(segment)
|
|
|
|
// Deprecated triggers stay registered for existing workflows but are
|
|
// excluded from generated documentation.
|
|
if (/\bdeprecated\s*:\s*true/.test(segment)) continue
|
|
|
|
if (idMatch && nameMatch) {
|
|
registry.set(idMatch[1], {
|
|
id: idMatch[1],
|
|
name: nameMatch[1],
|
|
description: descMatch?.[1] ?? '',
|
|
})
|
|
}
|
|
}
|
|
} catch {
|
|
// skip unreadable files silently
|
|
}
|
|
}
|
|
|
|
console.log(`✓ Loaded ${registry.size} trigger definitions`)
|
|
return registry
|
|
}
|
|
|
|
/**
|
|
* Write the icon mapping TypeScript file for the shared integrations data
|
|
* directory (`apps/sim/lib/integrations`). Mirrors `writeIconMapping` (the
|
|
* docs-app variant) but targets the sim app so it imports from
|
|
* `@/components/icons`. Unlike the docs variant, no bare-name aliasing is
|
|
* applied because consumers always look up by the canonical (possibly
|
|
* versioned) `integration.type` emitted into `integrations.json`.
|
|
*/
|
|
function writeIntegrationsIconMapping(iconMapping: Record<string, IconRef>): void {
|
|
try {
|
|
if (!fs.existsSync(INTEGRATIONS_DATA_PATH)) {
|
|
fs.mkdirSync(INTEGRATIONS_DATA_PATH, { recursive: true })
|
|
}
|
|
const iconMappingPath = path.join(INTEGRATIONS_DATA_PATH, 'icon-mapping.ts')
|
|
|
|
const imports = renderIconImports(Object.values(iconMapping))
|
|
const mappingEntries = Object.entries(iconMapping)
|
|
.sort(([a], [b]) => a.localeCompare(b))
|
|
.map(([blockType, iconRef]) => ` ${formatIconMapKey(blockType)}: ${iconRef.name},`)
|
|
.join('\n')
|
|
|
|
const content = `// Auto-generated file - do not edit manually
|
|
// Generated by scripts/generate-docs.ts
|
|
// Maps block types to their icon component references for the integrations page
|
|
|
|
import type { ComponentType, SVGProps } from 'react'
|
|
${imports}
|
|
|
|
type IconComponent = ComponentType<SVGProps<SVGSVGElement>>
|
|
|
|
export const blockTypeToIconMap: Record<string, IconComponent> = {
|
|
${mappingEntries}
|
|
}
|
|
`
|
|
fs.writeFileSync(iconMappingPath, content)
|
|
console.log('✓ Integration icon mapping written')
|
|
} catch (error) {
|
|
console.error('Error writing integration icon mapping:', error)
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Collect all integration entries from block definitions and write integrations.json
|
|
* to the shared integrations data directory (`apps/sim/lib/integrations`).
|
|
* Applies the same visibility filters as the docs generation pipeline.
|
|
*/
|
|
async function writeIntegrationsJson(iconMapping: Record<string, IconRef>): Promise<void> {
|
|
try {
|
|
if (!fs.existsSync(INTEGRATIONS_DATA_PATH)) {
|
|
fs.mkdirSync(INTEGRATIONS_DATA_PATH, { recursive: true })
|
|
}
|
|
|
|
const triggerRegistry = await buildTriggerRegistry()
|
|
const { desc: toolDescMap, name: toolNameMap } = await buildToolDescriptionMap()
|
|
|
|
// Hand-authored, integration-specific landing content (install walkthrough,
|
|
// privacy blurb), keyed by slug. Imported as pure data — its only import is
|
|
// type-only and erased at runtime — and baked into the entries below so the
|
|
// landing page reads a single source instead of augmenting at render time.
|
|
const landingContentModule = await import(
|
|
pathToFileURL(path.join(LANDING_INTEGRATIONS_DATA_PATH, 'landing-content.ts')).href
|
|
)
|
|
const landingContentMap = (landingContentModule.INTEGRATION_LANDING_CONTENT ?? {}) as Record<
|
|
string,
|
|
Record<string, unknown>
|
|
>
|
|
|
|
const integrations: IntegrationEntry[] = []
|
|
const seenBaseTypes = new Set<string>()
|
|
const blockFiles = (await glob(`${BLOCKS_PATH}/*.ts`)).sort()
|
|
|
|
for (const blockFile of blockFiles) {
|
|
const fileContent = fs.readFileSync(blockFile, 'utf-8')
|
|
const switchCaseMap = extractSwitchCaseToolMapping(fileContent)
|
|
const configs = extractAllBlockConfigs(fileContent)
|
|
|
|
for (const config of configs) {
|
|
const blockType = config.type
|
|
|
|
// Canonical integrations filter: only third-party tool blocks visible in the toolbar.
|
|
// `isIntegrationBlock` is the single source of truth for "is integration".
|
|
if (!isIntegrationBlock(config)) continue
|
|
|
|
// Every tools-category block MUST declare an `integrationType` from the canonical
|
|
// 16-value enum (apps/sim/blocks/types.ts). Fail loudly so the catalog never
|
|
// ships a tool without a category bucket.
|
|
if (!config.integrationType) {
|
|
throw new Error(
|
|
`Block "${blockType}" has \`category: 'tools'\` but is missing required \`integrationType\`. ` +
|
|
`Add one of the IntegrationType values from apps/sim/blocks/types.ts.`
|
|
)
|
|
}
|
|
if (!INTEGRATION_CATEGORY_VALUES.has(config.integrationType as IntegrationType)) {
|
|
throw new Error(
|
|
`Block "${blockType}" has unrecognised \`integrationType: "${config.integrationType}"\`. ` +
|
|
`Use one of: ${[...INTEGRATION_CATEGORY_VALUES].join(', ')}.`
|
|
)
|
|
}
|
|
const integrationType = config.integrationType as IntegrationType
|
|
|
|
// Deduplicate by stripped base type
|
|
const baseType = stripVersionSuffix(blockType)
|
|
if (seenBaseTypes.has(baseType)) continue
|
|
seenBaseTypes.add(baseType)
|
|
|
|
const iconName = (config as any).iconName || iconMapping[blockType]?.name || ''
|
|
const rawOps: { label: string; id: string }[] = (config as any).operations || []
|
|
|
|
// Enrich each operation with a description from the tool registry.
|
|
// Lookup order:
|
|
// 1. Derive toolId as `{baseType}_{operationId}` and check directly.
|
|
// 2. Check switch/case mapping parsed from tools.config.tool (handles
|
|
// cases where op IDs differ from tool IDs, e.g. get_carts → list_carts,
|
|
// or send_gmail → gmail_send).
|
|
// 3. Find the tool in tools.access whose name exactly matches the label.
|
|
const toolsAccess: string[] = (config as any).tools?.access || []
|
|
const operations: OperationInfo[] = rawOps.map(({ label, id }) => {
|
|
const toolId = `${baseType}_${id}`
|
|
let opDesc = toolDescMap.get(toolId) || toolDescMap.get(id) || ''
|
|
|
|
if (!opDesc) {
|
|
const switchMappedId = switchCaseMap.get(id)
|
|
if (switchMappedId) {
|
|
opDesc = toolDescMap.get(switchMappedId) || ''
|
|
// Also check versioned variants in tools.access (e.g. gmail_send_v2)
|
|
if (!opDesc) {
|
|
for (const tId of toolsAccess) {
|
|
if (tId === switchMappedId || tId.startsWith(`${switchMappedId}_v`)) {
|
|
opDesc = toolDescMap.get(tId) || ''
|
|
if (opDesc) break
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
if (!opDesc && toolsAccess.length > 0) {
|
|
for (const tId of toolsAccess) {
|
|
if (toolNameMap.get(tId)?.toLowerCase() === label.toLowerCase()) {
|
|
opDesc = toolDescMap.get(tId) || ''
|
|
if (opDesc) break
|
|
}
|
|
}
|
|
}
|
|
|
|
return { name: label, description: opDesc }
|
|
})
|
|
|
|
const triggerIds: string[] = (config as any).triggerIds || []
|
|
const triggers: TriggerInfo[] = triggerIds
|
|
.map((id) => triggerRegistry.get(id))
|
|
.filter((t): t is TriggerInfo => t !== undefined)
|
|
const docsUrl = (config as any).docsLink || `https://docs.sim.ai/integrations/${baseType}`
|
|
|
|
const slug = config.name
|
|
.toLowerCase()
|
|
.replace(/[^a-z0-9]+/g, '-')
|
|
.replace(/^-|-$/g, '')
|
|
|
|
const authType = extractAuthType(fileContent)
|
|
const oauthServiceId = authType === 'oauth' ? extractOAuthServiceId(fileContent) : undefined
|
|
// OAuth integrations resolve their connect UI through the service id
|
|
// (see `resolveOAuthServiceForIntegration`), so fail loudly rather than
|
|
// shipping a catalog entry that silently falls back to the API-key path.
|
|
if (authType === 'oauth' && !oauthServiceId) {
|
|
throw new Error(
|
|
`Block "${blockType}" is an OAuth integration but no \`serviceId\` could be ` +
|
|
`extracted from its \`oauth-input\` subBlock.`
|
|
)
|
|
}
|
|
|
|
integrations.push({
|
|
type: blockType,
|
|
slug,
|
|
name: config.name,
|
|
description: config.description,
|
|
longDescription: config.longDescription || '',
|
|
bgColor: config.bgColor || '#6B7280',
|
|
iconName,
|
|
docsUrl,
|
|
operations,
|
|
operationCount: operations.length,
|
|
triggers,
|
|
triggerCount: triggers.length,
|
|
authType,
|
|
...(oauthServiceId ? { oauthServiceId } : {}),
|
|
category: 'tools',
|
|
integrationType,
|
|
...(config.tags ? { tags: config.tags } : {}),
|
|
...(landingContentMap[slug] ? { landingContent: landingContentMap[slug] } : {}),
|
|
})
|
|
}
|
|
}
|
|
|
|
// Sort alphabetically by name for a predictable, crawl-friendly order
|
|
integrations.sort((a, b) => a.name.localeCompare(b.name))
|
|
|
|
const jsonPath = path.join(INTEGRATIONS_DATA_PATH, 'integrations.json')
|
|
// `JSON.stringify` always expands every array across multiple lines, but Biome's
|
|
// JSON formatter inlines short arrays of primitive strings. Pre-collapse those
|
|
// arrays here so the emitted file is already in Biome's canonical shape and
|
|
// `bun run check` does not churn it on every commit.
|
|
const serialize = (value: unknown) =>
|
|
JSON.stringify(value, null, 2).replace(
|
|
/\[\n(\s+"[^"\n]*"(?:,\n\s+"[^"\n]*")*)\n\s+\]/g,
|
|
(_match, inner) => {
|
|
const items = (inner as string).split(',\n').map((s: string) => s.trim())
|
|
return `[${items.join(', ')}]`
|
|
}
|
|
)
|
|
|
|
// `updatedAt` is re-stamped only when the integrations content actually
|
|
// changes, so sitemap/JSON-LD freshness never churns on no-op regens.
|
|
const previous = fs.existsSync(jsonPath)
|
|
? (JSON.parse(fs.readFileSync(jsonPath, 'utf-8')) as { integrations?: unknown })
|
|
: null
|
|
if (previous?.integrations && serialize(previous.integrations) === serialize(integrations)) {
|
|
console.log(`✓ Integration data unchanged: ${integrations.length} integrations → ${jsonPath}`)
|
|
return
|
|
}
|
|
|
|
const updatedAt = new Date().toISOString().slice(0, 10)
|
|
fs.writeFileSync(jsonPath, `${serialize({ updatedAt, integrations })}\n`)
|
|
console.log(`✓ Integration data written: ${integrations.length} integrations → ${jsonPath}`)
|
|
} catch (error) {
|
|
// Surface taxonomy violations (missing/invalid `integrationType`) loudly —
|
|
// they are programmer errors that must fail the generator, not be logged
|
|
// and silently swallowed.
|
|
console.error('Error writing integrations JSON:', error)
|
|
throw error
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Extract ALL block configs from a file, filtering out hidden blocks
|
|
*/
|
|
function extractAllBlockConfigs(fileContent: string): BlockConfig[] {
|
|
const configs: BlockConfig[] = []
|
|
|
|
// First, extract the primary icon from the file (for V2 blocks that inherit via spread)
|
|
const primaryIcon = extractIconNameFromContent(fileContent)
|
|
|
|
// Find all block exports in the file
|
|
const exportRegex = /export\s+const\s+(\w+)Block\s*:\s*BlockConfig[^=]*=\s*\{/g
|
|
let match
|
|
|
|
while ((match = exportRegex.exec(fileContent)) !== null) {
|
|
const blockName = match[1]
|
|
const startIndex = match.index + match[0].length - 1 // Position of opening brace
|
|
|
|
// Extract the block content by matching braces
|
|
const endIndex = findMatchingClose(fileContent, startIndex)
|
|
|
|
if (endIndex !== -1) {
|
|
const blockContent = fileContent.substring(startIndex, endIndex)
|
|
|
|
// Check if this block has hideFromToolbar: true
|
|
const hideFromToolbar = /hideFromToolbar\s*:\s*true/.test(stripSourceComments(blockContent))
|
|
if (hideFromToolbar) {
|
|
console.log(`Skipping ${blockName}Block - hideFromToolbar is true`)
|
|
continue
|
|
}
|
|
|
|
// Unreleased preview blocks stay out of every generated surface: docs
|
|
// .mdx pages, integrations.json (landing + workspace catalog + sitemap +
|
|
// OG images), and the icon mapping.
|
|
if (isPreviewSource(blockContent)) {
|
|
console.log(`Skipping ${blockName}Block - preview is true`)
|
|
continue
|
|
}
|
|
|
|
// Pass fileContent to enable spread inheritance resolution
|
|
const config = extractBlockConfigFromContent(blockContent, blockName, fileContent)
|
|
if (config) {
|
|
// For V2 blocks that don't have an explicit icon, use the primary icon from the file
|
|
if (!config.iconName && primaryIcon) {
|
|
;(config as any).iconName = primaryIcon
|
|
}
|
|
configs.push(config)
|
|
}
|
|
}
|
|
}
|
|
|
|
return configs
|
|
}
|
|
|
|
/**
|
|
* Extract the name of the spread base block (e.g., "GitHubBlock" from "...GitHubBlock")
|
|
*/
|
|
function extractSpreadBase(blockContent: string): string | null {
|
|
const spreadMatch = blockContent.match(/^\s*\.\.\.(\w+Block)\s*,/m)
|
|
return spreadMatch ? spreadMatch[1] : null
|
|
}
|
|
|
|
/**
|
|
* Extract block config from a specific block's content
|
|
* If the block uses spread inheritance (e.g., ...GitHubBlock), attempts to resolve
|
|
* missing properties from the base block in the file content.
|
|
*/
|
|
function extractBlockConfigFromContent(
|
|
blockContent: string,
|
|
blockName: string,
|
|
fileContent?: string
|
|
): BlockConfig | null {
|
|
try {
|
|
// Check for spread inheritance
|
|
const spreadBase = extractSpreadBase(blockContent)
|
|
let baseConfig: BlockConfig | null = null
|
|
|
|
if (spreadBase && fileContent) {
|
|
// Extract the base block's content from the file
|
|
const baseBlockRegex = new RegExp(
|
|
`export\\s+const\\s+${spreadBase}\\s*:\\s*BlockConfig[^=]*=\\s*\\{`,
|
|
'g'
|
|
)
|
|
const baseMatch = baseBlockRegex.exec(fileContent)
|
|
|
|
if (baseMatch) {
|
|
const startIndex = baseMatch.index + baseMatch[0].length - 1
|
|
const endIndex = findMatchingClose(fileContent, startIndex)
|
|
|
|
if (endIndex !== -1) {
|
|
const baseBlockContent = fileContent.substring(startIndex, endIndex)
|
|
// Recursively extract base config (but don't pass fileContent to avoid infinite loops)
|
|
baseConfig = extractBlockConfigFromContent(
|
|
baseBlockContent,
|
|
spreadBase.replace('Block', '')
|
|
)
|
|
}
|
|
}
|
|
}
|
|
|
|
// Extract properties from this block, using topLevelOnly=true for main properties
|
|
const blockType =
|
|
extractStringPropertyFromContent(blockContent, 'type', true) || blockName.toLowerCase()
|
|
const name =
|
|
extractStringPropertyFromContent(blockContent, 'name', true) ||
|
|
baseConfig?.name ||
|
|
`${blockName} Block`
|
|
const description =
|
|
extractStringPropertyFromContent(blockContent, 'description', true) ||
|
|
baseConfig?.description ||
|
|
''
|
|
const longDescription =
|
|
extractStringPropertyFromContent(blockContent, 'longDescription', true) ||
|
|
baseConfig?.longDescription ||
|
|
''
|
|
const category =
|
|
extractStringPropertyFromContent(blockContent, 'category', true) || baseConfig?.category || ''
|
|
const bgColor =
|
|
extractStringPropertyFromContent(blockContent, 'bgColor', true) ||
|
|
baseConfig?.bgColor ||
|
|
'#F5F5F5'
|
|
const iconName = extractIconNameFromContent(blockContent) || (baseConfig as any)?.iconName || ''
|
|
|
|
const outputs = extractOutputsFromContent(blockContent)
|
|
const toolsAccess = extractToolsAccessFromContent(blockContent)
|
|
|
|
// For tools.access, if not found directly, check if it's derived from base via map
|
|
let finalToolsAccess = toolsAccess
|
|
if (toolsAccess.length === 0 && baseConfig?.tools?.access) {
|
|
// Check if there's a map operation on base tools
|
|
// Pattern: access: (SomeBlock.tools?.access || []).map((toolId) => `${toolId}_v2`)
|
|
const mapMatch = blockContent.match(
|
|
/access\s*:\s*\(\s*\w+Block\.tools\?\.access\s*\|\|\s*\[\]\s*\)\.map\s*\(\s*\(\s*\w+\s*\)\s*=>\s*`\$\{\s*\w+\s*\}_v(\d+)`\s*\)/
|
|
)
|
|
if (mapMatch) {
|
|
// V2 block - append the version suffix to base tools
|
|
const versionSuffix = `_v${mapMatch[1]}`
|
|
finalToolsAccess = baseConfig.tools.access.map((tool) => `${tool}${versionSuffix}`)
|
|
}
|
|
}
|
|
|
|
const operations = extractOperationsFromContent(blockContent)
|
|
const triggerIds = extractTriggersAvailable(blockContent, fileContent)
|
|
const docsLink =
|
|
extractStringPropertyFromContent(blockContent, 'docsLink', true) ||
|
|
baseConfig?.docsLink ||
|
|
`https://docs.sim.ai/integrations/${stripVersionSuffix(blockType)}`
|
|
|
|
const integrationType =
|
|
extractEnumPropertyFromContent(blockContent, 'integrationType') ||
|
|
baseConfig?.integrationType ||
|
|
null
|
|
// Tags live on the block's `<BlockName>BlockMeta` export. For spread-inheriting
|
|
// blocks (e.g. `ConfluenceV2Block` extending `ConfluenceBlock`), also try the
|
|
// spread base's meta so V2 variants inherit tags.
|
|
const tags =
|
|
(fileContent ? extractTagsFromBlockMeta(fileContent, blockName) : null) ||
|
|
(fileContent && spreadBase
|
|
? extractTagsFromBlockMeta(fileContent, spreadBase.replace(/Block$/, ''))
|
|
: null)
|
|
|
|
return {
|
|
type: blockType,
|
|
name,
|
|
description,
|
|
longDescription,
|
|
category,
|
|
bgColor,
|
|
iconName,
|
|
outputs,
|
|
tools: {
|
|
access: finalToolsAccess.length > 0 ? finalToolsAccess : baseConfig?.tools?.access || [],
|
|
},
|
|
operations: operations.length > 0 ? operations : (baseConfig as any)?.operations || [],
|
|
triggerIds: triggerIds.length > 0 ? triggerIds : (baseConfig as any)?.triggerIds || [],
|
|
docsLink,
|
|
...(integrationType ? { integrationType } : {}),
|
|
...(tags ? { tags } : {}),
|
|
}
|
|
} catch (error) {
|
|
console.error(`Error extracting block configuration for ${blockName}:`, error)
|
|
return null
|
|
}
|
|
}
|
|
|
|
/**
|
|
* The single predicate that decides whether an extracted block config belongs
|
|
* in the integration surfaces emitted by this script — the integrations
|
|
* catalog (`integrations.json`) and the per-tool `/tools/*.mdx` docs. A block
|
|
* qualifies only when it is a third-party integration (`category: 'tools'`)
|
|
* that is currently surfaced in the toolbar (`hideFromToolbar` not set). Under
|
|
* the versioning upgrade paradigm only the latest version is visible, so this
|
|
* also naturally selects the canonical version. Recategorizing a block to
|
|
* `'blocks'` or `'triggers'` removes it from all integration surfaces.
|
|
*/
|
|
function isIntegrationBlock(config: {
|
|
category?: string
|
|
hideFromToolbar?: boolean
|
|
preview?: boolean
|
|
}): boolean {
|
|
return config.category === 'tools' && !config.hideFromToolbar && !config.preview
|
|
}
|
|
|
|
/**
|
|
* Block types that never belong in the integrations icon map regardless of
|
|
* category — core primitives, triggers, and webhook/feed plumbing.
|
|
*/
|
|
const ICON_MAP_EXCLUDED_TYPES = new Set([
|
|
'evaluator',
|
|
'number',
|
|
'webhook',
|
|
'schedule',
|
|
'mcp',
|
|
'generic_webhook',
|
|
'rss',
|
|
])
|
|
|
|
/**
|
|
* Extract a string property from block content.
|
|
* For top-level properties like 'description', only looks in the portion before nested objects
|
|
* to avoid matching properties inside nested structures like outputs.
|
|
*/
|
|
function extractStringPropertyFromContent(
|
|
content: string,
|
|
propName: string,
|
|
topLevelOnly = false
|
|
): string | null {
|
|
let searchContent = content
|
|
|
|
// For top-level properties, only search before nested objects like outputs, tools, inputs, subBlocks
|
|
if (topLevelOnly) {
|
|
const nestedObjectPatterns = [
|
|
/\boutputs\s*:\s*\{/,
|
|
/\btools\s*:\s*\{/,
|
|
/\binputs\s*:\s*\{/,
|
|
/\bsubBlocks\s*:\s*\[/,
|
|
/\btriggers\s*:\s*\{/,
|
|
]
|
|
|
|
let cutoffIndex = content.length
|
|
for (const pattern of nestedObjectPatterns) {
|
|
const match = content.match(pattern)
|
|
if (match && match.index !== undefined && match.index < cutoffIndex) {
|
|
cutoffIndex = match.index
|
|
}
|
|
}
|
|
searchContent = content.substring(0, cutoffIndex)
|
|
}
|
|
|
|
const singleQuoteMatch = searchContent.match(new RegExp(`${propName}\\s*:\\s*'([^']*)'`, 'm'))
|
|
if (singleQuoteMatch) return singleQuoteMatch[1]
|
|
|
|
const doubleQuoteMatch = searchContent.match(new RegExp(`${propName}\\s*:\\s*"([^"]*)"`, 'm'))
|
|
if (doubleQuoteMatch) return doubleQuoteMatch[1]
|
|
|
|
const templateMatch = searchContent.match(new RegExp(`${propName}\\s*:\\s*\`([^\`]+)\``, 's'))
|
|
if (templateMatch) {
|
|
let templateContent = templateMatch[1]
|
|
templateContent = templateContent.replace(/\$\{[^}]+\}/g, '')
|
|
templateContent = templateContent.replace(/\s+/g, ' ').trim()
|
|
return templateContent
|
|
}
|
|
|
|
return null
|
|
}
|
|
|
|
/**
|
|
* Extract an enum property value from block content. Maps an `IntegrationType`
|
|
* enum key (e.g. `Communication`) to its slug value (e.g. `'communication'`).
|
|
* Mirrors `apps/sim/blocks/types.ts → IntegrationType` — keep in sync.
|
|
*/
|
|
function extractEnumPropertyFromContent(content: string, propName: string): string | null {
|
|
const match = content.match(new RegExp(`${propName}\\s*:\\s*IntegrationType\\.(\\w+)`))
|
|
if (!match) return null
|
|
const enumKey = match[1]
|
|
const ENUM_MAP: Record<string, string> = {
|
|
AI: 'ai',
|
|
Analytics: 'analytics',
|
|
Commerce: 'commerce',
|
|
Communication: 'communication',
|
|
Databases: 'databases',
|
|
DevOps: 'devops',
|
|
Documents: 'documents',
|
|
Email: 'email',
|
|
HR: 'hr',
|
|
Marketing: 'marketing',
|
|
Observability: 'observability',
|
|
Productivity: 'productivity',
|
|
Sales: 'sales',
|
|
Search: 'search',
|
|
Security: 'security',
|
|
Support: 'support',
|
|
}
|
|
return ENUM_MAP[enumKey] || enumKey.toLowerCase()
|
|
}
|
|
|
|
/**
|
|
* Extract a string array property from block content.
|
|
* Matches patterns like `tags: ['api', 'oauth', 'webhooks']`
|
|
*/
|
|
function extractArrayPropertyFromContent(content: string, propName: string): string[] | null {
|
|
const match = content.match(new RegExp(`${propName}\\s*:\\s*\\[([^\\]]+)\\]`))
|
|
if (!match) return null
|
|
const items = match[1].match(/'([^']+)'|"([^"]+)"/g)
|
|
if (!items) return null
|
|
return items.map((item) => item.replace(/['"]/g, ''))
|
|
}
|
|
|
|
/**
|
|
* Extract `tags` from a `<BlockName>BlockMeta` literal in the source file.
|
|
* Looks for `export const <BlockName>BlockMeta = { ... tags: [...] ... }`
|
|
* at file scope and scans only the body of that literal. Returns null when
|
|
* no matching meta export exists or it contains no `tags` array.
|
|
*
|
|
* During the in-progress migration to per-block meta, some blocks declare
|
|
* `tags` on `BlockConfig` and others on `*BlockMeta`. The caller should
|
|
* try this extractor first and fall back to the `BlockConfig` extractor.
|
|
*/
|
|
function extractTagsFromBlockMeta(fileContent: string, blockName: string): string[] | null {
|
|
const headerRegex = new RegExp(`export\\s+const\\s+${blockName}BlockMeta\\s*(?::[^=]+)?=\\s*\\{`)
|
|
const metaHeaderMatch = fileContent.match(headerRegex)
|
|
if (!metaHeaderMatch || metaHeaderMatch.index === undefined) return null
|
|
const openBracePos = fileContent.indexOf('{', metaHeaderMatch.index)
|
|
if (openBracePos === -1) return null
|
|
const closeBracePos = findMatchingClose(fileContent, openBracePos)
|
|
if (closeBracePos === -1) return null
|
|
const metaBody = fileContent.substring(openBracePos + 1, closeBracePos)
|
|
return extractArrayPropertyFromContent(metaBody, 'tags')
|
|
}
|
|
|
|
/**
|
|
* Extract the component identifier assigned to a block's `icon` property.
|
|
* Most block icons are named `<Service>Icon`, but some reference a generic emcn
|
|
* icon (e.g. `icon: Library`) or use a lowercase identifier (`icon: xIcon`), so
|
|
* this matches any identifier rather than requiring the `Icon` suffix. Bare JS
|
|
* literals are excluded so `icon: undefined` never resolves to a component.
|
|
*/
|
|
function extractIconNameFromContent(content: string): string | null {
|
|
const iconMatch = content.match(
|
|
/(?:^|[\s{,])icon\s*:\s*(?!(?:undefined|null|true|false)\b)([A-Za-z_$][\w$]*)/
|
|
)
|
|
return iconMatch ? iconMatch[1] : null
|
|
}
|
|
|
|
/** Module an icon component is imported from, when the block file declares it. */
|
|
const DEFAULT_ICON_SOURCE = '@/components/icons'
|
|
|
|
/**
|
|
* Resolve the module a block's icon identifier is imported from by scanning the
|
|
* block file's import statements. Falls back to `@/components/icons`, which is
|
|
* where the overwhelming majority of block icons live and which both the sim
|
|
* app and the docs app (via the copied `icons.tsx`) resolve.
|
|
*/
|
|
function resolveIconSource(fileContent: string, iconName: string): string {
|
|
const importRegex = /import\s*(?:type\s*)?\{([^}]*)\}\s*from\s*['"]([^'"]+)['"]/g
|
|
let match: RegExpExecArray | null
|
|
while ((match = importRegex.exec(fileContent)) !== null) {
|
|
const named = match[1].split(',').map((entry) =>
|
|
entry
|
|
.trim()
|
|
.split(/\s+as\s+/)[0]
|
|
.trim()
|
|
)
|
|
if (named.includes(iconName)) return match[2]
|
|
}
|
|
return DEFAULT_ICON_SOURCE
|
|
}
|
|
|
|
/** Biome's configured `formatter.lineWidth` for this repo. */
|
|
const BIOME_LINE_WIDTH = 100
|
|
|
|
/**
|
|
* Quote an icon-map key that is not a bare JS identifier. Trigger providers use
|
|
* hyphenated ids (`google-drive`, `microsoft-teams`) that would otherwise emit
|
|
* as a subtraction expression and break the generated module.
|
|
*/
|
|
function formatIconMapKey(key: string): string {
|
|
return /^[A-Za-z_$][\w$]*$/.test(key) ? key : `'${key}'`
|
|
}
|
|
|
|
/**
|
|
* Render grouped `import { ... } from '...'` statements for every icon
|
|
* referenced by a mapping, so icons sourced outside `@/components/icons`
|
|
* (e.g. `@sim/emcn/icons`) resolve in the generated file. Output is emitted
|
|
* pre-formatted — package specifiers before `@/` aliases, specifiers sorted,
|
|
* and collapsed to one line when it fits — because these files are written
|
|
* verbatim and never passed through Biome.
|
|
*/
|
|
function renderIconImports(iconRefs: IconRef[]): string {
|
|
const bySource = new Map<string, Set<string>>()
|
|
for (const ref of iconRefs) {
|
|
let names = bySource.get(ref.source)
|
|
if (!names) {
|
|
names = new Set()
|
|
bySource.set(ref.source, names)
|
|
}
|
|
names.add(ref.name)
|
|
}
|
|
const isAlias = (source: string) => source.startsWith('@/') || source.startsWith('.')
|
|
return [...bySource.entries()]
|
|
.sort(([a], [b]) => Number(isAlias(a)) - Number(isAlias(b)) || biomeSortCompare(a, b))
|
|
.map(([source, names]) => {
|
|
const sorted = [...names].sort(biomeSortCompare)
|
|
const singleLine = `import { ${sorted.join(', ')} } from '${source}'`
|
|
if (singleLine.length <= BIOME_LINE_WIDTH) return singleLine
|
|
const specifiers = sorted.map((name) => ` ${name},`).join('\n')
|
|
return `import {\n${specifiers}\n} from '${source}'`
|
|
})
|
|
.join('\n')
|
|
}
|
|
|
|
function extractOutputsFromContent(content: string): Record<string, any> {
|
|
const outputsStart = content.search(/outputs\s*:\s*{/)
|
|
if (outputsStart === -1) return {}
|
|
|
|
const openBracePos = content.indexOf('{', outputsStart)
|
|
if (openBracePos === -1) return {}
|
|
|
|
const pos = findMatchingClose(content, openBracePos)
|
|
if (pos === -1) return {}
|
|
|
|
const outputsContent = content.substring(openBracePos + 1, pos - 1).trim()
|
|
const outputs: Record<string, any> = {}
|
|
|
|
const fieldRegex = /(\w+)\s*:\s*{/g
|
|
let match
|
|
const fieldPositions: Array<{ name: string; start: number }> = []
|
|
|
|
while ((match = fieldRegex.exec(outputsContent)) !== null) {
|
|
fieldPositions.push({
|
|
name: match[1],
|
|
start: match.index + match[0].length - 1,
|
|
})
|
|
}
|
|
|
|
fieldPositions.forEach((field) => {
|
|
const endPos = findMatchingClose(outputsContent, field.start)
|
|
|
|
if (endPos !== -1) {
|
|
const fieldContent = outputsContent.substring(field.start + 1, endPos - 1).trim()
|
|
|
|
const typeMatch = fieldContent.match(/type\s*:\s*['"](.*?)['"]/)
|
|
const description = extractDescription(fieldContent)
|
|
|
|
if (typeMatch) {
|
|
outputs[field.name] = {
|
|
type: typeMatch[1],
|
|
description: description || `${field.name} output from the block`,
|
|
}
|
|
}
|
|
}
|
|
})
|
|
|
|
return outputs
|
|
}
|
|
|
|
function extractToolsAccessFromContent(content: string): string[] {
|
|
const accessMatch = content.match(/access\s*:\s*\[\s*([^\]]+)\s*\]/)
|
|
if (!accessMatch) return []
|
|
return [...accessMatch[1].matchAll(/['"]([^'"]+)['"]/g)].map((m) => m[1])
|
|
}
|
|
|
|
/**
|
|
* Get the tool prefix (service name) from a tool name.
|
|
* e.g., "calcom_list_schedules" -> "calcom"
|
|
*/
|
|
function getToolPrefixFromName(toolName: string): string {
|
|
const parts = toolName.split('_')
|
|
|
|
// Try to find a valid tool directory
|
|
for (let i = parts.length - 1; i >= 1; i--) {
|
|
const possiblePrefix = parts.slice(0, i).join('_')
|
|
const toolDirPath = path.join(rootDir, `apps/sim/tools/${possiblePrefix}`)
|
|
|
|
if (fs.existsSync(toolDirPath) && fs.statSync(toolDirPath).isDirectory()) {
|
|
return possiblePrefix
|
|
}
|
|
}
|
|
|
|
return parts[0]
|
|
}
|
|
|
|
/**
|
|
* Resolve a const reference from a types file.
|
|
* Handles nested const references recursively.
|
|
*
|
|
* @param constName - The const name to resolve (e.g., "SCHEDULE_DATA_OUTPUT_PROPERTIES")
|
|
* @param toolPrefix - The tool prefix/service name (e.g., "calcom")
|
|
* @param depth - Recursion depth to prevent infinite loops
|
|
* @returns Resolved properties object or null if not found
|
|
*/
|
|
function resolveConstReference(
|
|
constName: string,
|
|
toolPrefix: string,
|
|
depth = 0
|
|
): Record<string, any> | null {
|
|
// Prevent infinite recursion
|
|
if (depth > 10) {
|
|
console.warn(`Max recursion depth reached resolving const: ${constName}`)
|
|
return null
|
|
}
|
|
|
|
// Check cache first
|
|
const cacheKey = `${toolPrefix}:${constName}`
|
|
if (constResolutionCache.has(cacheKey)) {
|
|
return constResolutionCache.get(cacheKey)!
|
|
}
|
|
|
|
// Read the types file for this tool
|
|
const typesFilePath = path.join(rootDir, `apps/sim/tools/${toolPrefix}/types.ts`)
|
|
if (!fs.existsSync(typesFilePath)) {
|
|
// Try to find const in the tool file itself
|
|
return null
|
|
}
|
|
|
|
const typesContent = fs.readFileSync(typesFilePath, 'utf-8')
|
|
|
|
// Find the const definition
|
|
// Pattern: export const CONST_NAME = { ... } as const
|
|
const constRegex = new RegExp(
|
|
`export\\s+const\\s+${constName}\\s*(?::\\s*[^=]+)?\\s*=\\s*\\{`,
|
|
'g'
|
|
)
|
|
const constMatch = constRegex.exec(typesContent)
|
|
|
|
if (!constMatch) {
|
|
return null
|
|
}
|
|
|
|
// Extract the const content
|
|
const startIndex = constMatch.index + constMatch[0].length - 1
|
|
const endIndex = findMatchingClose(typesContent, startIndex)
|
|
|
|
if (endIndex === -1) {
|
|
return null
|
|
}
|
|
|
|
const constContent = typesContent.substring(startIndex + 1, endIndex - 1).trim()
|
|
|
|
// Check if this const defines a complete output field (has type property)
|
|
// like EVENT_TYPE_OUTPUT = { type: 'object', description: '...', properties: {...} }
|
|
const typeMatch = constContent.match(/^\s*type\s*:\s*['"]([^'"]+)['"]/)
|
|
if (typeMatch) {
|
|
// This is a complete output definition - use parseConstFieldContent
|
|
const result = parseConstFieldContent(constContent, toolPrefix, typesContent, depth + 1)
|
|
if (result) {
|
|
constResolutionCache.set(cacheKey, result)
|
|
}
|
|
return result
|
|
}
|
|
|
|
// Otherwise, this is a properties object - use parseConstProperties
|
|
const properties = parseConstProperties(constContent, toolPrefix, typesContent, depth + 1)
|
|
|
|
// Cache the result
|
|
constResolutionCache.set(cacheKey, properties)
|
|
|
|
return properties
|
|
}
|
|
|
|
/**
|
|
* Parse properties from a const definition, resolving nested const references.
|
|
*/
|
|
function parseConstProperties(
|
|
content: string,
|
|
toolPrefix: string,
|
|
typesContent: string,
|
|
depth: number
|
|
): Record<string, any> {
|
|
const properties: Record<string, any> = {}
|
|
|
|
// First, handle spread operators (e.g., "...COMMENT_OUTPUT_PROPERTIES,")
|
|
const spreadRegex = /\.\.\.([A-Z][A-Z_0-9]+)\s*(?:,|$)/g
|
|
let spreadMatch
|
|
while ((spreadMatch = spreadRegex.exec(content)) !== null) {
|
|
const constName = spreadMatch[1]
|
|
|
|
// Check if at depth 0
|
|
const beforeMatch = content.substring(0, spreadMatch.index)
|
|
const openBraces = (beforeMatch.match(/\{/g) || []).length
|
|
const closeBraces = (beforeMatch.match(/\}/g) || []).length
|
|
if (openBraces !== closeBraces) {
|
|
continue
|
|
}
|
|
|
|
const resolvedConst = resolveConstFromTypesContent(constName, typesContent, toolPrefix, depth)
|
|
if (resolvedConst && typeof resolvedConst === 'object') {
|
|
// Spread all properties from the resolved const
|
|
Object.assign(properties, resolvedConst)
|
|
}
|
|
}
|
|
|
|
// Find all top-level property definitions
|
|
const propRegex = /(\w+)\s*:\s*(?:\{|([A-Z][A-Z_0-9]+)(?:\s*,|\s*$))/g
|
|
let match
|
|
|
|
while ((match = propRegex.exec(content)) !== null) {
|
|
const propName = match[1]
|
|
const constRef = match[2]
|
|
|
|
// Skip 'items' keyword (always a nested structure, never a field name)
|
|
if (propName === 'items') {
|
|
continue
|
|
}
|
|
|
|
// Check if this match is at depth 0 (not inside nested braces)
|
|
const beforeMatch = content.substring(0, match.index)
|
|
const openBraces = (beforeMatch.match(/\{/g) || []).length
|
|
const closeBraces = (beforeMatch.match(/\}/g) || []).length
|
|
if (openBraces !== closeBraces) {
|
|
continue // Skip - this is a nested property
|
|
}
|
|
|
|
// For 'properties' or 'type', check if it's an output field definition vs a keyword
|
|
// Output field definitions have 'type:' inside (e.g., { type: 'string', description: '...' })
|
|
if ((propName === 'properties' || propName === 'type') && !constRef) {
|
|
// Peek at what's inside the braces
|
|
const startPos = match.index + match[0].length - 1
|
|
const endPos = findMatchingClose(content, startPos)
|
|
if (endPos !== -1) {
|
|
const propContent = content.substring(startPos + 1, endPos - 1).trim()
|
|
// If it starts with 'type:', it's an output field definition - process it
|
|
if (propContent.match(/^\s*type\s*:/)) {
|
|
const parsedProp = parseConstFieldContent(propContent, toolPrefix, typesContent, depth)
|
|
if (parsedProp) {
|
|
properties[propName] = parsedProp
|
|
}
|
|
}
|
|
// Otherwise, it's a keyword usage (nested properties block or type specifier) - skip it
|
|
}
|
|
continue
|
|
}
|
|
|
|
if (constRef) {
|
|
// This property references a const (e.g., "attendees: ATTENDEES_OUTPUT")
|
|
const resolvedConst = resolveConstFromTypesContent(constRef, typesContent, toolPrefix, depth)
|
|
if (resolvedConst) {
|
|
properties[propName] = resolvedConst
|
|
}
|
|
} else {
|
|
// This property has inline definition
|
|
const startPos = match.index + match[0].length - 1
|
|
const endPos = findMatchingClose(content, startPos)
|
|
|
|
if (endPos !== -1) {
|
|
const propContent = content.substring(startPos + 1, endPos - 1).trim()
|
|
const parsedProp = parseConstFieldContent(propContent, toolPrefix, typesContent, depth)
|
|
if (parsedProp) {
|
|
properties[propName] = parsedProp
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
return properties
|
|
}
|
|
|
|
/**
|
|
* Resolve a const from the types content (for nested references within the same file).
|
|
*/
|
|
function resolveConstFromTypesContent(
|
|
constName: string,
|
|
typesContent: string,
|
|
toolPrefix: string,
|
|
depth: number
|
|
): Record<string, any> | null {
|
|
if (depth > 10) return null
|
|
|
|
// Check cache
|
|
const cacheKey = `${toolPrefix}:${constName}`
|
|
if (constResolutionCache.has(cacheKey)) {
|
|
return constResolutionCache.get(cacheKey)!
|
|
}
|
|
|
|
// Find the const definition in typesContent
|
|
const constRegex = new RegExp(
|
|
`export\\s+const\\s+${constName}\\s*(?::\\s*[^=]+)?\\s*=\\s*\\{`,
|
|
'g'
|
|
)
|
|
const constMatch = constRegex.exec(typesContent)
|
|
|
|
if (!constMatch) {
|
|
return null
|
|
}
|
|
|
|
const startIndex = constMatch.index + constMatch[0].length - 1
|
|
const endIndex = findMatchingClose(typesContent, startIndex)
|
|
|
|
if (endIndex === -1) return null
|
|
|
|
const constContent = typesContent.substring(startIndex + 1, endIndex - 1).trim()
|
|
|
|
// Check if this const defines a complete output field (has type property)
|
|
const typeMatch = constContent.match(/^\s*type\s*:\s*['"]([^'"]+)['"]/)
|
|
if (typeMatch) {
|
|
// This is a complete output definition (like ATTENDEES_OUTPUT)
|
|
const result = parseConstFieldContent(constContent, toolPrefix, typesContent, depth)
|
|
if (result) {
|
|
constResolutionCache.set(cacheKey, result)
|
|
}
|
|
return result
|
|
}
|
|
|
|
// This is a properties object (like ATTENDEE_OUTPUT_PROPERTIES)
|
|
const properties = parseConstProperties(constContent, toolPrefix, typesContent, depth + 1)
|
|
constResolutionCache.set(cacheKey, properties)
|
|
return properties
|
|
}
|
|
|
|
/**
|
|
* Parse a field content from a const, resolving nested const references.
|
|
*/
|
|
/**
|
|
* Extract description from field content, handling quoted strings properly.
|
|
* Handles single quotes, double quotes, and backticks, preserving internal quotes.
|
|
*/
|
|
function extractDescription(fieldContent: string): string | null {
|
|
// Walk through all `description:` matches and return the first one at depth 0.
|
|
// This prevents accidentally picking up `description:` keys inside nested child objects.
|
|
const descRegex = /description\s*:\s*('([^']*)'|"([^"]*)"|`([^`]*)`)/g
|
|
let m: RegExpExecArray | null
|
|
while ((m = descRegex.exec(fieldContent)) !== null) {
|
|
if (isAtDepthZero(fieldContent, m.index)) {
|
|
return m[2] ?? m[3] ?? m[4] ?? null
|
|
}
|
|
}
|
|
return null
|
|
}
|
|
|
|
function parseConstFieldContent(
|
|
fieldContent: string,
|
|
toolPrefix: string,
|
|
typesContent: string,
|
|
depth: number
|
|
): any {
|
|
const typeMatch = fieldContent.match(/type\s*:\s*['"]([^'"]+)['"]/)
|
|
const description = extractDescription(fieldContent)
|
|
|
|
if (!typeMatch) return null
|
|
|
|
const fieldType = typeMatch[1]
|
|
|
|
const result: any = {
|
|
type: fieldType,
|
|
description: description || '',
|
|
}
|
|
|
|
// Check for properties - either inline or const reference
|
|
if (fieldType === 'object' || fieldType === 'json') {
|
|
// Check for const reference first
|
|
const propsConstMatch = fieldContent.match(/properties\s*:\s*([A-Z][A-Z_0-9]+)/)
|
|
if (propsConstMatch) {
|
|
const resolvedProps = resolveConstFromTypesContent(
|
|
propsConstMatch[1],
|
|
typesContent,
|
|
toolPrefix,
|
|
depth + 1
|
|
)
|
|
if (resolvedProps) {
|
|
result.properties = resolvedProps
|
|
}
|
|
} else {
|
|
// Check for inline properties
|
|
const propertiesStart = fieldContent.search(/properties\s*:\s*\{/)
|
|
if (propertiesStart !== -1) {
|
|
const braceStart = fieldContent.indexOf('{', propertiesStart)
|
|
const braceEnd = findMatchingClose(fieldContent, braceStart)
|
|
|
|
if (braceEnd !== -1) {
|
|
const propertiesContent = fieldContent.substring(braceStart + 1, braceEnd - 1).trim()
|
|
result.properties = parseConstProperties(
|
|
propertiesContent,
|
|
toolPrefix,
|
|
typesContent,
|
|
depth + 1
|
|
)
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
// Check for items (arrays)
|
|
const itemsConstMatch = fieldContent.match(/items\s*:\s*([A-Z][A-Z_0-9]+)/)
|
|
if (itemsConstMatch) {
|
|
const resolvedItems = resolveConstFromTypesContent(
|
|
itemsConstMatch[1],
|
|
typesContent,
|
|
toolPrefix,
|
|
depth + 1
|
|
)
|
|
if (resolvedItems) {
|
|
result.items = resolvedItems
|
|
}
|
|
} else {
|
|
const itemsStart = fieldContent.search(/items\s*:\s*\{/)
|
|
if (itemsStart !== -1) {
|
|
const braceStart = fieldContent.indexOf('{', itemsStart)
|
|
const braceEnd = findMatchingClose(fieldContent, braceStart)
|
|
|
|
if (braceEnd !== -1) {
|
|
const itemsContent = fieldContent.substring(braceStart + 1, braceEnd - 1).trim()
|
|
const itemsType = itemsContent.match(/type\s*:\s*['"]([^'"]+)['"]/)
|
|
const itemsDesc = extractDescription(itemsContent)
|
|
|
|
result.items = {
|
|
type: itemsType ? itemsType[1] : 'object',
|
|
description: itemsDesc || '',
|
|
}
|
|
|
|
// Check for properties in items - either inline or const reference
|
|
const itemsPropsConstMatch = itemsContent.match(/properties\s*:\s*([A-Z][A-Z_0-9]+)/)
|
|
if (itemsPropsConstMatch) {
|
|
const resolvedProps = resolveConstFromTypesContent(
|
|
itemsPropsConstMatch[1],
|
|
typesContent,
|
|
toolPrefix,
|
|
depth + 1
|
|
)
|
|
if (resolvedProps) {
|
|
result.items.properties = resolvedProps
|
|
}
|
|
} else {
|
|
const itemsPropsStart = itemsContent.search(/properties\s*:\s*\{/)
|
|
if (itemsPropsStart !== -1) {
|
|
const propsBraceStart = itemsContent.indexOf('{', itemsPropsStart)
|
|
let propsBraceCount = 1
|
|
let propsBraceEnd = propsBraceStart + 1
|
|
|
|
while (propsBraceEnd < itemsContent.length && propsBraceCount > 0) {
|
|
if (itemsContent[propsBraceEnd] === '{') propsBraceCount++
|
|
else if (itemsContent[propsBraceEnd] === '}') propsBraceCount--
|
|
propsBraceEnd++
|
|
}
|
|
|
|
if (propsBraceCount === 0) {
|
|
const itemsPropsContent = itemsContent
|
|
.substring(propsBraceStart + 1, propsBraceEnd - 1)
|
|
.trim()
|
|
result.items.properties = parseConstProperties(
|
|
itemsPropsContent,
|
|
toolPrefix,
|
|
typesContent,
|
|
depth + 1
|
|
)
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
return result
|
|
}
|
|
|
|
/**
|
|
* Extract outputs from a tool content block by trying:
|
|
* 1. Const reference (e.g., `outputs: GIT_REF_OUTPUT_PROPERTIES,`)
|
|
* 2. Inline object (e.g., `outputs: { id: { type: 'string', ... } }`)
|
|
*/
|
|
function extractOutputsFromToolContent(content: string, toolPrefix: string): Record<string, any> {
|
|
const constMatch = content.match(/(?<![a-zA-Z_])outputs\s*:\s*([A-Z][A-Z_0-9]+)\s*(?:,|\}|$)/)
|
|
if (constMatch) {
|
|
const resolved = resolveConstReference(constMatch[1], toolPrefix)
|
|
if (resolved && typeof resolved === 'object') {
|
|
return resolved
|
|
}
|
|
}
|
|
|
|
const outputsStart = content.search(/(?<![a-zA-Z_])outputs\s*:\s*{/)
|
|
if (outputsStart !== -1) {
|
|
const openBracePos = content.indexOf('{', outputsStart)
|
|
if (openBracePos !== -1) {
|
|
const closePos = findMatchingClose(content, openBracePos)
|
|
if (closePos !== -1) {
|
|
const outputsContent = content.substring(openBracePos + 1, closePos - 1).trim()
|
|
return parseToolOutputsField(outputsContent, toolPrefix)
|
|
}
|
|
}
|
|
}
|
|
|
|
return {}
|
|
}
|
|
|
|
/**
|
|
* Resolves the module a tool delegates its config to, for tools built by a
|
|
* factory instead of an inline object literal:
|
|
*
|
|
* export const embeddingsOpenAITool = createEmbeddingTool({ id, provider, ... })
|
|
*
|
|
* The `params` block lives in the factory's module, so a file-local search finds
|
|
* nothing and the docs page renders an empty Input table. This follows the
|
|
* factory's import the way {@link extractSpreadBase} follows a same-file spread.
|
|
* Returns null when the tool declares its own config, which is the common case.
|
|
*/
|
|
function resolveFactorySource(fileContent: string, toolFilePath: string, rootDir: string): string {
|
|
const factoryCall = fileContent.match(/=\s*(create\w+)\s*\(\s*\{/)
|
|
if (!factoryCall) return ''
|
|
|
|
const factoryName = factoryCall[1]
|
|
const importMatch = fileContent.match(
|
|
new RegExp(`import\\s*\\{[^}]*\\b${factoryName}\\b[^}]*\\}\\s*from\\s*['"]([^'"]+)['"]`)
|
|
)
|
|
if (!importMatch) return ''
|
|
|
|
const specifier = importMatch[1]
|
|
const resolved = specifier.startsWith('@/')
|
|
? path.join(rootDir, 'apps/sim', specifier.slice(2))
|
|
: path.resolve(path.dirname(toolFilePath), specifier)
|
|
|
|
for (const candidate of [`${resolved}.ts`, path.join(resolved, 'index.ts')]) {
|
|
if (fs.existsSync(candidate)) return fs.readFileSync(candidate, 'utf-8')
|
|
}
|
|
return ''
|
|
}
|
|
|
|
function extractToolInfo(
|
|
toolName: string,
|
|
fileContent: string,
|
|
factorySource = ''
|
|
): {
|
|
description: string
|
|
params: Array<{ name: string; type: string; required: boolean; description: string }>
|
|
outputs: Record<string, any>
|
|
} | null {
|
|
try {
|
|
// First, try to find the specific tool definition by its ID
|
|
// Look for: id: 'toolName' or id: "toolName"
|
|
const toolIdRegex = new RegExp(`id:\\s*['"]${toolName}['"]`)
|
|
const toolIdMatch = fileContent.match(toolIdRegex)
|
|
|
|
let toolContent = fileContent
|
|
if (toolIdMatch && toolIdMatch.index !== undefined) {
|
|
// Find the tool definition block that contains this ID
|
|
// Search backwards for 'export const' or start of object
|
|
const beforeId = fileContent.substring(0, toolIdMatch.index)
|
|
const exportMatch = beforeId.match(/export\s+const\s+\w+[^=]*=\s*\{[\s\S]*$/)
|
|
|
|
if (exportMatch && exportMatch.index !== undefined) {
|
|
const startIndex = exportMatch.index + exportMatch[0].length - 1
|
|
const endIndex = findMatchingClose(fileContent, startIndex)
|
|
|
|
if (endIndex !== -1) {
|
|
toolContent = fileContent.substring(startIndex, endIndex)
|
|
}
|
|
}
|
|
}
|
|
|
|
// Prefer the params block scoped to this specific tool so that files
|
|
// defining multiple tools (e.g. file_compress + file_decompress in
|
|
// compress.ts) don't all inherit the first tool's params. Fall back to the
|
|
// full file for tools that inherit params via spread from a base object.
|
|
const toolConfigRegex =
|
|
/params\s*:\s*{([\s\S]*?)},?\s*(?:outputs|oauth|hosting|request|directExecution|postProcess|transformResponse)\s*:/
|
|
const toolConfigMatch =
|
|
toolContent.match(toolConfigRegex) ??
|
|
fileContent.match(toolConfigRegex) ??
|
|
factorySource.match(toolConfigRegex)
|
|
|
|
// Description should come from the specific tool block if found
|
|
// Only search before nested objects (params, outputs, request, etc.) to avoid matching
|
|
// descriptions inside outputs or params
|
|
let descriptionSearchContent = toolContent
|
|
const nestedObjectPatterns = [
|
|
/\bparams\s*:\s*[{]/,
|
|
/\boutputs\s*:\s*\{/,
|
|
/\brequest\s*:\s*\{/,
|
|
/\boauth\s*:\s*\{/,
|
|
/\btransformResponse\s*:/,
|
|
]
|
|
let cutoffIndex = toolContent.length
|
|
for (const pattern of nestedObjectPatterns) {
|
|
const match = toolContent.match(pattern)
|
|
if (match && match.index !== undefined && match.index < cutoffIndex) {
|
|
cutoffIndex = match.index
|
|
}
|
|
}
|
|
descriptionSearchContent = toolContent.substring(0, cutoffIndex)
|
|
|
|
// Match against the actual opening quote so apostrophes inside a double-quoted
|
|
// description (e.g. "Find someone's email") are not treated as the closing quote.
|
|
const descriptionRegex = /description\s*:\s*(?:'([^']*)'|"([^"]*)"|`([^`]*)`)/
|
|
let descriptionMatch = descriptionSearchContent.match(descriptionRegex)
|
|
|
|
// If description isn't found as a literal (might be inherited like description: baseTool.description),
|
|
// try to find the referenced tool's description
|
|
if (!descriptionMatch) {
|
|
const inheritedDescMatch = descriptionSearchContent.match(
|
|
/description\s*:\s*(\w+)Tool\.description/
|
|
)
|
|
if (inheritedDescMatch) {
|
|
const baseTool = inheritedDescMatch[1]
|
|
// Try to find the base tool's description in the file
|
|
const baseToolDescRegex = new RegExp(
|
|
`export\\s+const\\s+${baseTool}Tool[^{]*\\{[\\s\\S]*?description\\s*:\\s*(?:'([^']+)'|"([^"]+)"|\`([^\`]+)\`)`,
|
|
'i'
|
|
)
|
|
const baseToolMatch = fileContent.match(baseToolDescRegex)
|
|
if (baseToolMatch) {
|
|
descriptionMatch = baseToolMatch
|
|
}
|
|
}
|
|
}
|
|
|
|
const description = descriptionMatch
|
|
? (descriptionMatch[1] ??
|
|
descriptionMatch[2] ??
|
|
descriptionMatch[3] ??
|
|
'No description available')
|
|
: 'No description available'
|
|
|
|
const params: Array<{ name: string; type: string; required: boolean; description: string }> = []
|
|
|
|
if (toolConfigMatch) {
|
|
const paramsContent = toolConfigMatch[1]
|
|
|
|
const paramBlocksRegex = /(\w+)\s*:\s*{/g
|
|
let paramMatch
|
|
const paramPositions: Array<{ name: string; start: number; content: string }> = []
|
|
|
|
/**
|
|
* Checks if a position in the string is inside a quoted string.
|
|
* This prevents matching patterns like "Example: {" inside description strings.
|
|
*/
|
|
const isInsideString = (content: string, position: number): boolean => {
|
|
let inSingleQuote = false
|
|
let inDoubleQuote = false
|
|
let inBacktick = false
|
|
|
|
for (let i = 0; i < position; i++) {
|
|
const char = content[i]
|
|
const prevChar = i > 0 ? content[i - 1] : ''
|
|
|
|
// Skip escaped quotes
|
|
if (prevChar === '\\') continue
|
|
|
|
if (char === "'" && !inDoubleQuote && !inBacktick) {
|
|
inSingleQuote = !inSingleQuote
|
|
} else if (char === '"' && !inSingleQuote && !inBacktick) {
|
|
inDoubleQuote = !inDoubleQuote
|
|
} else if (char === '`' && !inSingleQuote && !inDoubleQuote) {
|
|
inBacktick = !inBacktick
|
|
}
|
|
}
|
|
|
|
return inSingleQuote || inDoubleQuote || inBacktick
|
|
}
|
|
|
|
while ((paramMatch = paramBlocksRegex.exec(paramsContent)) !== null) {
|
|
const paramName = paramMatch[1]
|
|
const startPos = paramMatch.index + paramMatch[0].length - 1
|
|
|
|
// Skip matches that are inside string literals (e.g., "Example: {" in descriptions)
|
|
if (isInsideString(paramsContent, paramMatch.index)) {
|
|
continue
|
|
}
|
|
|
|
const endPos = findMatchingClose(paramsContent, startPos)
|
|
|
|
if (endPos !== -1) {
|
|
const paramBlock = paramsContent.substring(startPos + 1, endPos - 1).trim()
|
|
paramPositions.push({ name: paramName, start: startPos, content: paramBlock })
|
|
// Resume scanning after this param's block so nested descriptors
|
|
// (e.g. an array param's `items: {...}`) are not parsed as params.
|
|
paramBlocksRegex.lastIndex = endPos
|
|
}
|
|
}
|
|
|
|
for (const param of paramPositions) {
|
|
const paramName = param.name
|
|
const paramBlock = param.content
|
|
|
|
if (paramName === 'accessToken' || paramName === 'params' || paramName === 'tools') {
|
|
continue
|
|
}
|
|
|
|
const typeMatch = paramBlock.match(/type\s*:\s*['"]([^'"]+)['"]/)
|
|
const requiredMatch = paramBlock.match(/required\s*:\s*(true|false)/)
|
|
|
|
let descriptionMatch = paramBlock.match(/description\s*:\s*'(.*?)'(?=\s*[,}])/s)
|
|
if (!descriptionMatch) {
|
|
descriptionMatch = paramBlock.match(/description\s*:\s*"(.*?)"(?=\s*[,}])/s)
|
|
}
|
|
if (!descriptionMatch) {
|
|
descriptionMatch = paramBlock.match(/description\s*:\s*`([^`]+)`/s)
|
|
}
|
|
if (!descriptionMatch) {
|
|
descriptionMatch = paramBlock.match(
|
|
/description\s*:\s*['"]([^'"]*(?:\n[^'"]*)*?)['"](?=\s*[,}])/s
|
|
)
|
|
}
|
|
|
|
params.push({
|
|
name: paramName,
|
|
type: typeMatch ? typeMatch[1] : 'string',
|
|
required: requiredMatch ? requiredMatch[1] === 'true' : false,
|
|
description: descriptionMatch ? descriptionMatch[1] : 'No description',
|
|
})
|
|
}
|
|
}
|
|
|
|
// Get the tool prefix for resolving const references
|
|
const toolPrefix = getToolPrefixFromName(toolName)
|
|
|
|
let outputs = extractOutputsFromToolContent(toolContent, toolPrefix)
|
|
|
|
// If no outputs found, check for spread inheritance (e.g., "...extendParserTool")
|
|
// toolContent may be narrowed past the spread line, so reconstruct the full block
|
|
if (Object.keys(outputs).length === 0) {
|
|
let fullToolBlock = toolContent
|
|
if (toolIdMatch && toolIdMatch.index !== undefined) {
|
|
const beforeId = fileContent.substring(0, toolIdMatch.index)
|
|
const exportRegex = /export\s+const\s+\w+[^=]*=\s*\{/g
|
|
let lastExportMatch: RegExpExecArray | null = null
|
|
let m: RegExpExecArray | null = null
|
|
while ((m = exportRegex.exec(beforeId)) !== null) {
|
|
lastExportMatch = m
|
|
}
|
|
if (lastExportMatch && lastExportMatch.index !== undefined) {
|
|
const bracePos = lastExportMatch.index + lastExportMatch[0].length - 1
|
|
const ep = findMatchingClose(fileContent, bracePos)
|
|
if (ep !== -1) {
|
|
fullToolBlock = fileContent.substring(bracePos, ep)
|
|
}
|
|
}
|
|
}
|
|
const spreadMatch = fullToolBlock.match(/\.\.\.(\w+(?:Tool|Base)\w*)/)
|
|
if (spreadMatch) {
|
|
const baseVarName = spreadMatch[1]
|
|
const baseToolRegex = new RegExp(
|
|
`export\\s+const\\s+${baseVarName}(?=[^a-zA-Z0-9_]|$)[^=]*=\\s*\\{`
|
|
)
|
|
const baseToolMatch = fileContent.match(baseToolRegex)
|
|
if (baseToolMatch && baseToolMatch.index !== undefined) {
|
|
const baseStart = baseToolMatch.index + baseToolMatch[0].length - 1
|
|
const endIdx = findMatchingClose(fileContent, baseStart)
|
|
if (endIdx !== -1) {
|
|
const baseToolContent = fileContent.substring(baseStart, endIdx)
|
|
outputs = extractOutputsFromToolContent(baseToolContent, toolPrefix)
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
return {
|
|
description,
|
|
params,
|
|
outputs,
|
|
}
|
|
} catch (error) {
|
|
console.error(`Error extracting info for tool ${toolName}:`, error)
|
|
return null
|
|
}
|
|
}
|
|
|
|
function formatOutputStructure(outputs: Record<string, any>, indentLevel = 0): string {
|
|
let result = ''
|
|
|
|
for (const [key, output] of Object.entries(outputs)) {
|
|
let type = 'unknown'
|
|
let description = `${key} output from the tool`
|
|
|
|
if (typeof output === 'object' && output !== null) {
|
|
if (output.type) {
|
|
type = output.type
|
|
}
|
|
|
|
if (output.description) {
|
|
description = output.description
|
|
}
|
|
}
|
|
|
|
const escapedDescription = description
|
|
.replace(/\|/g, '\\|')
|
|
.replace(/\{/g, '\\{')
|
|
.replace(/\}/g, '\\}')
|
|
.replace(/\(/g, '\\(')
|
|
.replace(/\)/g, '\\)')
|
|
.replace(/\[/g, '\\[')
|
|
.replace(/\]/g, '\\]')
|
|
.replace(/</g, '<')
|
|
.replace(/>/g, '>')
|
|
|
|
// Build prefix based on indent level - each level adds 2 spaces before the arrow
|
|
let prefix = ''
|
|
if (indentLevel > 0) {
|
|
const spaces = ' '.repeat(indentLevel)
|
|
prefix = `${spaces}↳ `
|
|
}
|
|
|
|
if (typeof output === 'object' && output !== null && output.type === 'array') {
|
|
result += `| ${prefix}\`${key}\` | ${type} | ${escapedDescription} |\n`
|
|
|
|
if (output.items?.properties) {
|
|
const arrayItemsResult = formatOutputStructure(output.items.properties, indentLevel + 1)
|
|
result += arrayItemsResult
|
|
}
|
|
} else if (
|
|
typeof output === 'object' &&
|
|
output !== null &&
|
|
output.properties &&
|
|
(output.type === 'object' || output.type === 'json')
|
|
) {
|
|
result += `| ${prefix}\`${key}\` | ${type} | ${escapedDescription} |\n`
|
|
|
|
const nestedResult = formatOutputStructure(output.properties, indentLevel + 1)
|
|
result += nestedResult
|
|
} else {
|
|
result += `| ${prefix}\`${key}\` | ${type} | ${escapedDescription} |\n`
|
|
}
|
|
}
|
|
|
|
return result
|
|
}
|
|
|
|
function parseToolOutputsField(outputsContent: string, toolPrefix?: string): Record<string, any> {
|
|
const outputs: Record<string, any> = {}
|
|
|
|
// First, handle top-level const references
|
|
// Patterns: "data: BOOKING_DATA_OUTPUT_PROPERTIES" or "pagination: PAGINATION_OUTPUT"
|
|
if (toolPrefix) {
|
|
// Pattern 1: Direct const reference
|
|
const constRefRegex = /(\w+)\s*:\s*([A-Z][A-Z_0-9]+)\s*(?:,|$)/g
|
|
let constMatch
|
|
while ((constMatch = constRefRegex.exec(outputsContent)) !== null) {
|
|
const propName = constMatch[1]
|
|
const constName = constMatch[2]
|
|
|
|
// Check if at depth 0
|
|
const beforeMatch = outputsContent.substring(0, constMatch.index)
|
|
const openBraces = (beforeMatch.match(/\{/g) || []).length
|
|
const closeBraces = (beforeMatch.match(/\}/g) || []).length
|
|
if (openBraces !== closeBraces) {
|
|
continue
|
|
}
|
|
|
|
const resolvedConst = resolveConstReference(constName, toolPrefix)
|
|
if (resolvedConst) {
|
|
outputs[propName] = resolvedConst
|
|
}
|
|
}
|
|
|
|
// Pattern 2: Property access on const (e.g., "status: BOOKING_DATA_OUTPUT_PROPERTIES.status,")
|
|
const propAccessRegex = /(\w+)\s*:\s*([A-Z][A-Z_0-9]+)\.(\w+)\s*(?:,|$)/g
|
|
let propAccessMatch
|
|
while ((propAccessMatch = propAccessRegex.exec(outputsContent)) !== null) {
|
|
const propName = propAccessMatch[1]
|
|
const constName = propAccessMatch[2]
|
|
const accessedProp = propAccessMatch[3]
|
|
|
|
// Skip if already resolved
|
|
if (outputs[propName]) {
|
|
continue
|
|
}
|
|
|
|
// Check if at depth 0
|
|
const beforeMatch = outputsContent.substring(0, propAccessMatch.index)
|
|
const openBraces = (beforeMatch.match(/\{/g) || []).length
|
|
const closeBraces = (beforeMatch.match(/\}/g) || []).length
|
|
if (openBraces !== closeBraces) {
|
|
continue
|
|
}
|
|
|
|
const resolvedConst = resolveConstReference(constName, toolPrefix)
|
|
if (resolvedConst?.[accessedProp]) {
|
|
outputs[propName] = resolvedConst[accessedProp]
|
|
}
|
|
}
|
|
|
|
// Pattern 3: Spread operator (e.g., "...COMMENT_OUTPUT_PROPERTIES,")
|
|
const spreadRegex = /\.\.\.([A-Z][A-Z_0-9]+)\s*(?:,|$)/g
|
|
let spreadMatch
|
|
while ((spreadMatch = spreadRegex.exec(outputsContent)) !== null) {
|
|
const constName = spreadMatch[1]
|
|
|
|
// Check if at depth 0 (not inside nested braces)
|
|
const beforeMatch = outputsContent.substring(0, spreadMatch.index)
|
|
const openBraces = (beforeMatch.match(/\{/g) || []).length
|
|
const closeBraces = (beforeMatch.match(/\}/g) || []).length
|
|
if (openBraces !== closeBraces) {
|
|
continue
|
|
}
|
|
|
|
const resolvedConst = resolveConstReference(constName, toolPrefix)
|
|
if (resolvedConst && typeof resolvedConst === 'object') {
|
|
// Spread all properties from the resolved const
|
|
Object.assign(outputs, resolvedConst)
|
|
}
|
|
}
|
|
}
|
|
|
|
const braces: Array<{ type: 'open' | 'close'; pos: number; level: number }> = []
|
|
for (let i = 0; i < outputsContent.length; i++) {
|
|
if (outputsContent[i] === '{') {
|
|
braces.push({ type: 'open', pos: i, level: 0 })
|
|
} else if (outputsContent[i] === '}') {
|
|
braces.push({ type: 'close', pos: i, level: 0 })
|
|
}
|
|
}
|
|
|
|
let currentLevel = 0
|
|
for (const brace of braces) {
|
|
if (brace.type === 'open') {
|
|
brace.level = currentLevel
|
|
currentLevel++
|
|
} else {
|
|
currentLevel--
|
|
brace.level = currentLevel
|
|
}
|
|
}
|
|
|
|
const fieldStartRegex = /(\w+)\s*:\s*{/g
|
|
let match
|
|
const fieldPositions: Array<{ name: string; start: number; end: number; level: number }> = []
|
|
|
|
while ((match = fieldStartRegex.exec(outputsContent)) !== null) {
|
|
const fieldName = match[1]
|
|
const bracePos = match.index + match[0].length - 1
|
|
|
|
// Skip if already resolved as const reference
|
|
if (outputs[fieldName]) {
|
|
continue
|
|
}
|
|
|
|
const openBrace = braces.find((b) => b.type === 'open' && b.pos === bracePos)
|
|
if (openBrace) {
|
|
const endPos = findMatchingClose(outputsContent, bracePos)
|
|
if (endPos !== -1) {
|
|
fieldPositions.push({
|
|
name: fieldName,
|
|
start: bracePos,
|
|
end: endPos,
|
|
level: openBrace.level,
|
|
})
|
|
}
|
|
}
|
|
}
|
|
|
|
const topLevelFields = fieldPositions.filter((f) => f.level === 0)
|
|
|
|
topLevelFields.forEach((field) => {
|
|
const fieldContent = outputsContent.substring(field.start + 1, field.end - 1).trim()
|
|
|
|
const parsedField = parseFieldContent(fieldContent, toolPrefix)
|
|
if (parsedField) {
|
|
outputs[field.name] = parsedField
|
|
}
|
|
})
|
|
|
|
return outputs
|
|
}
|
|
|
|
/**
|
|
* Returns true if the regex match at `matchIndex` within `content` is at brace depth 0.
|
|
* Used to distinguish top-level keys from keys nested inside child objects.
|
|
*/
|
|
function isAtDepthZero(content: string, matchIndex: number): boolean {
|
|
let depth = 0
|
|
for (let i = 0; i < matchIndex; i++) {
|
|
if (content[i] === '{') depth++
|
|
else if (content[i] === '}') depth--
|
|
}
|
|
return depth === 0
|
|
}
|
|
|
|
function parseFieldContent(fieldContent: string, toolPrefix?: string): any {
|
|
// Only match `type:` that is at the top level of fieldContent (depth 0).
|
|
// Child objects like `title: { type: 'string', ... }` also contain `type:` but at depth 1.
|
|
const typeRegex = /type\s*:\s*['"]([^'"]+)['"]/g
|
|
let typeMatch: RegExpExecArray | null = null
|
|
let m: RegExpExecArray | null
|
|
while ((m = typeRegex.exec(fieldContent)) !== null) {
|
|
if (isAtDepthZero(fieldContent, m.index)) {
|
|
typeMatch = m
|
|
break
|
|
}
|
|
}
|
|
const description = extractDescription(fieldContent)
|
|
|
|
// Check for spread operator at the start of field content (e.g., ...SUBSCRIPTION_OUTPUT)
|
|
// This pattern is used when a field spreads a complete output definition and optionally overrides properties
|
|
const spreadMatch = fieldContent.match(/^\s*\.\.\.([A-Z][A-Z_0-9]+)\s*,/)
|
|
if (spreadMatch && toolPrefix && !typeMatch) {
|
|
const constName = spreadMatch[1]
|
|
const resolvedConst = resolveConstReference(constName, toolPrefix)
|
|
if (resolvedConst && typeof resolvedConst === 'object') {
|
|
// Start with the resolved const and override with inline properties
|
|
const result: any = { ...resolvedConst }
|
|
// Override description if provided inline
|
|
if (description) {
|
|
result.description = description
|
|
}
|
|
return result
|
|
}
|
|
}
|
|
|
|
if (!typeMatch) {
|
|
// No top-level `type` key — check if the content contains named child fields that each
|
|
// have their own `type` property. This is the "implicit object" pattern used in trigger
|
|
// outputs (e.g., Cal.com's `payload`, Linear's `data`).
|
|
const properties = parsePropertiesContent(fieldContent, toolPrefix)
|
|
if (Object.keys(properties).length > 0) {
|
|
return {
|
|
type: 'object',
|
|
description: description || '',
|
|
properties,
|
|
}
|
|
}
|
|
return null
|
|
}
|
|
|
|
const fieldType = typeMatch[1]
|
|
|
|
const result: any = {
|
|
type: fieldType,
|
|
description: description || '',
|
|
}
|
|
|
|
if (fieldType === 'object' || fieldType === 'json') {
|
|
// Check for const reference first (e.g., properties: SCHEDULE_DATA_OUTPUT_PROPERTIES)
|
|
const propsConstMatch = fieldContent.match(/properties\s*:\s*([A-Z][A-Z_0-9]+)/)
|
|
if (propsConstMatch && toolPrefix) {
|
|
const resolvedProps = resolveConstReference(propsConstMatch[1], toolPrefix)
|
|
if (resolvedProps) {
|
|
result.properties = resolvedProps
|
|
}
|
|
} else {
|
|
// Check for inline properties
|
|
const propertiesRegex = /properties\s*:\s*{/
|
|
const propertiesStart = fieldContent.search(propertiesRegex)
|
|
|
|
if (propertiesStart !== -1) {
|
|
const braceStart = fieldContent.indexOf('{', propertiesStart)
|
|
const braceEnd = findMatchingClose(fieldContent, braceStart)
|
|
|
|
if (braceEnd !== -1) {
|
|
const propertiesContent = fieldContent.substring(braceStart + 1, braceEnd - 1).trim()
|
|
result.properties = parsePropertiesContent(propertiesContent, toolPrefix)
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
// Check for items const reference (e.g., items: ATTENDEES_OUTPUT)
|
|
const itemsConstMatch = fieldContent.match(/items\s*:\s*([A-Z][A-Z_0-9]+)/)
|
|
if (itemsConstMatch && toolPrefix) {
|
|
const resolvedItems = resolveConstReference(itemsConstMatch[1], toolPrefix)
|
|
if (resolvedItems) {
|
|
result.items = resolvedItems
|
|
}
|
|
} else {
|
|
const itemsRegex = /items\s*:\s*{/
|
|
const itemsStart = fieldContent.search(itemsRegex)
|
|
|
|
if (itemsStart !== -1) {
|
|
const braceStart = fieldContent.indexOf('{', itemsStart)
|
|
const braceEnd = findMatchingClose(fieldContent, braceStart)
|
|
|
|
if (braceEnd !== -1) {
|
|
const itemsContent = fieldContent.substring(braceStart + 1, braceEnd - 1).trim()
|
|
const itemsType = itemsContent.match(/type\s*:\s*['"]([^'"]+)['"]/)
|
|
|
|
// Check for inline properties FIRST (properties: {), then const reference
|
|
const propertiesInlineStart = itemsContent.search(/properties\s*:\s*{/)
|
|
// Only match const reference if it's at the TOP level (before any {)
|
|
const itemsPropsConstMatch =
|
|
propertiesInlineStart === -1
|
|
? itemsContent.match(/properties\s*:\s*([A-Z][A-Z_0-9]+)/)
|
|
: null
|
|
const searchContent =
|
|
propertiesInlineStart >= 0
|
|
? itemsContent.substring(0, propertiesInlineStart)
|
|
: itemsContent
|
|
const itemsDesc = extractDescription(searchContent)
|
|
|
|
result.items = {
|
|
type: itemsType ? itemsType[1] : 'object',
|
|
description: itemsDesc || '',
|
|
}
|
|
|
|
if (itemsPropsConstMatch && toolPrefix) {
|
|
const resolvedProps = resolveConstReference(itemsPropsConstMatch[1], toolPrefix)
|
|
if (resolvedProps) {
|
|
result.items.properties = resolvedProps
|
|
}
|
|
} else if (propertiesInlineStart !== -1) {
|
|
const itemsPropertiesRegex = /properties\s*:\s*{/
|
|
const itemsPropsStart = itemsContent.search(itemsPropertiesRegex)
|
|
|
|
if (itemsPropsStart !== -1) {
|
|
const propsBraceStart = itemsContent.indexOf('{', itemsPropsStart)
|
|
let propsBraceCount = 1
|
|
let propsBraceEnd = propsBraceStart + 1
|
|
|
|
while (propsBraceEnd < itemsContent.length && propsBraceCount > 0) {
|
|
if (itemsContent[propsBraceEnd] === '{') propsBraceCount++
|
|
else if (itemsContent[propsBraceEnd] === '}') propsBraceCount--
|
|
propsBraceEnd++
|
|
}
|
|
|
|
if (propsBraceCount === 0) {
|
|
const itemsPropsContent = itemsContent
|
|
.substring(propsBraceStart + 1, propsBraceEnd - 1)
|
|
.trim()
|
|
result.items.properties = parsePropertiesContent(itemsPropsContent, toolPrefix)
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
return result
|
|
}
|
|
|
|
function parsePropertiesContent(
|
|
propertiesContent: string,
|
|
toolPrefix?: string
|
|
): Record<string, any> {
|
|
const properties: Record<string, any> = {}
|
|
|
|
// First, handle const references at the property level
|
|
// Patterns: "attendees: ATTENDEES_OUTPUT" or "id: BOOKING_DATA_OUTPUT_PROPERTIES.id"
|
|
if (toolPrefix) {
|
|
// Pattern 1: Direct const reference (e.g., "eventType: EVENT_TYPE_OUTPUT,")
|
|
const constRefRegex = /(\w+)\s*:\s*([A-Z][A-Z_0-9]+)\s*(?:,|$)/g
|
|
let constMatch
|
|
while ((constMatch = constRefRegex.exec(propertiesContent)) !== null) {
|
|
const propName = constMatch[1]
|
|
const constName = constMatch[2]
|
|
|
|
// Skip keywords
|
|
if (propName === 'items' || propName === 'properties' || propName === 'type') {
|
|
continue
|
|
}
|
|
|
|
// Check if at depth 0
|
|
const beforeMatch = propertiesContent.substring(0, constMatch.index)
|
|
const openBraces = (beforeMatch.match(/\{/g) || []).length
|
|
const closeBraces = (beforeMatch.match(/\}/g) || []).length
|
|
if (openBraces !== closeBraces) {
|
|
continue
|
|
}
|
|
|
|
const resolvedConst = resolveConstReference(constName, toolPrefix)
|
|
if (resolvedConst) {
|
|
properties[propName] = resolvedConst
|
|
}
|
|
}
|
|
|
|
// Pattern 2: Property access on const (e.g., "id: BOOKING_DATA_OUTPUT_PROPERTIES.id,")
|
|
const propAccessRegex = /(\w+)\s*:\s*([A-Z][A-Z_0-9]+)\.(\w+)\s*(?:,|$)/g
|
|
let propAccessMatch
|
|
while ((propAccessMatch = propAccessRegex.exec(propertiesContent)) !== null) {
|
|
const propName = propAccessMatch[1]
|
|
const constName = propAccessMatch[2]
|
|
const accessedProp = propAccessMatch[3]
|
|
|
|
// Skip keywords
|
|
if (propName === 'items' || propName === 'properties' || propName === 'type') {
|
|
continue
|
|
}
|
|
|
|
// Skip if already resolved
|
|
if (properties[propName]) {
|
|
continue
|
|
}
|
|
|
|
// Check if at depth 0
|
|
const beforeMatch = propertiesContent.substring(0, propAccessMatch.index)
|
|
const openBraces = (beforeMatch.match(/\{/g) || []).length
|
|
const closeBraces = (beforeMatch.match(/\}/g) || []).length
|
|
if (openBraces !== closeBraces) {
|
|
continue
|
|
}
|
|
|
|
const resolvedConst = resolveConstReference(constName, toolPrefix)
|
|
if (resolvedConst?.[accessedProp]) {
|
|
properties[propName] = resolvedConst[accessedProp]
|
|
}
|
|
}
|
|
|
|
// Pattern 3: Spread operator (e.g., "...COMMENT_OUTPUT_PROPERTIES,")
|
|
const spreadRegex = /\.\.\.([A-Z][A-Z_0-9]+)\s*(?:,|$)/g
|
|
let spreadMatch
|
|
while ((spreadMatch = spreadRegex.exec(propertiesContent)) !== null) {
|
|
const constName = spreadMatch[1]
|
|
|
|
// Check if at depth 0
|
|
const beforeMatch = propertiesContent.substring(0, spreadMatch.index)
|
|
const openBraces = (beforeMatch.match(/\{/g) || []).length
|
|
const closeBraces = (beforeMatch.match(/\}/g) || []).length
|
|
if (openBraces !== closeBraces) {
|
|
continue
|
|
}
|
|
|
|
const resolvedConst = resolveConstReference(constName, toolPrefix)
|
|
if (resolvedConst && typeof resolvedConst === 'object') {
|
|
// Spread all properties from the resolved const
|
|
Object.assign(properties, resolvedConst)
|
|
}
|
|
}
|
|
}
|
|
|
|
const propStartRegex = /(\w+)\s*:\s*{/g
|
|
let match
|
|
const propPositions: Array<{ name: string; start: number; content: string }> = []
|
|
|
|
while ((match = propStartRegex.exec(propertiesContent)) !== null) {
|
|
const propName = match[1]
|
|
|
|
if (propName === 'items' || propName === 'properties') {
|
|
continue
|
|
}
|
|
|
|
// Skip if already resolved as const reference
|
|
if (properties[propName]) {
|
|
continue
|
|
}
|
|
|
|
// Check if this match is at depth 0 (not inside nested braces)
|
|
// Only process top-level properties, skip nested ones
|
|
const beforeMatch = propertiesContent.substring(0, match.index)
|
|
const openBraces = (beforeMatch.match(/{/g) || []).length
|
|
const closeBraces = (beforeMatch.match(/}/g) || []).length
|
|
if (openBraces !== closeBraces) {
|
|
continue // Skip - this is a nested property
|
|
}
|
|
|
|
const startPos = match.index + match[0].length - 1
|
|
|
|
const endPos = findMatchingClose(propertiesContent, startPos)
|
|
|
|
if (endPos !== -1) {
|
|
const propContent = propertiesContent.substring(startPos + 1, endPos - 1).trim()
|
|
|
|
const hasDescription = /description\s*:\s*/.test(propContent)
|
|
const hasProperties = /properties\s*:\s*[{A-Z]/.test(propContent)
|
|
const hasItems = /items\s*:\s*[{A-Z]/.test(propContent)
|
|
const isTypeOnly =
|
|
!hasDescription &&
|
|
!hasProperties &&
|
|
!hasItems &&
|
|
/^type\s*:\s*['"].*?['"]\s*,?\s*$/.test(propContent)
|
|
|
|
if (!isTypeOnly) {
|
|
propPositions.push({
|
|
name: propName,
|
|
start: startPos,
|
|
content: propContent,
|
|
})
|
|
}
|
|
}
|
|
}
|
|
|
|
propPositions.forEach((prop) => {
|
|
const parsedProp = parseFieldContent(prop.content, toolPrefix)
|
|
if (parsedProp) {
|
|
properties[prop.name] = parsedProp
|
|
}
|
|
})
|
|
|
|
return properties
|
|
}
|
|
|
|
async function getToolInfo(toolName: string): Promise<{
|
|
description: string
|
|
params: Array<{ name: string; type: string; required: boolean; description: string }>
|
|
outputs: Record<string, any>
|
|
} | null> {
|
|
try {
|
|
const parts = toolName.split('_')
|
|
|
|
let toolPrefix = ''
|
|
let toolSuffix = ''
|
|
|
|
for (let i = parts.length - 1; i >= 1; i--) {
|
|
const possiblePrefix = parts.slice(0, i).join('_')
|
|
const possibleSuffix = parts.slice(i).join('_')
|
|
|
|
const toolDirPath = path.join(rootDir, `apps/sim/tools/${possiblePrefix}`)
|
|
|
|
if (fs.existsSync(toolDirPath) && fs.statSync(toolDirPath).isDirectory()) {
|
|
toolPrefix = possiblePrefix
|
|
toolSuffix = possibleSuffix
|
|
break
|
|
}
|
|
}
|
|
|
|
if (!toolPrefix) {
|
|
toolPrefix = parts[0]
|
|
toolSuffix = parts.slice(1).join('_')
|
|
}
|
|
|
|
// Check if this is a versioned tool (e.g., _v2, _v3)
|
|
const isVersionedTool = isVersionedType(toolSuffix)
|
|
const strippedToolSuffix = stripVersionSuffix(toolSuffix)
|
|
|
|
const possibleLocations: Array<{ path: string; priority: 'exact' | 'fallback' }> = []
|
|
|
|
// For versioned tools, prioritize the exact versioned file first
|
|
// This handles cases like google_sheets where V2 is in a separate file (read_v2.ts)
|
|
if (isVersionedTool) {
|
|
// First priority: exact versioned file (e.g., read_v2.ts)
|
|
possibleLocations.push({
|
|
path: path.join(rootDir, `apps/sim/tools/${toolPrefix}/${toolSuffix}.ts`),
|
|
priority: 'exact',
|
|
})
|
|
// Second priority: stripped file that contains both V1 and V2 (e.g., pr.ts for github)
|
|
possibleLocations.push({
|
|
path: path.join(rootDir, `apps/sim/tools/${toolPrefix}/${strippedToolSuffix}.ts`),
|
|
priority: 'fallback',
|
|
})
|
|
} else {
|
|
// Non-versioned tool: try the direct file
|
|
possibleLocations.push({
|
|
path: path.join(rootDir, `apps/sim/tools/${toolPrefix}/${toolSuffix}.ts`),
|
|
priority: 'exact',
|
|
})
|
|
}
|
|
|
|
// Also try camelCase versions
|
|
const camelCaseSuffix = strippedToolSuffix
|
|
.split('_')
|
|
.map((part, i) => (i === 0 ? part : part.charAt(0).toUpperCase() + part.slice(1)))
|
|
.join('')
|
|
possibleLocations.push({
|
|
path: path.join(rootDir, `apps/sim/tools/${toolPrefix}/${camelCaseSuffix}.ts`),
|
|
priority: 'fallback',
|
|
})
|
|
|
|
// Fall back to index.ts
|
|
possibleLocations.push({
|
|
path: path.join(rootDir, `apps/sim/tools/${toolPrefix}/index.ts`),
|
|
priority: 'fallback',
|
|
})
|
|
|
|
let toolFileContent = ''
|
|
let foundFile = ''
|
|
let foundExactId = false
|
|
|
|
// Try to find a file that contains the exact tool ID
|
|
for (const location of possibleLocations) {
|
|
if (fs.existsSync(location.path)) {
|
|
const content = fs.readFileSync(location.path, 'utf-8')
|
|
|
|
// Check if this file contains the exact tool ID we're looking for
|
|
const toolIdRegex = new RegExp(`id:\\s*['"]${toolName}['"]`)
|
|
if (toolIdRegex.test(content)) {
|
|
toolFileContent = content
|
|
foundFile = location.path
|
|
foundExactId = true
|
|
break
|
|
}
|
|
|
|
// For fallback locations, store the content in case we don't find an exact match
|
|
if (location.priority === 'fallback' && !toolFileContent) {
|
|
toolFileContent = content
|
|
foundFile = location.path
|
|
}
|
|
}
|
|
}
|
|
|
|
// The named-file candidates above miss tools defined inside a sibling tool's
|
|
// file (e.g. file_decompress lives in compress.ts). Before accepting an
|
|
// arbitrary fallback file, scan the whole tool-prefix directory for the file
|
|
// that declares this exact tool ID.
|
|
if (!foundExactId) {
|
|
const prefixDir = path.join(rootDir, `apps/sim/tools/${toolPrefix}`)
|
|
if (fs.existsSync(prefixDir)) {
|
|
const dirFiles = await glob(`${prefixDir}/**/*.ts`)
|
|
const toolIdRegex = new RegExp(`id:\\s*['"]${toolName}['"]`)
|
|
for (const dirFile of dirFiles) {
|
|
if (dirFile.endsWith('.test.ts')) continue
|
|
const content = fs.readFileSync(dirFile, 'utf-8')
|
|
if (toolIdRegex.test(content)) {
|
|
toolFileContent = content
|
|
foundFile = dirFile
|
|
foundExactId = true
|
|
break
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
// If we didn't find a file with the exact ID, use the first available file
|
|
if (!toolFileContent) {
|
|
for (const location of possibleLocations) {
|
|
if (fs.existsSync(location.path)) {
|
|
toolFileContent = fs.readFileSync(location.path, 'utf-8')
|
|
foundFile = location.path
|
|
break
|
|
}
|
|
}
|
|
}
|
|
|
|
if (!toolFileContent) {
|
|
console.warn(`Could not find definition for tool: ${toolName}`)
|
|
return null
|
|
}
|
|
|
|
return extractToolInfo(
|
|
toolName,
|
|
toolFileContent,
|
|
resolveFactorySource(toolFileContent, foundFile, rootDir)
|
|
)
|
|
} catch (error) {
|
|
console.error(`Error getting info for tool ${toolName}:`, error)
|
|
return null
|
|
}
|
|
}
|
|
|
|
function extractManualContent(existingContent: string): Record<string, string> {
|
|
const manualSections: Record<string, string> = {}
|
|
const manualContentRegex =
|
|
/\{\/\*\s*MANUAL-CONTENT-START:(\w+)\s*\*\/\}([\s\S]*?)\{\/\*\s*MANUAL-CONTENT-END\s*\*\/\}/g
|
|
|
|
let match
|
|
while ((match = manualContentRegex.exec(existingContent)) !== null) {
|
|
const sectionName = match[1]
|
|
const content = match[2].trim()
|
|
manualSections[sectionName] = content
|
|
}
|
|
|
|
return manualSections
|
|
}
|
|
|
|
function mergeWithManualContent(
|
|
generatedMarkdown: string,
|
|
existingContent: string | null,
|
|
manualSections: Record<string, string>
|
|
): string {
|
|
if (!existingContent || Object.keys(manualSections).length === 0) {
|
|
return generatedMarkdown
|
|
}
|
|
|
|
let mergedContent = generatedMarkdown
|
|
|
|
Object.entries(manualSections).forEach(([sectionName, content]) => {
|
|
const insertionPoints: Record<string, { regex: RegExp }> = {
|
|
intro: {
|
|
regex: /<BlockInfoCard[\s\S]*?(\/>|<\/svg>`}\s*\/>)/,
|
|
},
|
|
usage: {
|
|
regex: /## Usage Instructions/,
|
|
},
|
|
outputs: {
|
|
regex: /## Outputs/,
|
|
},
|
|
notes: {
|
|
regex: /## Notes/,
|
|
},
|
|
}
|
|
|
|
const insertionPoint = insertionPoints[sectionName]
|
|
const wrapped = `{/* MANUAL-CONTENT-START:${sectionName} */}\n${content}\n{/* MANUAL-CONTENT-END */}`
|
|
|
|
const match = insertionPoint ? mergedContent.match(insertionPoint.regex) : null
|
|
if (match && match.index !== undefined) {
|
|
const insertPosition = match.index + match[0].length
|
|
mergedContent = `${mergedContent.slice(0, insertPosition)}\n\n${wrapped}\n${mergedContent.slice(insertPosition)}`
|
|
} else {
|
|
// Never drop manual content: when the anchor is missing (e.g. a `notes`
|
|
// section with no generated "## Notes" heading), append at the end.
|
|
console.log(`No insertion anchor for manual section "${sectionName}" — appending at end`)
|
|
mergedContent = `${mergedContent.replace(/\s*$/, '')}\n\n${wrapped}\n`
|
|
}
|
|
})
|
|
|
|
return mergedContent
|
|
}
|
|
|
|
async function generateBlockDoc(blockPath: string) {
|
|
try {
|
|
const blockFileName = path.basename(blockPath, '.ts')
|
|
if (blockFileName.endsWith('.test')) {
|
|
return
|
|
}
|
|
|
|
const fileContent = fs.readFileSync(blockPath, 'utf-8')
|
|
|
|
// Extract ALL block configs from the file (already filters out hideFromToolbar: true)
|
|
const blockConfigs = extractAllBlockConfigs(fileContent)
|
|
|
|
if (blockConfigs.length === 0) {
|
|
console.warn(`Skipping ${blockFileName} - no valid block configs found`)
|
|
return
|
|
}
|
|
|
|
// Process each block config
|
|
for (const blockConfig of blockConfigs) {
|
|
if (!blockConfig.type) {
|
|
continue
|
|
}
|
|
|
|
if (
|
|
blockConfig.type.includes('_trigger') ||
|
|
blockConfig.type.includes('_webhook') ||
|
|
blockConfig.type.includes('rss')
|
|
) {
|
|
console.log(`Skipping ${blockConfig.type} - contains '_trigger'`)
|
|
continue
|
|
}
|
|
|
|
if (
|
|
(blockConfig.category === 'blocks' &&
|
|
!NATIVE_RESOURCE_BLOCK_TYPES.has(stripVersionSuffix(blockConfig.type))) ||
|
|
blockConfig.type === 'sim_workspace_event' ||
|
|
blockConfig.type === 'evaluator' ||
|
|
blockConfig.type === 'number' ||
|
|
blockConfig.type === 'webhook' ||
|
|
blockConfig.type === 'schedule' ||
|
|
blockConfig.type === 'mcp' ||
|
|
blockConfig.type === 'generic_webhook' ||
|
|
blockConfig.type === 'rss'
|
|
) {
|
|
continue
|
|
}
|
|
|
|
// Use stripped type for file name (removes _v2, _v3 suffixes for cleaner URLs)
|
|
const displayType = stripVersionSuffix(blockConfig.type)
|
|
const outputFilePath = path.join(DOCS_OUTPUT_PATH, `${displayType}.mdx`)
|
|
|
|
let existingContent: string | null = null
|
|
if (fs.existsSync(outputFilePath)) {
|
|
existingContent = fs.readFileSync(outputFilePath, 'utf-8')
|
|
}
|
|
|
|
const manualSections = existingContent ? extractManualContent(existingContent) : {}
|
|
|
|
const markdown = await generateMarkdownForBlock(blockConfig, displayType)
|
|
|
|
let finalContent = markdown
|
|
if (Object.keys(manualSections).length > 0) {
|
|
finalContent = mergeWithManualContent(markdown, existingContent, manualSections)
|
|
}
|
|
|
|
fs.writeFileSync(outputFilePath, finalContent)
|
|
const logType =
|
|
displayType !== blockConfig.type ? `${displayType} (from ${blockConfig.type})` : displayType
|
|
console.log(`✓ Generated docs for ${logType}`)
|
|
}
|
|
} catch (error) {
|
|
console.error(`Error processing ${blockPath}:`, error)
|
|
}
|
|
}
|
|
|
|
async function generateMarkdownForBlock(
|
|
blockConfig: BlockConfig,
|
|
displayType?: string
|
|
): Promise<string> {
|
|
const {
|
|
type,
|
|
name,
|
|
description,
|
|
longDescription,
|
|
bgColor,
|
|
outputs = {},
|
|
tools = { access: [] },
|
|
} = blockConfig
|
|
|
|
let outputsSection = ''
|
|
|
|
if (outputs && Object.keys(outputs).length > 0) {
|
|
outputsSection = '## Outputs\n\n'
|
|
|
|
outputsSection += '| Output | Type | Description |\n'
|
|
outputsSection += '| ------ | ---- | ----------- |\n'
|
|
|
|
for (const outputKey in outputs) {
|
|
const output = outputs[outputKey]
|
|
|
|
const escapedDescription = output.description
|
|
? output.description
|
|
.replace(/\|/g, '\\|')
|
|
.replace(/\{/g, '\\{')
|
|
.replace(/\}/g, '\\}')
|
|
.replace(/\(/g, '\\(')
|
|
.replace(/\)/g, '\\)')
|
|
.replace(/\[/g, '\\[')
|
|
.replace(/\]/g, '\\]')
|
|
.replace(/</g, '<')
|
|
.replace(/>/g, '>')
|
|
: `Output from ${outputKey}`
|
|
|
|
if (typeof output.type === 'string') {
|
|
outputsSection += `| \`${outputKey}\` | ${output.type} | ${escapedDescription} |\n`
|
|
} else if (output.type && typeof output.type === 'object') {
|
|
outputsSection += `| \`${outputKey}\` | object | ${escapedDescription} |\n`
|
|
|
|
for (const propName in output.type) {
|
|
const propType = output.type[propName]
|
|
const commentMatch =
|
|
propName && output.type[propName]._comment
|
|
? output.type[propName]._comment
|
|
: `${propName} of the ${outputKey}`
|
|
|
|
outputsSection += `| ↳ \`${propName}\` | ${propType} | ${commentMatch} |\n`
|
|
}
|
|
} else if (output.properties) {
|
|
outputsSection += `| \`${outputKey}\` | object | ${escapedDescription} |\n`
|
|
|
|
for (const propName in output.properties) {
|
|
const prop = output.properties[propName]
|
|
const escapedPropertyDescription = prop.description
|
|
? prop.description
|
|
.replace(/\|/g, '\\|')
|
|
.replace(/\{/g, '\\{')
|
|
.replace(/\}/g, '\\}')
|
|
.replace(/\(/g, '\\(')
|
|
.replace(/\)/g, '\\)')
|
|
.replace(/\[/g, '\\[')
|
|
.replace(/\]/g, '\\]')
|
|
.replace(/</g, '<')
|
|
.replace(/>/g, '>')
|
|
: `The ${propName} of the ${outputKey}`
|
|
|
|
outputsSection += `| ↳ \`${propName}\` | ${prop.type} | ${escapedPropertyDescription} |\n`
|
|
}
|
|
}
|
|
}
|
|
} else {
|
|
outputsSection = 'This block does not produce any outputs.'
|
|
}
|
|
|
|
let toolsSection = ''
|
|
if (tools.access?.length) {
|
|
toolsSection = '## Actions\n\n'
|
|
|
|
for (const tool of tools.access) {
|
|
// Strip version suffix from tool name for display
|
|
const displayToolName = stripVersionSuffix(tool)
|
|
toolsSection += `### \`${displayToolName}\`\n\n`
|
|
|
|
console.log(`Getting info for tool: ${tool}`)
|
|
const toolInfo = await getToolInfo(tool)
|
|
|
|
if (toolInfo) {
|
|
if (toolInfo.description && toolInfo.description !== 'No description available') {
|
|
const escapedToolDescription = toolInfo.description
|
|
.replace(/\{/g, '\\{')
|
|
.replace(/\}/g, '\\}')
|
|
toolsSection += `${escapedToolDescription}\n\n`
|
|
}
|
|
|
|
toolsSection += '#### Input\n\n'
|
|
toolsSection += '| Parameter | Type | Required | Description |\n'
|
|
toolsSection += '| --------- | ---- | -------- | ----------- |\n'
|
|
|
|
if (toolInfo.params.length > 0) {
|
|
for (const param of toolInfo.params) {
|
|
const escapedDescription = param.description
|
|
? param.description
|
|
.replace(/\|/g, '\\|')
|
|
.replace(/\{/g, '\\{')
|
|
.replace(/\}/g, '\\}')
|
|
.replace(/\(/g, '\\(')
|
|
.replace(/\)/g, '\\)')
|
|
.replace(/\[/g, '\\[')
|
|
.replace(/\]/g, '\\]')
|
|
.replace(/</g, '<')
|
|
.replace(/>/g, '>')
|
|
: 'No description'
|
|
|
|
toolsSection += `| \`${param.name}\` | ${param.type} | ${param.required ? 'Yes' : 'No'} | ${escapedDescription} |\n`
|
|
}
|
|
}
|
|
|
|
toolsSection += '\n#### Output\n\n'
|
|
|
|
if (Object.keys(toolInfo.outputs).length > 0) {
|
|
toolsSection += '| Parameter | Type | Description |\n'
|
|
toolsSection += '| --------- | ---- | ----------- |\n'
|
|
|
|
toolsSection += formatOutputStructure(toolInfo.outputs)
|
|
} else if (Object.keys(outputs).length > 0) {
|
|
toolsSection += '| Parameter | Type | Description |\n'
|
|
toolsSection += '| --------- | ---- | ----------- |\n'
|
|
|
|
for (const [key, output] of Object.entries(outputs)) {
|
|
let type = 'string'
|
|
let description = `${key} output from the tool`
|
|
|
|
if (typeof output === 'string') {
|
|
type = output
|
|
} else if (typeof output === 'object' && output !== null) {
|
|
if ('type' in output && typeof output.type === 'string') {
|
|
type = output.type
|
|
}
|
|
if ('description' in output && typeof output.description === 'string') {
|
|
description = output.description
|
|
}
|
|
}
|
|
|
|
const escapedDescription = description
|
|
.replace(/\|/g, '\\|')
|
|
.replace(/\{/g, '\\{')
|
|
.replace(/\}/g, '\\}')
|
|
.replace(/\(/g, '\\(')
|
|
.replace(/\)/g, '\\)')
|
|
.replace(/\[/g, '\\[')
|
|
.replace(/\]/g, '\\]')
|
|
.replace(/</g, '<')
|
|
.replace(/>/g, '>')
|
|
|
|
toolsSection += `| \`${key}\` | ${type} | ${escapedDescription} |\n`
|
|
}
|
|
} else {
|
|
toolsSection += 'This tool does not produce any outputs.\n'
|
|
}
|
|
}
|
|
|
|
toolsSection += '\n'
|
|
}
|
|
}
|
|
|
|
let usageInstructions = ''
|
|
if (longDescription) {
|
|
usageInstructions = `## Usage Instructions\n\n${longDescription}\n\n`
|
|
}
|
|
|
|
return `---
|
|
title: ${name}
|
|
description: ${description}
|
|
---
|
|
|
|
import { BlockInfoCard } from "@/components/ui/block-info-card"
|
|
|
|
<BlockInfoCard
|
|
type="${type}"
|
|
color="${bgColor || '#F5F5F5'}"
|
|
/>
|
|
|
|
${usageInstructions}
|
|
|
|
${toolsSection}
|
|
`
|
|
}
|
|
|
|
/**
|
|
* Compute the canonical set of stripped block types that should have a
|
|
* `docs/tools/*.mdx` file — namely every visible `category: 'tools'` block
|
|
* (matching the writer filter at the top of this script). Any existing MDX
|
|
* not in this set is stale and gets cleaned up.
|
|
*
|
|
* Uses `extractAllBlockConfigs` so spread-inherited fields (e.g. a V2 that
|
|
* spreads `...GmailBlock` and inherits `category: 'tools'`) are resolved the
|
|
* same way the writer resolves them. `stripVersionSuffix` ensures V1 and V2
|
|
* map to the same doc filename — alphabetical glob order means the newest
|
|
* version naturally wins for both generation and cleanup.
|
|
*/
|
|
async function getCanonicalToolDocNames(): Promise<Set<string>> {
|
|
const validToolDocs = new Set<string>()
|
|
const blockFiles = (await glob(`${BLOCKS_PATH}/*.ts`)).sort()
|
|
|
|
for (const blockFile of blockFiles) {
|
|
const fileContent = fs.readFileSync(blockFile, 'utf-8')
|
|
const configs = extractAllBlockConfigs(fileContent)
|
|
|
|
for (const config of configs) {
|
|
// Match the writer filter: integration blocks, the documented
|
|
// native-resource blocks (category 'blocks'), and trigger-only service
|
|
// blocks (category 'triggers') whose pages the trigger pass writes.
|
|
const stripped = config.type ? stripVersionSuffix(config.type) : ''
|
|
const isDocumentedResource = NATIVE_RESOURCE_BLOCK_TYPES.has(stripped)
|
|
const isTriggerService =
|
|
config.category === 'triggers' &&
|
|
!config.hideFromToolbar &&
|
|
stripped !== 'sim_workspace_event'
|
|
if (!isIntegrationBlock(config) && !isDocumentedResource && !isTriggerService) continue
|
|
validToolDocs.add(stripped)
|
|
}
|
|
}
|
|
|
|
return validToolDocs
|
|
}
|
|
|
|
/**
|
|
* Remove any `docs/tools/*.mdx` that no longer corresponds to a visible
|
|
* `category: 'tools'` block — covers both hidden blocks and blocks that
|
|
* have been re-categorized to `'blocks'` / `'triggers'`. Keeps the
|
|
* tools/ docs directory in lockstep with the canonical block registry.
|
|
*/
|
|
function cleanupStaleToolDocs(validToolDocs: Set<string>): void {
|
|
console.log('Cleaning up stale tool docs...')
|
|
|
|
const existingDocs = fs
|
|
.readdirSync(DOCS_OUTPUT_PATH)
|
|
.filter((file: string) => file.endsWith('.mdx'))
|
|
|
|
let removedCount = 0
|
|
let keptForManualContent = 0
|
|
|
|
for (const docFile of existingDocs) {
|
|
const blockType = path.basename(docFile, '.mdx')
|
|
if (HANDWRITTEN_INTEGRATION_DOCS.has(blockType)) continue
|
|
if (validToolDocs.has(blockType)) continue
|
|
|
|
const docPath = path.join(DOCS_OUTPUT_PATH, docFile)
|
|
|
|
// Deleting a page that a later writer re-emits destroys its hand-written
|
|
// MANUAL-CONTENT blocks: the writer merges against the file on disk, and a
|
|
// deleted file reads as "no manual content". Whenever the two filters
|
|
// disagree, keep the prose and let the mismatch be fixed deliberately.
|
|
// Gate on what `extractManualContent` can actually recover — a stray or
|
|
// unterminated start marker preserves nothing, so it must not pin the page.
|
|
const manualSections = extractManualContent(fs.readFileSync(docPath, 'utf-8'))
|
|
if (Object.values(manualSections).some((section) => section.length > 0)) {
|
|
console.warn(
|
|
`⚠ Keeping ${blockType}.mdx: considered stale but holds MANUAL-CONTENT. ` +
|
|
`Add it to a doc-emitting set or delete it by hand once the content is migrated.`
|
|
)
|
|
keptForManualContent++
|
|
continue
|
|
}
|
|
|
|
fs.unlinkSync(docPath)
|
|
console.log(`✓ Removed stale tool doc: ${blockType}.mdx`)
|
|
removedCount++
|
|
}
|
|
|
|
if (keptForManualContent > 0) {
|
|
console.log(`⚠ Kept ${keptForManualContent} stale-looking doc(s) holding manual content`)
|
|
}
|
|
|
|
if (removedCount > 0) {
|
|
console.log(`✓ Cleaned up ${removedCount} stale tool doc files`)
|
|
} else {
|
|
console.log('✓ No stale tool docs to clean up')
|
|
}
|
|
}
|
|
|
|
// ============================================================================
|
|
// Trigger Documentation Generation
|
|
// ============================================================================
|
|
|
|
/**
|
|
* Format a trigger provider name for display, falling back to Title Case.
|
|
*/
|
|
function formatTriggerProviderName(provider: string): string {
|
|
if (TRIGGER_PROVIDER_DISPLAY_NAMES[provider]) {
|
|
return TRIGGER_PROVIDER_DISPLAY_NAMES[provider]
|
|
}
|
|
return provider.replace(/[-_]/g, ' ').replace(/\b\w/g, (c) => c.toUpperCase())
|
|
}
|
|
|
|
/**
|
|
* Escape text for use inside an MDX table cell.
|
|
*/
|
|
function escapeMdxCell(text: string): string {
|
|
return text
|
|
.replace(/\|/g, '\\|')
|
|
.replace(/\{/g, '\\{')
|
|
.replace(/\}/g, '\\}')
|
|
.replace(/\(/g, '\\(')
|
|
.replace(/\)/g, '\\)')
|
|
.replace(/\[/g, '\\[')
|
|
.replace(/\]/g, '\\]')
|
|
.replace(/</g, '<')
|
|
.replace(/>/g, '>')
|
|
}
|
|
|
|
/**
|
|
* Resolve a module-level `const varName = { ... }` declaration.
|
|
* Handles nested spreads of other const variables (but not property-access values).
|
|
* Used to expand variable spreads inside builder function return bodies.
|
|
*/
|
|
function resolveConstVariable(
|
|
varName: string,
|
|
primaryContent: string,
|
|
utilsContent: string,
|
|
depth = 0
|
|
): Record<string, any> {
|
|
if (depth > 8) return {}
|
|
|
|
// Match `const varName = {` (with optional type annotation)
|
|
const varRegex = new RegExp(`(?<![.\\w])const\\s+${varName}\\s*(?::[^=]+)?=\\s*\\{`)
|
|
|
|
for (const content of [primaryContent, utilsContent]) {
|
|
const varMatch = varRegex.exec(content)
|
|
if (!varMatch) continue
|
|
|
|
const openBrace = content.indexOf('{', varMatch.index + varMatch[0].length - 1)
|
|
if (openBrace === -1) continue
|
|
|
|
const closeBrace = findMatchingClose(content, openBrace)
|
|
if (closeBrace === -1) continue
|
|
|
|
const varBody = content.substring(openBrace + 1, closeBrace - 1).trim()
|
|
const result: Record<string, any> = {}
|
|
|
|
// Resolve nested variable spreads within this const (no parens = variable reference)
|
|
const nestedSpreadRegex = /\.\.\.\s*([a-zA-Z_]\w*)\b(?!\s*\()/g
|
|
let nestedMatch: RegExpExecArray | null
|
|
while ((nestedMatch = nestedSpreadRegex.exec(varBody)) !== null) {
|
|
const nested = resolveConstVariable(nestedMatch[1], primaryContent, utilsContent, depth + 1)
|
|
Object.assign(result, nested)
|
|
}
|
|
|
|
// Parse any inline `field: { type, description }` definitions
|
|
// (strip spread lines first; property-access values like `foo: bar.baz` are skipped by parser)
|
|
const bodyWithoutVarSpreads = varBody.replace(/\.\.\.\s*\w+\b(?!\s*\()\s*,?\s*/g, '')
|
|
const inlineOutputs = parseToolOutputsField(bodyWithoutVarSpreads)
|
|
Object.assign(result, inlineOutputs)
|
|
|
|
return result
|
|
}
|
|
|
|
return {}
|
|
}
|
|
|
|
/**
|
|
* Recursively resolve a trigger output builder function.
|
|
* Handles the common pattern where builders spread other builders:
|
|
* `return { ...buildBaseOutputs(), fieldA: { type: 'string', ... } }`
|
|
* Also handles variable spreads:
|
|
* `return { ...coreOutputs, ...deploymentOutputs }`
|
|
*
|
|
* Searches for the function definition in `primaryContent` first, then `utilsContent`.
|
|
* Recursion depth is capped to avoid infinite loops.
|
|
*/
|
|
function resolveTriggerBuilderFunction(
|
|
funcName: string,
|
|
primaryContent: string,
|
|
utilsContent: string,
|
|
depth = 0
|
|
): Record<string, any> {
|
|
if (depth > 8) return {}
|
|
|
|
const funcRegex = new RegExp(`(?:export\\s+)?function\\s+${funcName}\\s*\\(`)
|
|
let funcBody: string | null = null
|
|
|
|
for (const content of [primaryContent, utilsContent]) {
|
|
const funcMatch = funcRegex.exec(content)
|
|
if (!funcMatch) continue
|
|
|
|
const bodyStart = content.indexOf('{', funcMatch.index)
|
|
if (bodyStart === -1) continue
|
|
|
|
const bodyEnd = findMatchingClose(content, bodyStart)
|
|
if (bodyEnd === -1) continue
|
|
|
|
funcBody = content.substring(bodyStart + 1, bodyEnd - 1)
|
|
break
|
|
}
|
|
|
|
if (!funcBody) return {}
|
|
|
|
// Handle `return anotherFunc(...)` — full delegation to another builder,
|
|
// with or without arguments (argument values are ignored; only structure matters).
|
|
const returnFuncCallMatch = /\breturn\s+([a-z][a-zA-Z0-9_]*)\s*\(/.exec(funcBody.trim())
|
|
if (returnFuncCallMatch) {
|
|
return resolveTriggerBuilderFunction(
|
|
returnFuncCallMatch[1],
|
|
primaryContent,
|
|
utilsContent,
|
|
depth + 1
|
|
)
|
|
}
|
|
|
|
// Handle `return { ... }` — inline object literal
|
|
const returnMatch = /\breturn\s*\{/.exec(funcBody)
|
|
if (!returnMatch) return {}
|
|
|
|
const returnObjStart = funcBody.indexOf('{', returnMatch.index)
|
|
const returnObjEnd = findMatchingClose(funcBody, returnObjStart)
|
|
if (returnObjEnd === -1) return {}
|
|
|
|
const returnBody = funcBody.substring(returnObjStart + 1, returnObjEnd - 1).trim()
|
|
|
|
const result: Record<string, any> = {}
|
|
|
|
// Expand function-call spreads first: ...innerFuncName()
|
|
const spreadFuncRegex = /\.\.\.\s*(\w+)\s*\(\s*\)/g
|
|
let spreadMatch: RegExpExecArray | null
|
|
while ((spreadMatch = spreadFuncRegex.exec(returnBody)) !== null) {
|
|
const innerFuncName = spreadMatch[1]
|
|
const resolved = resolveTriggerBuilderFunction(
|
|
innerFuncName,
|
|
primaryContent,
|
|
utilsContent,
|
|
depth + 1
|
|
)
|
|
Object.assign(result, resolved)
|
|
}
|
|
|
|
// Expand variable spreads: ...varName (no parentheses — const references)
|
|
const spreadVarRegex = /\.\.\.\s*([a-zA-Z_]\w*)\b(?!\s*\()/g
|
|
let spreadVarMatch: RegExpExecArray | null
|
|
while ((spreadVarMatch = spreadVarRegex.exec(returnBody)) !== null) {
|
|
const varName = spreadVarMatch[1]
|
|
const resolved = resolveConstVariable(varName, primaryContent, utilsContent, depth + 1)
|
|
Object.assign(result, resolved)
|
|
}
|
|
|
|
// Then parse any inline field definitions (strip all spread lines first)
|
|
const bodyWithoutSpreads = returnBody
|
|
.replace(/\.\.\.\s*\w+\s*\(\s*\)\s*,?\s*/g, '') // function call spreads
|
|
.replace(/\.\.\.\s*\w+\b(?!\s*\()\s*,?\s*/g, '') // variable spreads
|
|
const inlineOutputs = parseToolOutputsField(bodyWithoutSpreads)
|
|
Object.assign(result, inlineOutputs)
|
|
|
|
return result
|
|
}
|
|
|
|
/**
|
|
* Read every sibling module of a trigger file so identifiers it references —
|
|
* builder functions and shared `outputs`/config constants alike — can be
|
|
* resolved from source text. Triggers keep these in `utils.ts` or `shared.ts`
|
|
* depending on the provider, so the whole directory is scanned rather than one
|
|
* hard-coded filename.
|
|
*/
|
|
function readTriggerSiblingModules(triggerFile: string): string {
|
|
const dir = path.dirname(triggerFile)
|
|
if (!fs.existsSync(dir)) return ''
|
|
return fs
|
|
.readdirSync(dir)
|
|
.filter((f) => f.endsWith('.ts') && !f.includes('.test.') && path.join(dir, f) !== triggerFile)
|
|
.map((f) => {
|
|
try {
|
|
return fs.readFileSync(path.join(dir, f), 'utf-8')
|
|
} catch {
|
|
return ''
|
|
}
|
|
})
|
|
.join('\n')
|
|
}
|
|
|
|
/**
|
|
* Resolve `outputs: SOME_CONSTANT` by locating the constant's object literal in
|
|
* the trigger file or one of its siblings. Without this the generated page
|
|
* silently loses the trigger's entire Output table the moment a provider
|
|
* factors its outputs out into a shared constant.
|
|
*/
|
|
function resolveTriggerOutputsConstant(
|
|
constName: string,
|
|
primaryContent: string,
|
|
siblingContent: string
|
|
): Record<string, any> {
|
|
const declRegex = new RegExp(`(?:export\\s+)?const\\s+${constName}\\s*(?::[^=]+)?=\\s*\\{`)
|
|
for (const content of [primaryContent, siblingContent]) {
|
|
const declMatch = declRegex.exec(content)
|
|
if (!declMatch) continue
|
|
const openPos = content.indexOf('{', declMatch.index)
|
|
if (openPos === -1) continue
|
|
const closePos = findMatchingClose(content, openPos)
|
|
if (closePos === -1) continue
|
|
return parseToolOutputsField(content.substring(openPos + 1, closePos - 1).trim())
|
|
}
|
|
return {}
|
|
}
|
|
|
|
/**
|
|
* Extract the outputs object from a TriggerConfig segment.
|
|
* Handles inline `outputs: { ... }`, function-call patterns like
|
|
* `outputs: buildIssueOutputs()`, and bare constant references like
|
|
* `outputs: SLACK_TRIGGER_OUTPUTS`, resolving each from the trigger file
|
|
* itself and its sibling modules.
|
|
*/
|
|
function extractTriggerOutputs(
|
|
segment: string,
|
|
fileContent: string,
|
|
utilsContent: string
|
|
): Record<string, any> {
|
|
// 1. Inline outputs: outputs: { ... }
|
|
const outputsMatch = /\boutputs\s*:\s*\{/.exec(segment)
|
|
if (outputsMatch) {
|
|
const openPos = segment.indexOf('{', outputsMatch.index + outputsMatch[0].length - 1)
|
|
if (openPos !== -1) {
|
|
const closePos = findMatchingClose(segment, openPos)
|
|
if (closePos !== -1) {
|
|
const outputsContent = segment.substring(openPos + 1, closePos - 1).trim()
|
|
return parseToolOutputsField(outputsContent)
|
|
}
|
|
}
|
|
}
|
|
|
|
// 2. Function-call outputs: outputs: buildFoo()
|
|
const funcCallMatch = /\boutputs\s*:\s*(\w+)\s*\(\s*\)/.exec(segment)
|
|
if (funcCallMatch) {
|
|
return resolveTriggerBuilderFunction(funcCallMatch[1], fileContent, utilsContent)
|
|
}
|
|
|
|
// 3. Constant reference: outputs: SLACK_TRIGGER_OUTPUTS
|
|
const constRefMatch = /\boutputs\s*:\s*([A-Za-z_$][\w$]*)\s*[,\n}]/.exec(segment)
|
|
if (constRefMatch) {
|
|
return resolveTriggerOutputsConstant(constRefMatch[1], fileContent, utilsContent)
|
|
}
|
|
|
|
return {}
|
|
}
|
|
|
|
/**
|
|
* Lazy-loaded cache of all TypeScript files in `lib/webhooks/providers/`.
|
|
* Used to resolve exported string constants that are imported by trigger utils files
|
|
* (e.g. `GONG_JWT_PUBLIC_KEY_CONFIG_KEY` from `lib/webhooks/providers/gong.ts`).
|
|
*/
|
|
let _webhookProviderConstantsCache: string | null = null
|
|
function getWebhookProviderConstants(): string {
|
|
if (_webhookProviderConstantsCache === null) {
|
|
const dir = path.join(rootDir, 'apps/sim/lib/webhooks/providers')
|
|
if (fs.existsSync(dir)) {
|
|
_webhookProviderConstantsCache = fs
|
|
.readdirSync(dir)
|
|
.filter((f) => f.endsWith('.ts'))
|
|
.map((f) => fs.readFileSync(path.join(dir, f), 'utf-8'))
|
|
.join('\n')
|
|
} else {
|
|
_webhookProviderConstantsCache = ''
|
|
}
|
|
}
|
|
return _webhookProviderConstantsCache
|
|
}
|
|
|
|
/**
|
|
* Try to resolve a SCREAMING_SNAKE_CASE constant to its string value by
|
|
* searching in the given content AND the webhook provider constants cache.
|
|
*/
|
|
function resolveConstStringValue(constName: string, content: string): string | null {
|
|
const pattern = new RegExp(`\\b${constName}\\s*=\\s*['"]([^'"]+)['"]`)
|
|
return pattern.exec(content)?.[1] ?? pattern.exec(getWebhookProviderConstants())?.[1] ?? null
|
|
}
|
|
|
|
/**
|
|
* Parse a single SubBlockConfig object literal into a TriggerConfigField.
|
|
* Returns null for blocks that should be skipped (UI-only IDs, text type, readOnly).
|
|
* Accepts optional `resolverContent` to resolve const-reference field IDs.
|
|
*/
|
|
/**
|
|
* Read a quoted string property, matching the opening quote to its own closing
|
|
* quote. A single `['"]…[^'"]+…['"]` character class ends the match at the
|
|
* first quote of *either* kind, so an apostrophe inside a double-quoted string
|
|
* ("Doesn't fire…") truncates the value mid-word.
|
|
*/
|
|
function matchQuotedProperty(content: string, propName: string): string | undefined {
|
|
const match = new RegExp(`\\b${propName}\\s*:\\s*(?:'([^']*)'|"([^"]*)"|\`([^\`]*)\`)`).exec(
|
|
content
|
|
)
|
|
if (!match) return undefined
|
|
return match[1] ?? match[2] ?? match[3]
|
|
}
|
|
|
|
function parseSubBlockObject(
|
|
obj: string,
|
|
uiOnlyIds: Set<string>,
|
|
resolverContent?: string
|
|
): TriggerConfigField | null {
|
|
let id: string | undefined = matchQuotedProperty(obj, 'id')
|
|
|
|
// Handle const-reference ids: `id: SCREAMING_CASE_IDENTIFIER`
|
|
if (!id) {
|
|
const constRefMatch = /\bid\s*:\s*([A-Z][A-Z0-9_]+)\b/.exec(obj)
|
|
if (constRefMatch) {
|
|
id = resolveConstStringValue(constRefMatch[1], resolverContent ?? '') ?? undefined
|
|
}
|
|
}
|
|
|
|
if (!id || uiOnlyIds.has(id)) return null
|
|
|
|
const type = matchQuotedProperty(obj, 'type')
|
|
if (type === 'text') return null
|
|
if (/\breadOnly\s*:\s*true/.test(obj)) return null
|
|
|
|
const title = matchQuotedProperty(obj, 'title')
|
|
const requiredMatch = /\brequired\s*:\s*(true)/.exec(obj)
|
|
const placeholder = matchQuotedProperty(obj, 'placeholder')
|
|
|
|
// Use title as description fallback so oauth-input and other fields without
|
|
// an explicit description still show something meaningful in the docs table.
|
|
const description = matchQuotedProperty(obj, 'description') ?? title
|
|
|
|
return {
|
|
id,
|
|
title: title ?? id,
|
|
type: type ?? 'short-input',
|
|
required: Boolean(requiredMatch),
|
|
placeholder,
|
|
description,
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Resolve a SubBlockConfig builder function to its field definitions.
|
|
* Handles `return [...]`, `return {...}`, and `blocks.push(...)` patterns.
|
|
* Searches `utilsContent` first, then `primaryContent`.
|
|
*/
|
|
function resolveSubBlockBuilderFunction(
|
|
funcName: string,
|
|
utilsContent: string,
|
|
primaryContent?: string
|
|
): TriggerConfigField[] {
|
|
const UI_ONLY_IDS = new Set(['webhookUrlDisplay', 'triggerInstructions', 'selectedTriggerId'])
|
|
|
|
for (const content of [utilsContent, primaryContent ?? '']) {
|
|
if (!content) continue
|
|
const funcRegex = new RegExp(`(?:export\\s+)?function\\s+${funcName}\\s*\\(`)
|
|
const funcMatch = funcRegex.exec(content)
|
|
if (!funcMatch) continue
|
|
|
|
// Find the closing ')' of the parameter list, then the '{' that opens the function body.
|
|
// Using just indexOf('{') would pick up '{' inside object-type parameters.
|
|
const openParen = content.indexOf('(', funcMatch.index)
|
|
if (openParen === -1) continue
|
|
const closeParen = findMatchingClose(content, openParen, '(', ')')
|
|
if (closeParen === -1) continue
|
|
const bodyStart = content.indexOf('{', closeParen)
|
|
if (bodyStart === -1) continue
|
|
const bodyEnd = findMatchingClose(content, bodyStart)
|
|
if (bodyEnd === -1) continue
|
|
|
|
const funcBody = content.substring(bodyStart + 1, bodyEnd - 1)
|
|
|
|
// Pattern 1: `return [...]`
|
|
const returnArrayMatch = /\breturn\s*\[/.exec(funcBody)
|
|
if (returnArrayMatch) {
|
|
const arrayStart = funcBody.indexOf('[', returnArrayMatch.index)
|
|
const arrayEnd = findMatchingClose(funcBody, arrayStart, '[', ']')
|
|
if (arrayEnd !== -1) {
|
|
return parseSubBlockArrayContent(
|
|
funcBody.substring(arrayStart + 1, arrayEnd - 1),
|
|
UI_ONLY_IDS,
|
|
content
|
|
)
|
|
}
|
|
}
|
|
|
|
// Pattern 2: `return { ... }` (single object)
|
|
const returnObjMatch = /\breturn\s*\{/.exec(funcBody)
|
|
if (returnObjMatch) {
|
|
const objStart = funcBody.indexOf('{', returnObjMatch.index)
|
|
const objEnd = findMatchingClose(funcBody, objStart)
|
|
if (objEnd !== -1) {
|
|
const field = parseSubBlockObject(
|
|
funcBody.substring(objStart, objEnd),
|
|
UI_ONLY_IDS,
|
|
content
|
|
)
|
|
return field ? [field] : []
|
|
}
|
|
}
|
|
|
|
// Pattern 3: `blocks.push({...})`
|
|
const pushFields: TriggerConfigField[] = []
|
|
const pushRegex = /\bblocks\.push\s*\(/g
|
|
let pushMatch: RegExpExecArray | null
|
|
while ((pushMatch = pushRegex.exec(funcBody)) !== null) {
|
|
const parenStart = pushMatch.index + pushMatch[0].length - 1
|
|
const parenEnd = findMatchingClose(funcBody, parenStart, '(', ')')
|
|
if (parenEnd === -1) continue
|
|
const pushArg = funcBody.substring(parenStart + 1, parenEnd - 1).trim()
|
|
if (pushArg.startsWith('{')) {
|
|
const field = parseSubBlockObject(pushArg, UI_ONLY_IDS, content)
|
|
if (field) pushFields.push(field)
|
|
}
|
|
}
|
|
if (pushFields.length > 0) return pushFields
|
|
}
|
|
|
|
return []
|
|
}
|
|
|
|
/**
|
|
* Parse SubBlockConfig items from within an array body (between the brackets).
|
|
* Handles inline `{...}` objects and function calls `funcName(...)`.
|
|
*/
|
|
function parseSubBlockArrayContent(
|
|
arrayContent: string,
|
|
uiOnlyIds: Set<string>,
|
|
utilsContent: string
|
|
): TriggerConfigField[] {
|
|
const fields: TriggerConfigField[] = []
|
|
let i = 0
|
|
|
|
while (i < arrayContent.length) {
|
|
if (arrayContent[i] === '{') {
|
|
const j = findMatchingClose(arrayContent, i)
|
|
if (j === -1) break
|
|
const field = parseSubBlockObject(arrayContent.substring(i, j), uiOnlyIds, utilsContent)
|
|
if (field) fields.push(field)
|
|
i = j
|
|
} else if (/[a-zA-Z_]/.test(arrayContent[i])) {
|
|
// Possible function call: funcName(args)
|
|
const funcCallMatch = /^(\w+)\s*\(/.exec(arrayContent.substring(i))
|
|
if (funcCallMatch && utilsContent) {
|
|
const funcName = funcCallMatch[1]
|
|
if (funcName !== 'true' && funcName !== 'false' && funcName !== 'null') {
|
|
fields.push(...resolveSubBlockBuilderFunction(funcName, utilsContent))
|
|
}
|
|
// Advance past the function call's closing paren
|
|
const openIdx = arrayContent.indexOf('(', i + funcName.length)
|
|
if (openIdx !== -1) {
|
|
const closeIdx = findMatchingClose(arrayContent, openIdx, '(', ')')
|
|
i = closeIdx !== -1 ? closeIdx : openIdx + 1
|
|
} else {
|
|
i += funcName.length
|
|
}
|
|
} else {
|
|
i++
|
|
}
|
|
} else {
|
|
i++
|
|
}
|
|
}
|
|
|
|
return fields
|
|
}
|
|
|
|
/**
|
|
* Extract user-facing configuration fields from a TriggerConfig subBlocks definition.
|
|
* Handles both inline arrays (`subBlocks: [...]`) and builder function calls
|
|
* (`subBlocks: buildXSubBlocks({...})`), resolving them from the trigger file and utils.ts.
|
|
*/
|
|
function extractTriggerConfigFields(
|
|
segment: string,
|
|
primaryContent?: string,
|
|
utilsContent?: string
|
|
): TriggerConfigField[] {
|
|
const UI_ONLY_IDS = new Set(['webhookUrlDisplay', 'triggerInstructions', 'selectedTriggerId'])
|
|
const allContent = utilsContent || primaryContent || ''
|
|
|
|
// Case 1: Inline subBlocks: [...]
|
|
const subBlocksMatch = /\bsubBlocks\s*:\s*\[/.exec(segment)
|
|
if (subBlocksMatch) {
|
|
const arrayStart = subBlocksMatch.index + subBlocksMatch[0].length - 1
|
|
const arrayEnd = findMatchingClose(segment, arrayStart, '[', ']')
|
|
if (arrayEnd === -1) return []
|
|
return parseSubBlockArrayContent(
|
|
segment.substring(arrayStart + 1, arrayEnd - 1),
|
|
UI_ONLY_IDS,
|
|
allContent
|
|
)
|
|
}
|
|
|
|
// Case 2: Builder function call — subBlocks: buildXFunc(...)
|
|
if (!allContent) return []
|
|
const builderCallMatch = /\bsubBlocks\s*:\s*(\w+)\s*\(/.exec(segment)
|
|
if (!builderCallMatch) return []
|
|
|
|
const funcName = builderCallMatch[1]
|
|
|
|
// Special case: buildTriggerSubBlocks — user config lives in the `extraFields` parameter
|
|
if (funcName === 'buildTriggerSubBlocks') {
|
|
const openParen = builderCallMatch.index + builderCallMatch[0].length - 1
|
|
const closeParen = findMatchingClose(segment, openParen, '(', ')')
|
|
if (closeParen === -1) return []
|
|
|
|
const argsBody = segment.substring(openParen + 1, closeParen - 1)
|
|
const extraFieldsMatch = /\bextraFields\s*:\s*/.exec(argsBody)
|
|
if (!extraFieldsMatch) return []
|
|
|
|
// Find first non-whitespace char after "extraFields:"
|
|
let valuePos = extraFieldsMatch.index + extraFieldsMatch[0].length
|
|
while (valuePos < argsBody.length && /\s/.test(argsBody[valuePos])) valuePos++
|
|
|
|
if (argsBody[valuePos] === '[') {
|
|
// extraFields: [...] — inline array, may contain function calls
|
|
const arrayEnd = findMatchingClose(argsBody, valuePos, '[', ']')
|
|
if (arrayEnd === -1) return []
|
|
return parseSubBlockArrayContent(
|
|
argsBody.substring(valuePos + 1, arrayEnd - 1),
|
|
UI_ONLY_IDS,
|
|
allContent
|
|
)
|
|
}
|
|
|
|
// extraFields: buildXFunc(args) — resolve the builder function
|
|
const extraFuncMatch = /^(\w+)\s*\(/.exec(argsBody.substring(valuePos))
|
|
if (!extraFuncMatch) return []
|
|
return resolveSubBlockBuilderFunction(extraFuncMatch[1], allContent)
|
|
}
|
|
|
|
// For all other builders, resolve the function body directly
|
|
return resolveSubBlockBuilderFunction(funcName, allContent)
|
|
}
|
|
|
|
/**
|
|
* Build the full trigger registry: id → TriggerFullInfo.
|
|
* Parses every trigger source file for config fields and output schemas.
|
|
*/
|
|
async function buildFullTriggerRegistry(): Promise<Map<string, TriggerFullInfo>> {
|
|
const registry = new Map<string, TriggerFullInfo>()
|
|
const SKIP = new Set(['index.ts', 'registry.ts', 'types.ts', 'constants.ts', 'utils.ts'])
|
|
|
|
const triggerFiles = (await glob(`${TRIGGERS_PATH}/**/*.ts`)).filter(
|
|
(f) => !SKIP.has(path.basename(f)) && !f.includes('.test.')
|
|
)
|
|
|
|
for (const file of triggerFiles) {
|
|
try {
|
|
const content = fs.readFileSync(file, 'utf-8')
|
|
|
|
// Load sibling modules (utils.ts, shared.ts, …) so builder functions and
|
|
// shared output/config constants referenced by name still resolve.
|
|
const utilsContent = readTriggerSiblingModules(file)
|
|
|
|
const exportRegex = /export\s+const\s+\w+\s*:\s*TriggerConfig\s*=\s*\{/g
|
|
let exportMatch: RegExpExecArray | null
|
|
const exportStarts: number[] = []
|
|
while ((exportMatch = exportRegex.exec(content)) !== null) {
|
|
exportStarts.push(exportMatch.index)
|
|
}
|
|
|
|
const segments =
|
|
exportStarts.length > 0
|
|
? exportStarts.map((start, i) => content.substring(start, exportStarts[i + 1]))
|
|
: [content]
|
|
|
|
for (const segment of segments) {
|
|
const idMatch = /\bid\s*:\s*['"]([^'"]+)['"]/.exec(segment)
|
|
const nameMatch = /\bname\s*:\s*['"]([^'"]+)['"]/.exec(segment)
|
|
const descMatch = /\bdescription\s*:\s*['"]([^'"]+)['"]/.exec(segment)
|
|
const providerMatch = /\bprovider\s*:\s*['"]([^'"]+)['"]/.exec(segment)
|
|
|
|
if (!idMatch || !nameMatch || !providerMatch) continue
|
|
|
|
// Deprecated triggers stay registered for existing workflows but are
|
|
// excluded from generated documentation.
|
|
if (/\bdeprecated\s*:\s*true/.test(segment)) continue
|
|
|
|
const polling = /\bpolling\s*:\s*true/.test(segment)
|
|
|
|
registry.set(idMatch[1], {
|
|
id: idMatch[1],
|
|
name: nameMatch[1],
|
|
description: descMatch?.[1] ?? '',
|
|
provider: providerMatch[1],
|
|
polling,
|
|
outputs: extractTriggerOutputs(segment, content, utilsContent),
|
|
configFields: extractTriggerConfigFields(segment, content, utilsContent),
|
|
})
|
|
}
|
|
} catch {
|
|
// skip unreadable files silently
|
|
}
|
|
}
|
|
|
|
console.log(`✓ Loaded full config for ${registry.size} triggers`)
|
|
return registry
|
|
}
|
|
|
|
/**
|
|
* Return the numeric version suffix of a trigger ID (e.g. `_v2` → 2, none → 1).
|
|
* Used to prefer the latest version when the same trigger name has v1 and v2 variants.
|
|
*/
|
|
function triggerVersionOrdinal(id: string): number {
|
|
const m = /_v(\d+)$/.exec(id)
|
|
return m ? Number.parseInt(m[1], 10) : 1
|
|
}
|
|
|
|
/**
|
|
* Group triggers by provider; triggers within each group are sorted alphabetically.
|
|
* When multiple triggers share the same display name (e.g. v1 + v2 of the same event),
|
|
* only the highest-version variant is kept so docs don't show duplicate sections.
|
|
*/
|
|
function groupTriggersByProvider(
|
|
registry: Map<string, TriggerFullInfo>
|
|
): Map<string, TriggerFullInfo[]> {
|
|
const groups = new Map<string, TriggerFullInfo[]>()
|
|
for (const trigger of registry.values()) {
|
|
const bucket = groups.get(trigger.provider) ?? []
|
|
bucket.push(trigger)
|
|
groups.set(trigger.provider, bucket)
|
|
}
|
|
for (const [provider, triggers] of groups) {
|
|
// Deduplicate by name: keep the highest-versioned trigger for each display name
|
|
const byName = new Map<string, TriggerFullInfo>()
|
|
for (const trigger of triggers) {
|
|
const existing = byName.get(trigger.name)
|
|
if (!existing || triggerVersionOrdinal(trigger.id) > triggerVersionOrdinal(existing.id)) {
|
|
byName.set(trigger.name, trigger)
|
|
}
|
|
}
|
|
groups.set(
|
|
provider,
|
|
[...byName.values()].sort((a, b) => a.name.localeCompare(b.name))
|
|
)
|
|
}
|
|
return groups
|
|
}
|
|
|
|
/**
|
|
* Map subBlock UI type identifiers to semantic data types for documentation.
|
|
* Users care about the data type (string/boolean/number), not the UI widget.
|
|
*/
|
|
const SUBBLOCK_TYPE_TO_SEMANTIC: Record<string, string> = {
|
|
'short-input': 'string',
|
|
'long-input': 'string',
|
|
dropdown: 'string',
|
|
switch: 'boolean',
|
|
slider: 'number',
|
|
'oauth-input': 'string',
|
|
code: 'string',
|
|
'file-upload': 'string',
|
|
text: 'string',
|
|
}
|
|
|
|
function toSemanticType(uiType: string): string {
|
|
return SUBBLOCK_TYPE_TO_SEMANTIC[uiType] ?? uiType
|
|
}
|
|
|
|
/**
|
|
* Generate MDX content for a single trigger provider page.
|
|
* Matches the structure of tool docs: ## Triggers, ### `trigger_id`, #### Configuration / Output.
|
|
*/
|
|
/**
|
|
* Build the "## Triggers" section for an integration page. A trigger is a block
|
|
* that starts a workflow, so this is appended to the service's actions page (or
|
|
* used as the body of a trigger-only service page).
|
|
*/
|
|
function buildTriggersSection(triggers: TriggerFullInfo[]): string {
|
|
const allPolling = triggers.every((t) => t.polling)
|
|
const mixedTypes = triggers.some((t) => t.polling) && triggers.some((t) => !t.polling)
|
|
|
|
let typeNote = ''
|
|
if (allPolling) {
|
|
typeNote =
|
|
'\nThese run on a schedule \\(**polling-based**\\) — they check for new data rather than receiving push notifications.\n'
|
|
} else if (mixedTypes) {
|
|
typeNote =
|
|
'\nSome of these are **polling-based** \\(checked on a schedule\\) while others are push-based webhooks.\n'
|
|
}
|
|
|
|
let triggersSection = ''
|
|
for (let i = 0; i < triggers.length; i++) {
|
|
const trigger = triggers[i]
|
|
|
|
// Configuration table
|
|
let configSection = ''
|
|
if (trigger.configFields.length > 0) {
|
|
configSection = '#### Configuration\n\n'
|
|
configSection += '| Parameter | Type | Required | Description |\n'
|
|
configSection += '| --------- | ---- | -------- | ----------- |\n'
|
|
for (const field of trigger.configFields) {
|
|
const type = toSemanticType(field.type)
|
|
const desc = escapeMdxCell(field.description ?? field.placeholder ?? '')
|
|
configSection += `| \`${field.id}\` | ${type} | ${field.required ? 'Yes' : 'No'} | ${desc} |\n`
|
|
}
|
|
configSection += '\n'
|
|
}
|
|
|
|
// Output table
|
|
let outputSection = ''
|
|
if (Object.keys(trigger.outputs).length > 0) {
|
|
outputSection = '#### Output\n\n'
|
|
outputSection += '| Parameter | Type | Description |\n'
|
|
outputSection += '| --------- | ---- | ----------- |\n'
|
|
outputSection += formatOutputStructure(trigger.outputs)
|
|
outputSection += '\n'
|
|
}
|
|
|
|
const separator = i < triggers.length - 1 ? '\n---\n\n' : ''
|
|
|
|
triggersSection += `### ${trigger.name}\n\n`
|
|
const escapedTriggerDescription = trigger.description
|
|
.replace(/\{/g, '\\{')
|
|
.replace(/\}/g, '\\}')
|
|
triggersSection += `${escapedTriggerDescription}\n\n`
|
|
triggersSection += configSection
|
|
triggersSection += outputSection
|
|
triggersSection += separator
|
|
}
|
|
|
|
return `## Triggers
|
|
|
|
A **Trigger** is a block that starts a workflow when an event happens in this service.
|
|
${typeNote}
|
|
${triggersSection}`
|
|
}
|
|
|
|
/** Standalone page for a trigger-only service (no actions block). */
|
|
function generateTriggerProviderDoc(
|
|
provider: string,
|
|
triggers: TriggerFullInfo[],
|
|
blockType: string,
|
|
providerColor: string
|
|
): string {
|
|
const providerName = formatTriggerProviderName(provider)
|
|
return `---
|
|
title: ${providerName}
|
|
description: ${providerName} triggers for automating workflows
|
|
---
|
|
|
|
import { BlockInfoCard } from "@/components/ui/block-info-card"
|
|
|
|
<BlockInfoCard
|
|
type="${blockType}"
|
|
color="${providerColor}"
|
|
/>
|
|
|
|
${buildTriggersSection(triggers)}`
|
|
}
|
|
|
|
/**
|
|
* Build a map of block-type → bgColor from all block definitions.
|
|
* Used to pick provider colours for the BlockInfoCard on trigger pages.
|
|
*/
|
|
async function buildProviderColorMap(): Promise<Map<string, string>> {
|
|
const colorMap = new Map<string, string>()
|
|
const blockFiles = (await glob(`${BLOCKS_PATH}/*.ts`)).sort()
|
|
|
|
for (const blockFile of blockFiles) {
|
|
const fileContent = fs.readFileSync(blockFile, 'utf-8')
|
|
const configs = extractAllBlockConfigs(fileContent)
|
|
for (const config of configs) {
|
|
if (config.bgColor && config.type) {
|
|
const baseType = stripVersionSuffix(config.type)
|
|
if (!colorMap.has(baseType)) colorMap.set(baseType, config.bgColor)
|
|
}
|
|
}
|
|
}
|
|
|
|
return colorMap
|
|
}
|
|
|
|
/**
|
|
* Generate one MDX file per trigger provider and update the sidebar meta.json.
|
|
* Hand-written docs (HANDWRITTEN_TRIGGER_DOCS) are never touched.
|
|
*/
|
|
/**
|
|
* Trigger ids that every hosting block gates behind `preview: true`.
|
|
*
|
|
* Blocks declare the triggers they expose via `triggers.available`. A trigger
|
|
* listed only by preview blocks inherits their gate — `slack_oauth` is reachable
|
|
* solely through the preview-gated `slack_v2` block, so documenting it would
|
|
* publish an unreleased surface under its own `slack_app` page. Triggers no
|
|
* block claims are left alone: standalone webhook providers are legitimately
|
|
* unlisted and must keep their pages.
|
|
*/
|
|
async function collectPreviewOnlyTriggerIds(): Promise<Set<string>> {
|
|
const listedByReleased = new Set<string>()
|
|
const listedByPreview = new Set<string>()
|
|
|
|
const blockFiles = (await glob(`${BLOCKS_PATH}/*.ts`)).sort()
|
|
for (const blockFile of blockFiles) {
|
|
const fileContent = fs.readFileSync(blockFile, 'utf-8')
|
|
const exportRegex = /export\s+const\s+(\w+)Block\s*:\s*BlockConfig[^=]*=\s*\{/g
|
|
let match: RegExpExecArray | null
|
|
|
|
while ((match = exportRegex.exec(fileContent)) !== null) {
|
|
const startIndex = match.index + match[0].length - 1
|
|
const endIndex = findMatchingClose(fileContent, startIndex)
|
|
if (endIndex === -1) continue
|
|
|
|
const blockContent = fileContent.substring(startIndex, endIndex)
|
|
const available = extractArrayPropertyFromContent(blockContent, 'available')
|
|
if (!available?.length) continue
|
|
|
|
const target = isPreviewSource(blockContent) ? listedByPreview : listedByReleased
|
|
for (const triggerId of available) target.add(triggerId)
|
|
}
|
|
}
|
|
|
|
return new Set([...listedByPreview].filter((id) => !listedByReleased.has(id)))
|
|
}
|
|
|
|
async function generateAllTriggerDocs(): Promise<void> {
|
|
try {
|
|
console.log('Generating trigger documentation...')
|
|
|
|
if (!fs.existsSync(TRIGGER_DOCS_OUTPUT_PATH)) {
|
|
fs.mkdirSync(TRIGGER_DOCS_OUTPUT_PATH, { recursive: true })
|
|
}
|
|
|
|
const fullRegistry = await buildFullTriggerRegistry()
|
|
|
|
// A trigger reachable only through a preview-gated block is as unreleased
|
|
// as that block — documenting it publishes an unshipped surface (and, when
|
|
// its provider has no released block, mints a whole page for it).
|
|
const previewOnly = await collectPreviewOnlyTriggerIds()
|
|
for (const triggerId of previewOnly) {
|
|
if (fullRegistry.delete(triggerId)) {
|
|
console.log(`Skipping trigger ${triggerId} — only hosted by a preview-gated block`)
|
|
}
|
|
}
|
|
|
|
const grouped = groupTriggersByProvider(fullRegistry)
|
|
const colorMap = await buildProviderColorMap()
|
|
|
|
const generatedProviders: string[] = []
|
|
|
|
for (const [provider, triggers] of grouped) {
|
|
if (SKIP_TRIGGER_PROVIDERS.has(provider)) {
|
|
console.log(`Skipping trigger provider: ${provider} (covered by hand-written docs)`)
|
|
continue
|
|
}
|
|
|
|
// The trigger lives on the same per-service integration page as the
|
|
// service's actions (provider ≠ block type for a few services).
|
|
const blockType = PROVIDER_TO_BLOCK_TYPE[provider] ?? provider
|
|
const outputFilePath = path.join(DOCS_OUTPUT_PATH, `${blockType}.mdx`)
|
|
const baseName = path.basename(outputFilePath, '.mdx')
|
|
|
|
if (HANDWRITTEN_INTEGRATION_DOCS.has(baseName) || HANDWRITTEN_TRIGGER_DOCS.has(baseName)) {
|
|
console.log(`Skipping ${provider} — hand-written page`)
|
|
continue
|
|
}
|
|
|
|
const existing = fs.existsSync(outputFilePath)
|
|
? fs.readFileSync(outputFilePath, 'utf-8')
|
|
: null
|
|
|
|
if (existing?.includes('\n## Actions')) {
|
|
// Actions page generated this run by the block pass — append the Triggers section.
|
|
if (!existing.includes('\n## Triggers')) {
|
|
fs.appendFileSync(outputFilePath, `\n${buildTriggersSection(triggers)}`)
|
|
}
|
|
} else {
|
|
// Trigger-only service (no actions block) — (re)write the standalone page,
|
|
// preserving manual content from the previous run. Cleanup spares these
|
|
// pages (category 'triggers' blocks are in the canonical set).
|
|
const providerColor = colorMap.get(blockType) ?? '#6B7280'
|
|
const markdown = generateTriggerProviderDoc(provider, triggers, blockType, providerColor)
|
|
const rawSections = existing ? extractManualContent(existing) : {}
|
|
const manualSections = Object.fromEntries(
|
|
Object.entries(rawSections).filter(([, v]) => v.length > 0)
|
|
)
|
|
const finalContent =
|
|
Object.keys(manualSections).length > 0
|
|
? mergeWithManualContent(markdown, existing, manualSections)
|
|
: markdown
|
|
fs.writeFileSync(outputFilePath, finalContent)
|
|
}
|
|
|
|
generatedProviders.push(blockType)
|
|
console.log(
|
|
`✓ Triggers for ${formatTriggerProviderName(provider)} (${triggers.length} trigger${triggers.length === 1 ? '' : 's'})`
|
|
)
|
|
}
|
|
|
|
console.log(`✓ Trigger sections merged into ${generatedProviders.length} integration pages`)
|
|
} catch (error) {
|
|
console.error('Error generating trigger documentation:', error)
|
|
}
|
|
}
|
|
|
|
async function generateAllBlockDocs() {
|
|
try {
|
|
// Copy icons from sim app to docs app
|
|
copyIconsFile()
|
|
|
|
// Generate icon mappings from block definitions
|
|
const docsIconMapping = await generateIconMapping({ includeHidden: true })
|
|
const visibleIconMapping = await generateIconMapping({ includeHidden: false })
|
|
writeIconMapping(docsIconMapping)
|
|
|
|
// Generate landing integrations page data (JSON + icon mapping)
|
|
await writeIntegrationsJson(visibleIconMapping)
|
|
writeIntegrationsIconMapping(visibleIconMapping)
|
|
|
|
// Compute the canonical set of tool docs and clean up anything stale —
|
|
// covers hidden blocks AND blocks re-categorized away from `'tools'`.
|
|
const validToolDocs = await getCanonicalToolDocNames()
|
|
cleanupStaleToolDocs(validToolDocs)
|
|
|
|
const blockFiles = (await glob(`${BLOCKS_PATH}/*.ts`)).sort()
|
|
|
|
for (const blockFile of blockFiles) {
|
|
await generateBlockDoc(blockFile)
|
|
}
|
|
|
|
// Merge trigger sections into the per-service pages (and write trigger-only pages)
|
|
await generateAllTriggerDocs()
|
|
|
|
// Write the integrations meta after both passes so trigger-only pages are included
|
|
updateMetaJson()
|
|
|
|
return true
|
|
} catch (error) {
|
|
console.error('Error generating documentation:', error)
|
|
return false
|
|
}
|
|
}
|
|
|
|
function updateMetaJson() {
|
|
const metaJsonPath = path.join(DOCS_OUTPUT_PATH, 'meta.json')
|
|
|
|
const blockFiles = fs
|
|
.readdirSync(DOCS_OUTPUT_PATH)
|
|
.filter((file: string) => file.endsWith('.mdx'))
|
|
.map((file: string) => path.basename(file, '.mdx'))
|
|
|
|
const items = [
|
|
...(blockFiles.includes('index') ? ['index'] : []),
|
|
...blockFiles.filter((file: string) => file !== 'index').sort(),
|
|
]
|
|
|
|
const metaJson = {
|
|
pages: items,
|
|
}
|
|
|
|
fs.writeFileSync(metaJsonPath, `${JSON.stringify(metaJson, null, 2)}\n`)
|
|
console.log(`Updated meta.json with ${items.length} entries`)
|
|
}
|
|
|
|
generateAllBlockDocs()
|
|
.then((success) => {
|
|
if (success) {
|
|
console.log('Documentation generation completed successfully')
|
|
process.exit(0)
|
|
} else {
|
|
console.error('Documentation generation failed')
|
|
process.exit(1)
|
|
}
|
|
})
|
|
.catch((error) => {
|
|
console.error('Fatal error:', error)
|
|
process.exit(1)
|
|
})
|