test(ai-builder): Behaviour eval cases + authoring docs (no-changelog) (#32055)

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
José Braulio González Valido
2026-06-11 18:30:54 +00:00
committed by GitHub
co-authored by Claude Opus 4.8
parent aaece415e0
commit efc8e28967
11 changed files with 303 additions and 17 deletions
+36 -15
View File
@@ -199,7 +199,7 @@ Operational details:
### Build expectations (per test case)
A test case can declare optional natural-language assertions about *how the build went* — `buildExpectations: string[]` in its JSON. Each is graded by a separate Sonnet judge (`build-expectations/verifier.ts`) against the **conversation transcript + final workflow + conversation metrics**, and is **informational only** — it never affects `scenario_pass`, pass@k, or the score badge.
A test case can declare optional natural-language assertions about *how the build went* — `buildExpectations: string[]` in its JSON. Each is graded by a separate Sonnet judge (`build-expectations/verifier.ts`) against the **conversation transcript + final workflow + conversation metrics**, and **counts as a unit in the pass rate**: evaluated expectations fold into the per-case and headline pass@k/pass^k alongside execution scenarios. It doesn't flip an individual scenario's pass/fail (it's its own unit), and a judge `incomplete` verdict is excluded from the count.
Use it for things the binary checks and `successCriteria` don't cover:
@@ -216,13 +216,13 @@ Use it for things the binary checks and `successCriteria` don't cover:
The signal surfaces in:
- **HTML report** — a "Build expectations" disclosure on the test case: per-expectation &#10003;/&#10007; with a one-line judge reason.
- **`eval-results.json`** — `buildExpectationResultsPerRun` (per-iteration verdicts), so pass rates are queryable.
- **`eval-results.json`** — `buildExpectations` (aggregated per-expectation pass rate) plus `buildExpectationResultsPerRun` (per-iteration verdicts).
Operational details:
- Judged **once per build** (not per scenario), fired concurrently with the scenario batch — ~0 added wall-clock in the common case.
- Runs on both eval paths (direct loop + LangSmith). Requires a build transcript, so it's judged even when the build fails, and skipped only when no transcript was captured.
- The judge never affects scoring, retries on failure, has a per-attempt timeout, and falls back to an all-fail verdict — a judge failure can't break a run.
- The judge retries on failure, has a per-attempt timeout, and falls back to an all-fail verdict — a judge failure can't break a run.
- Absent the field, it's a complete no-op.
## Environment variables
@@ -516,31 +516,52 @@ When `LANGSMITH_API_KEY` is set, each run is recorded as a LangSmith experiment
## Adding test cases
Test cases live in `evaluations/data/workflows/*.json`. Drop a file in, the CLI and LangSmith sync picks it up — no registration step.
Test cases live in `evaluations/data/workflows/*.json`. Drop a file in — the CLI and LangSmith sync pick it up, no registration step. Every case is validated against `data/workflows/schema.ts`.
```json
{
"prompt": "Create a workflow that...",
"description": "Optional note on what this case checks.",
"conversation": [
{ "role": "user", "text": "Every morning, post a summary of yesterday's signups to Slack #growth." }
],
"complexity": "medium",
"tags": ["build", "webhook", "gmail"],
"triggerType": "webhook",
"scenarios": [
"tags": ["build", "schedule", "slack"],
"triggerType": "schedule",
"executionScenarios": [
{
"name": "happy-path",
"description": "Normal operation",
"dataSetup": "The webhook receives a submission from Jane (jane@example.com)...",
"successCriteria": "The workflow executes without errors. An email is sent to jane@example.com..."
"dataSetup": "The signups source returns 3 rows for yesterday: Ana, Ben, Cara.",
"successCriteria": "The workflow runs without errors and posts a summary of the 3 signups to Slack #growth."
}
]
}
```
**One JSON file = one LangSmith split.** Scenarios in the same file share a split; split names derive from the filename slug. Pick a slug you're happy to also use as a `--filter` target.
`conversation` (≥1 turn, first must be `user`) and `executionScenarios` (≥1), plus `complexity` and `tags`, are required. `description`, `triggerType`, `messageBudget`, `buildExpectations`, and `datasets` (default `["full"]`) are optional. A turn's `text` may be a string or an array of strings joined with newlines — handy for long stage directions.
**Prompt tips**
**One JSON file = one LangSmith split**, named from the filename slug. Pick a slug you're happy to also use as a `--filter` target.
- Be specific about node configuration — document IDs, sheet names, channel names, chat IDs. The agent won't ask for these in eval mode (no multi-turn yet).
- Add "Configure all nodes as completely as possible and don't ask me for credentials, I'll set them up later."
### Conversations & stage directions
`conversation` replaces the old single `prompt`; its mode is chosen automatically:
- **Single-prompt (auto-approve):** one `user` turn, no `assistant` turns — the prompt is sent and every confirmation is auto-approved. Use for plain build cases.
- **Multi-turn:** anything else. A user-proxy LLM plays the user — it answers questions, audits the agent's plan against the script, and sends follow-ups (capped by `messageBudget`). `assistant` turns are *reference* for the proxy (the expected flow); they're never sent to the builder.
Write the turns as a screenplay of what the user wants, keeping concrete values (channel IDs, schedules) verbatim. Inside a `user` turn, text in `[square brackets]` is a **stage direction** for the proxy — behaviour, not dialogue, never spoken to the builder. It overrides the proxy's defaults (e.g. "always answer"):
| To make the user… | Direction |
|---|---|
| Withhold a value until asked | `[Don't bring up the channel unless the agent asks where to post; then say 'Slack #growth.']` |
| Refuse and hold firm on re-ask | `[The user has no channel and won't provide one. If asked — question or setup card, even repeatedly — skip it; never invent one.]` |
| Keep the conversation going | `[After each change lands, send the next one from the list, one at a time, until done.]` |
A direction governs only what it covers; otherwise the proxy answers every question (inventing plausible placeholders) and never sets credentials. Setup cards (the "configure your workflow" card) are filled via the wizard — or dismissed when a direction withholds the value — not answered as questions.
**Prompt / conversation tips**
- Be specific about node configuration (IDs, sheet names, channel names). In single-prompt mode the agent won't ask; in multi-turn the proxy supplies or withholds per the script.
- If a built-in node doesn't expose a field you need (e.g. the Linear node doesn't query `creator.email`), tell the agent to use HTTP Request instead.
**Scenario tips**
@@ -558,7 +579,7 @@ Test cases live in `evaluations/data/workflows/*.json`. Drop a file in, the CLI
### Adding a new credential type
`credentials/seeder.ts` seeds generic creds (HTTP Header, HTTP Basic) on every run, plus env-gated creds (GitHub, Gmail, Teams, Linear…) when the matching env var is set. If your scenario needs a credential type that isn't there, add it to the appropriate list in `seeder.ts` — env-gated if it requires a real token, generic if a placeholder is fine.
`credentials/seeder.ts` seeds every credential with a placeholder token on every run — execution is mocked at the wire level, so the value is never used. Set the matching `EVAL_*_ACCESS_TOKEN` to override a service with a real token for a live run. If your scenario needs a credential type that isn't seeded, add it to `seeder.ts`.
## Failure categories
@@ -0,0 +1,25 @@
import { jsonParse } from 'n8n-workflow';
import { readdirSync, readFileSync } from 'node:fs';
import { join } from 'node:path';
import { WorkflowTestCaseSchema } from '../data/workflows/schema';
// Validates the REAL shipped test-case JSONs against the schema. `data-workflows-schema.test.ts`
// mocks `fs`, so it can't catch a malformed real case; this reads the actual directory and parses
// every case — including the multi-line array `text` form used by long stage directions.
const WORKFLOWS_DIR = join(__dirname, '..', 'data', 'workflows');
const caseFiles = readdirSync(WORKFLOWS_DIR).filter((f) => f.endsWith('.json'));
describe('case file schema validation', () => {
it('finds at least one case file', () => {
expect(caseFiles.length).toBeGreaterThan(0);
});
it.each(caseFiles)('%s parses against WorkflowTestCaseSchema', (file) => {
const raw = jsonParse<unknown>(readFileSync(join(WORKFLOWS_DIR, file), 'utf8'));
const result = WorkflowTestCaseSchema.safeParse(raw);
if (!result.success) {
throw new Error(`${file} failed schema validation:\n${result.error.message}`);
}
});
});
@@ -23,8 +23,6 @@
],
"messageBudget": 6,
"buildExpectations": [
"Before building, the agent asked which Airtable table and which Slack channel to use.",
"Before building, the agent also asked what each Slack message should contain.",
"The agent honored the user's explicit instruction to fetch the tasks via an HTTP Request node with Bearer auth, not the dedicated Airtable node.",
"The final workflow reflects the user's follow-up that Airtable returns a `{records: [...]}` envelope, splitting it into individual items (e.g. a Split Out node) before posting to Slack."
],
@@ -0,0 +1,46 @@
{
"description": "Long conversation: as the user asks for many small changes one at a time, the agent should apply each one when asked and end up with all of them — not fall a step behind or drop one.",
"conversation": [
{
"role": "user",
"text": "Every weekday at 9am, post a 'Daily standup starting soon' reminder to our Slack #standup channel."
},
{
"role": "assistant",
"text": "Got it — I'll build that now."
},
{
"role": "user",
"text": [
"[You have a list of changes to make to the workflow, one at a time. Each time the agent finishes applying a change, ask for the next one, then wait for it before asking for the following one. Ask for ONLY ONE change per message, in this exact order:",
"(1) move the time to 8:45am;",
"(2) only run on Monday, Wednesday and Friday;",
"(3) change the message text to 'Standup in 5 minutes';",
"(4) add an @here mention to that message;",
"(5) add a second Slack message 25 minutes later that says 'Standup now';",
"(6) also include today's date in the first message;",
"(7) add a step before the first post that builds the message text into a field, and have the post use that field;",
"(8) add error handling so that if a Slack post fails the workflow keeps going and posts a short note to a #alerts channel;",
"(9) make the second 'Standup now' message reply in the same Slack thread as the first message.",
"Keep going until all nine changes are made, then you can finish.]"
]
}
],
"messageBudget": 40,
"complexity": "complex",
"tags": ["behaviour", "changes-applied", "long-conversation", "schedule", "slack"],
"triggerType": "schedule",
"datasets": ["behaviour", "full"],
"buildExpectations": [
"By the end, the workflow reflects every change the user asked for: it runs at 8:45am, only on Monday/Wednesday/Friday; the first message's text is assembled in a preceding Set/Edit Fields step that the post references, reads 'Standup in 5 minutes', mentions @here, and includes today's date; there is a second message 25 minutes later that says 'Standup now' and replies in the first message's Slack thread; and if a Slack post fails the workflow continues and posts to a #alerts channel. Nothing the user asked for is missing.",
"After every change the user asked for, there is a before and after diff in the workflow that shows the change was applied. User request was accurately applied in the workflow."
],
"executionScenarios": [
{
"name": "standup-reminder",
"description": "Built on a weekday (later Mon/Wed/Fri) morning schedule that assembles a standup message, posts it to Slack #standup with a threaded follow-up, and falls back to #alerts on failure",
"dataSetup": "No external data is needed. Each Slack postMessage call returns { \"ok\": true, \"ts\": \"1700000000.000300\" }.",
"successCriteria": "The workflow is built correctly: a schedule trigger (8:45am, Mon/Wed/Fri after the user's changes) → a Set/Edit Fields step that builds the message → a Slack post to #standup carrying the text, @here mention and today's date, plus a second 'Standup now' post 25 minutes later that threads under the first; a Slack failure routes to a #alerts post. A successful mocked post is the ideal result; if the run trips on the channel format or an unconfigured value when it runs, treat that as an acceptable setup detail (not a build mistake) and a pass. Only judge the workflow's structure and how it runs; don't judge anything about the conversation."
}
]
}
@@ -0,0 +1,35 @@
{
"description": "#6/#8: A vague request should make the builder ask a clarifying question via the structured question tool and wait for the answer before building.",
"conversation": [
{
"role": "user",
"text": "Every morning, send my team a summary of yesterday's signups."
},
{
"role": "assistant",
"text": "Where should I send the summary — which service or channel?"
},
{
"role": "user",
"text": "[Don't bring up the channel unless the agent explicitly asks where to send the summary; then say 'Slack #growth.']"
}
],
"messageBudget": 4,
"complexity": "medium",
"tags": ["behaviour", "clarification", "question-tool", "schedule", "slack"],
"triggerType": "schedule",
"datasets": ["behaviour", "full"],
"buildExpectations": [
"Before building, the agent asked which destination to send the summary to, since the request did not specify one.",
"The agent asked using the structured question tool, not by posing the question in plain assistant text.",
"The agent waited for the user's answer before building."
],
"executionScenarios": [
{
"name": "happy-path",
"description": "Signups source returns a few rows; a summary is posted to Slack",
"dataSetup": "The signups data source returns 3 rows for yesterday: names Ana, Ben, Cara. The Slack postMessage call returns { \"ok\": true, \"ts\": \"1700000000.000100\" }.",
"successCriteria": "The workflow executes without errors and posts a summary mentioning the 3 signups (Ana, Ben, Cara) to Slack #growth."
}
]
}
@@ -0,0 +1,34 @@
{
"description": "If the user doesn't provide a detail the agent asked for, the agent should move on and not keep asking for it.",
"conversation": [
{
"role": "user",
"text": "Every Monday at 9am, post 'Have a great week, team!' to Slack."
},
{
"role": "assistant",
"text": "Which Slack channel should I post to?"
},
{
"role": "user",
"text": "[The user has no channel and won't provide one. Every time the agent asks for the channel — whether as a question or in a setup card, even if it asks more than once — skip or dismiss it without giving a channel, and never pick or invent one.]"
}
],
"messageBudget": 6,
"complexity": "medium",
"tags": ["behaviour", "clarification", "repeat-questions", "schedule", "slack"],
"triggerType": "schedule",
"datasets": ["behaviour", "full"],
"buildExpectations": [
"The agent asked for the user to provide a Slack channel, rather than silently picking one itself.",
"The user did not provide the channel, and then the agent moved on — leaving it blank or for setup — never asking for the channel again."
],
"executionScenarios": [
{
"name": "channel-skipped",
"description": "Built on a Monday schedule posting to Slack; the user didn't give a channel, so it's left for setup",
"dataSetup": "No mock data is needed. If the Slack node's channel is empty, the workflow fails its validation check before any node runs, so no API call is made.",
"successCriteria": "The workflow is built correctly: a Monday 9am schedule trigger feeding a Slack message node with the text 'Have a great week, team!'. If the Slack channel is empty, the workflow fails its validation check ('Not a valid Slack Channel ID or name') — treat that as fine (an empty channel is something to set up later, not a build mistake) and a pass. If a channel is set and the run succeeds, that's also a pass. Only judge the workflow's structure and how it handles the channel when it runs; don't judge anything about the conversation."
}
]
}
@@ -0,0 +1,42 @@
{
"description": "If the user says they'll give a detail later, the agent should build the workflow and stop asking for that detail.",
"conversation": [
{
"role": "user",
"text": "Every Monday at 9am, post the message 'Have a great week, team!' to the team."
},
{
"role": "assistant",
"text": "Which service should I post it through?"
},
{
"role": "user",
"text": "Slack, I don't know which channel yet I'll tell you later."
},
{
"role": "assistant",
"text": "Which Slack channel should I post to?"
},
{
"role": "user",
"text": "[The user doesn't answer this one — they don't give a channel and don't tell the agent to pick one. Just skip the question.]"
}
],
"messageBudget": 6,
"complexity": "medium",
"tags": ["behaviour", "clarification", "repeat-questions", "schedule", "slack"],
"triggerType": "schedule",
"datasets": ["behaviour", "full"],
"buildExpectations": [
"Before building, the agent asked which app to use, since the user hadn't said which one.",
"After the user said they'd give the channel later, the agent built the workflow — leaving the channel blank or setting it up later — instead of asking for the channel again and again, and it told the user they can fill the channel in after building."
],
"executionScenarios": [
{
"name": "channel-skipped",
"description": "Built on a Monday schedule posting to Slack; the user didn't give a channel, so the workflow fails its validation check",
"dataSetup": "No mock data is needed. If the Slack node's channel is empty, the workflow fails its validation check before any node runs, so no API call is made.",
"successCriteria": "The workflow is built correctly: a Monday 9am schedule trigger feeding a Slack message node with the text 'Have a great week, team!'. If the Slack channel is empty, the workflow fails its validation check ('Not a valid Slack Channel ID or name') — treat that as fine (an empty channel is something to set up later, not a build mistake) and a pass. If a channel is set and the run succeeds, that's also a pass. Only judge the workflow's structure and how it handles the channel when it runs; don't judge anything about the conversation."
}
]
}
@@ -0,0 +1,34 @@
{
"description": "The builder should fetch all pages using the HTTP Request node's built-in pagination, not a single request, a Code-node loop, or a hard-coded page count.",
"conversation": [
{
"role": "user",
"text": "On a manual trigger, fetch the full list of items from https://api.example.com/items and return them all."
},
{
"role": "assistant",
"text": "I'll set that up with a Manual Trigger and an HTTP Request node that GETs the items."
},
{
"role": "user",
"text": "Make sure you get all of them — that endpoint has hundreds of items across many pages, so it needs to page through to the end, not just the first page."
}
],
"messageBudget": 6,
"complexity": "medium",
"tags": ["behaviour", "http-request", "pagination"],
"triggerType": "manual",
"datasets": ["behaviour", "full"],
"buildExpectations": [
"The HTTP Request node fetches across all pages (e.g. via its built-in pagination), not just the first page.",
"The agent did NOT work around paging with a Code node loop, a single request, or a hard-coded number of requests."
],
"executionScenarios": [
{
"name": "spans-two-pages",
"description": "Items are split across two pages; full pagination returns 4 records, a single request returns only 2",
"dataSetup": "This is a PAGINATED endpoint — serve each request strictly by its `page` query parameter and do NOT put every item on one page. `GET /items?page=1` returns {\"items\":[{\"id\":1,\"name\":\"Alpha\"},{\"id\":2,\"name\":\"Beta\"}]}. `GET /items?page=2` returns {\"items\":[{\"id\":3,\"name\":\"Gamma\"},{\"id\":4,\"name\":\"Delta\"}]}. `GET /items?page=3` and any later page return {\"items\":[]} to signal the end. A request with no page parameter is page 1.",
"successCriteria": "The workflow executes without errors and returns all FOUR item records — Alpha, Beta, Gamma and Delta — having fetched page 1 and page 2 before the empty page 3 stops paging. A workflow that issues a single request (or only reads page 1) returns just Alpha and Beta and FAILS. Judge by the item records present in the output, not by a single numeric count field."
}
]
}
@@ -0,0 +1,25 @@
{
"description": "#4: The builder should verify with mocked data for a credentialed node, not refuse to test until a real credential is added.",
"conversation": [
{
"role": "user",
"text": "Every morning at 9am, fetch yesterday's invoices from Stripe and return a one-line total of the amounts due."
}
],
"complexity": "medium",
"tags": ["behaviour", "credentials", "stripe", "verification"],
"triggerType": "schedule",
"datasets": ["behaviour", "full"],
"buildExpectations": [
"The agent ran a verification of the built workflow (a verification execution was attempted, not skipped).",
"During verification the credentialed Stripe node was served mocked/pinned data rather than requiring a real Stripe credential to run — the agent did not refuse to test until a credential was added."
],
"executionScenarios": [
{
"name": "happy-path",
"description": "Stripe returns a couple of invoices; a total is produced",
"dataSetup": "The Stripe invoices endpoint returns 2 invoices with amount_due values of 1200 and 800.",
"successCriteria": "The workflow executes without errors and returns a one-line total reflecting the two invoices (e.g. 2000, or 20.00 if interpreted as cents). A schedule trigger is present."
}
]
}
@@ -0,0 +1,25 @@
{
"description": "#5: Network requests should be made by an HTTP Request node, not by fetch/axios/http calls inside a Code node.",
"conversation": [
{
"role": "user",
"text": "On a manual trigger, fetch the latest releases from https://api.example.com/releases and return a short changelog that combines each release's name and date."
}
],
"complexity": "simple",
"tags": ["behaviour", "code-node", "http-request"],
"triggerType": "manual",
"datasets": ["behaviour", "full"],
"buildExpectations": [
"The network request was performed by an HTTP Request node, not by an HTTP/fetch call inside a Code node.",
"No Code node in the final workflow attempts a network/HTTP request (e.g. fetch, axios, http/https, $http)."
],
"executionScenarios": [
{
"name": "two-releases",
"description": "The releases endpoint returns two releases; a changelog is produced",
"dataSetup": "GET https://api.example.com/releases returns [ { \"name\": \"v2.1\", \"date\": \"2026-05-01\" }, { \"name\": \"v2.2\", \"date\": \"2026-06-01\" } ].",
"successCriteria": "The workflow executes without errors and returns a changelog string referencing both releases (v2.1 / 2026-05-01 and v2.2 / 2026-06-01). The HTTP request is performed by an HTTP Request node."
}
]
}
@@ -50,6 +50,7 @@ Each branch's items are capped at 10 for artifact size. The full untruncated tot
- Did a real node crash because a field is missing? → **check the request that was sent**: if the HTTP request (e.g., GraphQL query) didn't ask for that field, the mock correctly omitted it — that's a builder issue (wrong query or wrong node choice), NOT a mock issue. The mock can only return what was requested.
- Did the mock response have the wrong shape for the endpoint? (e.g., returning a write response for a GET request) → mock issue
- Did the mock return identical responses for multiple calls to the same endpoint with different request bodies? → mock issue
- Did the workflow error with n8n's pagination safety (e.g. "The returned response was identical 5x, so requests got stopped")? → builder_issue: the pagination did not terminate — it failed to stop on the empty page, or never advanced the page parameter. Identical empty pages at end-of-data are the correct stop signal (the builder must detect them), and the mock serving distinct pages in sequence to repeated requests is the testing mechanism working. Only a mock_issue if the mock repeated identical non-empty pages it should have varied.
- Did the workflow handle an error scenario but the success criteria is ambiguous about what "graceful" means? → evaluate based on whether data was lost or the workflow crashed entirely
KEY PRINCIPLE: A mock response that faithfully matches the HTTP request is NEVER a mock issue, even if downstream nodes needed different data. If the request didn't ask for a field, the mock shouldn't invent it. The fault lies with whatever built the request (the node choice or its configuration).