fix(wand): stop markdown code fences landing in generated code (#6289)

* fix(wand): stop markdown code fences landing in generated code

Strip fences from wand output for raw-value generation types, reset the
conversation history when a Function block switches language, and give the
Python prompt the worked example the JavaScript one already had.

* fix(wand): preserve nested fences and retire stale history on reset

Slice only the outermost fence delimiters so a fenced body containing
line-leading backticks keeps every interior line, and skip the history
append when a language reset retired the request mid-flight.

* refactor(wand): drop unreachable guard in fence stripper

The trimStart check made the -1 branch dead and stated "opens with a
fence" twice. Derive it once from the first fence line's position.

* chore(wand): remove unused onGenerationComplete callback

No call site ever passed it, so the branch never ran. The props interface
makes the removal compile-time verified.

* fix(wand): never treat an interior fence line as the closer

A truncated response whose body embeds line-leading backticks lost every
line after the first embedded delimiter. Only the opening line and a
final fence line are removed now. Sync the history epoch in a layout
effect so a request settling before the passive flush cannot append to
already-reset history.

* test(wand): record the trailing-fence ambiguity as a decision

A bare fence on the last line closes the wrapper in every well-formed
response and is content only when generation stopped exactly on an
embedded delimiter. Nothing separates the two, so assert the chosen
behavior instead of leaving it implicit.
This commit is contained in:
Waleed
2026-08-05 11:29:50 -07:00
committed by GitHub
parent 0583e2085f
commit 2d90c729c5
5 changed files with 309 additions and 15 deletions
+7
View File
@@ -328,6 +328,13 @@ export const POST = withRouteHandler(async (req: NextRequest) => {
'\n\nIMPORTANT: Return ONLY the raw cron expression (e.g., "0 9 * * 1-5"). Do NOT wrap it in markdown code blocks, backticks, or quotes. Do NOT include any explanation or text before or after the expression.'
}
// Both the JavaScript and Python function-body prompts share this type, so
// the reinforcement stays language-neutral.
if (generationType === 'javascript-function-body') {
finalSystemPrompt +=
'\n\nIMPORTANT: Return ONLY the raw function body. Do NOT wrap it in markdown code blocks (no ```javascript, no ```python, no ```). Do NOT include any explanation before or after the code.'
}
if (generationType === 'json-object') {
finalSystemPrompt +=
'\n\nIMPORTANT: Return ONLY the raw JSON object. Do NOT wrap it in markdown code blocks (no ```json or ```). Do NOT include any explanation or text before or after the JSON. The response must start with { and end with }.'
@@ -67,9 +67,32 @@ IMPORTANT FORMATTING RULES:
1. Reference Environment Variables: Use the exact syntax {{VARIABLE_NAME}}. Do NOT wrap it in quotes.
2. Reference Input Parameters/Workflow Variables: Use the exact syntax <variable_name>. Do NOT wrap it in quotes.
3. Function Body ONLY: Do NOT include the function signature (e.g., 'def my_func(...)') or surrounding braces. Return the final value with 'return'.
4. Imports: You may add imports as needed (standard library or pip-installed packages) without comments.
4. Imports: The Python standard library is always available. Third-party packages are available ONLY when the block has a sandbox selected — the sandbox's package list is appended below when one is. Never import a package that is not on that list.
5. No Markdown: Do NOT include backticks, code fences, or any markdown.
6. Clarity: Write clean, readable Python code.`
6. Clarity: Write clean, readable Python code.
7. No Explanations: Output the raw Python code only — no prose before or after it.
Example Scenario:
User Prompt: "Fetch user data from an API. Use the User ID passed in as 'userId' and an API Key stored as the 'SERVICE_API_KEY' environment variable."
Generated Code:
import json
import urllib.error
import urllib.request
user_id = <userId> # Correct: accessing an input parameter without quotes
api_key = {{SERVICE_API_KEY}} # Correct: accessing an environment variable without quotes
url = f"https://api.example.com/users/{user_id}"
request = urllib.request.Request(url, headers={"Authorization": f"Bearer {api_key}"})
try:
with urllib.request.urlopen(request) as response:
# Return the fetched data, which becomes the block's output
return json.loads(response.read().decode())
except urllib.error.HTTPError as error:
# Raising marks the block execution as failed
raise Exception(f"API request failed with status {error.code}: {error.read().decode()}")`
/**
* Line height constant for consistent rendering.
@@ -330,6 +353,9 @@ export const Code = memo(function Code({
tableId: typeof tableIdValue === 'string' ? tableIdValue : null,
sandboxId: typeof sandboxIdValue === 'string' ? sandboxIdValue : null,
},
// Keyed off the same value that swaps the prompt below, so history from the
// previous language cannot steer the next generation back to it.
historyResetKey: typeof languageValue === 'string' ? languageValue : undefined,
onStreamStart: () => handleStreamStartRef.current?.(),
onStreamChunk: (chunk: string) => handleStreamChunkRef.current?.(chunk),
onGeneratedContent: (content: string) => handleGeneratedContentRef.current?.(content),
@@ -1,4 +1,4 @@
import { useCallback, useRef, useState } from 'react'
import { useCallback, useLayoutEffect, useRef, useState } from 'react'
import { toast } from '@sim/emcn'
import { createLogger } from '@sim/logger'
import { filterUndefined } from '@sim/utils/object'
@@ -8,6 +8,7 @@ import { requestRaw } from '@/lib/api/client'
import { isApiClientError } from '@/lib/api/client/errors'
import { wandGenerateStreamContract } from '@/lib/api/contracts'
import { readSSEStream } from '@/lib/core/utils/sse'
import { shouldStripCodeFences, stripCodeFences } from '@/lib/wand/strip-code-fences'
import type { GenerationType } from '@/blocks/types'
import { subscriptionKeys } from '@/hooks/queries/subscription'
import { useSettingsNavigation } from '@/hooks/use-settings-navigation'
@@ -100,20 +101,26 @@ interface UseWandProps {
wandConfig?: WandConfig
currentValue?: string
contextParams?: WandContextParams
/**
* Clears the conversation history whenever this value changes. Pass anything
* that invalidates prior turns — a Function block switching language rewrites
* `wandConfig.prompt`, but replayed history would keep steering the model back
* to the previous language.
*/
historyResetKey?: string
onGeneratedContent: (content: string) => void
onStreamChunk?: (chunk: string) => void
onStreamStart?: () => void
onGenerationComplete?: (prompt: string, generatedContent: string) => void
}
export function useWand({
wandConfig,
currentValue,
contextParams,
historyResetKey,
onGeneratedContent,
onStreamChunk,
onStreamStart,
onGenerationComplete,
}: UseWandProps) {
const queryClient = useQueryClient()
const { navigateToSettings } = useSettingsNavigation()
@@ -127,6 +134,35 @@ export function useWand({
const [conversationHistory, setConversationHistory] = useState<ChatMessage[]>([])
/**
* Adjusted during render rather than in an effect so a generation started in
* the same commit as the change can never send the stale history. History is
* already empty on mount, so seeding the tracker with the current key
* correctly makes the first render a no-op.
*/
const [prevHistoryResetKey, setPrevHistoryResetKey] = useState(historyResetKey)
const [historyEpoch, setHistoryEpoch] = useState(0)
if (prevHistoryResetKey !== historyResetKey) {
setPrevHistoryResetKey(historyResetKey)
setConversationHistory([])
setHistoryEpoch((epoch) => epoch + 1)
}
/**
* Mirrors {@link historyEpoch} for the in-flight request to read on completion.
* A request that started before a reset must not append its turn to the fresh
* history — its prompt and reply belong to the superseded context.
*
* Synced in a layout effect, not a passive one: passive effects flush in a later
* task, so a request settling between the reset's commit and that flush would
* still read the old epoch and append anyway. Layout effects run synchronously
* during commit, before any promise continuation can observe the ref.
*/
const historyEpochRef = useRef(historyEpoch)
useLayoutEffect(() => {
historyEpochRef.current = historyEpoch
}, [historyEpoch])
const abortControllerRef = useRef<AbortController | null>(null)
const showPromptInline = useCallback(() => {
@@ -171,6 +207,9 @@ export function useWand({
setError(null)
setPromptInputValue('')
/** The context this request belongs to; a reset while it streams retires it. */
const startedHistoryEpoch = historyEpochRef.current
abortControllerRef.current = new AbortController()
if (onStreamStart) {
@@ -224,25 +263,37 @@ export function useWand({
signal: abortControllerRef.current?.signal,
})
if (accumulatedContent) {
onGeneratedContent(accumulatedContent)
/**
* Sanitized once the full response is known, then written back over the
* streamed text. Doing it per-chunk would mean guessing whether a
* trailing backtick run opens a fence or is part of the code, so the
* editor may briefly show a fence that the final value does not.
*/
const generatedContent = shouldStripCodeFences(wandConfig?.generationType)
? stripCodeFences(accumulatedContent)
: accumulatedContent
if (wandConfig?.maintainHistory) {
if (generatedContent) {
onGeneratedContent(generatedContent)
/**
* The sanitized form goes into history so a single fenced reply cannot
* become the in-context example for every later turn. Skipped entirely
* when a reset retired this request's context mid-flight.
*/
if (wandConfig?.maintainHistory && historyEpochRef.current === startedHistoryEpoch) {
setConversationHistory((prev) => [
...prev,
{ role: 'user', content: currentPrompt },
{ role: 'assistant', content: accumulatedContent },
{ role: 'assistant', content: generatedContent },
])
}
if (onGenerationComplete) {
onGenerationComplete(currentPrompt, accumulatedContent)
}
}
logger.debug('Wand generation completed', {
prompt,
contentLength: accumulatedContent.length,
contentLength: generatedContent.length,
strippedFences: generatedContent !== accumulatedContent,
})
setTimeout(() => {
@@ -282,7 +333,6 @@ export function useWand({
onGeneratedContent,
onStreamChunk,
onStreamStart,
onGenerationComplete,
queryClient,
contextParams?.tableId,
contextParams?.sandboxId,
+110
View File
@@ -0,0 +1,110 @@
/**
* @vitest-environment node
*/
import { describe, expect, it } from 'vitest'
import { shouldStripCodeFences, stripCodeFences } from '@/lib/wand/strip-code-fences'
describe('stripCodeFences', () => {
it('leaves unfenced code untouched', () => {
const code = 'const total = <a> + <b>;\nreturn total;'
expect(stripCodeFences(code)).toBe(code)
})
it('unwraps a fully wrapped response', () => {
expect(stripCodeFences('```python\nresult = <num1> + <num2>\nreturn result\n```')).toBe(
'result = <num1> + <num2>\nreturn result'
)
})
it('unwraps a response with no closing fence', () => {
expect(stripCodeFences('```javascript\nconst x = 1;\nreturn x;')).toBe(
'const x = 1;\nreturn x;'
)
})
it('unwraps an untagged fence', () => {
expect(stripCodeFences('```\nreturn 1;\n```')).toBe('return 1;')
})
it('tolerates leading whitespace before the opening fence', () => {
expect(stripCodeFences('\n ```python\nreturn 1\n```')).toBe('return 1')
})
it('preserves indentation inside the fence', () => {
const fenced = '```python\nif <flag>:\n return "yes"\nreturn "no"\n```'
expect(stripCodeFences(fenced)).toBe('if <flag>:\n return "yes"\nreturn "no"')
})
it('preserves fence lines embedded inside the fenced body', () => {
const fenced = '```javascript\nconst md = `\n```\nhello\n```\n`;\nreturn md;\n```'
expect(stripCodeFences(fenced)).toBe('const md = `\n```\nhello\n```\n`;\nreturn md;')
})
it('keeps every line when a body with nested fences is truncated mid-response', () => {
const truncated = '```javascript\nconst md = `\n```\nhello\n`;\nreturn md;'
expect(stripCodeFences(truncated)).toBe('const md = `\n```\nhello\n`;\nreturn md;')
})
it('treats a trailing bare fence as the closer even when the body was truncated at one', () => {
// Irreducibly ambiguous: a trailing bare fence closes the wrapper in every
// well-formed response, and is content only when generation stopped exactly
// at an embedded delimiter. Declining to strip it would leave a stray fence
// in the common case, which is the bug this util exists to fix.
expect(stripCodeFences('```javascript\nconst md = `\n```')).toBe('const md = `')
})
it('preserves a fenced docstring inside a Python body', () => {
const fenced = '```python\ntemplate = """\n```sql\nSELECT 1\n```\n"""\nreturn template\n```'
expect(stripCodeFences(fenced)).toBe(
'template = """\n```sql\nSELECT 1\n```\n"""\nreturn template'
)
})
it('keeps everything between the outer delimiters for a multi-block answer', () => {
// Prose survives rather than risk dropping code between two delimiters that
// may be a nested literal instead of a block boundary.
const fenced = '```js\nconst a = 1;\n```\nThen send it:\n```js\nreturn a;\n```'
expect(stripCodeFences(fenced)).toBe('const a = 1;\n```\nThen send it:\n```js\nreturn a;')
})
it('does not touch code that merely contains a fence later', () => {
const code = 'const doc = `\n```json\n{"a":1}\n```\n`;\nreturn doc;'
expect(stripCodeFences(code)).toBe(code)
})
it('returns the original when stripping would leave nothing', () => {
const empty = '```python\n```'
expect(stripCodeFences(empty)).toBe(empty)
})
it('is idempotent', () => {
const once = stripCodeFences('```python\nreturn <x>\n```')
expect(stripCodeFences(once)).toBe(once)
})
it('handles an empty string', () => {
expect(stripCodeFences('')).toBe('')
})
})
describe('shouldStripCodeFences', () => {
it('strips for code and structured value types', () => {
expect(shouldStripCodeFences('javascript-function-body')).toBe(true)
expect(shouldStripCodeFences('custom-tool-schema')).toBe(true)
expect(shouldStripCodeFences('json-object')).toBe(true)
expect(shouldStripCodeFences('cron-expression')).toBe(true)
})
it('does not strip free-form prose', () => {
expect(shouldStripCodeFences('system-prompt')).toBe(false)
})
it('does not strip when no generation type is declared', () => {
expect(shouldStripCodeFences(undefined)).toBe(false)
expect(shouldStripCodeFences('')).toBe(false)
})
it('does not strip an unrecognized type', () => {
expect(shouldStripCodeFences('something-new')).toBe(false)
})
})
+101
View File
@@ -0,0 +1,101 @@
import type { GenerationType } from '@/blocks/types'
/** A markdown fence delimiter at the start of a line, ignoring indentation. */
const FENCE_LINE = /^\s*```/
/**
* Whether a wand generation's output is a raw machine value, where a leading
* markdown fence is always wrong and must be removed.
*
* Declared as a total `Record` so adding a `GenerationType` fails the build
* until the new type opts in or out deliberately — a silent default would let a
* prose type start stripping fences (or a code type stop) without review.
*
* `system-prompt` is the sole exclusion: it is free-form prose for a model, so a
* fenced example inside it is legitimate authored content, not a formatting slip.
*/
const STRIPS_CODE_FENCES: Record<GenerationType, boolean> = {
'javascript-function-body': true,
'typescript-function-body': true,
'json-schema': true,
'json-object': true,
'table-schema': true,
'system-prompt': false,
'custom-tool-schema': true,
'sql-query': true,
postgrest: true,
'mongodb-filter': true,
'mongodb-pipeline': true,
'mongodb-sort': true,
'mongodb-documents': true,
'mongodb-update': true,
'neo4j-cypher': true,
'neo4j-parameters': true,
timestamp: true,
timezone: true,
'cron-expression': true,
'odata-expression': true,
}
/**
* Whether generated content for this type should have markdown fences stripped.
*
* An absent type means the field's `wandConfig` never declared one, which is the
* case for free-form prose fields — those are left untouched.
*/
export function shouldStripCodeFences(generationType?: string): boolean {
if (!generationType) return false
return STRIPS_CODE_FENCES[generationType as GenerationType] === true
}
/**
* Removes the markdown code fences a model wrapped around a raw value.
*
* Applies only when the response *opens* with a fence. Content that merely
* contains a fence later is left untouched, because a backtick run inside a
* template literal or a docstring is valid code that must survive verbatim —
* a false positive here would corrupt working code, which is far worse than
* leaving a rare unwrapped response for the user to fix.
*
* Only two lines can ever be removed: the opening fence, and the final line when
* it is also a fence. An interior fence line is always treated as content, because
* a generated body may legitimately contain line-leading backticks (code that
* builds a markdown string) and there is no way to tell that apart from a
* delimiter. Scanning for the *last* fence anywhere would truncate such a body
* whenever the response is cut off before its closing fence.
*
* The cost is that a model which answers with several fenced blocks and prose
* between them keeps that prose — visibly wrong output the user can re-roll,
* rather than code quietly missing a chunk.
*
* Falls back to the original text if stripping would leave nothing.
*/
export function stripCodeFences(text: string): string {
const lines = text.split('\n')
const openingFence = lines.findIndex((line) => FENCE_LINE.test(line))
if (openingFence === -1) return text
// Anything non-blank ahead of the first fence means the response does not open
// with one, so the backticks belong to the content.
for (let index = 0; index < openingFence; index++) {
if (lines[index].trim() !== '') return text
}
const inner = lines.slice(openingFence + 1)
// Trim blank lines only — leading whitespace on a kept line is indentation,
// which is load-bearing in Python.
while (inner.length > 0 && inner[inner.length - 1].trim() === '') inner.pop()
// Only the very last line may close the wrapper. A truncated response simply
// has no closer, and every line after the opener survives. The one case this
// cannot get right is generation stopping exactly on an embedded delimiter,
// where that final line is content — indistinguishable from a real closer, and
// rarer than the wrap it would otherwise fail to strip.
if (inner.length > 0 && FENCE_LINE.test(inner[inner.length - 1])) inner.pop()
while (inner.length > 0 && inner[0].trim() === '') inner.shift()
while (inner.length > 0 && inner[inner.length - 1].trim() === '') inner.pop()
return inner.length > 0 ? inner.join('\n') : text
}