fix(providers): final regenerated stream must not re-call tools (empty chat answers) (#5915)

* fix(providers): stop the final regenerated stream from re-calling tools and clobbering the answer

After the silent tool loop settles, OpenAI Responses and Gemini re-issue a
streaming request purely to stream the answer as prose — but with tools still
attached and auto tool choice, a reasoning model can re-decide to call a tool
there. Streamed calls are never executed on this path, so the run ends with a
dead function call, an empty streamed answer, and the stream callback
clobbering the tool loop's settled text with '' (deployed chat rendered
{"content": ""}).

Force tool_choice 'none' / functionCallingConfig NONE on the regeneration and
keep the tool loop's settled answer whenever the stream ends without text.

* feat(openai): settled tool chips on the regenerated answer stream

The silent Responses tool loop has no live stream while tools run, so opted-in
consumers saw no tool chips at all for OpenAI. The loop now records each
executed call and prepends settled tool_call_start/end pairs (name + status
only) to the agent-events stream ahead of the regenerated answer. Runs without
a sink never see these events, so legacy output is unchanged.

* fix(providers): apply the regeneration guard fleet-wide

Audit of every silent tool loop for the same race fixed for OpenAI/Gemini
(final regenerated stream re-calls a tool that is never executed, ending with
an empty answer that clobbers the settled text):

- anthropic (both implementations): tool_choice {type:'none'} on the
  regeneration (tools must stay — history carries tool_use blocks) + keep the
  settled answer when the stream ends without text
- groq: was re-applying the ORIGINAL tool_choice, so forced-tool runs
  re-forced the tool on the regeneration — guaranteed dead call; now 'none'
  + fallback
- deepseek, mistral, cerebras, azure-openai (legacy chat path), openrouter,
  xai: 'auto' -> 'none' + fallback
- bedrock: fallback only — Bedrock's ToolChoice has no 'none' and toolConfig
  is required when history carries toolUse blocks

Already guarded (no change): meta, sakana, nvidia, vllm, litellm, baseten,
together, fireworks, kimi, zai, ollama.
This commit is contained in:
Vikhyath Mondreti
2026-07-23 20:20:25 -07:00
committed by GitHub
parent bee45621ec
commit 7120cdc901
13 changed files with 291 additions and 56 deletions
@@ -80,7 +80,7 @@ Per-model support is generated from the model registry on the [Agent block page]
|--------|----------|------------|
| Anthropic / Azure Anthropic | Yes (incl. redacted blocks in traces). The newest Claude generations omit full thinking; Sim requests summarized thinking for them on streaming runs | Yes |
| Gemini / Vertex | Yes when a thinking level is set (thought summaries requested on agent-events runs) | Yes |
| OpenAI Responses | Reasoning **summaries** when streamed (requires OpenAI organization verification; unverified orgs fall back to no summaries) | Silent tool loop today (answer streams; chips may be absent) |
| OpenAI Responses | Reasoning **summaries** when streamed (requires OpenAI organization verification; unverified orgs fall back to no summaries) | Silent tool loop; tool chips arrive settled (start + end together) once tools finish, ahead of the streamed answer — no live in-progress chips yet |
| OpenAI-compat (Groq, DeepSeek, …) | Only if vendor streams `reasoning` / `reasoning_content` | Live loop where wired (e.g. Groq, DeepSeek) |
| Bedrock | Not invented | Yes when streaming tool loop is used |