mirror of
https://github.com/coder/coder.git
synced 2026-09-24 06:47:27 +08:00
27ed052d86b5c8a630ccaff45ca39632da2c419c
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d6daf273aa |
fix: tighten single tool result byte budget (#26763)
## Problem #26637 caps each locally-executed tool result (built-in, global/deployment MCP, workspace MCP) at a per-result byte budget derived from the model's context window. The budget was `ContextLimit/2 * 4 bytes` — i.e. **half the window at an optimistic 4 bytes/token**. On large-context models that is far too generous. With a `1,000,000`-token `ContextLimit` the per-result cap is **~2 MB**. A user hit exactly this with a chatd (deployment-pinned) MCP tool: the result was truncated to **1,998,709 characters** and still overflowed the prompt. 2 MB of dense text (JSON/logs/code) is ~650k–1M tokens — most or all of the window for a *single* result — so the cap fired but didn't actually prevent the overflow. ## Fix Tighten the two budget constants in `tooltruncate.go`: | constant | before | after | | --- | --- | --- | | `toolResultContextDivisor` | `2` (½ window) | `3` (⅓ window) | | `bytesPerTokenEstimate` | `4` | `3` (conservative) | The budget becomes `ContextLimit/3 * 3 ≈ ContextLimit` bytes: | ContextLimit | before | after | | --- | --- | --- | | 1,000,000 | ~2 MB | ~1 MB | | 200,000 | ~400 KB | ~200 KB | | unknown (≤0) | 64 KB | 64 KB (unchanged) | The 16 KB floor and 64 KB unknown-window default are unchanged. A conservative bytes-per-token estimate is intentional: dense payloads run well under 4 B/tok, so a lower estimate yields a smaller byte budget that is less likely to underestimate the true token cost. No behavioral code paths change — only the two constants and their doc comments. The existing `tooltruncate_internal_test.go` cases derive their expectations from the constants (`LargeWindow`) or exercise the floor/default (`BelowFloor`, `Unknown`), so they remain green. <details> <summary>Investigation notes</summary> Global/deployment MCP tools (`mcpclient.ConnectAll`) are appended to `prepared.Tools` and execute locally via `ExecuteLocalTools → executeTools → executeSingleTool`, so the #26637 cap *does* apply to them for text results (`convertCallResult` joins text content into `resp.Content`). The cap was simply too large: `toolResultByteBudget(ContextLimit)` = `ContextLimit/2*4` ≈ 2 MB for a 1M-token window. Reverse-engineering the reported `1,998,709` truncated characters confirms a `ContextLimit` of ~1,000,000 tokens. Known gaps left for follow-ups (out of scope here): - **Per-step aggregate is unbounded.** MCP tools advertise `Parallel: true` and `executeSingleTool` caps each result independently, so N parallel calls in one step can sum to N × the per-result cap. - **Binary/media `Data` bypasses the cap.** Only the text payload is bounded; `image`/`media`/blob embedded-resource results are base64-encoded untouched in `executeSingleTool`. - **Compaction is reactive.** It is gated on the prior step's reported usage (`latestPromptUsage`), so it can't pre-empt a single large result appended on the current step. </details> --- Generated by Coder Agents on behalf of @kylecarbs. |
||
|
|
32217259b7 |
feat: cap tool output to fit the model context window (#26637)
## Problem
Local tool results were persisted and replayed to the model verbatim,
with no size cap. A single oversized result, most often a multi-megabyte
response from an MCP tool, overflows the prompt on the next request.
Every retry rebuilds the same history and fails the same way, leaving
the chat wedged in `error`. Auto-compaction is reactive (token usage is
only known after a response), so it can't catch a single result that
blows the very next request.
## Fix
Cap every locally-executed tool result at its single choke point,
`executeSingleTool` in `chatloop`, so the cap covers built-in tools,
**global (deployment-pinned) MCP**, **workspace MCP**, and provider
runners uniformly. Because this runs before the result is published to
the live stream and before it is committed, the SSE preview, the
persisted message, and the model replay all see the same bounded output.
The budget is token-aware: a single tool result may use at most half the
model's context window (`~4 bytes/token`), with a `16KB` floor and a
`64KB` default when the window is unknown. Truncation keeps the head and
tail of the output and replaces the middle with a marker telling the
model how much was removed and to narrow its query; it is UTF-8 safe and
never exceeds the budget. Binary media `Data` is passed through
untouched (only the text payload is bounded).
A `coderd_chatd_tool_result_truncated_total{provider,model,tool_name}`
counter and a warning log record each truncation.
## Out of scope
- Provider-executed results (e.g. web search) arrive via the stream, not
`executeSingleTool`.
- Dynamic/external tool results submitted through the `/tool-results`
API are validated as JSON elsewhere.
- Cumulative growth across many results is still handled by context
compaction; this change only bounds any single result.
<details>
<summary>Implementation notes</summary>
- New `coderd/x/chatd/chatloop/tooltruncate.go`:
`toolResultByteBudget(contextLimitTokens)` and
`truncateToolResultText(text, maxBytes)` (pure, unit-tested).
- `chatloop.go`: added `ContextLimit` to `ExecuteLocalToolsOptions`;
threaded a computed byte budget through `executeTools` into
`executeSingleTool`, where `resp.Content` is capped for the text,
media-text, and error branches.
- `generation.go`: passes `ContextLimit: prepared.ContextLimitFallback`
(the model's configured context limit).
- `metrics.go`: new `ToolResultTruncatedTotal` counter +
`RecordToolResultTruncated`.
- Tunable knobs live as constants in `tooltruncate.go`
(`toolResultContextDivisor = 2`, `bytesPerTokenEstimate`,
`minToolResultBytes`, `defaultToolResultBytes`).
Verified: `go build ./coderd/x/chatd/...`, `go test
./coderd/x/chatd/chatloop/...`, and the `chatd` test binary compiles.
</details>
---
Resolves CODAGT-678
Generated by Coder Agents on behalf of @kylecarbs.
|