When converting a Chat Completions stream into Responses events, the first
tool_call delta chunk was copied wholesale into stream state (including
function.arguments), then the same chunk's arguments were accumulated again by
the shared `+=` block. For OpenAI this is harmless because its first tool_call
chunk carries empty arguments, but upstreams that pack id+name+arguments into a
single chunk (e.g. GLM/Zhipu) end up with doubled arguments such as
{"cmd":"ls"}{"cmd":"ls"}. Codex then fails to parse the tool call with
"trailing characters", breaking every tool invocation.
Reset the copied arguments so the shared accumulator counts them exactly once,
keeping the emitted delta and the final done/arguments consistent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Locks in that Claude Code detection keys on the billing block prefix +
cc_entrypoint=cli, not on the cch field that the new CLI (and now our own
mimicry) no longer sends:
- BillingBlockRecognizedWithoutCCH: an identity-prose-less sub-request whose
system block is `x-anthropic-billing-header: cc_version=...; cc_entrypoint=cli;`
(no cch) is still detected as Claude Code.
- NoCCHBlockStillRequiresClaudeCodeUA: dropping cch did not loosen detection —
a non-claude-cli UA is still rejected, so ClaudeCodeOnly groups can't be spoofed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Recent Claude Code CLI versions no longer emit the cch=... signature field in
their x-anthropic-billing-header system block (issue #3358). sub2api still
injected cch=00000 when mimicking Claude Code for OAuth accounts and optionally
signed it, so mimicked requests now diverge from real CLI traffic — the opposite
of what the mimicry is for.
- buildBillingAttributionText emits the block without the cch=00000 segment;
cc_version + cc_entrypoint=cli are kept (detection and Anthropic's first-party
signal rely on the block, not on cch).
- Retire signing: remove the two enableCCH signBillingHeaderCCH call sites in
buildUpstreamRequest / buildCountTokensRequest and delete the now-dead
signBillingHeaderCCH, cchPlaceholderRe, cchSeed, xxHash64Seeded helpers.
- enable_cch_signing is now a documented no-op (kept for backward compat).
- Drop the obsolete signing tests (TestSignBillingHeaderCCH, TestXXHash64Seeded,
TestSanitizeMustBeBeforeCCHSigning_HashConsistency) and update the prompt test
to assert the injected block no longer carries cch=.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Covers the #3358 fix:
- StripsUnsupportedClaudeCodeTokens reproduces the prod 400 — the four Vertex-
rejected tokens (advisor-tool, prompt-caching-scope, redact-thinking,
thinking-token-count) plus the identity betas are stripped while whitelisted
tokens survive. Fails before the builder fix, passes after.
- DropsHeaderWhenAllUnsupported: no anthropic-beta header is sent when every
client token is filtered out.
- BodySanitizeKeysOnFinalBeta: body.context_management is stripped based on the
final beta, not the raw client value.
- BlocksViaBetaPolicy: an admin block rule on a Vertex account returns BetaBlockedError.
- TestFilterVertexBetaTokens unit-tests whitelist/drop-set/dedupe/empty.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Vertex AI's Anthropic endpoint rejects unknown anthropic-beta tokens with
HTTP 400. buildUpstreamRequestAnthropicVertex forwarded the client header
verbatim via the allowedHeaders whitelist, so recent Claude Code CLIs that
send advisor-tool-2026-03-01, prompt-caching-scope-2026-01-05,
redact-thinking-2026-02-12 and thinking-token-count-2026-05-13 broke every
Vertex service_account request, even though plain account-test requests passed.
This is the only upstream builder that bypassed beta filtering: the
OAuth/API-key path uses computeFinalAnthropicBeta and the Bedrock path uses
filterBedrockBetaTokens. Close the gap with a Vertex-specific whitelist
(vertexSupportedBetaTokens) mirroring bedrockSupportedBetaTokens, plus the
existing BetaPolicy block check:
- evaluateBetaPolicy block check (symmetric to resolveBedrockBetaTokensForRequest)
- filterVertexBetaTokens strips policy-filtered + non-whitelisted tokens
- body context_management sanitize now keys on the final beta, not the raw client value
- overwrite the anthropic-beta header after the whitelist copy loop with the final value
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds a use-it-or-lose-it scheduling strategy: prefer accounts whose
session window resets soonest, so near-reset accounts get drained first
instead of accounts whose reset is still far away.
Both schedulers, opt-in, default behavior unchanged:
- Anthropic (gateway_service.go): new GatewaySchedulingConfig
.PreferSoonestReset flag. When on, the layered load-aware selection
inserts a filterBySoonestReset stage (priority -> soonest-reset ->
load -> LRU). Accounts with no active SessionWindowEnd are treated as
lowest priority; ties fall through to LRU.
- OpenAI/Codex (openai_account_scheduler.go): new "reset" score weight
in GatewayOpenAIWSSchedulerScoreWeights. Soonest-reset accounts score
higher; weight defaults to 0 (no effect).
SessionWindowEnd (upstream 5h/quota ResetsAt) is already carried in the
scheduler snapshot, so no snapshot changes are needed.
Documented in deploy/config.example.yaml. Adds unit tests for the
Anthropic filter and the OpenAI reset factor.
DeriveUpstreamEndpoint maps every OpenAI-platform request to /v1/responses, but
API-key accounts whose upstream only speaks Chat Completions
(!ShouldUseResponsesAPI) are forwarded directly to /v1/chat/completions. The
messages, responses and cyber-policy recording sites derived the endpoint via the
bare GetUpstreamEndpoint, so usage/ops records mislabeled those requests as
/v1/responses. Generalize the existing resolveRawCCUpstreamEndpoint into
resolveOpenAIUpstreamEndpoint and use it at every OpenAI recording site, matching
the already-correct chat-completions client path.
When an Anthropic upstream returned HTTP 200 but then emitted an SSE
`event: error` frame (overloaded_error / rate_limit_error / api_error /
etc.), Forward's stream branch matched on `err.Error() == "have error in
stream"` and returned `UpstreamFailoverError{StatusCode: 403}` with no
ResponseBody. That dropped three pieces of evidence:
- handleFailoverExhausted → ExtractUpstreamErrorMessage(nil) = "" →
ops_error_logs.upstream_error_message was empty.
- errorPassthroughService.MatchRule(_, 403, nil) could only match rules
without keywords, so keyword-based passthrough rules silently never
fired.
- upstream_errors carried no stream_error record, leaving ops looking at
a generic 403 with no clue whether the upstream was throttled,
overloaded, or rejecting the request. ping-during-slot-wait amplified
this by skipping failover (writerSizeBeforeForward guard), so 403s
ballooned in the ops view well past the upstream's actual rate.
Fix:
- Introduce *sseStreamErrorEventError that carries the SSE data line.
Error() still returns "have error in stream" so existing log searches
keep working.
- Forward extracts via errors.As, appends an OpsUpstreamErrorEvent
(kind="stream_error", with the sanitized message and a truncated raw
body honoring LogUpstreamErrorBody*), and returns
UpstreamFailoverError{StatusCode: 403, ResponseBody: rawJSON}.
StatusCode 403 is preserved verbatim: mapUpstreamError, failover
decisions (shouldFailoverUpstreamError(403)=true), client-visible message,
RetryableOnSameAccount, and rateLimitService side-effects (this path
already didn't invoke them) all match prior behavior. OAuth and API Key
accounts share this path; the API-Key passthrough branch is independent
and already forwards SSE error frames untouched, so it's unaffected.
Adds four unit tests: typed-error contract + RawData, empty data line,
event:error after partial stream output (streamStarted=true), and
non-JSON data line.