Locks in that Claude Code detection keys on the billing block prefix +
cc_entrypoint=cli, not on the cch field that the new CLI (and now our own
mimicry) no longer sends:
- BillingBlockRecognizedWithoutCCH: an identity-prose-less sub-request whose
system block is `x-anthropic-billing-header: cc_version=...; cc_entrypoint=cli;`
(no cch) is still detected as Claude Code.
- NoCCHBlockStillRequiresClaudeCodeUA: dropping cch did not loosen detection —
a non-claude-cli UA is still rejected, so ClaudeCodeOnly groups can't be spoofed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Recent Claude Code CLI versions no longer emit the cch=... signature field in
their x-anthropic-billing-header system block (issue #3358). sub2api still
injected cch=00000 when mimicking Claude Code for OAuth accounts and optionally
signed it, so mimicked requests now diverge from real CLI traffic — the opposite
of what the mimicry is for.
- buildBillingAttributionText emits the block without the cch=00000 segment;
cc_version + cc_entrypoint=cli are kept (detection and Anthropic's first-party
signal rely on the block, not on cch).
- Retire signing: remove the two enableCCH signBillingHeaderCCH call sites in
buildUpstreamRequest / buildCountTokensRequest and delete the now-dead
signBillingHeaderCCH, cchPlaceholderRe, cchSeed, xxHash64Seeded helpers.
- enable_cch_signing is now a documented no-op (kept for backward compat).
- Drop the obsolete signing tests (TestSignBillingHeaderCCH, TestXXHash64Seeded,
TestSanitizeMustBeBeforeCCHSigning_HashConsistency) and update the prompt test
to assert the injected block no longer carries cch=.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Covers the #3358 fix:
- StripsUnsupportedClaudeCodeTokens reproduces the prod 400 — the four Vertex-
rejected tokens (advisor-tool, prompt-caching-scope, redact-thinking,
thinking-token-count) plus the identity betas are stripped while whitelisted
tokens survive. Fails before the builder fix, passes after.
- DropsHeaderWhenAllUnsupported: no anthropic-beta header is sent when every
client token is filtered out.
- BodySanitizeKeysOnFinalBeta: body.context_management is stripped based on the
final beta, not the raw client value.
- BlocksViaBetaPolicy: an admin block rule on a Vertex account returns BetaBlockedError.
- TestFilterVertexBetaTokens unit-tests whitelist/drop-set/dedupe/empty.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Vertex AI's Anthropic endpoint rejects unknown anthropic-beta tokens with
HTTP 400. buildUpstreamRequestAnthropicVertex forwarded the client header
verbatim via the allowedHeaders whitelist, so recent Claude Code CLIs that
send advisor-tool-2026-03-01, prompt-caching-scope-2026-01-05,
redact-thinking-2026-02-12 and thinking-token-count-2026-05-13 broke every
Vertex service_account request, even though plain account-test requests passed.
This is the only upstream builder that bypassed beta filtering: the
OAuth/API-key path uses computeFinalAnthropicBeta and the Bedrock path uses
filterBedrockBetaTokens. Close the gap with a Vertex-specific whitelist
(vertexSupportedBetaTokens) mirroring bedrockSupportedBetaTokens, plus the
existing BetaPolicy block check:
- evaluateBetaPolicy block check (symmetric to resolveBedrockBetaTokensForRequest)
- filterVertexBetaTokens strips policy-filtered + non-whitelisted tokens
- body context_management sanitize now keys on the final beta, not the raw client value
- overwrite the anthropic-beta header after the whitelist copy loop with the final value
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds a use-it-or-lose-it scheduling strategy: prefer accounts whose
session window resets soonest, so near-reset accounts get drained first
instead of accounts whose reset is still far away.
Both schedulers, opt-in, default behavior unchanged:
- Anthropic (gateway_service.go): new GatewaySchedulingConfig
.PreferSoonestReset flag. When on, the layered load-aware selection
inserts a filterBySoonestReset stage (priority -> soonest-reset ->
load -> LRU). Accounts with no active SessionWindowEnd are treated as
lowest priority; ties fall through to LRU.
- OpenAI/Codex (openai_account_scheduler.go): new "reset" score weight
in GatewayOpenAIWSSchedulerScoreWeights. Soonest-reset accounts score
higher; weight defaults to 0 (no effect).
SessionWindowEnd (upstream 5h/quota ResetsAt) is already carried in the
scheduler snapshot, so no snapshot changes are needed.
Documented in deploy/config.example.yaml. Adds unit tests for the
Anthropic filter and the OpenAI reset factor.
DeriveUpstreamEndpoint maps every OpenAI-platform request to /v1/responses, but
API-key accounts whose upstream only speaks Chat Completions
(!ShouldUseResponsesAPI) are forwarded directly to /v1/chat/completions. The
messages, responses and cyber-policy recording sites derived the endpoint via the
bare GetUpstreamEndpoint, so usage/ops records mislabeled those requests as
/v1/responses. Generalize the existing resolveRawCCUpstreamEndpoint into
resolveOpenAIUpstreamEndpoint and use it at every OpenAI recording site, matching
the already-correct chat-completions client path.
Add the :Z (private unshared) SELinux label to all bind-mounted
directories in docker-compose.local.yml and docker-compose.dev.yml.
On systems with SELinux in Enforcing mode (e.g. Fedora, RHEL, CentOS),
containers using bind mounts are denied access to host directories
because the default 'user_home_t' context is not accessible to
container processes. The :Z label tells the container runtime to
relabel the mount point with 'container_file_t' so the container
can read/write it.
Named volumes in docker-compose.yml are not affected because the
runtime already handles their labels automatically.
Fixes permission-denied errors on:
- ./data:/app/data
- ./postgres_data:/var/lib/postgresql/data
- ./redis_data:/data
When an Anthropic upstream returned HTTP 200 but then emitted an SSE
`event: error` frame (overloaded_error / rate_limit_error / api_error /
etc.), Forward's stream branch matched on `err.Error() == "have error in
stream"` and returned `UpstreamFailoverError{StatusCode: 403}` with no
ResponseBody. That dropped three pieces of evidence:
- handleFailoverExhausted → ExtractUpstreamErrorMessage(nil) = "" →
ops_error_logs.upstream_error_message was empty.
- errorPassthroughService.MatchRule(_, 403, nil) could only match rules
without keywords, so keyword-based passthrough rules silently never
fired.
- upstream_errors carried no stream_error record, leaving ops looking at
a generic 403 with no clue whether the upstream was throttled,
overloaded, or rejecting the request. ping-during-slot-wait amplified
this by skipping failover (writerSizeBeforeForward guard), so 403s
ballooned in the ops view well past the upstream's actual rate.
Fix:
- Introduce *sseStreamErrorEventError that carries the SSE data line.
Error() still returns "have error in stream" so existing log searches
keep working.
- Forward extracts via errors.As, appends an OpsUpstreamErrorEvent
(kind="stream_error", with the sanitized message and a truncated raw
body honoring LogUpstreamErrorBody*), and returns
UpstreamFailoverError{StatusCode: 403, ResponseBody: rawJSON}.
StatusCode 403 is preserved verbatim: mapUpstreamError, failover
decisions (shouldFailoverUpstreamError(403)=true), client-visible message,
RetryableOnSameAccount, and rateLimitService side-effects (this path
already didn't invoke them) all match prior behavior. OAuth and API Key
accounts share this path; the API-Key passthrough branch is independent
and already forwards SSE error frames untouched, so it's unaffected.
Adds four unit tests: typed-error contract + RawData, empty data line,
event:error after partial stream output (streamStarted=true), and
non-JSON data line.
The function signature was changed to require a mappedModel parameter
for protocol-aware thinking-block filtering, but this test call site
was not updated. Without the fix, the unit test build fails on CI.
The gateway's thinking-block handling was designed for Anthropic's strict
semantics (drop blocks with missing/invalid signature), but third-party
Anthropic-compatible upstreams have INVERTED semantics:
* DeepSeek `/anthropic`, Kimi `/coding`, GLM, Moonshot, qwen-*-thinking
require ALL historical thinking blocks to round-trip verbatim.
* Stripping any of them produces:
400 "The content[].thinking in the thinking mode must be passed back
to the API"
Without this fix, every multi-turn request from a thinking-capable client
(Claude Code, pi, etc.) to such upstreams loses its thinking blocks and
fails. This becomes especially painful when an account's model_mapping
maps `claude-sonnet-4-6 → deepseek-v4-pro` — `reqModel` looks Anthropic
but the upstream contract is the opposite.
Approach
--------
Branch all thinking-block transforms by the *mapped* model id (after
account model_mapping is applied), classifying into three families:
* `anthropic-strict` claude-/opus-/sonnet-/haiku- → existing behaviour
* `passback-required` deepseek-/kimi-/moonshot-/glm-/ → preserve verbatim
qwen-*-thinking
* `unknown` other models → conservative
(preserve, no retry)
Affected entry points (all guarded):
* Pre-filter on outbound: `FilterThinkingBlocks`
Previously dropped blocks with missing/invalid signature; now skips
entirely for non-strict families. Pre-filter is needed because the
post-error retry path can run out of budget on long conversations
(maxRetryElapsed = 10 s).
* 400 retry rectifier: `FilterThinkingBlocksForRetry`
Disables top-level thinking and converts thinking → text. Now skips
for passback-required (those 400s aren't signature errors and any
transformation breaks the round-trip contract).
* 400 retry rectifier (tools): `FilterSignatureSensitiveBlocksForRetry`
Same family-aware short-circuit.
* 400 detector: `shouldRectifySignatureError`
Returns false for passback-required, so the retry path doesn't even
fire.
Tests
-----
* `thinking_protocol_test.go` — classifier across all known vendor
prefixes plus edge cases (empty, case, qwen non-thinking).
* `thinking_protocol_filter_integration_test.go` — locks in that the
three filter entry points return the body byte-for-byte unchanged
when the model id is passback-required or unknown, and still strip
invalid blocks for anthropic-strict.
This PR supersedes #1350 (which only added the pre-filter without the
upstream-family awareness, and would have made third-party upstreams
worse). Once merged, please close#1350.
Reference issues:
- NousResearch/hermes-agent#16748 — DeepSeek /anthropic strip behaviour
- NousResearch/hermes-agent#15700 — DeepSeek thinking:disabled requirement
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
pi maps defaultThinkingLevel 'xhigh' → 'max' for deepseek-v4-pro via
thinkingLevelMap, but normalizeOpenAIReasoningEffort did not recognize
'max', causing all reasoning_effort to be dropped (0% fill rate since
traffic switched from claude to deepseek on May 1).
Add 'max' → 'xhigh' mapping to match Claude's NormalizeClaudeOutputEffort
behavior.