3969 Commits
Author SHA1 Message Date
shaw 6936687870 fix(lint): check WriteString return in summarizeOpenAIImagesNoOutputBody
修复 PR #3381 引入的 errcheck lint 错误,satisfy golangci-lint
对 strings.Builder.WriteString 返回值的检查要求。
v0.1.138
2026-06-21 21:43:58 +08:00
Wesley Liddick 8105846f33 Merge pull request #3326 from tairan/fix/selinux-bind-mount-labels
fix(deploy): add :Z SELinux labels to bind mounts in compose files
2026-06-21 21:23:27 +08:00
Wesley Liddick f079742bab Merge pull request #3371 from cugxuan/feat/subscription-affiliate-rebate
feat: apply affiliate rebate to subscription payments
2026-06-21 21:18:50 +08:00
Wesley Liddick a560d2f91e Merge pull request #3363 from kangjwme/feat/prefer-soonest-reset-scheduling
feat(scheduling): opt-in "prefer soonest reset" account selection
2026-06-21 21:17:44 +08:00
Wesley Liddick d2ff9c0c4c Merge pull request #3369 from skyswordw/codex/default-ccswitch-openai-gpt-55
fix(ccswitch): default OpenAI import model to gpt-5.5
2026-06-21 21:16:07 +08:00
Wesley Liddick 92ce5bfe24 Merge pull request #3347 from SavitarC/fix/usage-cache-token-tooltip
fix(usage): 显示缓存 Token 明细
2026-06-21 21:15:45 +08:00
Wesley Liddick febfe538b7 Merge pull request #3344 from Milesians/main
fix(frontend): refresh custom page document title
2026-06-21 21:15:17 +08:00
Wesley Liddick 6e27b82335 Merge pull request #3381 from 404QAQ/fix/images-incomplete-failover
fix(images): 识别 response.incomplete 触发 failover + 记录软失败上游响应
2026-06-21 21:14:48 +08:00
Wesley Liddick f6d734ec23 Merge pull request #3308 from a0yark/fix/gemini-tool-schema-cleanup
fix(gemini): clean unsupported tool schema fields
2026-06-21 21:14:09 +08:00
Wesley Liddick af1a032d35 Merge pull request #3375 from StarryKira/fix/3358-vertex-beta-and-cch
fix(gateway): filter anthropic-beta on the Vertex Anthropic path + drop cch sign (#3358)
2026-06-21 21:13:14 +08:00
Wesley Liddick 93de19665b Merge pull request #3359 from alfadb/fix/glm-effort-mapping
修复 GLM 推理强度映射
2026-06-21 21:12:46 +08:00
Wesley Liddick bed12b7016 Merge pull request #3360 from feitianbubu/fix/auto-mode-cc-entrypoint-ide
fix(gateway): in auto mode recognize Claude Code IDE clients via any cc_entrypoint
2026-06-21 21:12:33 +08:00
Wesley Liddick f597e926da Merge pull request #3335 from FjlI5/fix/openai-upstream-endpoint-logging
fix: 修正 chat-only API-key 账号上游端点被误记为 /v1/responses
2026-06-21 21:12:15 +08:00
Wesley Liddick 7b5fbe5197 Merge pull request #3364 from feitianbubu/fix/promo-clear-expiry
fix(promo): allow clearing promo code expiry on edit
2026-06-21 21:12:02 +08:00
Wesley Liddick 04494fd0d9 Merge pull request #3312 from 315944211/codex/update-pnpm-action-setup-node24
chore: update pnpm action setup
2026-06-21 21:11:43 +08:00
Wesley Liddick 945b9b208b Merge pull request #3374 from wucm667/fix/email-binding-enforce-suffix-whitelist
fix(auth): 邮箱身份绑定流程同样校验注册后缀白名单,堵住绕过
2026-06-21 12:56:53 +08:00
404QAQ b0d5592ae2 fix(images): 识别 response.incomplete + 记录软失败上游响应
修复 gpt-image-2 大图/编辑请求 502 报错(社区 issue #2232/#3135/#2516,
现象:upstream did not return image output)。根因是上游生成超时/截断时返回
response.incomplete,旧逻辑只认 error/response.failed,导致:
1) 软失败报成模糊 502 且不触发 failover 换账号重试
2) 上游真实响应未记录,ops_error_logs 里上游信息全空,无法排查

- openAIImagesUpstreamErrorFromSSEPayload 增加 response.incomplete 识别:
  生成超时/截断(max_output_tokens 等) → 可重试 502 触发 failover;
  content_filter/moderation → 400 不重试
- 软失败兜底(无图无标准错误)记录上游诊断摘要到 ops(last_event/status/
  incomplete_reason/body 片段),非流式+流式两条路径一致
- summarizeOpenAIImagesNoOutputBody 提取诊断信息,body 截断上限 1KB

基于官方 v0.1.137 净分支。测试: 5 个新单测 + 图片回归通过 (-tags unit)
2026-06-21 08:34:12 +08:00
harukaandClaude Opus 4.8 5cb8cdd36c test(claude-code): detection recognizes the new-CLI billing block (no cch)
Locks in that Claude Code detection keys on the billing block prefix +
cc_entrypoint=cli, not on the cch field that the new CLI (and now our own
mimicry) no longer sends:

- BillingBlockRecognizedWithoutCCH: an identity-prose-less sub-request whose
  system block is `x-anthropic-billing-header: cc_version=...; cc_entrypoint=cli;`
  (no cch) is still detected as Claude Code.
- NoCCHBlockStillRequiresClaudeCodeUA: dropping cch did not loosen detection —
  a non-claude-cli UA is still rejected, so ClaudeCodeOnly groups can't be spoofed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 08:02:44 -07:00
harukaandClaude Opus 4.8 6cfb7898df fix(claude-mimicry): drop the cch sign to match new Claude Code CLI
Recent Claude Code CLI versions no longer emit the cch=... signature field in
their x-anthropic-billing-header system block (issue #3358). sub2api still
injected cch=00000 when mimicking Claude Code for OAuth accounts and optionally
signed it, so mimicked requests now diverge from real CLI traffic — the opposite
of what the mimicry is for.

- buildBillingAttributionText emits the block without the cch=00000 segment;
  cc_version + cc_entrypoint=cli are kept (detection and Anthropic's first-party
  signal rely on the block, not on cch).
- Retire signing: remove the two enableCCH signBillingHeaderCCH call sites in
  buildUpstreamRequest / buildCountTokensRequest and delete the now-dead
  signBillingHeaderCCH, cchPlaceholderRe, cchSeed, xxHash64Seeded helpers.
- enable_cch_signing is now a documented no-op (kept for backward compat).
- Drop the obsolete signing tests (TestSignBillingHeaderCCH, TestXXHash64Seeded,
  TestSanitizeMustBeBeforeCCHSigning_HashConsistency) and update the prompt test
  to assert the injected block no longer carries cch=.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 08:01:29 -07:00
harukaandClaude Opus 4.8 efffd5d791 test(gateway): Vertex anthropic-beta filtering
Covers the #3358 fix:
- StripsUnsupportedClaudeCodeTokens reproduces the prod 400 — the four Vertex-
  rejected tokens (advisor-tool, prompt-caching-scope, redact-thinking,
  thinking-token-count) plus the identity betas are stripped while whitelisted
  tokens survive. Fails before the builder fix, passes after.
- DropsHeaderWhenAllUnsupported: no anthropic-beta header is sent when every
  client token is filtered out.
- BodySanitizeKeysOnFinalBeta: body.context_management is stripped based on the
  final beta, not the raw client value.
- BlocksViaBetaPolicy: an admin block rule on a Vertex account returns BetaBlockedError.
- TestFilterVertexBetaTokens unit-tests whitelist/drop-set/dedupe/empty.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 07:53:32 -07:00
harukaandClaude Opus 4.8 40e1cc14b3 fix(gateway): filter anthropic-beta on the Vertex Anthropic path (#3358)
Vertex AI's Anthropic endpoint rejects unknown anthropic-beta tokens with
HTTP 400. buildUpstreamRequestAnthropicVertex forwarded the client header
verbatim via the allowedHeaders whitelist, so recent Claude Code CLIs that
send advisor-tool-2026-03-01, prompt-caching-scope-2026-01-05,
redact-thinking-2026-02-12 and thinking-token-count-2026-05-13 broke every
Vertex service_account request, even though plain account-test requests passed.

This is the only upstream builder that bypassed beta filtering: the
OAuth/API-key path uses computeFinalAnthropicBeta and the Bedrock path uses
filterBedrockBetaTokens. Close the gap with a Vertex-specific whitelist
(vertexSupportedBetaTokens) mirroring bedrockSupportedBetaTokens, plus the
existing BetaPolicy block check:

- evaluateBetaPolicy block check (symmetric to resolveBedrockBetaTokensForRequest)
- filterVertexBetaTokens strips policy-filtered + non-whitelisted tokens
- body context_management sanitize now keys on the final beta, not the raw client value
- overwrite the anthropic-beta header after the whitelist copy loop with the final value

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 07:51:24 -07:00
wucm667 ecedc7c8d3 fix(auth): enforce email bind suffix whitelist 2026-06-19 21:17:45 +08:00
cugxuan 0fa604ba71 feat: apply affiliate rebate to subscription payments 2026-06-19 17:47:10 +08:00
skywalker d3dfa28f8e Update CC Switch OpenAI default model 2026-06-19 17:11:25 +08:00
feitianbubu 2dc1387b59 fix(promo): allow clearing promo code expiry on edit 2026-06-18 23:07:25 +08:00
kangjwme 510adf703c feat(scheduling): add opt-in "prefer soonest reset" account selection
Adds a use-it-or-lose-it scheduling strategy: prefer accounts whose
session window resets soonest, so near-reset accounts get drained first
instead of accounts whose reset is still far away.

Both schedulers, opt-in, default behavior unchanged:

- Anthropic (gateway_service.go): new GatewaySchedulingConfig
  .PreferSoonestReset flag. When on, the layered load-aware selection
  inserts a filterBySoonestReset stage (priority -> soonest-reset ->
  load -> LRU). Accounts with no active SessionWindowEnd are treated as
  lowest priority; ties fall through to LRU.

- OpenAI/Codex (openai_account_scheduler.go): new "reset" score weight
  in GatewayOpenAIWSSchedulerScoreWeights. Soonest-reset accounts score
  higher; weight defaults to 0 (no effect).

SessionWindowEnd (upstream 5h/quota ResetsAt) is already carried in the
scheduler snapshot, so no snapshot changes are needed.

Documented in deploy/config.example.yaml. Adds unit tests for the
Anthropic filter and the OpenAI reset factor.
2026-06-18 22:50:46 +08:00
feitianbubu e3e31bd4c5 fix(gateway): auto mode recognize Claude Code IDE clients via any cc_entrypoint 2026-06-18 19:22:10 +08:00
alfadb 89cfe24a0c fix(openai): normalize glm reasoning effort 2026-06-18 19:07:08 +08:00
SavitarC 51d722906c fix(usage): 显示缓存 Token 明细 2026-06-18 14:34:13 +08:00
milesians 952be8711b fix(frontend): refresh custom page document title 2026-06-18 12:30:21 +08:00
FjlI5 bab8a9a93e fix(openai): log /v1/chat/completions upstream endpoint for chat-only API-key accounts
DeriveUpstreamEndpoint maps every OpenAI-platform request to /v1/responses, but
API-key accounts whose upstream only speaks Chat Completions
(!ShouldUseResponsesAPI) are forwarded directly to /v1/chat/completions. The
messages, responses and cyber-policy recording sites derived the endpoint via the
bare GetUpstreamEndpoint, so usage/ops records mislabeled those requests as
/v1/responses. Generalize the existing resolveRawCCUpstreamEndpoint into
resolveOpenAIUpstreamEndpoint and use it at every OpenAI recording site, matching
the already-correct chat-completions client path.
2026-06-18 00:29:04 +08:00
WANG TAIRAN 31640363fc fix(deploy): add :Z SELinux labels to bind mounts
Add the :Z (private unshared) SELinux label to all bind-mounted
directories in docker-compose.local.yml and docker-compose.dev.yml.

On systems with SELinux in Enforcing mode (e.g. Fedora, RHEL, CentOS),
containers using bind mounts are denied access to host directories
because the default 'user_home_t' context is not accessible to
container processes. The :Z label tells the container runtime to
relabel the mount point with 'container_file_t' so the container
can read/write it.

Named volumes in docker-compose.yml are not affected because the
runtime already handles their labels automatically.

Fixes permission-denied errors on:
- ./data:/app/data
- ./postgres_data:/var/lib/postgresql/data
- ./redis_data:/data
2026-06-17 12:38:40 +08:00
a0yarkandCursor 8c4a43cf72 fix(gemini): satisfy schema cleanup test lint
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-16 23:19:45 +08:00
315944211 369f53a7a4 chore: force node24 for cla action 2026-06-16 21:12:09 +08:00
315944211 abc203a37b chore: update pnpm action setup 2026-06-16 21:12:09 +08:00
github-actions[bot] 4a5665da5b chore: sync VERSION to 0.1.137 [skip ci] 2026-06-16 12:59:59 +00:00
Wesley Liddick eba9bea959 Merge pull request #3278 from 315944211/feat/show-account-id-in-admin-list
feat: show account id in account list
v0.1.137
2026-06-16 20:35:55 +08:00
Wesley Liddick de38d62360 Merge pull request #3250 from WesleyZiwen/codex/anthropic-429-window-reset
fix: preserve Anthropic window cooldowns
2026-06-16 20:34:36 +08:00
Wesley Liddick 9e9e154f57 Merge pull request #3306 from bestony/fix/ip-acl-denial-message
fix(auth): include client ip in acl denial message
2026-06-16 20:32:50 +08:00
Wesley Liddick 7fb3e1aacd Merge pull request #3247 from alfadb/fix/reasoning-and-thinking-protocol
fix(gateway): 整合推理强度与思考协议处理(替代 #2155 / #2136 / #3246)
2026-06-16 20:29:21 +08:00
Wesley Liddick f204216918 Merge pull request #3243 from alfadb/feature/chinese-llm-fallback-pricing
feat(billing): 国产 LLM 兜底定价 (GLM / Kimi / MiniMax) + 收编 DeepSeek V4
2026-06-16 20:28:49 +08:00
Wesley Liddick 03ec90a25a Merge pull request #3223 from feitianbubu/fix/intercept-streaming-haiku-probe
fix: 修复CC Switch改为流式测试后请求失败的问题
2026-06-16 20:27:33 +08:00
shaw 6c7203d83b fix(gateway): preserve SSE event:error body so ops logs reflect real upstream errors
When an Anthropic upstream returned HTTP 200 but then emitted an SSE
`event: error` frame (overloaded_error / rate_limit_error / api_error /
etc.), Forward's stream branch matched on `err.Error() == "have error in
stream"` and returned `UpstreamFailoverError{StatusCode: 403}` with no
ResponseBody. That dropped three pieces of evidence:

- handleFailoverExhausted → ExtractUpstreamErrorMessage(nil) = "" →
  ops_error_logs.upstream_error_message was empty.
- errorPassthroughService.MatchRule(_, 403, nil) could only match rules
  without keywords, so keyword-based passthrough rules silently never
  fired.
- upstream_errors carried no stream_error record, leaving ops looking at
  a generic 403 with no clue whether the upstream was throttled,
  overloaded, or rejecting the request. ping-during-slot-wait amplified
  this by skipping failover (writerSizeBeforeForward guard), so 403s
  ballooned in the ops view well past the upstream's actual rate.

Fix:
- Introduce *sseStreamErrorEventError that carries the SSE data line.
  Error() still returns "have error in stream" so existing log searches
  keep working.
- Forward extracts via errors.As, appends an OpsUpstreamErrorEvent
  (kind="stream_error", with the sanitized message and a truncated raw
  body honoring LogUpstreamErrorBody*), and returns
  UpstreamFailoverError{StatusCode: 403, ResponseBody: rawJSON}.

StatusCode 403 is preserved verbatim: mapUpstreamError, failover
decisions (shouldFailoverUpstreamError(403)=true), client-visible message,
RetryableOnSameAccount, and rateLimitService side-effects (this path
already didn't invoke them) all match prior behavior. OAuth and API Key
accounts share this path; the API-Key passthrough branch is independent
and already forwards SSE error frames untouched, so it's unaffected.

Adds four unit tests: typed-error contract + RawData, empty data line,
event:error after partial stream output (streamStarted=true), and
non-JSON data line.
2026-06-16 20:25:09 +08:00
alfadb a05d9e87c0 feat(billing): 国产模型 thinking-enabled 自动填充 reasoning_effort 默认值
问题:Kimi/GLM/MiniMax 等国产 LLM 协议层只有 thinking on/off 开关,没有
reasoning_effort 档位概念。客户端启用 thinking 后 usage_log.reasoning_effort
长期为 NULL,无法在用量分析里区分 'thinking 开启' 与 'thinking 关闭'。

方案:仅在 'thinking 启用 + 上游属于 passback-required 国产模型 + 客户端
未明确指定 effort' 三者同时成立时,给 usage_log.reasoning_effort 写默认值
'high'(与 DeepSeek thinking-enabled 默认 effort 一致)。

设计原则:
1. **白名单**:仅 ResolveThinkingProtocol == PassbackRequired 集合内,
   且排除原生支持 effort 的 DeepSeek(避免覆盖客户端意图)。
2. **fail-open**:客户端显式传 effort 时永远不覆盖。
3. **未来兼容**:如 Kimi 后续加入真 effort 档位,客户端开始发 effort,
   guard (3) 自动让出,本逻辑变 no-op。

实现:
- gateway_request.go: 加 DefaultEffortForThinkingEnabled (按模型白名单)
  + OpenAIBodyHasThinkingEnabled (检测 OpenAI 协议 body 里的 thinking.type)
  + ApplyThinkingEnabledFallback (包装现有 extractor 的 nil-then-default 逻辑)
- gateway_handler.go: Anthropic 路径两处(主 + retry)对称补充
- OpenAI 路径全覆盖:openai_gateway_service.go (passthrough + non-passthrough)
  + openai_gateway_chat_completions_raw.go + openai_gateway_responses_chat_fallback.go
  + openai_ws_http_bridge.go + openai_ws_forwarder.go
- 跨协议路径:gateway_forward_as_chat_completions.go (CC client → Anthropic upstream)
  + gateway_forward_as_responses.go (Responses client → Anthropic upstream)
  + gemini_chat_completions_compat_service.go (一致性保持)

未覆盖:openai_ws_v2_passthrough_adapter.go 两处。原因:该 adapter 持有的是
session-level 客户端原始 model,没有 *Account 句柄无法走 GetMappedModel。
WS v2 当前对国产模型场景不重要(pi 调用 Kimi/GLM/MiniMax 走 sync HTTP),
留待后续如果出现 WS v2 + 国产模型用例时单独处理。

测试:
- TestDefaultEffortForThinkingEnabled (14 用例):覆盖 Kimi/GLM/MiniMax 大小写、
  Qwen thinking 变体、DeepSeek 排除、Claude/GPT/Gemini 不命中。
- TestOpenAIBodyHasThinkingEnabled (8 用例):covers enabled/adaptive/disabled、
  大小写、空 body、缺字段、invalid JSON fail-safe。
- TestApplyThinkingEnabledFallback (9 用例):现有 effort 不覆盖、nil + 启用 +
  passback → high、nil + disabled → nil、nil + 启用 + 排除模型 → nil。
2026-06-16 19:37:31 +08:00
alfadb 5c5283979b doc(thinking-protocol): clarify mappedModel vs originalModel semantics per call path
回应 PR #3247 Copilot review:

1. NormalizeChineseLLMThinking 文档写 'MiniMax M3 / M3.x' 但实现覆盖
   minimax-m* (M2.x 也命中)。改成 'MiniMax M-series (M2.x/M3.x)'。

2. Copilot 质疑 gemini_messages_compat 传 originalModel 不是
   mappedModel 是误用。实际是有意为之的跨协议场景:
   - 上游是 Gemini,但被剥离的 body 是 Anthropic 格式
   - 剥离逻辑要按客户端请求的 Anthropic 子协议族判定
   - 传 mappedModel (gemini-3.1-pro) 会被判为 Unknown→不剥离→retry 死循环

   未改代码逻辑,改为更新文档:
   - thinking_protocol.go ResolveThinkingProtocol 添加「调用路径语义」
     说明 Anthropic gateway 与 Gemini compat 两条路径的参数语义差异。
   - gemini_messages_compat_service.go 调用点加 inline 注释解释
     路径语义为什么该传 originalModel。
2026-06-16 19:37:31 +08:00
alfadb 56c6325d15 fix(gateway): rewrite thinking.type=enabled to adaptive for MiniMax M-series
\u95ee\u9898\uff1aAnthropic-SDK \u5ba2\u6237\u7aef\uff08\u5982 pi-ai\uff09\u5728 budget-based thinking \u8def\u5f84\u4e0a\u9ed8\u8ba4\u53d1
`thinking: { type: 'enabled', budget_tokens: ... }` \u7ed9\u4efb\u4f55 Anthropic \u517c\u5bb9\u4e0a\u6e38\uff0c\u5305\u62ec
MiniMax M3 / M2.x \u3002\u4f46 MiniMax \u5b98\u65b9\u6587\u6863\u660e\u786e\u5217\u51fa\uff1athinking.type \u53ea\u63a5\u53d7
"adaptive" \u6216 "disabled"\uff0c\u4e0d\u63a5\u53d7 "enabled"\u3002\u9020\u6210\u4e0a\u6e38\u53ef\u80fd 400 / \u9759\u9ed8\u5ffd\u7565\u3002

\u4fee\u590d\uff1a\u5728 Anthropic forward \u8def\u5f84\u4e0a\u3001StripEmptyTextBlocks \u4e4b\u540e\u3001\u91cd\u8bd5\u5faa\u73af\u4e4b\u524d\uff0c
\u52a0\u4e00\u4e2a NormalizeChineseLLMThinking \u9884\u6539\u5199\u6b65\u9aa4\uff1a
- \u4ec5\u5bf9 ResolveThinkingProtocol == PassbackRequired \u4e0a\u6e38\u751f\u6548\uff08claude-* \u4e0d\u4f1a\u8fdb\u6765\uff09
- \u4ec5\u5339\u914d minimax-m* \u524d\u7f00\u7684\u6620\u5c04\u540e\u6a21\u578b ID
- \u4ec5\u5f53 thinking.type == "enabled" \u65f6\u6539\u5199\u4e3a "adaptive"
- Kimi/GLM/DeepSeek \u7684 enabled \u4e0d\u53d7\u5f71\u54cd\uff08\u8fd9\u4e9b\u4e0a\u6e38\u63a5\u53d7 enabled\uff09

\u8bbe\u8ba1\u539f\u5219\uff1a
1. "\u767d\u540d\u5355\u8bed\u4e49"\uff1a\u53ea\u5bf9\u5df2\u77e5\u9700\u8981\u5904\u7406\u7684\u4e0a\u6e38\u6539\u5199\uff0c\u4e0d\u505a\u9690\u5f0f\u5e7f\u8c31\u8865\u4e01\u3002
2. "\u4f4e\u4fb5\u5165\u6027"\uff1aMiniMax \u4e4b\u5916\u7684\u6a21\u578b\u96f6\u6539\u52a8\uff0c\u73b0\u6709 4 \u4e2a\u4e0a\u6e38\u7684\u884c\u4e3a\u4e0d\u53d8\u3002
3. "\u53ef\u8bca\u65ad"\uff1a\u6539\u5199\u65f6 LegacyPrintf \u8f93\u51fa\u4e00\u884c\u65e5\u5fd7\uff0c\u4fbf\u4e8e\u8c03\u8bd5\u3002

\u6d4b\u8bd5\uff1a9 \u4e2a\u65b0\u7528\u4f8b\u8986\u76d6 MiniMax M3/M2.7\u3001adaptive/disabled \u4e0d\u53d8\u3001\u65e0 thinking
\u5b57\u6bb5\u3001Kimi/GLM/DeepSeek/Claude \u4e0d\u53d7\u5f71\u54cd\u3001invalid JSON fail-safe\u3002

\u4f9d\u8d56\uff1a\u672c\u5206\u652f\u57fa\u4e8e fix/thinking-block-protocol-aware-filter\uff0c\u7528\u5176
ResolveThinkingProtocol \u4f5c\u4e3a passback-required \u5224\u522b\u6e90\u3002\u4e0a\u6e38 PR \u987a\u5e8f\u9700\u5148\u5408
fix/thinking-block-protocol-aware-filter\uff08PR #2136\uff09\u518d\u5408\u672c PR\u3002
2026-06-16 19:37:31 +08:00
alfadb efbf6d2092 fix(test): update FilterThinkingBlocksForRetry call to use mappedModel param
The function signature was changed to require a mappedModel parameter
for protocol-aware thinking-block filtering, but this test call site
was not updated. Without the fix, the unit test build fails on CI.
2026-06-16 19:37:31 +08:00
alfadbandClaude Opus 4.7 6baf00d784 fix(gateway): protocol-aware thinking-block filtering for Anthropic-compatible upstreams
The gateway's thinking-block handling was designed for Anthropic's strict
semantics (drop blocks with missing/invalid signature), but third-party
Anthropic-compatible upstreams have INVERTED semantics:

  * DeepSeek `/anthropic`, Kimi `/coding`, GLM, Moonshot, qwen-*-thinking
    require ALL historical thinking blocks to round-trip verbatim.
  * Stripping any of them produces:
    400 "The content[].thinking in the thinking mode must be passed back
    to the API"

Without this fix, every multi-turn request from a thinking-capable client
(Claude Code, pi, etc.) to such upstreams loses its thinking blocks and
fails. This becomes especially painful when an account's model_mapping
maps `claude-sonnet-4-6 → deepseek-v4-pro` — `reqModel` looks Anthropic
but the upstream contract is the opposite.

Approach
--------

Branch all thinking-block transforms by the *mapped* model id (after
account model_mapping is applied), classifying into three families:

  * `anthropic-strict`     claude-/opus-/sonnet-/haiku-      → existing behaviour
  * `passback-required`    deepseek-/kimi-/moonshot-/glm-/   → preserve verbatim
                           qwen-*-thinking
  * `unknown`              other models                       → conservative
                                                               (preserve, no retry)

Affected entry points (all guarded):

  * Pre-filter on outbound:   `FilterThinkingBlocks`
    Previously dropped blocks with missing/invalid signature; now skips
    entirely for non-strict families. Pre-filter is needed because the
    post-error retry path can run out of budget on long conversations
    (maxRetryElapsed = 10 s).
  * 400 retry rectifier:      `FilterThinkingBlocksForRetry`
    Disables top-level thinking and converts thinking → text. Now skips
    for passback-required (those 400s aren't signature errors and any
    transformation breaks the round-trip contract).
  * 400 retry rectifier (tools): `FilterSignatureSensitiveBlocksForRetry`
    Same family-aware short-circuit.
  * 400 detector:             `shouldRectifySignatureError`
    Returns false for passback-required, so the retry path doesn't even
    fire.

Tests
-----

  * `thinking_protocol_test.go` — classifier across all known vendor
    prefixes plus edge cases (empty, case, qwen non-thinking).
  * `thinking_protocol_filter_integration_test.go` — locks in that the
    three filter entry points return the body byte-for-byte unchanged
    when the model id is passback-required or unknown, and still strip
    invalid blocks for anthropic-strict.

This PR supersedes #1350 (which only added the pre-filter without the
upstream-family awareness, and would have made third-party upstreams
worse). Once merged, please close #1350.

Reference issues:
  - NousResearch/hermes-agent#16748 — DeepSeek /anthropic strip behaviour
  - NousResearch/hermes-agent#15700 — DeepSeek thinking:disabled requirement

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-16 19:37:31 +08:00
alfadb 34b1e56e29 test: add 'max' → 'xhigh' test cases for reasoning effort normalization
Cover the DeepSeek reasoning_effort 'max' alias in:
- TestExtractOpenAIReasoningEffortFromBody
- TestExtractCCReasoningEffortFromBody
- TestExtractResponsesReasoningEffortFromBody
2026-06-16 19:37:31 +08:00
alfadb 142d8c3618 fix(gateway): normalize DeepSeek reasoning_effort 'max' to 'xhigh'
pi maps defaultThinkingLevel 'xhigh' → 'max' for deepseek-v4-pro via
thinkingLevelMap, but normalizeOpenAIReasoningEffort did not recognize
'max', causing all reasoning_effort to be dropped (0% fill rate since
traffic switched from claude to deepseek on May 1).

Add 'max' → 'xhigh' mapping to match Claude's NormalizeClaudeOutputEffort
behavior.
2026-06-16 19:37:31 +08:00