Commit Graph
2941 Commits
Author SHA1 Message Date
Wesley Liddick 6db9e42352 Merge pull request #3476 from zeyugao/main
Detect OpenAI overloaded errors by structured code
2026-06-26 09:42:22 +08:00
Wesley Liddick fd0da2570d Merge pull request #3481 from visa2/fix/gateway-glm-codex-tool-args-doubled
fix(apicompat): avoid doubling tool_call arguments from single-chunk upstreams (Codex + GLM)
2026-06-26 09:42:13 +08:00
Wesley Liddick 50c3e52799 Merge pull request #3460 from StarryKira/codex/fix-responses-cache-input
[codex] fix responses cache input for chat completions bridge
2026-06-26 09:42:04 +08:00
Wesley Liddick db1813f7a3 Merge pull request #3434 from feitianbubu/fix/codex-cli-only-chat-completions
fix(gateway): enforce codex_cli_only restriction on /v1/chat/completions
2026-06-26 09:41:46 +08:00
Wesley Liddick 6d50934de6 Merge pull request #3457 from feitianbubu/fix/admin-usage-cache-breakdown
fix(admin/usage): populate cache creation/read token breakdown in stats
2026-06-26 09:41:36 +08:00
Wesley Liddick 27278f62ed Merge pull request #3467 from jianjianai/fix/invalid-refresh-token-nonretryable
fix(token-refresh): treat refresh_token_invalidated as non-retryable / 添加 refresh_token_invalidated 到不可重试列表
2026-06-26 09:40:04 +08:00
Wesley Liddick 38355ed204 Merge pull request #3436 from feitianbubu/fix/email-identity-error-shadowing
fix(auth): stop swallowing email auth identity create error via shadowed err
2026-06-26 09:38:49 +08:00
shaw c9f42e1f77 fix(lint): gofmt token_refresh_service.go after refresh_token_invalidated addition 2026-06-26 09:34:20 +08:00
visa2andClaude Opus 4.8 29122e3051 fix(apicompat): avoid doubling tool_call arguments from single-chunk upstreams
When converting a Chat Completions stream into Responses events, the first
tool_call delta chunk was copied wholesale into stream state (including
function.arguments), then the same chunk's arguments were accumulated again by
the shared `+=` block. For OpenAI this is harmless because its first tool_call
chunk carries empty arguments, but upstreams that pack id+name+arguments into a
single chunk (e.g. GLM/Zhipu) end up with doubled arguments such as
{"cmd":"ls"}{"cmd":"ls"}. Codex then fails to parse the tool call with
"trailing characters", breaking every tool invocation.

Reset the copied arguments so the shared accumulator counts them exactly once,
keeping the emitted delta and the final done/arguments consistent.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-26 00:58:38 +08:00
Elsa cc7612bdbd Detect OpenAI overloaded error codes 2026-06-25 21:18:39 +08:00
DaydreamCodingandClaude Opus 4.8 0112782049 fix(gateway-openai): codex spark 剥离 image_generation 工具,修复 502
gpt-5.3-codex-spark 为 text-only,不支持 image_generation 工具;codex CLI 默认
在 /responses 的 tools[] 携带该工具,导致 ChatGPT 后端返回 400
invalid_request_error(param=tools),被 mapUpstreamError 兜底成 502
"Upstream request failed"(真实 400 见 ops_error_logs.upstream_errors)。

新增 stripCodexSparkImageGenerationTools,对 spark 剥离客户端携带的
image_generation 工具(tools 清空则删键),接入三条路径:
- applyCodexOAuthTransformWithOptions:OAuth 全入口(responses/messages/chat-completions)
- buildUpstreamRequestOpenAIPassthrough 无条件 spark 块:覆盖 OAuth+APIKey
  /responses,不受 image-generation 开关影响
- WS stripCodexSparkImageGenerationToolFromRawPayload:覆盖 WebSocket /responses

bridge 注入(ensureOpenAIResponsesImageGenerationTool)本就跳过 spark;真正的缺口
在 normalizeOpenAIResponsesImageGenerationTools 会规范化并保留客户端工具。
临时绕过=换非 spark 模型;关 bridge/allow_image_generation 无效(工具是客户端带的)。

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 17:27:19 +08:00
简简aw 0a97a5f461 fix(token-refresh): treat refresh_token_invalidated as non-retryable 2026-06-25 15:12:21 +08:00
shaw 00d68ff6df feat(openai): add GPT-5.5 codex instructions and use as latest fallback 2026-06-25 11:44:42 +08:00
haruka dbdbfb1122 fix: avoid default codex instructions for chat bridge 2026-06-25 02:32:15 +08:00
feitianbubu 4567f6582b test(admin/usage): update sqlmock rows for cache breakdown columns 2026-06-24 23:59:22 +08:00
feitianbubu 063454ae96 fix(admin/usage): populate cache creation/read token breakdown in stats 2026-06-24 23:43:03 +08:00
feitianbubu 82576e0a38 fix(auth): stop swallowing email auth identity create error via shadowed err 2026-06-23 19:29:17 +08:00
feitianbubu ae5e980dd1 fix(gateway): enforce codex_cli_only restriction on /v1/chat/completions 2026-06-23 18:38:52 +08:00
github-actions[bot] d430343f51 chore: sync VERSION to 0.1.138 [skip ci] 2026-06-22 01:24:13 +00:00
shaw 6936687870 fix(lint): check WriteString return in summarizeOpenAIImagesNoOutputBody
修复 PR #3381 引入的 errcheck lint 错误,satisfy golangci-lint
对 strings.Builder.WriteString 返回值的检查要求。
2026-06-21 21:43:58 +08:00
Wesley Liddick f079742bab Merge pull request #3371 from cugxuan/feat/subscription-affiliate-rebate
feat: apply affiliate rebate to subscription payments
2026-06-21 21:18:50 +08:00
Wesley Liddick a560d2f91e Merge pull request #3363 from kangjwme/feat/prefer-soonest-reset-scheduling
feat(scheduling): opt-in "prefer soonest reset" account selection
2026-06-21 21:17:44 +08:00
Wesley Liddick 6e27b82335 Merge pull request #3381 from 404QAQ/fix/images-incomplete-failover
fix(images): 识别 response.incomplete 触发 failover + 记录软失败上游响应
2026-06-21 21:14:48 +08:00
Wesley Liddick f6d734ec23 Merge pull request #3308 from a0yark/fix/gemini-tool-schema-cleanup
fix(gemini): clean unsupported tool schema fields
2026-06-21 21:14:09 +08:00
Wesley Liddick af1a032d35 Merge pull request #3375 from StarryKira/fix/3358-vertex-beta-and-cch
fix(gateway): filter anthropic-beta on the Vertex Anthropic path + drop cch sign (#3358)
2026-06-21 21:13:14 +08:00
Wesley Liddick 93de19665b Merge pull request #3359 from alfadb/fix/glm-effort-mapping
修复 GLM 推理强度映射
2026-06-21 21:12:46 +08:00
Wesley Liddick bed12b7016 Merge pull request #3360 from feitianbubu/fix/auto-mode-cc-entrypoint-ide
fix(gateway): in auto mode recognize Claude Code IDE clients via any cc_entrypoint
2026-06-21 21:12:33 +08:00
Wesley Liddick f597e926da Merge pull request #3335 from FjlI5/fix/openai-upstream-endpoint-logging
fix: 修正 chat-only API-key 账号上游端点被误记为 /v1/responses
2026-06-21 21:12:15 +08:00
Wesley Liddick 7b5fbe5197 Merge pull request #3364 from feitianbubu/fix/promo-clear-expiry
fix(promo): allow clearing promo code expiry on edit
2026-06-21 21:12:02 +08:00
404QAQ b0d5592ae2 fix(images): 识别 response.incomplete + 记录软失败上游响应
修复 gpt-image-2 大图/编辑请求 502 报错(社区 issue #2232/#3135/#2516,
现象:upstream did not return image output)。根因是上游生成超时/截断时返回
response.incomplete,旧逻辑只认 error/response.failed,导致:
1) 软失败报成模糊 502 且不触发 failover 换账号重试
2) 上游真实响应未记录,ops_error_logs 里上游信息全空,无法排查

- openAIImagesUpstreamErrorFromSSEPayload 增加 response.incomplete 识别:
  生成超时/截断(max_output_tokens 等) → 可重试 502 触发 failover;
  content_filter/moderation → 400 不重试
- 软失败兜底(无图无标准错误)记录上游诊断摘要到 ops(last_event/status/
  incomplete_reason/body 片段),非流式+流式两条路径一致
- summarizeOpenAIImagesNoOutputBody 提取诊断信息,body 截断上限 1KB

基于官方 v0.1.137 净分支。测试: 5 个新单测 + 图片回归通过 (-tags unit)
2026-06-21 08:34:12 +08:00
harukaandClaude Opus 4.8 5cb8cdd36c test(claude-code): detection recognizes the new-CLI billing block (no cch)
Locks in that Claude Code detection keys on the billing block prefix +
cc_entrypoint=cli, not on the cch field that the new CLI (and now our own
mimicry) no longer sends:

- BillingBlockRecognizedWithoutCCH: an identity-prose-less sub-request whose
  system block is `x-anthropic-billing-header: cc_version=...; cc_entrypoint=cli;`
  (no cch) is still detected as Claude Code.
- NoCCHBlockStillRequiresClaudeCodeUA: dropping cch did not loosen detection —
  a non-claude-cli UA is still rejected, so ClaudeCodeOnly groups can't be spoofed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 08:02:44 -07:00
harukaandClaude Opus 4.8 6cfb7898df fix(claude-mimicry): drop the cch sign to match new Claude Code CLI
Recent Claude Code CLI versions no longer emit the cch=... signature field in
their x-anthropic-billing-header system block (issue #3358). sub2api still
injected cch=00000 when mimicking Claude Code for OAuth accounts and optionally
signed it, so mimicked requests now diverge from real CLI traffic — the opposite
of what the mimicry is for.

- buildBillingAttributionText emits the block without the cch=00000 segment;
  cc_version + cc_entrypoint=cli are kept (detection and Anthropic's first-party
  signal rely on the block, not on cch).
- Retire signing: remove the two enableCCH signBillingHeaderCCH call sites in
  buildUpstreamRequest / buildCountTokensRequest and delete the now-dead
  signBillingHeaderCCH, cchPlaceholderRe, cchSeed, xxHash64Seeded helpers.
- enable_cch_signing is now a documented no-op (kept for backward compat).
- Drop the obsolete signing tests (TestSignBillingHeaderCCH, TestXXHash64Seeded,
  TestSanitizeMustBeBeforeCCHSigning_HashConsistency) and update the prompt test
  to assert the injected block no longer carries cch=.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 08:01:29 -07:00
harukaandClaude Opus 4.8 efffd5d791 test(gateway): Vertex anthropic-beta filtering
Covers the #3358 fix:
- StripsUnsupportedClaudeCodeTokens reproduces the prod 400 — the four Vertex-
  rejected tokens (advisor-tool, prompt-caching-scope, redact-thinking,
  thinking-token-count) plus the identity betas are stripped while whitelisted
  tokens survive. Fails before the builder fix, passes after.
- DropsHeaderWhenAllUnsupported: no anthropic-beta header is sent when every
  client token is filtered out.
- BodySanitizeKeysOnFinalBeta: body.context_management is stripped based on the
  final beta, not the raw client value.
- BlocksViaBetaPolicy: an admin block rule on a Vertex account returns BetaBlockedError.
- TestFilterVertexBetaTokens unit-tests whitelist/drop-set/dedupe/empty.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 07:53:32 -07:00
harukaandClaude Opus 4.8 40e1cc14b3 fix(gateway): filter anthropic-beta on the Vertex Anthropic path (#3358)
Vertex AI's Anthropic endpoint rejects unknown anthropic-beta tokens with
HTTP 400. buildUpstreamRequestAnthropicVertex forwarded the client header
verbatim via the allowedHeaders whitelist, so recent Claude Code CLIs that
send advisor-tool-2026-03-01, prompt-caching-scope-2026-01-05,
redact-thinking-2026-02-12 and thinking-token-count-2026-05-13 broke every
Vertex service_account request, even though plain account-test requests passed.

This is the only upstream builder that bypassed beta filtering: the
OAuth/API-key path uses computeFinalAnthropicBeta and the Bedrock path uses
filterBedrockBetaTokens. Close the gap with a Vertex-specific whitelist
(vertexSupportedBetaTokens) mirroring bedrockSupportedBetaTokens, plus the
existing BetaPolicy block check:

- evaluateBetaPolicy block check (symmetric to resolveBedrockBetaTokensForRequest)
- filterVertexBetaTokens strips policy-filtered + non-whitelisted tokens
- body context_management sanitize now keys on the final beta, not the raw client value
- overwrite the anthropic-beta header after the whitelist copy loop with the final value

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 07:51:24 -07:00
wucm667 ecedc7c8d3 fix(auth): enforce email bind suffix whitelist 2026-06-19 21:17:45 +08:00
cugxuan 0fa604ba71 feat: apply affiliate rebate to subscription payments 2026-06-19 17:47:10 +08:00
feitianbubu 2dc1387b59 fix(promo): allow clearing promo code expiry on edit 2026-06-18 23:07:25 +08:00
kangjwme 510adf703c feat(scheduling): add opt-in "prefer soonest reset" account selection
Adds a use-it-or-lose-it scheduling strategy: prefer accounts whose
session window resets soonest, so near-reset accounts get drained first
instead of accounts whose reset is still far away.

Both schedulers, opt-in, default behavior unchanged:

- Anthropic (gateway_service.go): new GatewaySchedulingConfig
  .PreferSoonestReset flag. When on, the layered load-aware selection
  inserts a filterBySoonestReset stage (priority -> soonest-reset ->
  load -> LRU). Accounts with no active SessionWindowEnd are treated as
  lowest priority; ties fall through to LRU.

- OpenAI/Codex (openai_account_scheduler.go): new "reset" score weight
  in GatewayOpenAIWSSchedulerScoreWeights. Soonest-reset accounts score
  higher; weight defaults to 0 (no effect).

SessionWindowEnd (upstream 5h/quota ResetsAt) is already carried in the
scheduler snapshot, so no snapshot changes are needed.

Documented in deploy/config.example.yaml. Adds unit tests for the
Anthropic filter and the OpenAI reset factor.
2026-06-18 22:50:46 +08:00
feitianbubu e3e31bd4c5 fix(gateway): auto mode recognize Claude Code IDE clients via any cc_entrypoint 2026-06-18 19:22:10 +08:00
alfadb 89cfe24a0c fix(openai): normalize glm reasoning effort 2026-06-18 19:07:08 +08:00
FjlI5 bab8a9a93e fix(openai): log /v1/chat/completions upstream endpoint for chat-only API-key accounts
DeriveUpstreamEndpoint maps every OpenAI-platform request to /v1/responses, but
API-key accounts whose upstream only speaks Chat Completions
(!ShouldUseResponsesAPI) are forwarded directly to /v1/chat/completions. The
messages, responses and cyber-policy recording sites derived the endpoint via the
bare GetUpstreamEndpoint, so usage/ops records mislabeled those requests as
/v1/responses. Generalize the existing resolveRawCCUpstreamEndpoint into
resolveOpenAIUpstreamEndpoint and use it at every OpenAI recording site, matching
the already-correct chat-completions client path.
2026-06-18 00:29:04 +08:00
a0yarkandCursor 8c4a43cf72 fix(gemini): satisfy schema cleanup test lint
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-16 23:19:45 +08:00
github-actions[bot] 4a5665da5b chore: sync VERSION to 0.1.137 [skip ci] 2026-06-16 12:59:59 +00:00
Wesley Liddick de38d62360 Merge pull request #3250 from WesleyZiwen/codex/anthropic-429-window-reset
fix: preserve Anthropic window cooldowns
2026-06-16 20:34:36 +08:00
Wesley Liddick 9e9e154f57 Merge pull request #3306 from bestony/fix/ip-acl-denial-message
fix(auth): include client ip in acl denial message
2026-06-16 20:32:50 +08:00
Wesley Liddick 7fb3e1aacd Merge pull request #3247 from alfadb/fix/reasoning-and-thinking-protocol
fix(gateway): 整合推理强度与思考协议处理(替代 #2155 / #2136 / #3246)
2026-06-16 20:29:21 +08:00
Wesley Liddick f204216918 Merge pull request #3243 from alfadb/feature/chinese-llm-fallback-pricing
feat(billing): 国产 LLM 兜底定价 (GLM / Kimi / MiniMax) + 收编 DeepSeek V4
2026-06-16 20:28:49 +08:00
Wesley Liddick 03ec90a25a Merge pull request #3223 from feitianbubu/fix/intercept-streaming-haiku-probe
fix: 修复CC Switch改为流式测试后请求失败的问题
2026-06-16 20:27:33 +08:00
shaw 6c7203d83b fix(gateway): preserve SSE event:error body so ops logs reflect real upstream errors
When an Anthropic upstream returned HTTP 200 but then emitted an SSE
`event: error` frame (overloaded_error / rate_limit_error / api_error /
etc.), Forward's stream branch matched on `err.Error() == "have error in
stream"` and returned `UpstreamFailoverError{StatusCode: 403}` with no
ResponseBody. That dropped three pieces of evidence:

- handleFailoverExhausted → ExtractUpstreamErrorMessage(nil) = "" →
  ops_error_logs.upstream_error_message was empty.
- errorPassthroughService.MatchRule(_, 403, nil) could only match rules
  without keywords, so keyword-based passthrough rules silently never
  fired.
- upstream_errors carried no stream_error record, leaving ops looking at
  a generic 403 with no clue whether the upstream was throttled,
  overloaded, or rejecting the request. ping-during-slot-wait amplified
  this by skipping failover (writerSizeBeforeForward guard), so 403s
  ballooned in the ops view well past the upstream's actual rate.

Fix:
- Introduce *sseStreamErrorEventError that carries the SSE data line.
  Error() still returns "have error in stream" so existing log searches
  keep working.
- Forward extracts via errors.As, appends an OpsUpstreamErrorEvent
  (kind="stream_error", with the sanitized message and a truncated raw
  body honoring LogUpstreamErrorBody*), and returns
  UpstreamFailoverError{StatusCode: 403, ResponseBody: rawJSON}.

StatusCode 403 is preserved verbatim: mapUpstreamError, failover
decisions (shouldFailoverUpstreamError(403)=true), client-visible message,
RetryableOnSameAccount, and rateLimitService side-effects (this path
already didn't invoke them) all match prior behavior. OAuth and API Key
accounts share this path; the API-Key passthrough branch is independent
and already forwards SSE error frames untouched, so it's unaffected.

Adds four unit tests: typed-error contract + RawData, empty data line,
event:error after partial stream output (streamStarted=true), and
non-JSON data line.
2026-06-16 20:25:09 +08:00
alfadb a05d9e87c0 feat(billing): 国产模型 thinking-enabled 自动填充 reasoning_effort 默认值
问题:Kimi/GLM/MiniMax 等国产 LLM 协议层只有 thinking on/off 开关,没有
reasoning_effort 档位概念。客户端启用 thinking 后 usage_log.reasoning_effort
长期为 NULL,无法在用量分析里区分 'thinking 开启' 与 'thinking 关闭'。

方案:仅在 'thinking 启用 + 上游属于 passback-required 国产模型 + 客户端
未明确指定 effort' 三者同时成立时,给 usage_log.reasoning_effort 写默认值
'high'(与 DeepSeek thinking-enabled 默认 effort 一致)。

设计原则:
1. **白名单**:仅 ResolveThinkingProtocol == PassbackRequired 集合内,
   且排除原生支持 effort 的 DeepSeek(避免覆盖客户端意图)。
2. **fail-open**:客户端显式传 effort 时永远不覆盖。
3. **未来兼容**:如 Kimi 后续加入真 effort 档位,客户端开始发 effort,
   guard (3) 自动让出,本逻辑变 no-op。

实现:
- gateway_request.go: 加 DefaultEffortForThinkingEnabled (按模型白名单)
  + OpenAIBodyHasThinkingEnabled (检测 OpenAI 协议 body 里的 thinking.type)
  + ApplyThinkingEnabledFallback (包装现有 extractor 的 nil-then-default 逻辑)
- gateway_handler.go: Anthropic 路径两处(主 + retry)对称补充
- OpenAI 路径全覆盖:openai_gateway_service.go (passthrough + non-passthrough)
  + openai_gateway_chat_completions_raw.go + openai_gateway_responses_chat_fallback.go
  + openai_ws_http_bridge.go + openai_ws_forwarder.go
- 跨协议路径:gateway_forward_as_chat_completions.go (CC client → Anthropic upstream)
  + gateway_forward_as_responses.go (Responses client → Anthropic upstream)
  + gemini_chat_completions_compat_service.go (一致性保持)

未覆盖:openai_ws_v2_passthrough_adapter.go 两处。原因:该 adapter 持有的是
session-level 客户端原始 model,没有 *Account 句柄无法走 GetMappedModel。
WS v2 当前对国产模型场景不重要(pi 调用 Kimi/GLM/MiniMax 走 sync HTTP),
留待后续如果出现 WS v2 + 国产模型用例时单独处理。

测试:
- TestDefaultEffortForThinkingEnabled (14 用例):覆盖 Kimi/GLM/MiniMax 大小写、
  Qwen thinking 变体、DeepSeek 排除、Claude/GPT/Gemini 不命中。
- TestOpenAIBodyHasThinkingEnabled (8 用例):covers enabled/adaptive/disabled、
  大小写、空 body、缺字段、invalid JSON fail-safe。
- TestApplyThinkingEnabledFallback (9 用例):现有 effort 不覆盖、nil + 启用 +
  passback → high、nil + disabled → nil、nil + 启用 + 排除模型 → nil。
2026-06-16 19:37:31 +08:00