3932 Commits
Author SHA1 Message Date
Wesley Liddick eba9bea959 Merge pull request #3278 from 315944211/feat/show-account-id-in-admin-list
feat: show account id in account list
v0.1.137
2026-06-16 20:35:55 +08:00
Wesley Liddick de38d62360 Merge pull request #3250 from WesleyZiwen/codex/anthropic-429-window-reset
fix: preserve Anthropic window cooldowns
2026-06-16 20:34:36 +08:00
Wesley Liddick 9e9e154f57 Merge pull request #3306 from bestony/fix/ip-acl-denial-message
fix(auth): include client ip in acl denial message
2026-06-16 20:32:50 +08:00
Wesley Liddick 7fb3e1aacd Merge pull request #3247 from alfadb/fix/reasoning-and-thinking-protocol
fix(gateway): 整合推理强度与思考协议处理(替代 #2155 / #2136 / #3246)
2026-06-16 20:29:21 +08:00
Wesley Liddick f204216918 Merge pull request #3243 from alfadb/feature/chinese-llm-fallback-pricing
feat(billing): 国产 LLM 兜底定价 (GLM / Kimi / MiniMax) + 收编 DeepSeek V4
2026-06-16 20:28:49 +08:00
Wesley Liddick 03ec90a25a Merge pull request #3223 from feitianbubu/fix/intercept-streaming-haiku-probe
fix: 修复CC Switch改为流式测试后请求失败的问题
2026-06-16 20:27:33 +08:00
shaw 6c7203d83b fix(gateway): preserve SSE event:error body so ops logs reflect real upstream errors
When an Anthropic upstream returned HTTP 200 but then emitted an SSE
`event: error` frame (overloaded_error / rate_limit_error / api_error /
etc.), Forward's stream branch matched on `err.Error() == "have error in
stream"` and returned `UpstreamFailoverError{StatusCode: 403}` with no
ResponseBody. That dropped three pieces of evidence:

- handleFailoverExhausted → ExtractUpstreamErrorMessage(nil) = "" →
  ops_error_logs.upstream_error_message was empty.
- errorPassthroughService.MatchRule(_, 403, nil) could only match rules
  without keywords, so keyword-based passthrough rules silently never
  fired.
- upstream_errors carried no stream_error record, leaving ops looking at
  a generic 403 with no clue whether the upstream was throttled,
  overloaded, or rejecting the request. ping-during-slot-wait amplified
  this by skipping failover (writerSizeBeforeForward guard), so 403s
  ballooned in the ops view well past the upstream's actual rate.

Fix:
- Introduce *sseStreamErrorEventError that carries the SSE data line.
  Error() still returns "have error in stream" so existing log searches
  keep working.
- Forward extracts via errors.As, appends an OpsUpstreamErrorEvent
  (kind="stream_error", with the sanitized message and a truncated raw
  body honoring LogUpstreamErrorBody*), and returns
  UpstreamFailoverError{StatusCode: 403, ResponseBody: rawJSON}.

StatusCode 403 is preserved verbatim: mapUpstreamError, failover
decisions (shouldFailoverUpstreamError(403)=true), client-visible message,
RetryableOnSameAccount, and rateLimitService side-effects (this path
already didn't invoke them) all match prior behavior. OAuth and API Key
accounts share this path; the API-Key passthrough branch is independent
and already forwards SSE error frames untouched, so it's unaffected.

Adds four unit tests: typed-error contract + RawData, empty data line,
event:error after partial stream output (streamStarted=true), and
non-JSON data line.
2026-06-16 20:25:09 +08:00
alfadb a05d9e87c0 feat(billing): 国产模型 thinking-enabled 自动填充 reasoning_effort 默认值
问题:Kimi/GLM/MiniMax 等国产 LLM 协议层只有 thinking on/off 开关,没有
reasoning_effort 档位概念。客户端启用 thinking 后 usage_log.reasoning_effort
长期为 NULL,无法在用量分析里区分 'thinking 开启' 与 'thinking 关闭'。

方案:仅在 'thinking 启用 + 上游属于 passback-required 国产模型 + 客户端
未明确指定 effort' 三者同时成立时,给 usage_log.reasoning_effort 写默认值
'high'(与 DeepSeek thinking-enabled 默认 effort 一致)。

设计原则:
1. **白名单**:仅 ResolveThinkingProtocol == PassbackRequired 集合内,
   且排除原生支持 effort 的 DeepSeek(避免覆盖客户端意图)。
2. **fail-open**:客户端显式传 effort 时永远不覆盖。
3. **未来兼容**:如 Kimi 后续加入真 effort 档位,客户端开始发 effort,
   guard (3) 自动让出,本逻辑变 no-op。

实现:
- gateway_request.go: 加 DefaultEffortForThinkingEnabled (按模型白名单)
  + OpenAIBodyHasThinkingEnabled (检测 OpenAI 协议 body 里的 thinking.type)
  + ApplyThinkingEnabledFallback (包装现有 extractor 的 nil-then-default 逻辑)
- gateway_handler.go: Anthropic 路径两处(主 + retry)对称补充
- OpenAI 路径全覆盖:openai_gateway_service.go (passthrough + non-passthrough)
  + openai_gateway_chat_completions_raw.go + openai_gateway_responses_chat_fallback.go
  + openai_ws_http_bridge.go + openai_ws_forwarder.go
- 跨协议路径:gateway_forward_as_chat_completions.go (CC client → Anthropic upstream)
  + gateway_forward_as_responses.go (Responses client → Anthropic upstream)
  + gemini_chat_completions_compat_service.go (一致性保持)

未覆盖:openai_ws_v2_passthrough_adapter.go 两处。原因:该 adapter 持有的是
session-level 客户端原始 model,没有 *Account 句柄无法走 GetMappedModel。
WS v2 当前对国产模型场景不重要(pi 调用 Kimi/GLM/MiniMax 走 sync HTTP),
留待后续如果出现 WS v2 + 国产模型用例时单独处理。

测试:
- TestDefaultEffortForThinkingEnabled (14 用例):覆盖 Kimi/GLM/MiniMax 大小写、
  Qwen thinking 变体、DeepSeek 排除、Claude/GPT/Gemini 不命中。
- TestOpenAIBodyHasThinkingEnabled (8 用例):covers enabled/adaptive/disabled、
  大小写、空 body、缺字段、invalid JSON fail-safe。
- TestApplyThinkingEnabledFallback (9 用例):现有 effort 不覆盖、nil + 启用 +
  passback → high、nil + disabled → nil、nil + 启用 + 排除模型 → nil。
2026-06-16 19:37:31 +08:00
alfadb 5c5283979b doc(thinking-protocol): clarify mappedModel vs originalModel semantics per call path
回应 PR #3247 Copilot review:

1. NormalizeChineseLLMThinking 文档写 'MiniMax M3 / M3.x' 但实现覆盖
   minimax-m* (M2.x 也命中)。改成 'MiniMax M-series (M2.x/M3.x)'。

2. Copilot 质疑 gemini_messages_compat 传 originalModel 不是
   mappedModel 是误用。实际是有意为之的跨协议场景:
   - 上游是 Gemini,但被剥离的 body 是 Anthropic 格式
   - 剥离逻辑要按客户端请求的 Anthropic 子协议族判定
   - 传 mappedModel (gemini-3.1-pro) 会被判为 Unknown→不剥离→retry 死循环

   未改代码逻辑,改为更新文档:
   - thinking_protocol.go ResolveThinkingProtocol 添加「调用路径语义」
     说明 Anthropic gateway 与 Gemini compat 两条路径的参数语义差异。
   - gemini_messages_compat_service.go 调用点加 inline 注释解释
     路径语义为什么该传 originalModel。
2026-06-16 19:37:31 +08:00
alfadb 56c6325d15 fix(gateway): rewrite thinking.type=enabled to adaptive for MiniMax M-series
\u95ee\u9898\uff1aAnthropic-SDK \u5ba2\u6237\u7aef\uff08\u5982 pi-ai\uff09\u5728 budget-based thinking \u8def\u5f84\u4e0a\u9ed8\u8ba4\u53d1
`thinking: { type: 'enabled', budget_tokens: ... }` \u7ed9\u4efb\u4f55 Anthropic \u517c\u5bb9\u4e0a\u6e38\uff0c\u5305\u62ec
MiniMax M3 / M2.x \u3002\u4f46 MiniMax \u5b98\u65b9\u6587\u6863\u660e\u786e\u5217\u51fa\uff1athinking.type \u53ea\u63a5\u53d7
"adaptive" \u6216 "disabled"\uff0c\u4e0d\u63a5\u53d7 "enabled"\u3002\u9020\u6210\u4e0a\u6e38\u53ef\u80fd 400 / \u9759\u9ed8\u5ffd\u7565\u3002

\u4fee\u590d\uff1a\u5728 Anthropic forward \u8def\u5f84\u4e0a\u3001StripEmptyTextBlocks \u4e4b\u540e\u3001\u91cd\u8bd5\u5faa\u73af\u4e4b\u524d\uff0c
\u52a0\u4e00\u4e2a NormalizeChineseLLMThinking \u9884\u6539\u5199\u6b65\u9aa4\uff1a
- \u4ec5\u5bf9 ResolveThinkingProtocol == PassbackRequired \u4e0a\u6e38\u751f\u6548\uff08claude-* \u4e0d\u4f1a\u8fdb\u6765\uff09
- \u4ec5\u5339\u914d minimax-m* \u524d\u7f00\u7684\u6620\u5c04\u540e\u6a21\u578b ID
- \u4ec5\u5f53 thinking.type == "enabled" \u65f6\u6539\u5199\u4e3a "adaptive"
- Kimi/GLM/DeepSeek \u7684 enabled \u4e0d\u53d7\u5f71\u54cd\uff08\u8fd9\u4e9b\u4e0a\u6e38\u63a5\u53d7 enabled\uff09

\u8bbe\u8ba1\u539f\u5219\uff1a
1. "\u767d\u540d\u5355\u8bed\u4e49"\uff1a\u53ea\u5bf9\u5df2\u77e5\u9700\u8981\u5904\u7406\u7684\u4e0a\u6e38\u6539\u5199\uff0c\u4e0d\u505a\u9690\u5f0f\u5e7f\u8c31\u8865\u4e01\u3002
2. "\u4f4e\u4fb5\u5165\u6027"\uff1aMiniMax \u4e4b\u5916\u7684\u6a21\u578b\u96f6\u6539\u52a8\uff0c\u73b0\u6709 4 \u4e2a\u4e0a\u6e38\u7684\u884c\u4e3a\u4e0d\u53d8\u3002
3. "\u53ef\u8bca\u65ad"\uff1a\u6539\u5199\u65f6 LegacyPrintf \u8f93\u51fa\u4e00\u884c\u65e5\u5fd7\uff0c\u4fbf\u4e8e\u8c03\u8bd5\u3002

\u6d4b\u8bd5\uff1a9 \u4e2a\u65b0\u7528\u4f8b\u8986\u76d6 MiniMax M3/M2.7\u3001adaptive/disabled \u4e0d\u53d8\u3001\u65e0 thinking
\u5b57\u6bb5\u3001Kimi/GLM/DeepSeek/Claude \u4e0d\u53d7\u5f71\u54cd\u3001invalid JSON fail-safe\u3002

\u4f9d\u8d56\uff1a\u672c\u5206\u652f\u57fa\u4e8e fix/thinking-block-protocol-aware-filter\uff0c\u7528\u5176
ResolveThinkingProtocol \u4f5c\u4e3a passback-required \u5224\u522b\u6e90\u3002\u4e0a\u6e38 PR \u987a\u5e8f\u9700\u5148\u5408
fix/thinking-block-protocol-aware-filter\uff08PR #2136\uff09\u518d\u5408\u672c PR\u3002
2026-06-16 19:37:31 +08:00
alfadb efbf6d2092 fix(test): update FilterThinkingBlocksForRetry call to use mappedModel param
The function signature was changed to require a mappedModel parameter
for protocol-aware thinking-block filtering, but this test call site
was not updated. Without the fix, the unit test build fails on CI.
2026-06-16 19:37:31 +08:00
alfadbandClaude Opus 4.7 6baf00d784 fix(gateway): protocol-aware thinking-block filtering for Anthropic-compatible upstreams
The gateway's thinking-block handling was designed for Anthropic's strict
semantics (drop blocks with missing/invalid signature), but third-party
Anthropic-compatible upstreams have INVERTED semantics:

  * DeepSeek `/anthropic`, Kimi `/coding`, GLM, Moonshot, qwen-*-thinking
    require ALL historical thinking blocks to round-trip verbatim.
  * Stripping any of them produces:
    400 "The content[].thinking in the thinking mode must be passed back
    to the API"

Without this fix, every multi-turn request from a thinking-capable client
(Claude Code, pi, etc.) to such upstreams loses its thinking blocks and
fails. This becomes especially painful when an account's model_mapping
maps `claude-sonnet-4-6 → deepseek-v4-pro` — `reqModel` looks Anthropic
but the upstream contract is the opposite.

Approach
--------

Branch all thinking-block transforms by the *mapped* model id (after
account model_mapping is applied), classifying into three families:

  * `anthropic-strict`     claude-/opus-/sonnet-/haiku-      → existing behaviour
  * `passback-required`    deepseek-/kimi-/moonshot-/glm-/   → preserve verbatim
                           qwen-*-thinking
  * `unknown`              other models                       → conservative
                                                               (preserve, no retry)

Affected entry points (all guarded):

  * Pre-filter on outbound:   `FilterThinkingBlocks`
    Previously dropped blocks with missing/invalid signature; now skips
    entirely for non-strict families. Pre-filter is needed because the
    post-error retry path can run out of budget on long conversations
    (maxRetryElapsed = 10 s).
  * 400 retry rectifier:      `FilterThinkingBlocksForRetry`
    Disables top-level thinking and converts thinking → text. Now skips
    for passback-required (those 400s aren't signature errors and any
    transformation breaks the round-trip contract).
  * 400 retry rectifier (tools): `FilterSignatureSensitiveBlocksForRetry`
    Same family-aware short-circuit.
  * 400 detector:             `shouldRectifySignatureError`
    Returns false for passback-required, so the retry path doesn't even
    fire.

Tests
-----

  * `thinking_protocol_test.go` — classifier across all known vendor
    prefixes plus edge cases (empty, case, qwen non-thinking).
  * `thinking_protocol_filter_integration_test.go` — locks in that the
    three filter entry points return the body byte-for-byte unchanged
    when the model id is passback-required or unknown, and still strip
    invalid blocks for anthropic-strict.

This PR supersedes #1350 (which only added the pre-filter without the
upstream-family awareness, and would have made third-party upstreams
worse). Once merged, please close #1350.

Reference issues:
  - NousResearch/hermes-agent#16748 — DeepSeek /anthropic strip behaviour
  - NousResearch/hermes-agent#15700 — DeepSeek thinking:disabled requirement

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-16 19:37:31 +08:00
alfadb 34b1e56e29 test: add 'max' → 'xhigh' test cases for reasoning effort normalization
Cover the DeepSeek reasoning_effort 'max' alias in:
- TestExtractOpenAIReasoningEffortFromBody
- TestExtractCCReasoningEffortFromBody
- TestExtractResponsesReasoningEffortFromBody
2026-06-16 19:37:31 +08:00
alfadb 142d8c3618 fix(gateway): normalize DeepSeek reasoning_effort 'max' to 'xhigh'
pi maps defaultThinkingLevel 'xhigh' → 'max' for deepseek-v4-pro via
thinkingLevelMap, but normalizeOpenAIReasoningEffort did not recognize
'max', causing all reasoning_effort to be dropped (0% fill rate since
traffic switched from claude to deepseek on May 1).

Add 'max' → 'xhigh' mapping to match Claude's NormalizeClaudeOutputEffort
behavior.
2026-06-16 19:37:31 +08:00
alfadb 262fe1230d feat(billing): 为 doubao-embedding-vision 添加图文差别兜底定价
火山方舟 doubao-embedding-vision 多模态向量化按量付费官方价为文本 ¥0.7/MTok、图片 ¥1.8/MTok,二者不同价。原计费引擎对 embedding 仅记单一 input token、无图片输入档位,无法表达该差别。

变更:
- ModelPricing 新增 ImageInputPricePerToken 字段
- OpenAIUsage / UsageTokens 新增 ImageInputTokens 字段
- extractOpenAIEmbeddingsUsage 解析 usage.prompt_tokens_details.image_tokens
- CalculateCost 拆分文本/图片输入计费;ImageInputTokens 为 0 时走原单价路径,存量 chat/vision 流量行为不变
- getFallbackPricing 新增 doubao-embedding-vision 分支(most-specific-first,覆盖带版本后缀别名),fallback 表填入 $0.098/$0.252 per MTok(汇率 ÷7.14)

测试:新增图文混合/纯文本/图片 token 超额三类用例及定价回退断言;service 包单测与 golangci-lint 全通过。
2026-06-16 19:37:02 +08:00
alfadb 4f5f2788e1 fix(billing): add kimi-for-coding fallback pricing 2026-06-16 19:37:02 +08:00
alfadb c90089c814 fix(billing): address Copilot review feedback
应 PR #3243 review 修正三项问题:

1. 删除 DeepSeek V4 fallback 定价的重复块(cherry-pick 时
   的旧位置 + 新国产 LLM 分段都写了一遍,后写覆盖前写,
   当前值相同所以无 bug,但冗余,未来改一处易漏另一处)。
2. 删除 minimax-m3 匹配中的死代码——modelLower 已经
   strings.ToLower(),'MiniMax-M3' 字面量永远不可能匹中。
3. 测试结构改 *float64 替代 0 sentinel,让 free-tier
   模型(GLM-4.5-Flash / GLM-4.7-Flash)的 0 价能被
   真正断言而非默默跳过;定义内联 floatPtr 辅助函数。
2026-06-16 19:37:02 +08:00
alfadb a4ce73391b feat(billing): add GLM / Kimi / MiniMax fallback pricing for Chinese LLM providers
将 fix/deepseek-fallback-pricing 的 DeepSeek V4 Pro/Flash 定价收编,
并扩展国产 LLM 兜底定价覆盖:

- 智谱 GLM (z.ai 公开 SKU 13 个): glm-5.1 / glm-5 / glm-5-turbo / glm-4.7 /
  glm-4.7-flashx / glm-4.6 / glm-4.5 / glm-4.5-x / glm-4.5-air / glm-4.5-airx /
  glm-4-32b-0414-128k / glm-4.5-flash / glm-4.7-flash
- 月之暗面 Kimi K 系列 4 个: kimi-k2.6 / kimi-k2.5 / kimi-k2-thinking /
  kimi-k2 (K2-0905/K2-0711 官方未保留定价,隐性回退到 kimi-k2)
- MiniMax M 系列 6 个: minimax-m3 (≤512K) / minimax-m2.7 / minimax-m2.7-highspeed /
  minimax-m2.5 / minimax-m2.1 / minimax-m2
- DeepSeek V4 沿用原 fix/deepseek-fallback-pricing 的 entry 移植过来

所有定价数据来自各家官方定价页 (USD/MTok 口径),与现有
Claude/GPT 风格保持一致 (硬编码在 Go 代码里,admin channel 优先)。

匹配策略:长 key 优先 (glm-5.1 优先于 glm-5, k2.6 优先于 k2);
未列出的国产厂商 alias (qwen/doubao/hunyuan) 一律不返回兜底价,
避免误计价——这与 fix/deepseek-fallback-pricing 的白名单语义一致。

数据源注释 (Source URL) 内嵌在每个 initFallbackPricing 分组开头,
便于后续 PR 同步上游价格变化。
2026-06-16 19:37:02 +08:00
alfadb f597d98bdf test(openai): use unpriced model in usage test 2026-06-16 19:37:02 +08:00
alfadb 5a593a511e test(billing): tighten DeepSeek V4 fallback assertions; clarify branch comments
Address copilot-pull-request-reviewer feedback on #2157:

- billing_service_test.go: extend TestGetFallbackPricing_FamilyMatching
  with optional expectedOutput / expectedCacheRead fields and assert
  full Input/Output/CacheRead pricing for all 4 DeepSeek cases
  (v4-pro, v4-flash, deepseek-chat→flash, deepseek-reasoner→flash),
  preventing silent regression to 0 output/cache cost.
- billing_service.go: rewrite the DeepSeek block comment to explicitly
  describe its scope (V4 Pro/Flash + chat/reasoner aliases, no
  unknown-deepseek fallback) and tighten the OpenAI comment to make
  it unambiguous that it only describes the OpenAI/Codex branch
  immediately below it.
2026-06-16 19:37:02 +08:00
alfadb 27e26a3a90 chore: fix gofmt alignment 2026-06-16 19:37:02 +08:00
alfadb c906bf000e feat(billing): add DeepSeek V4 Pro / Flash fallback pricing
Pricing source: https://api-docs.deepseek.com/quick_start/pricing

deepseek-v4-pro:  input $0.435/M, output $0.87/M, cache_read $0.003625/M
deepseek-v4-flash: input $0.14/M,  output $0.28/M, cache_read $0.0028/M

Also map legacy aliases deepseek-chat / deepseek-reasoner → v4-flash.
2026-06-16 19:37:02 +08:00
shaw b8a482e127 fix(ci): unblock main after recent merges
Three independent CI blockers landed on main from concurrent PR merges:

- openai_quota_service.go (introduced by b8169492): const block spacing
  not gofmt-compliant + trailing blank line. golangci-lint v2.9 flagged it
  on every push after the merge.
- openai_images_failover_test.go (introduced by PR #3155, da30c599):
  NewOpenAIGatewayHandler call missing the opsService argument added by
  PR #3230 (b62b573f). Test was authored before #3230 and merged without
  rebase, causing "not enough arguments" compile error.
- account_quota_reset_test.go: TestIsFixedDailyPeriodExpired_NotExpired
  and TestIsFixedWeeklyPeriodExpired_NotExpired used time.Now()-1min as
  periodStart, which crosses the 09:00 UTC reset boundary when CI runs in
  the 09:00:00-09:00:59 window. Anchoring periodStart to today's 12:00
  UTC removes the race.
2026-06-16 17:59:06 +08:00
Wesley Liddick 44f5791008 Merge pull request #3154 from wucm667/fix/responses-session-hash-input-field
fix(openai): Responses API session hash fallback 纳入 input 内容,避免 sticky 绑死账号
2026-06-16 17:01:16 +08:00
Wesley Liddick 9c2c8ab3e5 Merge pull request #3155 from wucm667/fix/openai-images-server-error-failover
fix(openai): 图像接口上游 server_error 触发 failover 换号,不再 5xx 透传
2026-06-16 17:00:40 +08:00
Bestony@Homelab 56c62c59c8 fix(auth): include client ip in acl denial message 2026-06-16 17:00:35 +08:00
Wesley Liddick e4ccb75d0f Merge pull request #3220 from wucm667/fix/oauth-signup-apply-promo-code
fix(auth): OAuth 注册支持应用 URL 上的 promo_code 优惠码
2026-06-16 16:59:38 +08:00
Wesley Liddick 2ce8788929 Merge pull request #3299 from alfadb/fix/openai-responses-probe-tool-capability
fix(openai-probe): /responses 能力探测增加工具调用校验
2026-06-16 16:58:48 +08:00
Wesley Liddick 2e0ff1cfd5 Merge pull request #3258 from bwliangc/feat/channel-monitor-jitter
feat(渠道监控): 检测间隔支持正负随机抖动配置
2026-06-16 16:58:23 +08:00
Wesley Liddick 16765bde69 Merge pull request #3230 from DaydreamCoding/feat/openai-cyber-policy-passthrough
feat(openai): cyber_policy 硬阻断全链路透传、审计与计费
2026-06-16 16:55:51 +08:00
shaw b816949291 feat(openai-quota): query + reset rate-limit credits for OpenAI accounts
Adds an admin-side action that mirrors the Codex Desktop "rate-limit reset"
flow against chatgpt.com upstream for OpenAI OAuth accounts.

Backend
- OpenAIQuotaService.QueryUsage / ResetCredit hit /wham/usage and
  /wham/rate-limit-reset-credits/consume with the Codex Desktop header set,
  reusing OpenAITokenProvider for refreshed tokens and PrivacyClientFactory
  for the impersonated Chrome TLS fingerprint.
- Honors the account's configured proxy by reading the eager-loaded
  account.Proxy directly (falls back to proxyRepo only when missing).
- GET /api/v1/admin/openai/accounts/:id/quota
  POST /api/v1/admin/openai/accounts/:id/reset-quota
- Wire DI for the new service + handler dependency.

Frontend
- OpenAIQuotaResetCell renders a single action row in AccountUsageCell's
  OpenAI section: the existing local "查询" (active sampling) is injected
  via #pre-actions, alongside a "次数 N" button that doubles as the
  upstream query trigger and the available-credit indicator, and a "重置"
  button that consumes one credit.
- No duplicate 5h/7d window display; the local UsageProgressBar owns those
  bars to avoid confusion.
2026-06-16 16:55:07 +08:00
alfadb b88f8e4c04 fix(openai-probe): /responses 能力探测增加工具调用校验
原探测仅以 HTTP 状态码判定 /v1/responses 端点是否存在(404/405 视为不存在,其余视为存在),无法识别"端点存在、基础补全可用、但工具调用不可用"的上游。火山方舟 coding/v3 的 kimi-k2.6 即属此类:携带 tools 的请求仅返回 reasoning、不产出 function_call,导致账号被误判 openai_responses_supported=true。网关据此走 /responses 转换路径后工具调用结果丢失,带工具的请求确定性失败。

改动:
- 探测请求携带一个工具并以 tool_choice=required 强制调用,仅当响应 output 数组包含 function_call 项时判定为支持;2xx 但无 function_call 判定为不支持,使网关改走 /v1/chat/completions 直转路径(同一模型在该路径下工具调用正常)。
- 探测模型改用账号 model_mapping 中的真实上游模型;占位模型在第三方上游会返回 400 model-not-found,无法判定能力。
- 探测超时 8s 调整为 15s(探测在后台异步执行,为推理型模型先推理再产出工具调用预留时间)。
- 非 2xx(404/405 除外)仍保守判定为支持,不改变既有账号行为;运行时回退函数 isResponsesEndpointSupportedByStatus 保持不变。
- 新增单元测试覆盖判定逻辑与探测模型选择。
2026-06-16 15:59:00 +08:00
shaw f069c9ae00 fix(outbox-dedup): buildSchedulerGroupPayload typed-nil broke dedup_key consistency
#3255 introduced payload-aware dedup_key (sha256 over event_type, account_id,
group_id, payload_json). enqueueSchedulerOutbox checks "if payload != nil"
before json.Marshal — but Go interface containing a typed-nil map (returned
by buildSchedulerGroupPayload(empty)) is NOT == nil at the interface level.

So an ungrouped account's account_changed event went through this path:
  payload := buildSchedulerGroupPayload(account.GroupIDs)  // typed-nil map
  enqueueSchedulerOutbox(..., payload)                     // interface != nil
    → json.Marshal(typedNilMap) = "null"
    → dedup_key hash = sha256(... + "null")

While other call sites pass literal nil:
  enqueueSchedulerOutbox(..., nil)                         // interface == nil
    → payloadJSON stays empty
    → dedup_key hash = sha256(... + "")

The two dedup_keys differ for what should be the same logical event,
silently degrading dedup effectiveness in bursts on ungrouped accounts.

Fix: change buildSchedulerGroupPayload return type from map[string]any to
any so empty input returns true untyped-nil. All call sites pass the
result straight to enqueueSchedulerOutbox(payload any) — no inspection,
no breakage.

Adds regression test TestEnqueueSchedulerOutbox_UngroupedAccountDedupesWithLiteralNilPayload
asserting (1) typed-nil regression doesn't sneak back, (2) dedup_key for
empty-groups payload matches the literal-nil-payload key.
2026-06-16 14:21:09 +08:00
shaw acaffe29ec fix(account-repo): refresh candidates SQL excluded healthy accounts; fix CI build
Post-merge audit of #3272 found two regressions:

1. ListOAuthRefreshCandidates used "AND NOT (a AND b)" which, under PG
   3-valued logic, evaluates to NULL when both temp_unschedulable_until
   and temp_unschedulable_reason are NULL — i.e., the common healthy
   account state. Such rows were silently excluded from the background
   token refresh worker, so their OAuth access tokens would never get
   refreshed and eventually start returning 401.

   Verified empirically against PostgreSQL: only 3 of 5 test rows
   matched before the fix; after switching to "(a AND b) IS NOT TRUE"
   the expected 4 rows match.

2. The new ListOAuthRefreshCandidates method on AccountRepository was
   not implemented on stubAccountRepo in api_contract_test.go (build
   tag "unit"), breaking "make test-unit" which CI runs in
   .github/workflows/backend-ci.yml.

Tests:
- Added IS NOT TRUE and "AND NOT (" assertions to the SQL-shape unit
  test so the predicate can't regress to the broken form again.
- "go test -tags=unit ./internal/..." now passes cleanly.
2026-06-16 14:08:50 +08:00
shaw 2ba52bf4aa Merge pull request #3274 from jianjianai/fix/scheduler-outbox-cleanup
fix: cleanup consumed scheduler outbox rows / 添加 scheduler_outbox 表的清理代码,避免表过大

Additional hardening: WHERE clause adds a 10s grace via
created_at < NOW() - INTERVAL '10 seconds' to defend against the PG
sequence-id vs commit race (id assigned in tx, commit delayed past
watermark advance). Without it, slow committers could lose their
outbox row before the snapshot poller reads it.
2026-06-16 11:57:43 +08:00
shaw 31dc8913ac chore(outbox-cleanup): add 10s grace to defend against id-vs-commit race
PG sequences advance outside transactions, so a slow committer can hold
an id that gets surpassed by the watermark before its commit becomes
visible. Without the grace period, cleanup would delete such rows before
the snapshot poller ever sees them. 10s is a comfortable upper bound
on realistic enqueueSchedulerOutbox commit latency.
2026-06-16 11:57:32 +08:00
jjaw cb14935e9a fix: cleanup consumed scheduler outbox rows 2026-06-16 11:56:40 +08:00
shaw 4254dcb352 Merge pull request #3255 from jianjianai/fix1/outbox-scheduler-snapshot-coalesce
降低高频账号状态变更下的 outbox 写入压力 / safely coalesce scheduler outbox events

Migrations renumbered 151/152 -> 152/153 to avoid collision with #3232.
2026-06-16 11:50:35 +08:00
shaw 1fdbe52f96 chore(migrations): renumber scheduler outbox dedup migrations 151/152 -> 152/153
#3232 (already merged) occupies 151. Renumber #3255's dedup migrations to
avoid collision and keep linear ordering. Updated runner constant + tests.
2026-06-16 11:50:25 +08:00
jjaw b3ec6288ad fix: release scheduler outbox dedup on claim 2026-06-16 11:48:13 +08:00
jjaw 60cf89ae26 fix: recover scheduler outbox invalid dedup index 2026-06-16 11:48:13 +08:00
jjaw 3ef70b045d fix: safely coalesce scheduler outbox events 2026-06-16 11:48:12 +08:00
jjaw 34e66ec0a5 fix: outbox scheduler snapshot coalesce 2026-06-16 11:48:12 +08:00
shaw 45f3b0dd74 Merge pull request #3272 from jianjianai/fix/token-refresh-retry-backoff
fix: token refresh 筛选避免拉取全部账号导致数据库爆炸 / reduce token refresh retry amplification
2026-06-16 11:47:07 +08:00
jjaw 9b270f11d7 refactor: inline token refresh retry reason prefix 2026-06-16 11:44:39 +08:00
jjaw 74199b6a6e fix: reduce token refresh retry amplification 2026-06-16 11:44:39 +08:00
shaw b56c501d30 Merge pull request #3273 from jianjianai/fix/redis-wait-queue-hotpath
fix: move user wait queue accounting off hot path / 用户槽位立即成功的请求不调用等待队列
2026-06-16 11:43:07 +08:00
jjaw b0579c4891 fix: move user wait queue accounting off hot path 2026-06-16 11:41:54 +08:00
shaw 87d0f8da66 Merge pull request #3132 from jianjianai/main
修复 token refresh 拉取大量账号时触发 PostgreSQL 参数上限
2026-06-16 11:40:47 +08:00
jjaw 8b698ff4c1 fix account list parameter limit 2026-06-16 11:39:15 +08:00