Follow-up fixes to #3569 based on code audit:
- Expose server_timezone / server_utc_offset in public settings (and the
__APP_CONFIG__ injection payload) and label every peak-window display
with the server UTC offset, so users don't misread the billing window
as browser-local time
- Unify CreateGroup/UpdateGroup peak-config sanitization via a single
NormalizePeakRateConfig chokepoint: non-subscription groups always get
peak fields cleared; unparseable window strings and negative
multipliers are scrubbed when peak is disabled
- Replace hot-path time.Parse in PeakMultiplierAt with a manual HH:MM
parser (accept set verified byte-for-byte identical to
time.Parse("15:04") by exhaustive fuzzing) and reuse it in validation
- Revert the zero-behavior CalculateCost indirection churn in
billing_service/gateway_service introduced by #3569
- Remove dead GetGroupPlatformMap and the duplicate deref helper in the
admin handler package
- Share frontend peak formatting via utils/peak-rate.ts, unify the ×N
label format, and move hardcoded Chinese tooltips to i18n keys
When converting a Chat Completions stream into Responses events, the first
tool_call delta chunk was copied wholesale into stream state (including
function.arguments), then the same chunk's arguments were accumulated again by
the shared `+=` block. For OpenAI this is harmless because its first tool_call
chunk carries empty arguments, but upstreams that pack id+name+arguments into a
single chunk (e.g. GLM/Zhipu) end up with doubled arguments such as
{"cmd":"ls"}{"cmd":"ls"}. Codex then fails to parse the tool call with
"trailing characters", breaking every tool invocation.
Reset the copied arguments so the shared accumulator counts them exactly once,
keeping the emitted delta and the final done/arguments consistent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
PR #3016's merge appended a verbatim second copy of four
TestStream_Reasoning* functions into
chatcompletions_responses_stream_lifecycle_test.go, causing
'redeclared in this block' build failures that broke both the
test and golangci-lint CI jobs.
Remove the duplicate block; each test now appears once.
- bump impersonated CLI version 2.1.92 -> 2.1.161; derive User-Agent from
CLICurrentVersion so the two hardcoded copies can no longer drift apart
- fix x-stainless headers to real 2.1.161 values: package-version
0.70.0 -> 0.94.0, runtime-version v24.13.0 -> v24.3.0 (verified against
the installed Bun-compiled binary)
- expand the disguise-path system prompt from a 2-block identity skeleton
to a 3-block layout (billing + identity + tool-agnostic prose), matching
real CC's multi-block shape; cache breakpoint moved to the last static
block. Deliberately excludes # Doing tasks / # Using your tools /
# Executing actions to avoid polluting proxied-client behavior
- stabilize the synthesized metadata.user_id session_id across conversation
turns: derive it from (account + client discriminator + first user
message) instead of a per-turn content/body hash. Sticky-routing
GenerateSessionHash is intentionally left untouched; remove now-dead
hashBodyForSessionSeed
Tests: update the 3-block system assertions in gateway_prompt_test and
gateway_anthropic_apikey_passthrough_test; add a session_id cross-turn
stability test in gateway_oauth_metadata_test.
When an OpenAI Chat Completions client targets an Anthropic-platform group,
ForwardAsChatCompletions converts the request CC → Responses → Anthropic
(ChatCompletionsToResponses → ResponsesToAnthropicRequest) before forwarding it
upstream. The Responses→Anthropic converter emits each function_call as its own
assistant message and each function_call_output as its own user message and
relies solely on mergeConsecutiveMessages to alternate roles. That is not enough
to satisfy Anthropic's tool-pairing invariants, so a trimmed or partial tool
history produces an upstream 400, e.g.:
tool_use_id found in tool_result blocks: call_00_...
Each tool_result block must have a corresponding tool_use block in the
previous message.
The failures this leaves unrepaired:
- orphan tool_result — a client that does sliding-window context management
keeps a recent tool result but drops the assistant tool_calls message that
announced it, so the tool_result has no matching tool_use;
- unanswered/dangling tool_use — a parallel call whose sibling result never
came back, or a call left dangling, which Anthropic also rejects.
Add normalizeAnthropicToolPairing, run between two merge passes: the first merge
groups parallel calls and their results; the pairing pass indexes every
tool_result by its tool_use id, keeps only answered tool_use blocks (dropping
unanswered/dangling calls, and the assistant message entirely when nothing else
remains) and re-emits the matching tool_result blocks as the immediately
following user message; standalone/orphan tool_results are dropped from their
original position; the second merge restores alternation. This mirrors
normalizeChatMessages on the Responses→Chat path.
Tested two ways: responses_to_anthropic_tool_pairing_test.go covers the repair
on direct Responses input (developer message between call and output, parallel
both-answered kept grouped, parallel one-unanswered dropped, orphan tool_result,
dangling call, single-call baseline); responses_to_anthropic_cc_chain_test.go
drives the real ChatCompletionsToResponses → ResponsesToAnthropicRequest chain
and reproduces the production 400 (orphan and unanswered-parallel) — both fail
without the repair and pass with it. The full apicompat suite stays green.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Include codex-auto-review in the OpenAI fallback models list so /v1/models exposes it when no account mapping is configured. Keep the entry aligned with the existing default model catalog.
errcheck (check-type-assertions) flagged unchecked single-value type
assertions; switch to the comma-ok form so golangci-lint passes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Codex CLI speaks the OpenAI Responses protocol (streaming, store:false), while
many upstreams (e.g. DeepSeek in thinking mode) only expose Chat Completions.
The bridge that translates between the two had grown field by field and leaned
on Go's serialization defaults, which both the Responses client (Codex) and the
Chat upstream reject in ways the official OpenAI endpoints tolerate.
Problems this fixes (all observed running Codex CLI against a DeepSeek upstream):
- Streaming reasoning was never shown in the Codex TUI (the answer appeared
with no visible thinking): reasoning deltas were emitted before the reasoning
item was opened, so the strict client discarded them.
- A tool-using turn could wedge the session into a "no response" state: the
function_call stream was never closed (no function_call_arguments.done /
output_item.done), so Codex never saw the tool call complete.
- Parallel tool calls were rejected upstream (400/502): each function_call
became its own assistant message, producing consecutive assistant messages
with mismatched tool replies.
- A tool turn was rejected with "reasoning_content in the thinking mode must be
passed back": the reasoning that produced the tool call was dropped instead
of being returned on the assistant message.
- Items with no Chat equivalent (web_search_call, ...) and Codex's
command-approval notice landed between an assistant tool_calls message and
its tool reply, triggering "An assistant message with 'tool_calls' must be
followed by tool messages responding to each 'tool_call_id'".
- Interrupt/reconnect left an unanswered or dangling tool_call in the history,
triggering the same 400.
The shared root cause is reliance on serialization defaults — omitempty dropping
protocol-required zero values, and unrecognized item types falling through a
generic path — rather than deliberately reproducing the target protocol. The
bridge is reworked into two explicit layers.
Request direction (Responses input -> Chat messages): a parse -> build ->
normalize pipeline.
- reasoning_content is carried back on the assistant message that produced a
tool call (DeepSeek thinking mode requires it to continue the same thought)
- consecutive function_call items (parallel tool calls) are merged into a
single assistant message's tool_calls array
- item types with no Chat equivalent are skipped instead of leaking through a
generic path
- normalizeChatMessages is the single invariant gate: it guarantees every
assistant tool_calls message is immediately followed by one tool reply per
tool_call_id — reordering any intervening message (such as a command-approval
notice) to after the replies, dropping unanswered tool_calls and orphan tool
replies, and preserving bare passthrough tool messages.
Response direction (Chat SSE -> Responses SSE): ResponsesStreamEvent.MarshalJSON
constructs each streamed event explicitly so protocol-required fields are always
present (output_index/content_index/summary_index at 0, message content:[],
reasoning summary:[], function_call call_id/name/arguments, output_text part
text/annotations/logprobs). This is a single source of truth that removes any
post-hoc JSON patching. Reasoning is emitted as its own output item, opened
before its deltas, and tool calls are fully closed
(function_call_arguments.done + output_item.done with complete arguments).
Tests cover request-direction message invariants against golden Codex request
shapes (parallel calls, unknown items, intervening messages, partial/dangling
calls), per-event wire completeness, and streaming lifecycle ordering.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>