mirror of
https://github.com/coder/coder.git
synced 2026-09-24 15:04:27 +08:00
Chat message ordering was derived from `created_at`, which is `now()` and therefore the transaction start time. That makes it unusable as an append-order column for two independent reasons: every row in one `InsertChatMessages` batch shares a single timestamp, and two concurrent transactions can commit in the opposite order to the one they started in. This PR gives `chat_messages.id` a real append-order guarantee and moves the history reads onto it. ## Changes **`InsertChatMessages` had no input-order guarantee.** Callers index the returned slice by input position. That only worked because PostgreSQL happens to evaluate the `BIGSERIAL` default in row order. Ids are now allocated up front and the k-th smallest is assigned to input index k, so the pairing does not depend on where the column default is evaluated. Returned rows are explicitly `ORDER BY id`. **Three history reads now order by `id`.** | Query | Was | Now | |---|---|---| | `GetChatMessagesByChatID` | `created_at ASC` | `id ASC` | | `GetChatMessagesByRevisionForStream` | `created_at ASC, id ASC` | `id ASC` | | `GetLastChatMessageByRole` | `created_at DESC, id DESC` | `id DESC` | `GetChatMessagesByChatID` paginated by `id` while ordering by `created_at`, which is incoherent on its own terms. The other two matter because of who consumes them. The stream query supplies incremental updates on the same socket that emits a full `GetChatMessagesByChatID` snapshot on history reset, so once that snapshot moved to `id` the two disagreed under timestamp skew. `GetLastChatMessageByRole` returns an id that is then used as an id cursor, both as `AfterID` when synthesizing tool cancellations and as `chats.last_read_message_id`, where a stale anchor leaves later assistant messages permanently unread. A tie-breaker would not have fixed either one. It only resolves equal timestamps; leading with `created_at` is the actual defect. **`GetLastChatMessageByRole` loses its index, so this adds one.** `ORDER BY created_at DESC, id DESC` could take an ordered scan of `idx_chat_messages_chat_created`. Nothing in the schema can supply `ORDER BY id DESC LIMIT 1` for a given `chat_id` and `role`, so the planner switches to a backward scan of the primary key and filters every newer row in the table, scanning all of it when the chat has no message in that role, which is the routine case for a fresh chat. Migration `000559` adds `(chat_id, role, id DESC) WHERE deleted = false`, the same shape as the existing `idx_chat_messages_user_prompts`. This matters because the query is hot: it runs on every stream connect and disconnect, and once per turn when synthesizing tool cancellations. `GetChatMessagesForPromptByChatID` has the same defect and is fixed in the stacked PR, because its compaction boundary change is semantic and deserves a separate review. Auto-archive stays timestamp-based deliberately: it measures activity, not order. Wrapping the insert in a CTE (needed because `INSERT` cannot take `ORDER BY`) makes sqlc synthesize `InsertChatMessagesRow`. It is structurally identical to `ChatMessage`, so the call sites use a direct struct conversion that stops compiling if the two ever diverge. ## Testing Behavior tests write `created_at` values inverted against id order, so a reader that leads with `created_at` returns the batch backwards. All three queries were verified red by reverting the `ORDER BY` and regenerating: the stream query returned `[3,2,1]` for `[1,2,3]`, and `GetLastChatMessageByRole` picked id 1 instead of id 3. `TestInsertChatMessagesOrderContract` asserts against the generated SQL, covering what a behavior test cannot: PostgreSQL evaluates the id default in row order anyway, so a batch still looks ordered once the guarantee is removed. `TestChatMessagesSequenceCacheIsOne` guards the cross-batch half of the invariant. Ids follow chat row lock order only while the sequence hands out one value at a time; sequence cache blocks are per session, so with a cache above one a backend holding stale cached values can lock second and still commit lower ids. Bumping a sequence cache is an ordinary throughput tweak, and it would silently corrupt history order. The index was checked on a 200k row fixture. Without it, the zero-match lookup filters all 200,000 rows over 2763 buffers; with it, the plan is an index scan with both `chat_id` and `role` in the index condition, no sort node, and 3 buffers. Note that the within-batch mapping does not depend on the cache size. It is established by `ROW_NUMBER() OVER (ORDER BY id)` over the allocated ids, so it holds regardless of `nextval` evaluation order. ## Note on the deleted subagent hand-sort The subagent history reader's hand-sort stays deleted, but calling it redundant was imprecise. It sorted by `created_at` then `id`, so it is only equivalent to `id` ordering when the two agree. When they disagree the old code selected a different "latest assistant". This is a deliberate behavior change to match the new invariant, not dead-code removal. > Opened by Mux on behalf of Mike.
122 lines
4.2 KiB
Go
122 lines
4.2 KiB
Go
package chatstate
|
|
|
|
import (
|
|
"database/sql"
|
|
|
|
"github.com/google/uuid"
|
|
"github.com/sqlc-dev/pqtype"
|
|
|
|
"github.com/coder/coder/v2/coderd/database"
|
|
)
|
|
|
|
// Message is the durable message input shape used by chatstate
|
|
// transitions. It is intentionally lower level than the SDK message
|
|
// request types: callers must produce a fully materialized message
|
|
// (parsed parts, calculated cost, resolved model config) before
|
|
// passing it in.
|
|
//
|
|
// The state machine never reshapes a Message except to attach the
|
|
// runtime `chat_id`.
|
|
type Message struct {
|
|
Role database.ChatMessageRole
|
|
Content pqtype.NullRawMessage
|
|
Visibility database.ChatMessageVisibility
|
|
ModelConfigID uuid.NullUUID
|
|
ReasoningEffort database.NullChatReasoningEffort
|
|
CreatedBy uuid.NullUUID
|
|
ContentVersion int16
|
|
Compressed bool
|
|
InputTokens sql.NullInt64
|
|
OutputTokens sql.NullInt64
|
|
TotalTokens sql.NullInt64
|
|
ReasoningTokens sql.NullInt64
|
|
CacheCreationTokens sql.NullInt64
|
|
CacheReadTokens sql.NullInt64
|
|
ContextLimit sql.NullInt64
|
|
TotalCostMicros sql.NullInt64
|
|
RuntimeMs sql.NullInt64
|
|
}
|
|
|
|
// toInsertParams converts a batch of Messages into the parallel-array
|
|
// shape required by `InsertChatMessages`. The returned struct has all
|
|
// arrays sized to len(messages).
|
|
//
|
|
// The chat ID is supplied by the caller because Message itself does
|
|
// not carry one (the chat machine already knows the chat).
|
|
func toInsertParams(chatID uuid.UUID, messages []Message) database.InsertChatMessagesParams {
|
|
n := len(messages)
|
|
params := database.InsertChatMessagesParams{
|
|
ChatID: chatID,
|
|
CreatedBy: make([]uuid.UUID, n),
|
|
ModelConfigID: make([]uuid.UUID, n),
|
|
ReasoningEffort: make([]string, n),
|
|
Role: make([]database.ChatMessageRole, n),
|
|
Content: make([]string, n),
|
|
ContentVersion: make([]int16, n),
|
|
Visibility: make([]database.ChatMessageVisibility, n),
|
|
InputTokens: make([]int64, n),
|
|
OutputTokens: make([]int64, n),
|
|
TotalTokens: make([]int64, n),
|
|
ReasoningTokens: make([]int64, n),
|
|
CacheCreationTokens: make([]int64, n),
|
|
CacheReadTokens: make([]int64, n),
|
|
ContextLimit: make([]int64, n),
|
|
Compressed: make([]bool, n),
|
|
TotalCostMicros: make([]int64, n),
|
|
RuntimeMs: make([]int64, n),
|
|
}
|
|
for i, m := range messages {
|
|
params.CreatedBy[i] = nullUUIDOrNil(m.CreatedBy)
|
|
params.ModelConfigID[i] = nullUUIDOrNil(m.ModelConfigID)
|
|
if m.ReasoningEffort.Valid {
|
|
params.ReasoningEffort[i] = string(m.ReasoningEffort.ChatReasoningEffort)
|
|
}
|
|
params.Role[i] = m.Role
|
|
if m.Content.Valid {
|
|
params.Content[i] = string(m.Content.RawMessage)
|
|
} else {
|
|
// Use the JSON null literal; UNNEST + ::jsonb requires a
|
|
// valid JSON value and the trigger leaves it untouched.
|
|
params.Content[i] = "null"
|
|
}
|
|
params.ContentVersion[i] = m.ContentVersion
|
|
params.Visibility[i] = m.Visibility
|
|
params.InputTokens[i] = nullInt64Or(m.InputTokens, 0)
|
|
params.OutputTokens[i] = nullInt64Or(m.OutputTokens, 0)
|
|
params.TotalTokens[i] = nullInt64Or(m.TotalTokens, 0)
|
|
params.ReasoningTokens[i] = nullInt64Or(m.ReasoningTokens, 0)
|
|
params.CacheCreationTokens[i] = nullInt64Or(m.CacheCreationTokens, 0)
|
|
params.CacheReadTokens[i] = nullInt64Or(m.CacheReadTokens, 0)
|
|
params.ContextLimit[i] = nullInt64Or(m.ContextLimit, 0)
|
|
params.Compressed[i] = m.Compressed
|
|
params.TotalCostMicros[i] = nullInt64Or(m.TotalCostMicros, 0)
|
|
params.RuntimeMs[i] = nullInt64Or(m.RuntimeMs, 0)
|
|
}
|
|
return params
|
|
}
|
|
|
|
// fromInsertedRows converts the rows returned by `InsertChatMessages`, which
|
|
// sqlc types separately because the query wraps the insert in a CTE. The
|
|
// conversion stops compiling if the row ever stops matching ChatMessage.
|
|
func fromInsertedRows(rows []database.InsertChatMessagesRow) []database.ChatMessage {
|
|
messages := make([]database.ChatMessage, len(rows))
|
|
for i, row := range rows {
|
|
messages[i] = database.ChatMessage(row)
|
|
}
|
|
return messages
|
|
}
|
|
|
|
func nullUUIDOrNil(u uuid.NullUUID) uuid.UUID {
|
|
if u.Valid {
|
|
return u.UUID
|
|
}
|
|
return uuid.Nil
|
|
}
|
|
|
|
func nullInt64Or(v sql.NullInt64, fallback int64) int64 {
|
|
if v.Valid {
|
|
return v.Int64
|
|
}
|
|
return fallback
|
|
}
|