mirror of
https://github.com/coder/coder.git
synced 2026-09-23 05:43:53 +08:00
Closes https://linear.app/codercom/issue/CODAGT-268 ## Problem The chat UI collapses large pastes (>=10 lines or >=1000 chars) into a synthetic `pasted-text-*.txt` attachment. A chat created with only such an attachment had no title input anywhere: the create path derived `titleSource` only from text and file-reference parts (so the chat was named "New Chat"), async auto-titling extracted text the same way and silently skipped generation, and the manual propose/regenerate paths returned an empty title for the same reason. The regular prompt path already inlines these files for the model; only the title paths were blind. ## Fix Add a single title-input derivation in `chatprompt` and use it everywhere: - `chatprompt.TitleText` joins text and file-reference parts (unchanged formatting), and falls back to synthetic pasted-text attachment content (truncated to a 16 KiB title budget) when they yield nothing. - `chatprompt.SyntheticPasteFileIDs` identifies paste attachments; `chatprompt.FallbackTitle` consolidates the previously duplicated `chatTitleFromMessage` / `fallbackChatTitle`. - Chat creation captures paste blob references while validating file parts (the file row was already loaded there) and derives `titleSource` via `TitleText`. Only the create path derives titles; message send and edit reuse the same validation without copying any blob data. - `GenerateChatTitleAsync` and the manual propose/regenerate paths resolve paste content via `titlePasteText`, which only queries when a visible user message has no other title text, so chats with typed text never incur a file fetch. - Title-path paste fetches are bounded: a new `GetChatFileDataPrefixesByIDs` query returns only a `substr` prefix (`chatprompt.TitlePasteBytePrefix`, 64 KiB = 4 bytes x the 16 Ki-rune title budget) so full blobs (up to 10 MiB each) never leave the database for titling, and `chatprompt.TitlePasteText` applies the same bound to the create path which already holds the loaded row. Deliberate side effect: because generation-time extraction now matches create-time `titleSource` exactly, file-reference-only chats also become eligible for AI titles. They were previously skipped by the same derivation mismatch. Non-goals: no frontend changes (attachment chip UX stays as is), and non-synthetic user-uploaded `.txt` files still yield "New Chat". ## Testing - Unit tests for `TitleText`, `TitlePasteText`, `SyntheticPasteFileIDs`, `FallbackTitle`, `titleInput`, `titlePasteText`, and paste-aware `extractManualTitleTurns`. - Real-database test for `GetChatFileDataPrefixesByIDs` (prefix shorter and longer than stored data) plus dbauthz coverage for the new query. - Integration tests: paste-only create gets a fallback title from the paste content, async title generation fires with the paste content as input, and `RegenerateChatTitle` works on a paste-only chat. > This PR was written by [Mux](https://mux.coder.com) on Mike's behalf.
63 lines
2.5 KiB
SQL
63 lines
2.5 KiB
SQL
-- name: InsertChatFile :one
|
|
INSERT INTO chat_files (owner_id, organization_id, name, mimetype, data)
|
|
VALUES (@owner_id::uuid, @organization_id::uuid, @name::text, @mimetype::text, @data::bytea)
|
|
RETURNING id, owner_id, organization_id, created_at, name, mimetype;
|
|
|
|
-- name: GetChatFileByID :one
|
|
SELECT * FROM chat_files WHERE id = @id::uuid;
|
|
|
|
-- name: GetChatFilesByIDs :many
|
|
SELECT * FROM chat_files WHERE id = ANY(@ids::uuid[]);
|
|
|
|
-- name: GetChatFileDataPrefixesByIDs :many
|
|
-- GetChatFileDataPrefixesByIDs returns a bounded prefix of each
|
|
-- file's content, keeping full blobs out of server memory. Owner and
|
|
-- organization columns support row-level authorization.
|
|
SELECT id, owner_id, organization_id, substr(data, 1, @prefix_bytes::int) AS data_prefix
|
|
FROM chat_files
|
|
WHERE id = ANY(@ids::uuid[]);
|
|
|
|
-- name: GetChatFileMetadataByChatID :many
|
|
-- GetChatFileMetadataByChatID returns lightweight file metadata for
|
|
-- all files linked to a chat. The data column is excluded to avoid
|
|
-- loading file content.
|
|
SELECT cf.id, cf.owner_id, cf.organization_id, cf.name, cf.mimetype, cf.created_at
|
|
FROM chat_files cf
|
|
JOIN chat_file_links cfl ON cfl.file_id = cf.id
|
|
WHERE cfl.chat_id = @chat_id::uuid
|
|
ORDER BY cf.created_at ASC;
|
|
|
|
-- TODO(cian): Add indexes on chats(archived, updated_at) and
|
|
-- chat_files(created_at) for purge query performance.
|
|
-- See: https://github.com/coder/internal/issues/1438
|
|
-- name: DeleteOldChatFiles :execrows
|
|
-- Deletes chat files that are older than the given threshold and are
|
|
-- not referenced by any chat that is still active or was archived
|
|
-- within the same threshold window. This covers two cases:
|
|
-- 1. Orphaned files not linked to any chat.
|
|
-- 2. Files whose every referencing chat has been archived for longer
|
|
-- than the retention period.
|
|
WITH kept_file_ids AS (
|
|
-- NOTE: This uses updated_at as a proxy for archive time
|
|
-- because there is no archived_at column. Correctness
|
|
-- requires that updated_at is never backdated on archived
|
|
-- chats. See ArchiveChatByID.
|
|
SELECT DISTINCT cfl.file_id
|
|
FROM chat_file_links cfl
|
|
JOIN chats c ON c.id = cfl.chat_id
|
|
WHERE c.archived = false
|
|
OR c.updated_at >= @before_time::timestamptz
|
|
),
|
|
deletable AS (
|
|
SELECT cf.id
|
|
FROM chat_files cf
|
|
LEFT JOIN kept_file_ids k ON cf.id = k.file_id
|
|
WHERE cf.created_at < @before_time::timestamptz
|
|
AND k.file_id IS NULL
|
|
ORDER BY cf.created_at ASC
|
|
LIMIT @limit_count
|
|
)
|
|
DELETE FROM chat_files
|
|
USING deletable
|
|
WHERE chat_files.id = deletable.id;
|