mirror of
https://github.com/coder/coder.git
synced 2026-09-22 21:22:17 +08:00
closes CODAGT-839 closes CODAGT-843 closes CODAGT-773 ## Summary Adds a new heartbeat usage event type, `hb_agent_runtime_v1`, measuring the total agent-loop runtime of Coder Agents (chats) per UTC hour, plus a reconciler that generates one event per hour with self-healing backfill over a trailing 7-day window. Events flow to Tallyman through the existing publisher unchanged. This measures the new Coder Agents (the `chats` tables), not the deprecated Tasks counted by `dc_managed_agents_v1`. Independent of #27508, which fixes the dead ai-seats cron registration. Both PRs carry the identical `usage_event` create permission hunk for the usage-publisher subject (this feature's generator and the ai-seats cron each need it for heartbeat inserts), so they can land in either order and the overlap merges cleanly. > [!WARNING] > **Do not include this in a release until Tallyman accepts `hb_agent_runtime_v1`.** The publisher marks permanently rejected events as done-forever, and the generator then sees those buckets as complete locally, so their usage would be silently and permanently lost. ## Details Each event's payload is `{"runtime_ms": N}`: the sum of `chat_messages.runtime_ms` for messages created in the hour bucket `[H, H+1)`, across all chats (sub-agents, API-created, archived, and soft-deleted messages included). Events use deterministic IDs (`hb_agent_runtime_v1:<bucket start>`) with `created_at` set to the bucket start, so concurrent replicas race safely via `ON CONFLICT (id) DO NOTHING` without locking, and daily rollups attribute backfilled hours to the correct day. Idle hours produce zero-valued events. A bucket becomes eligible 5 minutes after it closes; hours missing for longer than the 7-day window are forfeited, which can only undercount. Note that this makes `usage_events.created_at` explicitly the *event occurrence time* rather than the row insertion time; the two only diverge for backfilled events. It already behaved as the occurrence timestamp (it drives the daily rollup day and is shipped to Tallyman/Metronome as the event timestamp), and the migration now documents this with a `COMMENT ON COLUMN`, which also surfaces as a Go doc comment on `UsageEvent.CreatedAt`. The new `usage.Generator` runs unconditionally in enterprise builds; the `publish_usage_data` license flag continues to gate egress only, so air-gapped deployments still fill their local ledger. The `aggregate_usage_event()` trigger sums `runtime_ms` per day into `usage_events_daily` (unlike `hb_ai_seats_v1`, which takes the daily max). `InsertHeartbeatUsageEvent` now takes an explicit `createdAt` so generators can backfill historical buckets; the cron passes `clock.Now()` to preserve its existing behavior. ## Tallyman follow-up <details> <summary>Prompt for the Tallyman-repo agent</summary> > **Task**: Add support for the new Coder usage event type `hb_agent_runtime_v1` so Tallyman accepts, validates, and forwards it to Metronome. > > **Background**: coder/coder PR (this PR) adds hourly heartbeat events measuring Coder Agent runtime. Events arrive via the existing `/api/v1/events/ingest` endpoint with: `event_type: "hb_agent_runtime_v1"`, `event_data: {"runtime_ms": <int64 >= 0>}`, deterministic `id` of the form `hb_agent_runtime_v1:2026-07-15_14:00:00` (UTC hour bucket start), and `created_at` set to the bucket start (may be up to ~8 days in the past due to backfill; within Metronome's 34-day dedup window). Zero-value events are normal (idle hours). > > **Work**: > 1. Update Tallyman's vendored/imported `coderd/usage/usagetypes` (or equivalent) to the coder/coder commit that adds `UsageEventTypeHBAgentRuntimeV1` and `HBAgentRuntime`. > 2. Ensure ingestion validation accepts the type (`Valid()` switches) and rejects negative `runtime_ms`. > 3. Ensure Metronome forwarding maps the event with transaction ID derived from the event `id` as for existing types, passing `runtime_ms` through as the property for a SUM-aggregated billable metric ("Coder Agent Hours" = `SUM(runtime_ms) / 3,600,000`). > 4. Do NOT permanently reject unknown-but-well-formed future `hb_*` types if avoidable; at minimum confirm current behavior for unknown types (temporary vs permanent rejection) and report it. > 5. Tests: ingest accept/validate, dedup by ID, Metronome payload mapping. > > **Constraint**: this must be deployed to tallyman-prod **before** any coder/coder release containing the event generator; coderd treats permanent rejections as terminal per event. </details>
120 lines
4.7 KiB
SQL
120 lines
4.7 KiB
SQL
-- name: InsertUsageEvent :exec
|
|
-- Duplicate events are ignored intentionally to allow for multiple replicas to
|
|
-- publish heartbeat events.
|
|
INSERT INTO
|
|
usage_events (
|
|
id,
|
|
event_type,
|
|
event_data,
|
|
created_at,
|
|
publish_started_at,
|
|
published_at,
|
|
failure_message
|
|
)
|
|
VALUES
|
|
(@id, @event_type, @event_data, @created_at, NULL, NULL, NULL)
|
|
ON CONFLICT (id) DO NOTHING;
|
|
|
|
-- name: UsageEventExistsByID :one
|
|
SELECT EXISTS(
|
|
SELECT 1 FROM usage_events WHERE id = @id
|
|
)::bool;
|
|
|
|
-- name: ListUsageEventCreatedAtsByTypeSince :many
|
|
-- Used by the usage generator to find missing heartbeat buckets.
|
|
SELECT created_at
|
|
FROM usage_events
|
|
WHERE event_type = @event_type
|
|
AND created_at >= @since::timestamptz;
|
|
|
|
-- name: SelectUsageEventsForPublishing :many
|
|
WITH usage_events AS (
|
|
UPDATE
|
|
usage_events
|
|
SET
|
|
publish_started_at = @now::timestamptz
|
|
WHERE
|
|
id IN (
|
|
SELECT
|
|
potential_event.id
|
|
FROM
|
|
usage_events potential_event
|
|
WHERE
|
|
-- Do not publish events that have already been published or
|
|
-- have permanently failed to publish.
|
|
potential_event.published_at IS NULL
|
|
-- Do not publish events that are already being published by
|
|
-- another replica.
|
|
AND (
|
|
potential_event.publish_started_at IS NULL
|
|
-- If the event has publish_started_at set, it must be older
|
|
-- than an hour ago. This is so we can retry publishing
|
|
-- events where the replica exited or couldn't update the
|
|
-- row.
|
|
-- The parentheses around @now::timestamptz are necessary to
|
|
-- avoid sqlc from generating an extra argument.
|
|
OR potential_event.publish_started_at < (@now::timestamptz) - INTERVAL '1 hour'
|
|
)
|
|
-- Do not publish events older than 30 days. Tallyman will
|
|
-- always permanently reject these events anyways. This is to
|
|
-- avoid duplicate events being billed to customers, as
|
|
-- Metronome will only deduplicate events within 34 days.
|
|
-- Also, the same parentheses thing here as above.
|
|
AND potential_event.created_at > (@now::timestamptz) - INTERVAL '30 days'
|
|
ORDER BY potential_event.created_at ASC
|
|
FOR UPDATE SKIP LOCKED
|
|
LIMIT 100
|
|
)
|
|
RETURNING *
|
|
)
|
|
SELECT *
|
|
-- Note that this selects from the CTE, not the original table. The CTE is named
|
|
-- the same as the original table to trick sqlc into reusing the existing struct
|
|
-- for the table.
|
|
FROM usage_events
|
|
-- The CTE and the reorder is required because UPDATE doesn't guarantee order.
|
|
ORDER BY created_at ASC;
|
|
|
|
-- name: UpdateUsageEventsPostPublish :exec
|
|
UPDATE
|
|
usage_events
|
|
SET
|
|
publish_started_at = NULL,
|
|
published_at = CASE WHEN input.set_published_at THEN @now::timestamptz ELSE NULL END,
|
|
failure_message = NULLIF(input.failure_message, '')
|
|
FROM (
|
|
SELECT
|
|
UNNEST(@ids::text[]) AS id,
|
|
UNNEST(@failure_messages::text[]) AS failure_message,
|
|
UNNEST(@set_published_ats::boolean[]) AS set_published_at
|
|
) input
|
|
WHERE
|
|
input.id = usage_events.id
|
|
-- If the number of ids, failure messages, and set published ats are not the
|
|
-- same, do not do anything. Unfortunately you can't really throw from a
|
|
-- query without writing a function or doing some jank like dividing by
|
|
-- zero, so this is the best we can do.
|
|
AND cardinality(@ids::text[]) = cardinality(@failure_messages::text[])
|
|
AND cardinality(@ids::text[]) = cardinality(@set_published_ats::boolean[]);
|
|
|
|
-- name: GetTotalUsageDCManagedAgentsV1 :one
|
|
-- Gets the total number of managed agents created between two dates. Uses the
|
|
-- aggregate table to avoid large scans or a complex index on the usage_events
|
|
-- table.
|
|
--
|
|
-- This has the trade off that we can't count accurately between two exact
|
|
-- timestamps. The provided timestamps will be converted to UTC and truncated to
|
|
-- the events that happened on and between the two dates. Both dates are
|
|
-- inclusive.
|
|
SELECT
|
|
-- The first cast is necessary since you can't sum strings, and the second
|
|
-- cast is necessary to make sqlc happy.
|
|
COALESCE(SUM((usage_data->>'count')::bigint), 0)::bigint AS total_count
|
|
FROM
|
|
usage_events_daily
|
|
WHERE
|
|
event_type = 'dc_managed_agents_v1'
|
|
-- Parentheses are necessary to avoid sqlc from generating an extra
|
|
-- argument.
|
|
AND day BETWEEN date_trunc('day', (@start_date::timestamptz) AT TIME ZONE 'UTC')::date AND date_trunc('day', (@end_date::timestamptz) AT TIME ZONE 'UTC')::date;
|