Files
sim/apps
Waleed 220da449c9 fix(mothership): stop inlining full execution traces for the logs context (#5353)
* fix(mothership): stop inlining full execution traces for the logs context

Tagging a run via "Troubleshoot in Chat" (or any @-mention of a logs
context) resolved through processExecutionLogFromDb, which materialized
the ENTIRE execution trace (every block's input/output, nested tool-call
spans) and inlined it directly into the prompt. For any non-trivial run
this repeatedly blew the context window, forcing multiple compactions and
eventually auto-stopping the agent before it could investigate anything.

Every other context resolver in this file already avoids this by sending
a lightweight pointer instead of a full inline dump (workflow/blocks/
workflow_block contexts point into the VFS). Logs contexts have no VFS
materialization to point at, but the equivalent lightweight mechanism
already exists as a tool: query_logs supports incremental disclosure
(overview for timing/cost, full for a scoped block's input/output, or
pattern to grep the trace) and is already registered for the mothership
agent.

Now processExecutionLogFromDb sends a compact summary (id, workflow,
level, trigger, timing, cost) plus a note pointing the model at
query_logs with the executionId, instead of materializing and embedding
the trace. Also drops the now-unused executionData column from the
select projection, so resolving a logs context no longer fetches a
potentially large JSONB blob it never reads.

* improvement(mothership): send a bounded block overview instead of a bare tool pointer

Follow-up to the previous commit's fix (stop inlining full execution
traces). A pure text pointer telling the model to call query_logs made
the agent's very first useful action against a tagged run contingent on
it noticing and correctly acting on prose in a JSON blob it may only
skim — every sibling resolver in this file instead returns a
deterministic mechanism (a VFS path) the model reads on demand.

There's no VFS materialization for individual execution logs, but the
same deterministic signal is available cheaply: toOverview() (the exact
projection query_logs's own "overview" view already returns) walks the
raw trace spans and produces a compact tree — block name/type/status/
timing/cost, no input or output — without touching large-value refs at
all. The summary now includes that tree, so the model can see which
block failed on the first turn, and the note narrows to what still
requires a tool call: a block's actual input/output/error, or a grep.

materializeExecutionData is still called, but it's a no-op for the
common inline case (it only unwraps a top-level object-storage pointer
for runs whose whole trace was offloaded as one blob) and was needed to
reach traceSpans at all for those heavier runs — exactly the runs most
worth an overview.

A serialized-size cap (mirroring query-logs.ts's own truncation
fallback, scaled down since this lands in the prompt unconditionally)
drops the overview if a pathological span count pushes it over budget,
falling back to the note alone.

Extends the tests: the happy path now asserts the overview tree is
present and that no raw input/output payload leaks into the serialized
summary, plus a new test for the size-cap fallback.
2026-07-01 22:50:26 -07:00
..