A provider can accept a request, return response headers, and then never
send a byte of body data. The connection-phase request timeout was cleared
as soon as headers arrived, so nothing bounded that wait and the agent turn
hung indefinitely after a tool call completed: step-finish:tool-calls was
recorded and the next step-start never arrived, with the HTTP server still
responsive.
Extend the same configured timeout deadline to the wait for the response
body's first byte. The connection-phase timer covers the fetch up to
headers; once headers arrive, the remaining deadline is handed to a
first-byte guard that aborts the request if no data arrives. After the
first byte the guard becomes a passthrough, so idle gaps inside an already
streaming response (reasoning, buffering, slow token generation) are never
touched and remain opt-in via chunkTimeout.
This is a transport-level signal (bytes on the wire, before any content)
rather than the absence of normalized AI SDK events, so it cannot fire on
long prompt processing or reasoning the way the reverted stream watchdog
did. timeout: false still disables the bound entirely.
Adds a hermetic regression test that injects a simulated stalled socket
through the provider's own fetch option via the plugin config hook, so the
SDK, Kilo's fetch wrapper, SSE parsing, the processor and the agent loop
all stay production code. The stalled request is transient, so the test
asserts the turn recovers by retrying and completing instead of freezing.
The test goes red without the fix (no retry, frozen at step-finish) and
green with it.
Refs #8656
* fix(cli): enforce permissions on shell commands the parser fails to scan
* fix(cli): fail closed on error chunks without command names, move pwsh execution test to kilo file
* test(cli): cover TUI startup outside package
* fix(cli): use native preload path in TUI test
---------
Co-authored-by: Johnny Eric Amancio <johnnyeric@gmail.com>
* fix(cli): drain session ingest queue on shutdown and flush terminal batches promptly
* fix(cli): drain the session ingest queue on process shutdown
* fix(cli): pin drain bound expiry, add changeset, conform to naming rule
* fix(cli): keep kilo-sessions out of the CLI startup import graph
* fix(cli): never let the ingest drain task reject the shutdown sequence
* test(cli): pin drain-before-dispose ordering on the KiloCli shutdown path
* fix(cli): make the guarded ingest drain non-rejecting and correct the lazy-import rationale
* test(cli): cover the retryable-status drain path under shutdown
* test(cli): decouple cli-shutdown drain assertions from declaration order
* fix(cli): advertise the instance from enableRemote so /remote registers as a spawn target
Enabling the remote relay from the TUI `/remote` slash command connected the
socket and mirrored sessions, but never advertised the instance, so the CLI
never appeared as a spawn target in the mobile "Run on" picker. Only the
explicit `kilo remote` command called setInstanceAdvertisement.
The advertisement now runs on every successful enableRemote() entry, before the
already-connected and coalescing early returns. That ordering matters: bootstrap
auto-enable frequently connects first, so `/remote` usually hits
`if (remote) return` and an advertisement placed in the connection-setup body
would leave the defect unfixed in the common case. `ingestDisabled` returns
before the advertisement and stays unadvertised.
The ensure helper is a no-op when an advertisement is already set, so it fires no
extra heartbeat, while explicit setInstanceAdvertisement keeps its existing
replace semantics. buildInstanceAdvertisement moves to a shared module so the
command path and the enable path derive it identically.
* fix(cli): report pending question and permission on the session heartbeat
The heartbeat built each session's status from SessionStatus.Service, whose
union is idle/retry/busy/offline and which never consults Question.Service or
Permission.Service. deriveStatus() already did consult both, but only fed the
ingest session_status sync. So a session genuinely blocked on a question was
advertised as busy on the heartbeat, and the mobile app — which takes live row
status from the heartbeat — showed no needs-input badge.
Extract the precedence (permission, then question, then SessionStatus) into a
shared helper used by both deriveStatus and the heartbeat, so the two channels
cannot drift.
The heartbeat runs on a ~10s timer across every session, and deriveStatus makes
service calls per session, so the permission and question lists are fetched once
per tick and indexed by session id rather than queried per session. A test pins
the call count.
Behaviour note beyond the strict fix: sharing the derivation also means a
SessionStatus of offline now reports as retry on the wire, matching what
deriveStatus has always sent to ingest. Nothing consumes offline from the
heartbeat — the transport forwards only idle and busy, and the mobile row treats
both as non-attention — so the effect is that the two channels now agree. The
detach fence test is parameterised accordingly; its assertion that the status
clears on detach is unchanged.
* chore(cli): widen the promise-facade allowlist for the heartbeat attention tests
The DEF-3 heartbeat tests raise and reply to real Question and Permission
requests through the global AppRuntime, which took kilo-sessions.test.ts from 4
classified references to 29 and failed the allowlist check.
Bumping the count rather than restructuring the tests is deliberate: the
heartbeat resolves attention status from the global Question.Service and
Permission.Service, so asserting it requires driving those same services.
Scoped layers cannot express that — the global-runtime coupling is the thing
under test — and it is the same integration pattern this entry already
sanctioned for the detach fence. The reason string records that.
* refactor(cli): run remote sessions in one process with safe per-session exit
Consolidate remote session handling into a single CLI process instead of
spawning one process per remote-created session (addresses the PR review):
- restore in-process create_session (accepts an absent sessionId and targets
the connection directory); remove the session spawner, the
KILO_REMOTE_ATTACH_SESSION attach-on-boot path, the child-advertisement gate,
and their tests
- retain instance advertisement and fire one immediate out-of-band heartbeat on
(re)connect when advertising, so a headless `kilo remote` host is discoverable
without delay
Make /exit (wire command exit_cli, unchanged for compatibility) detach only the
target session instead of terminating the CLI:
- AttachedState.detach with a presence-suppression tombstone; detach also clears
the target's SessionStatus so the negative-containment heartbeat fence resolves
deterministically for busy/retry/offline sessions
- exit_cli handler verifies ownership, cancels the active prompt, detaches and
awaits the detach heartbeat, then ACKs; the interactive RemoteExit callback is
invoked only after the ACK when the last owned session exits; a headless
`kilo remote` host stays alive and advertising at zero sessions
- add an optional canExitSession boolean to the list_commands v1 catalog
(always true, independent of exitAvailable) so clients can detect safe
session-exit semantics
History and stored sessions are preserved on exit.
* fix(cli): break module-load cycle in remote session prompt-cancel
The K1 in-process exit_cli seam added a static `import { SessionPrompt }`
to kilo-sessions.ts. @/session/prompt evaluates KiloSessionPrompt at module
load, so the new static edge raced that init and left the namespace in TDZ,
crashing unrelated test files with 'undefined is not an object (evaluating
KiloSessionPrompt.shouldAskPlanFollowup)'. Defer to a dynamic import at the
single call site, mirroring remote-command.ts.
* fix(cli): correct AttachedState announce/detach concurrency and rollback
Address review findings on the shared-process session lifecycle:
- announce/detach no longer join the OPPOSITE in-flight operation. Joining
detach's negative-containment fence made announce resolve success for a
detached id (and vice versa: detach joined announce and resolved success
while still attached, which exit_cli treats as license to ACK/close). Each
path now joins only a same-kind in-flight op and, when the opposite op is
in flight, awaits it to settle and then performs the real work.
- Failed-detach rollback now releases the suppression tombstone, so a
still-attached session is not dropped by the next setPresence (the tombstone
loop would otherwise remove the still-present id and never clear).
- Both catch/rollback branches now honor the lifecycle generation guard
(mirroring the success path); a stale in-flight op that rejects after
reset() no longer mutates the new lifecycle's presence/pending/suppressed
sets (reset clears the same Set instances).
Adds regression tests for each fix, plus AC6f covering the remote-ws
detachSessionId negative-containment waiter.