* fix(cli): advertise the instance from enableRemote so /remote registers as a spawn target
Enabling the remote relay from the TUI `/remote` slash command connected the
socket and mirrored sessions, but never advertised the instance, so the CLI
never appeared as a spawn target in the mobile "Run on" picker. Only the
explicit `kilo remote` command called setInstanceAdvertisement.
The advertisement now runs on every successful enableRemote() entry, before the
already-connected and coalescing early returns. That ordering matters: bootstrap
auto-enable frequently connects first, so `/remote` usually hits
`if (remote) return` and an advertisement placed in the connection-setup body
would leave the defect unfixed in the common case. `ingestDisabled` returns
before the advertisement and stays unadvertised.
The ensure helper is a no-op when an advertisement is already set, so it fires no
extra heartbeat, while explicit setInstanceAdvertisement keeps its existing
replace semantics. buildInstanceAdvertisement moves to a shared module so the
command path and the enable path derive it identically.
* fix(cli): report pending question and permission on the session heartbeat
The heartbeat built each session's status from SessionStatus.Service, whose
union is idle/retry/busy/offline and which never consults Question.Service or
Permission.Service. deriveStatus() already did consult both, but only fed the
ingest session_status sync. So a session genuinely blocked on a question was
advertised as busy on the heartbeat, and the mobile app — which takes live row
status from the heartbeat — showed no needs-input badge.
Extract the precedence (permission, then question, then SessionStatus) into a
shared helper used by both deriveStatus and the heartbeat, so the two channels
cannot drift.
The heartbeat runs on a ~10s timer across every session, and deriveStatus makes
service calls per session, so the permission and question lists are fetched once
per tick and indexed by session id rather than queried per session. A test pins
the call count.
Behaviour note beyond the strict fix: sharing the derivation also means a
SessionStatus of offline now reports as retry on the wire, matching what
deriveStatus has always sent to ingest. Nothing consumes offline from the
heartbeat — the transport forwards only idle and busy, and the mobile row treats
both as non-attention — so the effect is that the two channels now agree. The
detach fence test is parameterised accordingly; its assertion that the status
clears on detach is unchanged.
* chore(cli): widen the promise-facade allowlist for the heartbeat attention tests
The DEF-3 heartbeat tests raise and reply to real Question and Permission
requests through the global AppRuntime, which took kilo-sessions.test.ts from 4
classified references to 29 and failed the allowlist check.
Bumping the count rather than restructuring the tests is deliberate: the
heartbeat resolves attention status from the global Question.Service and
Permission.Service, so asserting it requires driving those same services.
Scoped layers cannot express that — the global-runtime coupling is the thing
under test — and it is the same integration pattern this entry already
sanctioned for the detach fence. The reason string records that.
* refactor(cli): run remote sessions in one process with safe per-session exit
Consolidate remote session handling into a single CLI process instead of
spawning one process per remote-created session (addresses the PR review):
- restore in-process create_session (accepts an absent sessionId and targets
the connection directory); remove the session spawner, the
KILO_REMOTE_ATTACH_SESSION attach-on-boot path, the child-advertisement gate,
and their tests
- retain instance advertisement and fire one immediate out-of-band heartbeat on
(re)connect when advertising, so a headless `kilo remote` host is discoverable
without delay
Make /exit (wire command exit_cli, unchanged for compatibility) detach only the
target session instead of terminating the CLI:
- AttachedState.detach with a presence-suppression tombstone; detach also clears
the target's SessionStatus so the negative-containment heartbeat fence resolves
deterministically for busy/retry/offline sessions
- exit_cli handler verifies ownership, cancels the active prompt, detaches and
awaits the detach heartbeat, then ACKs; the interactive RemoteExit callback is
invoked only after the ACK when the last owned session exits; a headless
`kilo remote` host stays alive and advertising at zero sessions
- add an optional canExitSession boolean to the list_commands v1 catalog
(always true, independent of exitAvailable) so clients can detect safe
session-exit semantics
History and stored sessions are preserved on exit.
* fix(cli): break module-load cycle in remote session prompt-cancel
The K1 in-process exit_cli seam added a static `import { SessionPrompt }`
to kilo-sessions.ts. @/session/prompt evaluates KiloSessionPrompt at module
load, so the new static edge raced that init and left the namespace in TDZ,
crashing unrelated test files with 'undefined is not an object (evaluating
KiloSessionPrompt.shouldAskPlanFollowup)'. Defer to a dynamic import at the
single call site, mirroring remote-command.ts.
* fix(cli): correct AttachedState announce/detach concurrency and rollback
Address review findings on the shared-process session lifecycle:
- announce/detach no longer join the OPPOSITE in-flight operation. Joining
detach's negative-containment fence made announce resolve success for a
detached id (and vice versa: detach joined announce and resolved success
while still attached, which exit_cli treats as license to ACK/close). Each
path now joins only a same-kind in-flight op and, when the opposite op is
in flight, awaits it to settle and then performs the real work.
- Failed-detach rollback now releases the suppression tombstone, so a
still-attached session is not dropped by the next setPresence (the tombstone
loop would otherwise remove the still-present id and never clear).
- Both catch/rollback branches now honor the lifecycle generation guard
(mirroring the success path); a stale in-flight op that rejects after
reset() no longer mutates the new lifecycle's presence/pending/suppressed
sets (reset clears the same Set instances).
Adds regression tests for each fix, plus AC6f covering the remote-ws
detachSessionId negative-containment waiter.
Introduce a `metadata` method on SessionProcessor.Handle that buffers
metadata emitted before tool-call registration, then applies it on the
running transition. This decouples metadata emission timing from
tool-call lifecycle.
Downgrade virtua from 0.49.1 to 0.42.3 and migrate the virtualizer API:
- Replace `findItemIndex(scrollOffset)` with `findStartIndex()`
- Replace `bufferSize` prop with `overscan` (count-based)
Additional changes:
- Change Permission.reply return type from Promise<boolean> to Promise<void>
- Make Ruleset type readonly and remove unnecessary array spreads
- Update nvidia provider headers to reference Kilo branding
- Reorder SDK event type definitions for installation events
- Reduce promise facade allowlist in check script
* feat(session-export): scaffold module with config constants
* feat(session-export): add zstd compression wrapper
* feat(session-export): event and envelope type definitions
* feat(session-export): eligibility check with kill-switch
* feat(session-export): org signal collector with auth resolver
* feat(session-export): worker SQLite schema and storage helpers
* feat(session-export): content-addressed chunker with zstd dedup
* feat(session-export): client-side light scrubber
* feat(session-export): IPC contract and inbox with back-pressure
* feat(session-export): persist scrubbed events with chunked payloads
* feat(session-export): worker entry point with inbox drain loop
* feat(session-export): main-thread capture module
* feat(session-export): workspace baseline and delta fibers
* feat(session-export): sync subscriber and tool io chunking
* feat(session-export): bootstrap wiring and compaction hook
* feat(session-export): wire capture hooks into sessions
* feat(session-export): uploader and buffer cap
* chore(session-export): annotate shared session hooks
* changeset(session-export): add release note
* fix(session-export): keep bootstrap non-blocking without instance context
* feat(session-export): respawn worker after failures
* test(session-export): add performance budget assertions
* test(session-export): add worker end-to-end smoke test
* fix(session-export): strip identity and high-risk baseline paths
* test(session-export): gate perf assertions for stable sweeps
* fix(session-export): bundle worker in single-file builds
* fix(session-export): make capture envelopes cloneable
* fix(session-export): drain worker on CLI shutdown
* feat(session-export): capture workspace baseline and deltas
* fix(session-export): preserve request metadata
* test(session-export): cover request metadata capture
* feat(cli): send indexed session export batches
* feat(cli): authorize session export uploads
* fix(cli): flush session export shutdown uploads
* feat(cli): optimize session export replay payloads
* fix(cli): include session export surface metadata
* fix(session-export): decrement chunk refs after upload
decRefChunks was never called, so chunk ref_count stayed at 1 (or
higher with dedup) and DELETE FROM chunk WHERE ref_count <= 0 never
matched. Chunks accumulated in the local SQLite buffer until the 50 GB
cap. Call decRefChunks alongside markUploaded so deleteUploaded can
reclaim the rows.
* fix(session-export): add periodic uploader flush timer
flushIntervalMs and retryBackoffMaxMs were both defined in config and
neither was referenced anywhere; scheduleFlush only ran on inbound
events or reconnect. After a 5xx or network failure the row was backed
off for 1 s, but if no further event arrived the retry never fired and
the events stranded until the CLI restarted. Drive a periodic flush
from a setInterval, unref the handle so it doesn't pin the process, and
clear it from a new dispose() hook the worker calls on shutdown.
* fix(session-export): stop emitting absolute workspace root in baseline
CaptureMetadata.root was the literal absolute filesystem path
(/Users/<name>/Projects/<repo>), shipped in every
workspace_baseline_completed event and not stripped by handlers'
identity filter. The field was set but never read anywhere downstream
— file paths in the baseline are already relative, so the root added
no replay signal. Remove the field outright.
* fix(session-export): cap pendingEvents result set
The SELECT had no LIMIT clause, so under a backlog (network outage,
crashed receiver) it would materialize the full pending table into a
JavaScript array before the byte-limit truncation applied. With a
50 GB buffer cap that is hundreds of MB of heap inside the worker per
drain. Add LIMIT 500 — the drain loop already re-queries until empty,
so no events are missed.
* fix(session-export): exponential retry backoff up to retryBackoffMaxMs
Both the 5xx and network-error branches always retried after the floor
delay regardless of how many attempts had already failed, and
retryBackoffMaxMs was unreferenced. During a sustained outage every
session re-tried at roughly 1 Hz against the dead receiver. Surface
upload_attempts on EventRow and compute the next delay as
min * 2^attempts capped at max.
* fix(session-export): evict superseded workspace snapshots on remember
createWorkspaceProvider retained every captured snapshot — in-memory
and inside the persisted state file — even after the session moved on
to a newer one. For a 1k-file repo over a 100-turn session that is
~2 GB of unreachable heap plus a state file that grows monotonically.
Drop the previous snapshot for the session on remember() when no other
session still references it.
* fix(session-export): anchor aws_secret_key scrubber to key name
The bare 40-char base64 pattern matched every 40-character hex string,
including all git commit SHAs. Tool outputs, diffs, and conversation
messages were silently rewritten as <<REDACTED:aws_secret_key>>,
destroying lineage information in training data. Require the key name
context — naked secrets in unstructured text are rare and the .env /
.aws/credentials high-risk path strip already covers the common case.
* fix(session-export): preserve root linkage in SyncSubscriber events
SyncSubscriber hardcoded rootSessionId = sessionId on every tool,
permission, and feedback event, so sub-agent sessions lost their root
linkage and a future training pipeline could not reconstruct the
agent topology from these side-channel events. Expose the rootSessionId
mapping from Capture and plumb it through the same pattern as
getTurnId.
* fix(session-export): atomic chunk GC after upload
markUploaded + decRefChunks + deleteUploaded were three separate SQL
statements; a crash between markUploaded and decRefChunks would leave
events flagged uploaded (never retried) and chunks with stale
ref_count (never reclaimed by deleteUploaded). Bundle the three calls
into a single transactional commitUploaded helper so either all three
land or none of them do.
* fix(session-export): wait on transient sqlite locks
* fix(session-export): snapshot current workspace directory
* fix(cli): preserve Kilo model export metadata
* fix(cli): send anon id for session export
* fix(cli): fallback to telemetry anon id
* chore(cli): annotate session export config test
* chore: remove session export docs
* fix(cli): harden session export uploads
* fix(cli): preserve stream lifecycle for session export
* fix(kilo-docs): exclude session export ingest link
* fix(cli): restrict session export workspace sync to git repos
* chore: remove session export changeset
* fix(cli): tighten session export payload types
* fix(cli): type session export model payloads
* fix(cli): simplify session export cleanup
* fix(cli): avoid session export shutdown race
* refactor(cli): use drizzle for session export storage
* fix(cli): finalize session export sqlite statements
* fix(cli): link compaction exports to root sessions
* fix(cli): stop exporting raw stream parts
* fix(cli): prune stale workspace snapshots
* refactor(cli): clarify chunk ref counting
* test(cli): clarify dropped upload assertions
* ci: avoid visual path filter action failure
* fix(cli): preserve in-flight workspace snapshots
* fix(cli): stop uploading baseline start events
* fix(cli): trim redundant export metadata
* fix(cli): trim workspace export bookkeeping
* fix(cli): fold terminal outcome into tool exports
* fix(cli): dedupe request context in export batches
* fix(cli): normalize compaction export payloads
* fix(cli): avoid duplicate tool result exports
* test(cli): align session export expectations
* fix(cli): run secretlint during session export scrubbing
* fix(cli): keep retried export batches contiguous
* fix(cli): include agent info in session exports
* fix(cli): flush pending session exports on serve startup
* fix(cli): drop exports when scrubbing fails
* fix(cli): narrow secretlint value extraction
* fix(cli): limit exported agent info
* test(cli): remove brittle agent export source assertion
* fix(cli): ignore corrupt session export workspace state
* fix(cli): keep session export close best effort
* fix(cli): avoid following session export symlinks
* fix(cli): preserve session export chunk ref counts
* fix(cli): retry transient session export uploads
* fix(cli): validate session export ingest endpoint
* fix(cli): tolerate missing export token details
* fix(cli): fail closed on session export org lookup
* fix(cli): decode session export permission replies
* fix(cli): finalize session export on stream close
* fix(cli): bound session export baseline timeout
* fix(cli): avoid persisting workspace file contents
* fix(cli): bound workspace snapshot capture
* fix(cli): validate session export worker messages
* fix(cli): revoke stale session export eligibility
* fix(cli): scope session export workspaces
* fix(cli): extend session export shutdown flush
* fix(cli): infer free Kilo models for export
* fix(cli): avoid duplicate chunk refs
* fix(cli): throttle session export uploads
* fix(cli): harden session export capture
* feat: disclose free model data collection (#10767)
* feat: disclose free model data collection
* chore(cli): document free model footer sorting
* fix(vscode): add free model data translations
* fix: simplify free model data label
* fix: simplify data collection badges
* fix: remove duplicate model info disclosure
* fix: restore composer data tooltip
* fix: align jetbrains data collection indicator
* fix: align free model data indicators
* fix(vscode): show data collection in model preview
* fix: limit data indicators to kilo gateway
* test(cli): relax prompt cancel timeout
* test(cli): annotate prompt cancel timeout
* test(cli): classify prompt queue runtime test
* fix(cli): defer session export startup
Migrate all Session.* promise-based helper usage (create, get, messages,
updateMessage, updatePart, setPermission, children, fork) to direct
Effect service consumption via Session.Service. This eliminates the
makeRuntime-backed legacy promise wrappers from session.ts and updates
all dependent modules and tests to use the Effect-based interface.
Key changes:
- Remove exported promise helpers (create, get, messages, etc.) from
session/session.ts along with the module-level makeRuntime instance
- Update kilo-sessions, remote-sender, allow-everything permission,
plan-followup, recall tool, and fork module to yield Session.Service
- Convert remapChildren from async function to Effect generator
- Add PlanFollowupRuntime.session helper for effectful session access
- Update all affected test files to use Effect.runPromise with
Session.Service.use pattern
- Remove session/session.ts from the promise facade allowlist and
update test allowlist counts accordingly