* feat(tui): expand a collapsed paste on a second identical paste
* chore: retrigger review
* chore(tui): add paste expansion changeset
* fix(cli): target paste changeset
* fix(tui): refresh autocomplete after expanding a paste
Call auto()?.onInput on the expand-placeholder path so open autocomplete
matches onContentChange when a second identical paste expands text.
* feat: add signal-to-noise controls to grep tool
Add configurable options to reduce noise in grep results:
- context: show N lines before/after each match
- limit: bound maximum matches (default 100) with early termination
- literal: treat pattern as plain text instead of regex
- ignoreCase: case-insensitive matching
These controls help models avoid overwhelming context windows with too
many matches and provide clear guidance when results are truncated.
Implementation preserves upstream ripgrep structure with minimal Kilo
hooks for additive behavior only. Shared-file changes reduced from 183
to 60 lines compared to initial implementation.
* test(core): tolerate PTY event publication race
* fix(cli): count grep matches independently of context
* feat(agent-manager): allow sessions to move their worktree between sections or ungroup
Add a new to the tool so a session can
reassign its own worktree to another section or ungroup it by passing
. The move operation validates the target section and
the session's worktree, then pushes refreshed state to the panel.
Also tighten the model-facing contract so the list-to-move workflow is
unambiguous: is required first, its result is the
source of truth for section and session IDs, and direct edits to
are rejected by , , and
with an explicit pointer to the tool.
- New / in the Agent Manager protocol
- New domain function with section/session validation
- New OpenAPI hook so stays
nullable in the generated SDK
- Tool schema descriptions now spell out the list-first workflow and
the no-direct-edit rule
- List output now includes an instruction block and is pretty-printed
- Optional nullable Task fields tolerate from models
- Focused CLI tests, bridge unit test, and protection test
* chore: remove local PR screenshot
* style(vscode): format agent manager orchestration
* fix(agent-manager): protect state paths on Windows
* feat(opencode): remote create_session fields, rename adoption, title sync
Extend create_session wire with optional agent/model/orgId (strict v1,
old-CLI degrade via client retry); claim org via session metadata
(metadata > KILO_ORG_ID > auth); adopt system session.renamed via
setTitle with consume-on-failure adoption marks; POST generation-aware
title changes through readiness (auto-titles marked by ensureTitle,
same-title Updated consumes pending adoptions).
* test(opencode): prove cancel→reprompt reaches idle; lock exit survivor
Item 14 CLI prove-it at SessionPrompt level: cancel-when-idle,
mid-stream, mid-tool, queued follow-up (deterministic queue wait), and
abortIntakes all settle to idle and reprompt completes — no production
hang found, no src change. Item 8: lock survivor session send_message
after sibling exit_cli.
* test(opencode): drop AppRuntime spy from create_session default test
Satisfies check-opencode-promise-facades while still proving the
production default forwards {agent, model, metadata} into
Session.Service.create.
* fix(opencode): bound rename marks, wire title report path, harden title tests
Kilobot review on #12704: adoption/auto-title maps now carry timestamps,
prune on write (60s TTL), and clear on Session.Event.Deleted (exported
clear/clearAll); the Updated watcher calls the interface
reportSessionTitle and fullSync passes preloaded info into meta();
ensureTitle's Kilo logic lives in kilocode/session/prompt.ts behind one
kilocode_change call site; title tests poll instead of sleeping and lock
mark-before-write plus clear-on-failure for real; meta() get-failure
org fallback covered via the _metaForTests seam.
* fix(kilo-sessions): mark bookkeeping before ingest sync, AppRuntime, test cleanup
Kilobot round 2 on #12704: consume rename/auto-title marks before the
ingest.sync network hop so the 60s TTL spans only the in-process hop;
call reportSessionTitle via AppRuntime.runPromise; auth cleanup back
under Effect.ensuring; restore the upstream blank line in prompt.ts so
the fork diff is only the kilocode_change call site.
* fix(kilo-sessions): keep title report self-healing if ingest.sync fails
Advance knownTitles only after successful sync; restore consumed rename/
auto-title marks on failure so the next Updated can re-POST. IIFE keeps
const-style outcome derivation.
* fix(kilo-sessions): optimistic knownTitles with full title-path rollback
Advance knownTitles before the network hop so concurrent Updated handlers
see sameTitle and cannot POST the same title with a wrong generated flag.
Restore prev + consumed marks when ingest.sync throws or reportSessionTitle
returns not-ok, so the next Updated retries the full self-healing path.
* style(kilo-sessions): prettier title Updated handler
* fix(kilo-sessions): preserve newer title state
* refactor(kilo-sessions): simplify title reporting tests
* fix(kilo-sessions): report unseeded title updates
* fix(kilo-sessions): consume unseeded title marks
* test(kilo-sessions): cover unseeded title marks
* test(kilo-sessions): unique ids for unseeded title tests
Thread a distinct session id through unseededMockSessionLayer so
session_share Storage records do not couple the three unseeded cases.
Applies the kilocode-merge-minimizer skill to the prior fix. The
hard-veto and headless-subagent DeniedError sites are reverted to
their exact pre-fix shape -- neither carries a specific rule anyway,
so wrapping their ruleset in a { rule, matches } object added shared
upstream diff for no benefit. Only the main deny path (which already
had the deciding rule in scope) still changes, and now passes the
bare rule instead of a wrapper object, shrinking that hunk from a
multi-line block to a single-line swap.
PermissionProvenance.classifyDenial now duck-types ruleset as a
possible bare Permission.Rule (checking action === "deny" and a
string pattern) instead of expecting a { rule } wrapper, so it still
reads the main deny path's rule directly while falling back to a
synthesized deny rule for the other paths, exactly as before.
Net shared-file diff across permission/index.ts, session/tools.ts, and
the TUI's routes/session/index.tsx for this whole feature is now 9
insertions / 12 deletions, down from ~50+ lines.
DeniedError.ruleset only carried the deny-permission subset, so
PermissionProvenance.classifyDenial had to guess the deciding rule via
findLast(action === "deny"). With two deny rules for different
patterns under the same permission (e.g. bash: { "git push *": deny,
"rm -rf *": deny }), this could attribute a denial to whichever rule
sorted last instead of the one that actually matched the request.
Permission.ask now embeds the exact rule resolve()/evaluate() matched
against the request's pattern directly on the error (ruleset: { rule,
matches }), so classifyDenial reads it instead of re-deriving it.
Some denials carry no rule at all (e.g. the headless-subagent policy
denial), where classify({ rule: undefined }) reports the same
{ source: "default" } shape as the *approval* fallback -- silently
rendering a refusal as an auto-approval in the TUI and kilo export.
classifyDenial now synthesizes an explicit deny rule for the request's
permission/pattern in that case, so rule.action always reflects the
real outcome.
Adds test/kilocode/permission/deny-provenance.test.ts covering both
regressions against the real Permission.Service, and updates the
existing session-tools.test.ts denial fixture to the new ruleset
shape.
* fix(cli): stabilize Windows CI tests and rebalance slow shards
Three Windows-only instabilities in the CLI unit suite:
1. httpapi-instance-route-auth.test.ts failed with an uncaught
"Invalid handle" error. The test's ConfigProvider.layer(
fromUnknown(...)) replaced the ambient config provider, blinding
KILO_EXPERIMENTAL_DISABLE_FILEWATCHER=true that CI/preload sets. With
the flag hidden, the @parcel/watcher Windows backend subscribed on the
temp repo's .git; the tmpdir fixture then deleted that directory while
the never-disposed per-test runtime still held the subscription, and
CreateFileW failed with the hardcoded "Invalid handle" (napi rejection
with no JS stack). Add the disable-filewatcher flag to every test
config map that boots instances via the HttpApi app (instance-route-auth,
cors, ui, exercise backend, kilo-edit, memory).
2. config-overlay.test.ts intermittently returned HTTP 500 on Windows.
Filesystem.write's atomic temp-file+rename had no retry for Windows
transient locked-file errors (EPERM/EACCES/EBUSY) from Defender/indexer
and the detached background plugin install racing the rename in the same
tmpdir. Mirror the proven cleanup.ts locked-error retry pattern with a
short backoff, Windows-only.
3. Windows shards were badly imbalanced: the sharder weighted files by
byte size, which concentrated every slow spawn/FS/lock-heavy file
(snapshot, prompt, provider, run-process, instance-bootstrap,
httpapi-session) into one shard (~612s vs ~356s siblings), and the
resulting contention forced whole-file retries that doubled cost. Add
TestShard.timedWeight and a committed test-timings.json seeded from CI
junit data so shards balance by measured runtime (spread collapses from
~200s to ~18s) and contention-prone files spread across shards.
Platforms without manifest entries fall back to size weighting.
* fix(cli): skip stale manifest entries in timed shard weighting
Bun.file().size returns 0 (never throws) for missing paths, so the
try/catch in timedWeight was dead code and stale/renamed manifest entries
added their time to the scale numerator with zero size, inflating the
size-to-time ratio that estimates unknown files. Skip entries with a
non-positive on-disk size instead of catching a throw that never happens.
* revert(cli): drop hardcoded test-timings manifest
The committed test-timings.json (482 entries) was a maintenance burden:
it goes stale as tests are added/renamed and no size-based heuristic can
replace it (slow subprocess outliers like run-process.test.ts are 7kb but
112s, 10x the runtime-per-byte of other files). Revert the timing-weighted
sharding to the prior size-based LPT. The Windows reliability fixes
(ConfigProvider filewatcher flag + Filesystem.write locked-file retry)
remain and are what eliminate the failures and the ~360s of retry overhead
that dominated the 12m50s shard. A maintainable runtime-based rebalance
(self-updating CI cache fed from the junit artifacts CI already uploads)
is a separate follow-up.
Auto-approval provenance was only recorded on the metadata of allowed
tool calls (state.metadata.approval), so denied calls had no
structured explanation of which rule/config/agent denied them. Since
'kilo export' serializes state.metadata verbatim into the JSON session
log, denials showed up with no provenance at all.
Add PermissionProvenance.classifyDenial, which reads the deciding deny
rule off a DeniedError's tagged ruleset and classifies it the same way
approvals are classified. Wire it into SessionTools' ctx.ask via
Effect.tapErrorTag so denials are recorded before the tool call fails,
reusing the existing carryApproval/failToolCall preservation so the
metadata survives onto the final error state.
The auth mode probes every protected route with valid credentials to
prove the route accepts them. Two routes intentionally block after a
valid request: tui.control.next waits for queued TUI input, and
instance.reload can outlive the one second probe. The probe won by
timeout, marked the scenario passed, and left the request alive inside
the cached web handler. Final disposeApps() then waited for those
requests forever, so any run that rebuilt the app (cold turbo cache)
never exited and CI cancelled the job at 15 minutes.
Add a scenario flag that keeps the missing-credentials check (the route
must still return 401) but skips the valid-credential probe for routes
whose valid requests intentionally block, and mark tui.control.next and
instance.reload with it.