--update-durations now disables batching so every file is measured
per-file (batched members' entries otherwise rot), and only passing
runs record durations (a timed-out file would write the kill deadline
and permanently skew LPT).
Timed-out work items are no longer retried: the deadline already
proves a hang, and a second 600s attempt could push a shard past the
45-minute job budget.
A flaky batch now annotates the member files that failed the first
attempt instead of the pseudo-file. Defender exclusions apply per
path instead of aborting on the first policy-locked one.
TURBO_FORCE covered only the CLI test step; the non-CLI packages step
still replayed cached results on same-commit dispatches, so a soak run
could report package tests green without executing them.
workflow_dispatch runs exist to validate real execution (soak passes,
timing measurements); a same-commit dispatch was replaying FULL TURBO
cache hits and executing nothing. Push and PR runs keep the cache.
The deadline exists to catch hangs, not to cap healthy runs. A
saturated Linux shard held cli/run/run-process past 300s with every
test inside passing, and the runner killed it twice. Windows already
ran with 600s for the same reason.
Measured with the Defender exclusions active: Setup Bun ran 83-102s
with cache restore vs ~86s with a fresh install — no benefit, plus a
save-step cost on non-PR runs. The original reasoning stands.
Measured on the integration branch: five shards were worse on both
wall-clock (slowest job 9.3m vs 6.7m) and machine minutes (43 vs 37).
With two workers per shard, packing more work per shard loses.
Shared-process batching cut per-shard test time enough that five shards
hold the previous wall-clock, saving one job's fixed setup and pool
pressure per run.
The Windows bun cache goes back on: the restore-slower-than-install
measurement predates the Defender exclusions, which cover both the bun
cache directory and node_modules writes.
Real-time scanning taxes every process spawn and temp-file write; the
unit suite spawns hundreds of test processes and creates temp git repos.
Excluding the workspace, temp dirs, and bun cache cuts that overhead.
Best effort: on images where Defender is absent or locked down the step
logs and the suite runs unchanged.
The worker queue sorted by raw file weight, which is 0 for batch
pseudo-files (not real paths) — batches started last, leaving one
worker running a full batch after the rest of the shard drained.
Sorting by shard weight starts the heaviest items first and shortens
the tail.
The batch-safety guard now reads member sources in parallel, and the
diagnostics-only artifact upload no longer fails a green job on a
runner-side network blip (seen as ECONNREFUSED on Blacksmith).
Windows CLI unit tests ran 4-at-a-time on the 4-vCPU runner (custom
test-runner default = min(4, cpus)), oversubscribing CPU so heavy
real-server test files blew their per-test timeouts.
- Cap KILO_TEST_CONCURRENCY=2 on Windows (each file gets ~2 vCPU).
- Grow Windows shards 4 -> 6 to absorb the lower per-shard parallelism.
- Shard by observed duration instead of file size, so the two heaviest
files no longer stack in one shard.
- Raise the Windows per-file kill deadline to 600s (KILO_TEST_FILE_TIMEOUT);
the heaviest file runs ~270s, only ~30s under the old 300s default.
Scoped to the @kilocode/cli custom runner on Windows only. Linux/macOS
unchanged; the non-CLI/core suite runs via plain 'bun test' and is not
affected (separate lever).
Three defects from the review of #12823.
- A throw after the learn step started left no prompt artifact, so triage and
edit ran with no learned rule at all. The step is continue-on-error, so the
run continued and the failure was silent. learn.mjs now writes both prompt
artifacts from the checked-out file before any fallible work, and replaces
them once the rolling branch copy loads.
- The direct marker PATCH sent the body read before the extraction call, so it
overwrote any body edit made in the minutes since. learn.mjs now re-reads the
body immediately before the PATCH.
- Two additions in one model response could carry one id or one rule text.
validateDelta now rejects a duplicate of an earlier accepted addition.
Tests: 10s pins the prompt artifacts across a failed API call and the removal of
a stale block. 10t drives learn.mjs against a stub GitHub API whose second read
returns a maintainer edit, and asserts the edit survives the PATCH. 10u pins
both duplicate rejections. DOCS_SYNC_API_BASE is the new selftest-only hook that
points lib.mjs at the stub server.
* fix(ci): pass --auto to the docs-sync kilo runs and keep full stderr logs
Headless kilo run auto-rejects every permission ask it receives, and the
GitHub runner has no user config granting bash — so without --auto the
docs-sync bot's triage, edit and verify-fix calls were silently crippled
whenever the agent reached for a non-allowlisted shell command (CI run
30306629290: 9 rejections, all 11 edit batches failed, exit 0).
- Pass --auto immediately after "run" at all three call sites
(triage.mjs, edit.mjs, docs-sync.yml Fix verify failures step)
- runKilo now always writes the child's full stderr to
docs-sync-out/kilo-stderr-<label>.log, on success as well as failure —
the blindness that hid the defect
- selftest asserts --auto non-vacuously (region-scoped source match +
a stub invocation that records argv) and proves the stderr log is
written on both the failure and the summary-writing success path
* fix(cli): exit nonzero when a headless run auto-rejects or its session errors
Non-interactive kilo run reported success for runs that accomplished
nothing — a caller cannot distinguish success from a dead run, which is
why the docs-sync bot had to stop trusting exit codes entirely.
- Plain headless run (neither --auto nor --dangerously-skip-permissions)
in which the CLI auto-rejected at least one permission ask now exits
non-zero, even when the session still reaches idle afterwards, with
the diagnostic: run ended with an auto-rejected permission; pass
--auto for autonomous use. Deliberate contract change: any
auto-rejected ask means the run was crippled, not successful.
- A mid-stream session error now writes its diagnostic to stderr before
emitting the json error event, so the cause is visible under
--format json as well (the emit-first shape skipped the stderr write).
- The old exit-0 contract lock-in test is replaced: its llm.fail fixture
never published a consumable session.error, so it locked in a false
premise. New tests cover both scenarios in both output formats;
happy-path and --format json runs still exit 0 unchanged.
* fix(ci): redact secret env values from persisted kilo stderr logs
The full-stderr capture added for observability lands in uploaded CI
artifacts, which are raw files (GitHub masks secrets in log streams
only) on a public repo, and the runner env holds a long-lived
KILO_API_KEY. Redact exact values of KEY|TOKEN|SECRET-named env vars
(len >= 8) once at capture, so both the console tail and the artifact
file are safe. Adds the changeset for the headless exit-code contract
change.
* fix(ci): redact secrets from persisted kilo stdout and harden redaction
The docs-sync workflow uploads docs-sync-out/ as a public 14-day artifact,
and kilo stdout was persisted raw there in two more places: triage-raw-*.txt
and the edit-log.txt tee. Redact at capture in runKilo for stdout as already
done for stderr, and pipe the verify-fix step's kilo stdout through a new
line-wise redact-stream.mjs filter before tee.
Also harden redactEnvSecrets: widen the name pattern to
KEY|TOKEN|SECRET|CREDENTIAL|PASSWORD|ORG_ID|_PAT (KILO_ORG_ID is a repo
secret), replace longer values first so a short secret that prefixes a
longer one cannot leak the remainder, and document the exact-substring
limitation. Selftest gains cases 2e (prefix ordering), 2f (stdout capture),
2g (stream filter) and a non-vacuous exact-line assertion in 2d.
* fix(cli): correct the headless-exit changeset's json-format claim
The auto-reject path adds a new error event to the --format json stream;
only the existing event shapes are unchanged. Also state that the exit-1
rule covers a plain non-interactive --attach run that auto-rejects an ask.
* fix(ci): document --auto security trade-off and deferred hardening
Update comments in triage.mjs and edit.mjs to accurately describe the
security implications of --auto (unrestricted bash for an agent steered
by external PR content) and note that a scoped permission.bash map via
KILO_CONFIG_CONTENT is the intended hardening, deferred until required
shell patterns are stable.
* fix(ci): raise docs-sync budgets so the backlog can actually drain
--auto fixes the batches the bot attempted; it does not fix the ones it
never started. In run 30306629290 (254 PRs collected, 51 docs-worthy),
the wall-clock budgets deferred 54 PRs untriaged and 31 unedited without
an attempt — 45 of the 60 pending rows on the rolling PR. Triage got 8 of
11 chunks in 35 min; edit got 4 of 11 batches in 50 min.
Both are ceilings, not costs. A caught-up run needs ~2 chunks and ~1
batch and finishes in ~25 min, so raising them spends nothing on a normal
day and drains the backlog on a bad one. 90/120 covers 20 chunks and 14
batches — 500 triaged and 70 edited PRs against a ~5 docs-worthy/day
inflow — inside a 240-minute job timeout.
Also strip ANSI CSI sequences in tailText, the shared path both triage
and edit route their pending causes through. kilo renders its TUI to
stderr, so every "Why" cell on the rolling PR currently reads
"^[[0m→ ^[[0mRead packages/..." instead of the diagnostic. The persisted
docs-sync-out/kilo-stderr-*.log stays raw as the debugging record.
selftest case 2h asserts the pending reason is escape-free and still
carries the diagnostic text; case 2i asserts each budget fits at least
two units and the job timeout outlasts both, so a future edit cannot
silently restore a budget too small to run anything. Both shown failing
on the unmodified code first.
* fix(ci): keep the docs-sync rebuild authoritative in the fix step
The "Fix verify failures" step runs under `set -o pipefail` and the
default `bash -e`. Once the CLI half of this PR ships, `kilo run` exits 1
on a mid-stream session error, which aborts the block before the rebuild
runs: verify2.log is never written and `Re-verify status` reports
VERIFIED=false even when the docs build fine. A transient provider error
would flip every rolling docs PR to "Verification: failing".
The agent's exit code was never the signal for this step — the rebuild
is. Guard the pipeline with `|| echo ::warning::` so a nonzero kilo run
is surfaced but the rebuild still decides the outcome.
Extends selftest case 2b (which already parses this step) rather than
adding a case: the guard must appear between the pipeline's tee and the
rebuild, so a comment elsewhere in the block cannot satisfy it. Shown
failing with the guard removed.
---------
Co-authored-by: kiloconnect[bot] <240665456+kiloconnect[bot]@users.noreply.github.com>
The generateOpenApiSpec Gradle task fetches pinned Kilo CLI release
metadata from the GitHub REST API to generate the backend OpenAPI
client. The test-jetbrains workflow did not expose a token to this
step, so the calls were unauthenticated and subject to the 60/hour
per-IP limit. On shared Blacksmith runners that budget is exhausted
across concurrent jobs, intermittently failing the build with
'API rate limit exceeded'.
Pass GITHUB_TOKEN to the test step so the task uses the authenticated
(1000/hour repo-scoped) limit.
* fix(ci): configure git identity before the docs-sync merge and classify merge failures
* fix(ci): hold the docs-sync watermark back until every PR has an outcome
* test(ci): self-check for the docs-sync failure paths
* fix(ci): isolate PR selftest concurrency from the daily docs-sync run
Passing the JETBRAINS_CERTIFICATE_CHAIN / JETBRAINS_PRIVATE_KEY multiline
secret content directly as certificateChain/privateKey Gradle properties
gets mishandled by the zip-signer CLI when signPlugin and
verifyPluginSignature run as separate Gradle invocations (#12567), causing
verifyPluginSignature to fail with 'Invalid argument: ***' as the masked
multiline content is split into extra CLI args.
Mirror script/build-version.sh: write the certificate chain and private
key to temp files under $RUNNER_TEMP and wire
certificateChainFile/privateKeyFile (file-based) into the intellij
signing extension instead of certificateChain/privateKey (raw content).
Co-authored-by: kiloconnect[bot] <240665456+kiloconnect[bot]@users.noreply.github.com>
Without KILO_ORG_ID the kilo provider sends no X-KiloCode-OrganizationId
header, so the gateway bills the key owner's personal balance instead of
the org. The dry run (run 30107542234) 402'd every triage chunk with
"Add credits to continue, or switch to a free model" and silently
classified 0/187 PRs.
Hoist KILO_API_KEY + KILO_ORG_ID to job-level env, matching smoke-test.yml.
* feat: daily docs-sync bot workflow (Kilo CLI)
Adds a scheduled workflow that keeps packages/kilo-docs in sync with PRs
merged to Kilo-Org/cloud and Kilo-Org/kilocode:
- watermark.mjs derives the processing window from the bot's own PR body
marker (self-healing, no external state; 72h fallback, 14d cap)
- collect.mjs queries merged PRs via the GitHub API and applies a
deterministic pre-filter (bots, chores, docs-only PRs)
- triage.mjs classifies PRs in chunks of 25 with kilo run; failed chunks
degrade to unclassified instead of failing the run
- edit.mjs updates docs in batches of 5 PRs with kilo run, bounded per
batch; failures surface as skipped entries in the PR body
- verify runs the kilo-docs build + test suite; one LLM fix pass on
failure; still-red becomes a draft PR
- upsert-pr.mjs maintains one rolling auto-docs PR (appends while open,
fresh branch after merge), with a 15-file draft cap and a
machine-readable processed-through watermark
Also adds docs-sync.yml to the workflow allowlist in
script/check-workflows.ts.
* fix: correct kilo run invocation and auth
- message positional must come before flags: --file is multi-value and
consumes a trailing message as a file path (File not found)
- authenticate via the existing KILO_API_KEY repo secret (the kilo
provider reads it natively); drop the DOCS_SYNC_KILO_CONFIG config
secret requirement
- fix default model IDs: gateway provider id is kilo/, not kilocode/
- include stderr tail in triage/edit failure logs
* fix: handle kilo run double-printed assistant output
kilo run prints the assistant message twice (streaming render + final
summary), so stdout can contain the same JSON array back-to-back. Parse
the largest valid trailing array instead of slicing first-to-last
bracket. Verified against real chunked triage output.
* fix: reviewer-pass robustness fixes
- edit.mjs: unambiguous summary file path in the batch prompt and a
fallback read when the agent drops the docs-sync-out/ prefix, so real
edits never report as skipped
- prepare-branch.mjs: use the open auto-docs PR's actual head.ref
instead of assuming docs/auto-sync
- upsert-pr.mjs: compute the 15-file draft cap on the cumulative PR
diff (origin/main...HEAD), not just the latest commit
* fix: address Kilobot review findings
Security:
- sanitize HTML-comment sequences out of agent-generated PR body values
so a crafted value cannot forge section markers or the watermark
- draft any PR whose diff touches non-content files in packages/kilo-docs
(outside pages/ and lib/nav/) — build-executable changes force human
review before merge
- on merge conflict, keep the conflicted rolling branch untouched
(preserving human commits) and continue on a fresh dated branch that
links the old PR
Resilience:
- retry GitHub API calls on network errors and 5xx, not just 403
rate limits
- isolate per-PR collect failures instead of aborting the run
- trust watermark markers only on bot-authored PRs and clamp future
dates loudly
- validate chunk triage entries belong to their chunk before the shared
dedupe
- use changed_files for files_total and skip docs-only classification
on truncated (300+) file lists
- pipe stderr in the edit pass so failure warnings carry the real CLI
error
* fix: address second Kilobot review round
- escape pipe characters in changeRow actions (same as skippedRow)
- sanitize agent-chosen file paths before they land in draftReasons
and the PR body (residual marker-forgery path via filenames)
- log expected fetch misses in prepare-branch instead of silent catches
* feat: keep bot-authored PRs in the docs-sync digest
Release and dependency bots ship user-facing changes (e.g. JetBrains
release PRs from kilo-maintainer[bot]). The auto-docs label check and
docs-only path filter remain as the loop guards.
Main's per-instance watcher lands without the unit-test disable env on
the exerciser job, so 289 scenario instance boots spawn inotify watchers
on Linux and the process dies with EBADF. Mirror the unit step's
KILO_EXPERIMENTAL_DISABLE_FILEWATCHER for the exerciser.
The CLI never builds the v2 location stack, so the .git watcher that
feeds Vcs branch updates never starts and the sidebar branch label goes
stale after a git switch outside Kilo.
Warm the LocationServiceMap entry for the instance directory during
kilocode bootstrap. This starts the same watcher the file/pty handlers
use, keyed by the same ref, so the stack is shared instead of
duplicated. The warm-up is forked off the bootstrap critical path and
failures only log a warning.
Fixes#12225