Picks up the shared button primitives (#13164), the one-row-per-tool-call
chat rendering (#13186), the refined session chat layout (#13205), and the
@pierre/diffs hunk renderer (#13201) that landed since 0.2.0-next.3.
* Render desktop diff view hunks with shared @pierre/diffs renderer
Replace DiffView's hand-rolled DiffHunk +/- line rows with ToolFileDiff
from @cline/ui (backed by @pierre/diffs), matching the chat tool rows.
Hunks carrying complete new contents (created files) render with real
line numbers; fragment hunks hide them, mirroring ToolCallRow. All of
DiffView's chrome (collapse, copy, open-in-editor, counts) is unchanged.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Make ToolFileDiff syntax palette follow the app theme, not browser preference
@pierre/diffs declares 'color-scheme: light dark' on its shadow :host, so
its light-dark() token colors resolve from the browser's preferred scheme.
Apps themed by the .dark class (desktop app) got the light palette's
near-black text on dark surfaces. Inline colorScheme: inherit on the host
wins over the :host rule and follows the app's color-scheme, which the
@cline/ui theme already flips with .dark. Skipped when a caller pins an
explicit themeType.
Also key diff-view hunks by index so repeated same-shaped hunks (a file
created twice with identical contents) don't collide.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
The Claude Code provider was unusable for agentic work (#13146):
- The claude-code manifest lacked the provider-tools capability, so the
gateway sent Cline's tool definitions (which the provider drops as
unbridgeable) while the CLI's native tools stayed enabled with no
approval plumbing - every write was refused and no prompt appeared.
- ai-sdk-provider-claude-code defaults settingSources to [], so the
spawned session read neither ~/.claude/settings.json nor project
settings, silently ignoring user-configured permission rules.
- No cwd was passed, so the session inherited the extension host's
cwd (/ on macOS) and refused writes outside it.
Changes:
- Mark claude-code with provider-tools (same treatment as the Codex
CLI provider): stop sending unbridgeable external tools and let the
CLI execute its own, tagged executionMode=provider for the runtime.
- Forward the session workspace cwd from @cline/core into the
claude-code gateway provider options and lift it into the agent
session settings.
- Default settingSources to [user, project] and permissionMode to
acceptEdits (file edits under cwd auto-approved; command execution
stays gated by the user's own Claude settings), all overridable via
explicit defaultSettings.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(desktop): don't let a stale queued send response wedge the composer
A fresh session is still busy while its interactive loop starts, so the
sidecar coerces the first send onto the pending-prompt queue and replies
{queued:true} with a queue snapshot taken at enqueue time. The turn itself
runs via the runtime's queue drain and completes through stream events
(chat_queued_prompt_start -> deltas -> chat_done). On cold/slow sidecars the
RPC response lands only after those events; the webview then applied the
stale snapshot and unconditionally set status back to "running", leaving
the composer on "Agent is working..." forever and resurrecting a phantom
queue entry.
Webview: capture the turn epoch at send dispatch; chat_queued_prompt_start
bumps it, so a mismatch when the queued response arrives means the stream
already advanced the turn lifecycle and the response is ignored. Aborts now
resolve the queued branch to "cancelled" like the direct path.
Sidecar: the queued send response no longer routes its enqueue-time snapshot
through applyPendingPrompts, which overwrote the event-maintained
session.promptsInQueue and rebroadcast the stale list to every webview.
Includes deterministic regression tests for the stale-response orderings
plus temporary [P0DBG] debug instrumentation (region-marked, to be removed
after runtime verification).
* fix(desktop): ignore stale hub 'running' status after turn settles
The sidecar core is hub-attached, so chat_session_status events are
asynchronous projections of the hub's session record. A stale 'running'
can trail the stream's chat_done and flip a settled turn back to busy,
wedging the composer on 'Agent is working…' with nothing left to
reconcile. Track the epoch at which the turn settled and drop 'running'
status events until a new turn bumps the epoch.
* chore: remove stray QA screenshot artifacts from repo root
* chore(desktop): remove P0 debug instrumentation and fault injection
Strips all [P0DBG] logging, the /p0dbg sidecar route, the webview log
mirror + heartbeat, and the P0DBG_STARTUP_BUSY_MS /
P0DBG_DELAY_QUEUED_RESPONSE_MS fault-injection paths used to reproduce
the stuck-composer P0. The two real fixes (stale queued-response epoch
guard + stale-running-after-settle guard in the webview, and the
sidecar's non-clobbering queued-send snapshot) and the regression tests
remain.
* refactor(desktop): replace turn-epoch guards with an explicit turn lifecycle
The stuck-composer fixes left the hook with two hand-rolled epoch refs
(turnEpochRef / turnSettledEpochRef) mutated and compared inline across
eight call sites. Extract the rules into a pure TurnLifecycle module that
is now the only writer of the session status:
- a settled turn cannot be reopened: stale hub 'running' projections and
stale queued-send acknowledgements are dropped by the lifecycle instead
of by inline epoch comparisons
- async work (send RPC responses, queue reconciliation) captures an opaque
token and the lifecycle decides whether the world moved on, instead of
handlers comparing counters
- every status write goes through a named operation (begin, turnStarted,
settle, projectStatus, apply, reset), so the state machine is explicit
and unit-testable in isolation
No behavior change: the 5 wedge regression tests and the full hook suite
pass unchanged, plus 10 new unit tests for the lifecycle module itself.
* Revert "refactor(desktop): replace turn-epoch guards with an explicit turn lifecycle"
This reverts commit c480aaabe8.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* fix(vscode): stop legacy-migration backlog telemetry spam, emit real migration outcomes
* refactor(vscode): slim migration telemetry fix to minimal surface
* fix(vscode): emit legacy migration outcome only after seeded session start settles
The completed event fired at in-memory conversion time, before the
seeded session start persisted the migration, so a start/persistence
failure was misreported as a successful migration and never produced an
error outcome. Conversion now records a pending migration; the followup
and compaction coordinators settle it after the session start resolves
(completed) or rejects (error/session_start_failed).
* fix(vscode): surface seeded-persistence failures in migration outcomes
LocalRuntimeHost.startSession deliberately swallows seeded-message
persistence failures (the in-memory session still works), so a resolved
start was not proof the legacy conversion became durable. The start
result now reports seededMessagesPersistence, and the resume/compaction
coordinators settle the migration from that result: completed only when
the seed write succeeded, error/seed_persistence_failed when the start
resolved but the write failed, error/session_start_failed when the
start rejected. durationMs now spans conversion through settlement.
Adds the core boundary test forcing persistSessionMessages to fail and
asserting the start still resolves with the failure visible on the
result, plus coordinator tests for both failure modes.
* refactor(vscode): drop per-task migration outcome events, keep volume fixes only
Scope the PR down to the zero-behavioral-risk telemetry fixes, per
review: keep the backlog event transition gating and the one-line
migratedSdkTaskCount fix (counting resumed legacy sessions via their
legacyTask metadata), and revert the per-task terminal outcome
plumbing (pending-migration settlement, coordinator hooks, and the
core StartSessionResult.seededMessagesPersistence field) along with
the success->completed outcome rename on the now-uncalled
captureLegacyTaskMigration. The per-task outcome events can land
separately on the observable persistence boundary.
* fix(llms): reject truncated tool-call JSON with unterminated strings
* fix(llms): scope truncation guard to jsonrepair only and handle single quotes
* fix(shared): keep jsonrepair ahead of bare-object repair for typed literals
Restores main's precedence for inputs both strategies can handle:
{"flag": True} must repair to a typed true, not the string "True".
The truncation guard now gates only the jsonrepair step, which is the
only strategy that can invent a string terminator.
---------
Co-authored-by: Mikołaj Kondratek <19799111+mkondratek@users.noreply.github.com>
A message whose content array held only empty text parts slipped past the
existing empty-content guards in formatMessagesForAiSdk (which cover
content: "" and content: []). The AI SDK then strips empty text parts,
producing {"role":"user","content":[]} on the wire, which strict
providers reject — seen in prod as Vercel 400s for kimi-k3:
"user message must have content".
* fix(telemetry): emit disjoint per-request token buckets in task.tokens
SDK usage events follow the AI SDK convention where inputTokens is the
full request input including cache reads/writes. task.tokens forwarded
that value as tokensIn while also reporting cacheReadTokens and
cacheWriteTokens, so every event re-counted the whole (mostly cached)
conversation context and per-task token sums inflated ~5x on
cache-heavy sessions relative to the legacy contract (tokensIn =
uncached input only, disjoint buckets).
task.tokens now subtracts the cache buckets from tokensIn at the
capture site (mirroring the webview's normalizeUsageEvent), defaults
the cache buckets to 0 instead of undefined, and stamps the provider
attribute for parity with the legacy event schema. Event and attribute
names are unchanged.
* fix(core): normalize registered ApiHandler usage to cache-inclusive inputTokens
Review follow-up: two producer contracts shared AgentUsage.inputTokens.
Native AI SDK usage reports the full cache-inclusive prompt size, but the
ApiHandler adapter forwarded classic disjoint chunk.inputTokens unchanged,
so the task.tokens cache subtraction would zero out real uncached input
for a cache-reporting registered handler.
Normalize at the adapter boundary (inputTokens + cacheReadTokens +
cacheWriteTokens) so every producer entering AgentUsage satisfies the
same cache-inclusive invariant, document that invariant on
AgentTokenUsage.inputTokens, and reframe the telemetry clamp as a
defensive guard rather than a supported producer shape. Adds an adapter
normalization test and a boundary test from an ApiStreamUsageChunk
through task.tokens asserting the disjoint buckets round-trip.
* Revert "fix(core): normalize registered ApiHandler usage to cache-inclusive inputTokens"
This reverts commit 9a9aff374a.
* fix(telemetry): report involuntary Cline logouts from the SDK auth service
The SDK auth service cleared credentials silently when a refresh token was
rejected (invalid grant), both mid-session and during startup restore, so
user.auth_logged_out never captured involuntary logouts on the next bundle.
Emit token_invalid at both credential-clearing sites and restore_error when
startup restore throws, matching the reason vocabulary the legacy bundle now
uses so the same warehouse query measures involuntary logouts across rollout
variants. Startup with no stored session still emits nothing.
* refactor(telemetry): trim logout-reason parity change to the minimum
* fix(telemetry): report Cline invalid-grant logouts as token_invalid in the SDK resolver
getValidClineCredentials is the single owner of the involuntary-logout
event for the Cline provider; normalize its reason to the legacy
extension's LogoutReason vocabulary (token_invalid) so warehouse queries
cover both bundles. The raw OAuth code stays in errorCode. Codex/OCA
paths keep emitting invalid_grant and are unaffected.
* fix(telemetry): let the SDK resolver own token_invalid; keep restore_error for real restore failures
Address review on the SDK-adapter half of the logout-reason split:
- drop both adapter-side token_invalid emissions - the SDK resolver
already emits user.auth_logged_out on the same telemetry instance, so
the adapter was double-counting the exact signal being measured
- transient failures refreshing the stored session on startup (resolver
throws: network/timeout/5xx) no longer book as restore_error; stored
credentials are kept and the SDK books auth_refresh_soft_failure, so
an offline startup is not a logout
- single-source LogoutReason in services/auth/types.ts and re-export it
from the SDK auth service instead of maintaining two parallel enums
- boundary test runs the real getValidClineCredentials and asserts
exactly one auth_logged_out (reason=token_invalid) total, so a
reintroduced adapter emission fails the suite
* feat(hub): prompt update and restart when another install replaces the shared Hub
* feat(hub): make managed Hub build-watch interval configurable via CLINE_HUB_BUILD_WATCH_INTERVAL_MS
* feat(hub): reuse newer managed Hub builds instead of retiring them
Embed a build epoch alongside the deterministic runtime fingerprint so
managed-Hub compatibility can order builds in time. When fingerprints
differ, a Hub produced after the client's own build is attached over the
compatible wire protocol (and the build-mismatch watcher prompts the user
to update) instead of being retired, so concurrent installations converge
on the newest build rather than replacing each other's daemons. Older,
unordered, or metadata-less Hubs are retired and replaced as before.
* refactor(hub): simplify mismatch status derivation and dedupe sidecar event encoding
* fix(cli): only watch for managed Hub build mismatches in hub-attached sessions
Yolo and sandbox sessions force the local backend and never attach to the
shared managed Hub, so a newer Hub owned by another installation must not
interrupt them with the blocking update dialog.
* fix(desktop): stage an app update before hub-mismatch restart
'Update and restart' previously invoked restart_to_apply_update directly,
which only relaunches the current bundle. With no update staged by the
background 2h updater loop, the app came back on the same version, hit the
same newer Hub, and re-prompted immediately.
Add a check_for_update_now Tauri command that runs one updater
check/download/stage cycle on demand and reports the resulting status. The
dialog now stages the update first and restarts only when the updater
reports 'ready'; otherwise it stays open and explains that no update is
downloadable yet (or that the check failed) instead of restarting into the
same version. Addresses the outstanding Greptile P1 on the dialog.
* fix(desktop): reset the no-update hint when a new hub mismatch arrives
Without this, a dialog for a fresh mismatch reopened pre-set to 'Try again'
with the previous prompt's stale hint.
* fix(desktop): serialize updater cycles so overlapping checks cannot clobber a staged update
The periodic update loop and the on-demand check_for_update_now command
run the same check/download/stage cycle against shared state. Without
exclusion, two overlapping cycles could download the same bundle
concurrently, and the later one could overwrite a freshly staged "ready"
status with "idle" or "error" decided from its stale pre-await
ready_version snapshot - making the update dialog deny that a staged
update exists. A tokio::sync::Mutex now serializes whole cycles; the
ready_version snapshot is read under the lock, so it stays authoritative
for the cycle that took it.
* fix(core): harden Hub daemon lifecycle
* fix(core): wait for Hub listener before replacement
* fix(core): recover after Hub cleanup errors
* fix(core): assign the close memo handle before socket termination re-enters beginClose
On every shutdown with a connected client, the daemon logged
'unhandledRejection: AggregateError: hub server close failed' and exited
with code 1 instead of 0. Root cause: beginClose() terminated the
tracked WebSockets before assigning closeHandle. terminate() fires close
events whose microtask continuations advance the daemon coordinator's
deferred cleanup into server.beginClose() while the first invocation is
still mid-body, so the memo guard passes twice and a second set of
wss.close()/server.close() calls runs against the already-closing server,
rejecting with 'Server is not running' and spuriously failing the close
aggregate. The rejection then rode the daemon's unhandledRejection
fatal path and escalated the exit code.
Construct the close promises and assign the memo handle first, and only
then terminate sockets and run detach handlers; a re-entrant call now
hits the memo guard. Also observe the /shutdown handler's
fire-and-forget closeServer() so a genuine close failure is reported
solely by the owner's own await on the same memoized promise instead of
the unhandledRejection path.
Verified with a real daemon: shutdown with zero clients stays graceful
(exit 0, ~26ms); shutdown with a held-open authenticated WebSocket now
exits 0 with no unhandled rejection, still bounded by the 2s coordinator
deadline for the genuine Bun listener-close stall, with the discovery
record cleaned. Hub suites (228) and the shutdown e2e (5) pass.
---------
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
* fix(llms): update AI SDK deps to fix streamed tool calls with non-zero indexes
LiteLLM's Anthropic passthrough emits chat-completions tool_call deltas
whose index mirrors the Anthropic content-block index (1 when a text
block precedes the tool call; see BerriAI/litellm#11580).
@ai-sdk/provider-utils 5.0.18 stored streamed tool calls in a sparse
array keyed by that index and crashed at stream flush with
"Cannot read properties of undefined (reading 'hasFinished')",
aborting the agent turn. Upstream fixed this in provider-utils 5.0.21
("Fix streamed tool calls with non-zero, non-contiguous, reused, or
missing indexes.").
Update the ai / @ai-sdk packages so every chat-completions streaming
path resolves @ai-sdk/provider-utils 5.0.25, and drop the root
">=4.0.0" override on @ai-sdk/provider-utils: with intersect semantics
it pinned the workspace to the already-locked 5.0.18 even after parents
began requiring 5.0.25, and it force-upgraded dify-ai-provider two
majors past its declared ^3 range. Each package now resolves the
version line it declares.
Fixes#13119
* test(llms): pin non-zero streamed tool_call index regression (#13119)
Wire-level regression test: an openai-compatible SSE stream whose only
tool_call delta carries index 1 (Anthropic content-block numbering via
LiteLLM) must complete and emit the tool-call part instead of throwing
at flush.
* fix: address review findings from merge-conflict resolution
- Restore apps/vscode/proto/cline/state.proto to main's version: the
merge commit's pre-commit hook regenerated it with a stale generator,
deleting auto_approve_all_toggled = 174 and moving a reserved line,
creating drift against the checked-in descriptor. The deletion was
never intended.
- Restore FeatureSettingsSection.tsx to main's version (the same hook
reformatted main's file during the merge).
- Regenerate bun.lock narrowly from main's lockfile without --force so
the diff contains only the @ai-sdk family and its direct transitives;
drop the spurious webview-ui-scoped @radix-ui duplicate entries the
previous install introduced (hoisted resolutions still satisfy
webview-ui's unchanged ranges; verified with --frozen-lockfile).
- Align @ai-sdk/provider to ^4.0.7 in @cline/llms to match the rest of
the AI SDK family and avoid parallel provider resolutions.
Revalidated: wire repro streams to finishReason=tool-calls, @cline/llms
suite passes incl. the index-1 regression test, all workspaces
typecheck, SDK builds clean.
* fix: restore FeatureSettingsSection.tsx to main's formatting
The branch's pre-commit biome hook (--semicolons=as-needed, --write
--staged) strips a blank line from this file whenever it is staged,
which is how the unintended diff appeared in the merge commit. Commit
with --no-verify to keep the file byte-identical to main; this PR does
not touch the VS Code webview.
* refactor(desktop): extract chat transcript logic to messages/
Verbatim moves out of chat-messages.tsx (2,415 -> ~1,300 lines), with no behavior changes.
- Extract shared constants, grouping and reasoning helpers, tool summaries, and tool icons into messages/.
- Add unit tests for the extracted pure logic.
- Update test:chat-ui to include tests under messages/.
chat-messages.test.tsx remains unchanged and continues to pass.
* refactor(desktop): extract chat message components to messages/
Moves MessageBubble, ReasoningBlock, ToolMessageBlock, ToolApprovalPanel
(+ formatApprovalTimestamp and the ToolApprovalRequestItem type), and the
image lightbox out of chat-messages.tsx into their own modules under
messages/. Memo wrappers, comparators, and prop contracts are unchanged;
chat-messages.tsx keeps only the ChatMessages orchestration (~780 lines).
chat-messages.test.tsx is untouched and still passes.
* refactor(desktop): extract chat transcript logic to messages/
Verbatim moves out of chat-messages.tsx (2,415 -> ~1,300 lines), with no behavior changes.
- Extract shared constants, grouping and reasoning helpers, tool summaries, and tool icons into messages/.
- Add unit tests for the extracted pure logic.
- Update test:chat-ui to include tests under messages/.
chat-messages.test.tsx remains unchanged and continues to pass.
* refactor(desktop): extract chat message components to messages/
Moves MessageBubble, ReasoningBlock, ToolMessageBlock, ToolApprovalPanel
(+ formatApprovalTimestamp and the ToolApprovalRequestItem type), and the
image lightbox out of chat-messages.tsx into their own modules under
messages/. Memo wrappers, comparators, and prop contracts are unchanged;
chat-messages.tsx keeps only the ChatMessages orchestration (~780 lines).
chat-messages.test.tsx is untouched and still passes.
* refactor(desktop): extract chat message components to messages/
Moves MessageBubble, ReasoningBlock, ToolMessageBlock, ToolApprovalPanel
(+ formatApprovalTimestamp and the ToolApprovalRequestItem type), and the
image lightbox out of chat-messages.tsx into their own modules under
messages/. Memo wrappers, comparators, and prop contracts are unchanged;
chat-messages.tsx keeps only the ChatMessages orchestration (~780 lines).
chat-messages.test.tsx is untouched and still passes.
* fix(telemetry): stop mirroring per-token stream deltas into telemetry
Gate assistant-text-delta, assistant-reasoning-delta, and tool-updated
runtime events out of the unconditional telemetry.capture mirror in
AgentRuntime.emit. These fire once per streamed token or tool progress
chunk and accounted for ~97% of all agent.* telemetry volume in the
field with no analytical value. Listeners, hooks.onEvent, and the
run-failed sdk.error reporting are unchanged; the gate is a static
Set lookup so no per-event allocation is added.
* refactor(telemetry): inline stream-delta telemetry gate as a switch
Replace the module-level Set constant with case labels directly at the
capture site; same behavior, less indirection.
* feat(ui): styled label parts, terminal-style commands, and patch fidelity fixes
Label segments: ToolSummary gains labelParts ({text, code?}[]) so
consumers can render code-ish segments (file names, commands, queries,
URLs) in a monospace face. Single commands now read like a terminal
prompt — '$ bun test' — and an untruncated single command no longer
duplicates itself as a detail line.
Review fixes folded in:
- apply_patch preserves hunk boundaries: per-hunk oldText/newText on
ApplyPatchFile and file items (re-diffing concatenated hunks let a
deletion in one hunk pair with an addition in another), plus action
metadata — Delete File labels as 'Deleted x' with no phantom diff,
'*** Move to:' renames display as 'old → new'.
- run_commands accepts every RunCommandsInputUnionSchema shape (single
entry, bare arrays, top-level {command,args}, {cmd}).
- makeUnifiedDiff treats empty text as zero lines, so creating an empty
file or deleting all content no longer reports a phantom +1.
- parseWebFetchInput drops non-string urls instead of stringifying
objects into labels.
- hoisted a double normalizeValue in the unknown-tool fallback.
* refactor(desktop): render each tool call as its own chat row
Drops the consecutive-call grouping ('Read 3 files · Ran 2 commands')
in favor of one row per tool call — each with its own icon, status,
disclosure, and treatment per kind:
- commands read like a terminal: '$ bun run test' in monospace, with
the captured output in a capped scrollable mono block on expand and
'$ '-prefixed detail lines for multi-command calls
- edit rows carry mono filenames, the +/- badge, and their pierre
diffs pre-expanded (one diff per hunk for multi-hunk patches),
keeping the user-toggle override from the grouped implementation
- reads/searches/fetches keep inline specifics with mono code segments
via the shared labelParts
Also fixes the test:chat-ui exit-1 regression flagged in review:
@pierre/diffs' custom element calls CSSStyleSheet.replaceSync, which
jsdom lacks — a prototype polyfill in the suite keeps the real
component in the test tree (and the pre-expand assertions meaningful)
while letting the run exit 0. This suite gates ui-publish.yml.
* fix(desktop): keep the thinking indicator up during quiet turn stretches
The indicator only covered the gap right after a user message, so the
turn looked frozen while the model composed its next step — most
noticeably while streaming tool-call arguments, when neither text nor
a tool row is on screen. It now shows whenever the turn is running and
nothing else is visibly active (no streaming text, no in-progress tool
row, no pending approval/question).
* feat(ui): action-first tool labels
Every row leads with the plain action phrase — 'Ran command',
'Read file', 'Edited file', 'Created file', 'Deleted file' — with the
specifics (command, file name, line range) following as a monospace
segment. The mono segment renders at full size; the previous 0.92em
downscale made it look smaller than the surrounding prose.
* fix(desktop): chat polish — indicator alignment, action spacing, no expanded fade
- The Thinking indicator now mirrors the tool-row trigger metrics
(min-h-7, py-1, gap-2, 16px icon, font-medium, 8px rhythm) so the
text no longer shifts when the indicator swaps with an arriving
tool row.
- The copy/fork/timestamp action row sat 4px up into the message text
above it (-translate-y-1); it now rests 2px below the message block.
- Expanded reasoning/tool panels rendered at 70% opacity with
hover-to-unfade; expanded content is what the user is reading, so it
now renders at full opacity.
* feat(ui): violet active rows, gray finished rows, no green hover
Tool-row colors follow activity: running/pending rows (and the row
spinner) carry the brand violet, finished rows settle into
muted-foreground gray, and hover brightens toward the foreground
instead of hue-shifting to the success green. Errors stay red.
Also: maxInlineChars default raised 60 → 200 so real commands stop
getting truncated (the cap is now only a guard against pathological
payloads; layout handles overflow), and expanded editor rows lead with
the fuller file path above the diff, matching read rows.
* refactor(ui): let layout own label overflow instead of char caps
maxInlineChars now defaults to unlimited — labels carry the full
command/task/question text (whitespace collapsed to one line) and
.cline-chat-tool-label ellipsizes at the container edge via CSS
(nowrap + text-overflow) instead of wrapping. The cap remains as an
opt-in for width-constrained surfaces like TUIs. Since the label can
now be visually cut by layout, single-command rows always carry the
full command in their expanded details.
* fix(ui): drop stale green base color on tool triggers
The redesign moved finished tool rows to muted gray and running rows to
brand violet, but a leftover .cline-chat-tool-trigger { color:
var(--success-text) } rule later in the sheet overrode the gray base, so
every settled row still rendered green.
* fix(desktop): give message actions clear separation from message text
2px below the text read as touching; 6px (translate-y-1.5) gives the
copy/fork/timestamp row visible breathing room.
* feat(ui): spinner replaces the tool icon while a call is in flight
The progress ring used to append to the right of the label, so running
rows sprouted chrome instead of reading as one glyph + label. It now
takes the icon slot and fills the same 1rem box, so the label never
shifts when the icon swaps back in on completion.
* fix(desktop): align thinking indicator with the tool row that replaces it
The indicator sits outside the message column, so it already inherits the
conversation gap; its own mt-2 stacked on top and rendered it 8px lower
than the tool row that swaps in.
* style(desktop): message actions match chat text scale in a lighter gray
Copy/edit/restore/fork icons go from 12-14px to the 16px the rest of the
chat chrome uses, the timestamp moves from 11px to text-sm, and the whole
row renders at 70% muted-foreground so it reads as secondary chrome;
hover still brightens to full foreground.
* style(ui): running tool rows share the thinking indicator's gray
Violet-on-running read as a different system than the muted thinking
state it replaces; the spinner alone now signals activity. The progress
ring draws in currentColor so it stays gray on normal rows and red on
error rows without extra rules.
* style(desktop): nudge message actions down 2px and scale them down a step
Actions row moves from 6px to 8px below the message text; icons go
16px -> 14px and the timestamp text-sm -> text-xs after the previous
bump overshot.
* fix(ui): don't unstick conversation follow when content grows
Stick-to-bottom flipped off whenever a scroll event landed between a
content-height jump (tall diff rows mounting) and the resize observer's
re-pin: the handler read the new distance-from-bottom as the user having
left the bottom. Sticking is now released only by an actual upward
scroll and always restored on reaching the bottom, so the transcript
keeps following while rows stream in.
* feat(ui): user message bubbles on a filled brand-violet surface
The card-colored bubble sat too close to the app background to read at
a glance. New brand-violet-surface tokens (deep enough for near-white
text in both themes) fill the user bubble.
* style(desktop): give the conversation bottom padding above the composer
The last message (and its hover actions hanging below) butted against
the composer border.
* fix(desktop): composer keeps its two-line height when unfocused
Collapsing to one row on blur made the input and the conversation above
it jump on every focus change; the focus-tracking state existed only to
drive that resize.
* refactor(ui): simplify AgentAskQuestion and move it to the brand accent
The 'Follow-up question' heading, intro sentence, and box-in-box nesting
made a one-question prompt read like a form. The question now leads the
card directly (icon + text + option buttons) and the accent shifts from
blue to brand violet, with the section still labelled for assistive
tech.
* fix(desktop): pending questions and approvals render at the end of the transcript
They rendered above the whole conversation like a banner, so a follow-up
question appeared at the top of the chat instead of where the
conversation actually is.
* fix(desktop): keep message actions reachable and make hover/focus feedback instant
The 8px offset under a message was a translated gap — dead space that
dropped the parent's :hover midway to the buttons, hiding them before
they could be clicked. The offset is now padding on the actions element
so the hover chain stays unbroken. Also removes the opacity fade on the
actions row and the composer's focus border transition: both read as lag
rather than polish.
* feat(ui): add shared tool-summary presentation module
Pure, framework-free tool-call presentation logic under
@cline/ui/components/agent-chat/tool-summary: buildToolSummary and
buildGroupedToolLabel turn raw {toolName, input, result} payloads into
rich row labels (file names with line ranges, inline commands, search
queries, URLs), per-item details, +/- diff counts, per-file unified
diffs (editor old/new text and apply_patch envelopes), team_* labels,
and MCP-aware output text extraction. Merges the desktop app's
buildToolSummary layer with the CLI's tool-parsing/diff utilities so
every @cline/ui consumer renders tool rows consistently.
Exports the new subpath from package.json, extends the packed-tarball
smoke test to cover it, documents the boundary change in ADOPTION.md,
and bumps the package to 0.2.0-next.3.
* refactor(desktop): adopt shared tool-summary for chat tool rows
Replaces ~950 lines of app-local tool extraction (buildToolSummary,
teamSummary, parsers, grouped-label logic) in chat-messages.tsx with
the @cline/ui tool-summary module. Desktop tool rows gain single-call
specifics inline (Read app.tsx (10-80), Ran bun test, Edited util.ts
with +/- badge), line ranges on reads, shortened paths with directory
context in expanded details, per-file unified diffs in the expanded
panel, and grouped labels joined with a middot. Detail keys switch to
index-based to fix duplicate-line key collisions. Icons re-key on the
shared ToolKind classification.
* fix(ui): stop fabricating line positions in fragment tool diffs
Editor str_replace payloads carry old_text/new_text as fragments of the
file, but makeUnifiedDiff treated them as whole files and emitted hunk
headers anchored at line 1, mislocating the change in expanded edit
rows (Greptile P1 on #13151). Fragment diffs now use a neutral
'@@ … @@' separator; only whole-file content (editor create,
apply_patch Add File sections) keeps real hunk positions.
File items also expose the raw oldText/newText (reconstructed from
hunks for apply_patch) plus a fragment flag, so rich diff renderers
can consume the texts directly instead of re-parsing unified output.
* feat(ui): render tool-row edit diffs with @pierre/diffs
Adds @cline/ui/components/agent-chat/tool-diff exporting ToolFileDiff,
a thin wrapper over @pierre/diffs (optional peer dependency) that
renders a tool-summary file item as a syntax-highlighted, theme-aware
unified diff. Fragment diffs hide line numbers instead of showing
misleading ones. The desktop chat renders edit diffs through it, and
tool groups containing an edit diff now open pre-expanded so the diff
is immediately visible.
ADOPTION.md reframes the shared-module story: extracting presentation
logic products would otherwise duplicate is the direction @cline/ui is
headed, with tool-summary and tool-diff as the first two modules. The
packed-tarball smoke test covers the new subpath in both consumers.
* fix(ui): blend tool diffs into the app surface
Two polish fixes to ToolFileDiff from design review:
- Normalize trailing newlines on both sides before diffing so tool
payload fragments (which rarely end in a newline) don't litter every
diff with 'No newline at end of file' markers.
- Map @pierre/diffs' background hooks (--diffs-light-bg/--diffs-dark-bg)
to the host app's --background token (stock white/black fallback),
so the diff surface and all its color-mixed tints (context lines,
gutters, separators) derive from the app background instead of
pierre's pure white/black. Overridable via a new background prop.
Storybook's ToolSummaries story now renders file items through
ToolFileDiff pre-expanded (matching the apps) and adds a multi-file
apply_patch fixture.
* fix(desktop): pre-expand tool groups when edit diffs arrive mid-stream
defaultOpen only applies at mount, but a streaming tool group mounts
with its first (often read) call and gains the edit later, so live runs
never saw the promised pre-expanded diff. Drive the disclosure with
controlled state that opens when a file diff first appears, unless the
user has toggled the row themselves. Covers the streaming path with a
rerender test.
* chore(ui): replace font dependencies
* refactor(ui): migrate shared typography tokens
* refactor(ui): adopt Inter and Geist Mono in apps
* fix(hub): preserve variable font weight tokens
* feat(ui): tune font weights for dark mode
* docs(ui): add font migration screenshots
* (chore)ui: misc typography adjustments
* fix(hub): make dark-mode font-weight overrides take effect
Tailwind's @theme inline bakes literal values into utilities, so the
.dark --font-weight-* overrides were dead code and dark mode rendered
the heavier light-mode weights. Declare the weights in :root instead so
font-* utilities keep their var() references, matching the @cline/ui
tokens approach. Also rewrap --font-mono to satisfy biome format.
* fix(ui): restore light-mode semibold to 640 and pin weight scales in test
The PR intent is a 480/560/640/640 light scale with 400/500/600/600
dark overrides, and the Hub already uses 640; tokens.css had drifted to
600 for light semibold. Regenerate scoped-tokens.css and assert both
the light and dark weight scales in the theme contract test.
* chore(desktop): remove stray double space in provider header class
---------
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
Verbatim moves out of chat-messages.tsx (2,415 -> ~1,300 lines), with no behavior changes.
- Extract shared constants, grouping and reasoning helpers, tool summaries, and tool icons into messages/.
- Add unit tests for the extracted pure logic.
- Update test:chat-ui to include tests under messages/.
chat-messages.test.tsx remains unchanged and continues to pass.
getAuthToken captured expiresAt before refreshing, then validated the new
token against that stale value. When the old token was already past expiry
(not just inside the 5-minute buffer), a successful refresh was thrown
away and null returned, so the first call after long idle failed despite
valid credentials. Re-read the expiry from the refreshed auth info.
* fix(hub): don't forward recoverable agent errors to dashboard peers
Recoverable error events are in-run notices, not turn outcomes: the
MistakeTracker emits one for every recorded mistake (e.g. a plan-mode
guard-blocked run_commands call) while the run continues. The hub
dashboard forwarded every error event to peers, so the webview dropped
out of the sending state and appended an error row mid-turn — the same
host bug fixed for VS Code and the CLI in #12953.
Gate the forward on recoverable, matching those hosts: the tool failure
is already shown inline via the failed tool_event, and the turn's
outcome stays decided by how it actually ends (turn_done or a
non-recoverable error). Recoverable errors are logged server-side.
* fix(hub): forward recoverable flag to peers instead of filtering server-side
Per review: the server is a translation layer between agent events and
the webview protocol, so it should not embed display policy or console
logging. Forward every agent error with its recoverable flag on the
peer message and let each peer decide — the webview keeps recoverable
errors out of the transcript and keeps the turn state, matching how the
CLI gates display on the same flag while the information stays
available to any peer that wants it.
* add custom model selection to the vertex provider
* fix race conditions from PR review
* fix linter warnings
* fix test failures
* refactor(vscode): drop Vertex global-endpoint picker filtering
The SDK catalog is live (models.dev), so a static host allowlist of
global-endpoint-capable models lags every model launch and silently hides
new models from users on vertexRegion=global. Remove the allowlist, the
host override that injected supportsGlobalEndpoint, and the picker filter;
show the full catalog for every region.
An unsupported pick now fails loudly at request time: map Vertex's
'model not available in region: global' (and Google's Publisher Model
locations/global not-found body) to recovery guidance in the error row.
Also drop Anthropic's universal pricing from the Vertex Fable 5 overlay —
Vertex bills region-dependently, so the copied price understated recorded
cost; the record now carries no pricing instead of a wrong one.
---------
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
* fix: remove stale Double-Check Completion feature tip
The rotating feature tips still told users to enable "Double-Check
Completion" in settings, but that toggle was removed in the new UI —
the Features section now offers Auto Compact, Feature Tips, Background
Edit, Checkpoints, Worktrees and Hooks. Following the tip sent users
searching the settings panel for something that isn't there.
Drop the tip. The remaining ten were checked against the current UI and
all still hold, including the "Settings → Features → Feature Tips" path.
* chore: remove dead CLI settings e2e page object and orphaned test
`page-objects/settings.ts` asserted the CLI settings Features tab shows
"Double-check completion" — the same removed setting behind the stale
feature tip. Nothing in the live tui-test suite (apps/cli/src/tests)
imported it; only chat.ts and auth.ts page objects are in use.
Its one importer, apps/vscode/tests/e2e/cli/interactive.test.ts, is a
leftover from the pre-2026-06-02 SDK migration squash: all three of its
imports resolve to files that don't exist, there's no tui-test config in
that tree, and no npm script runs it. It cannot execute.
* fix: respect user max output tokens in compaction summarizer requests
The compaction summarizer hardcoded max_tokens to 1024 and the VSCode host
never mirrored the user's Max Output Tokens onto providerConfig, so summary
requests were always capped at 1024 tokens. Reasoning models can spend that
entire budget thinking; the reasoning stream is discarded, so no summary
text arrives and compaction is skipped on every attempt.
- Mirror maxTokensPerTurn onto providerConfig.maxOutputTokens in the VSCode
session factory so consumers that build handlers straight from it (the
compaction summarizer) honor the user's setting, matching the CLI.
- Resolve the summarizer output budget from explicit config, then model
info, then knownModels, before the default; raise the default to 4096.
- Log a diagnostic warning (reasoning chars, incompleteReason, likely
cause) when the summarizer returns no summary text instead of silently
skipping.
* fix: clamp summarizer default output budget by model metadata instead of adopting it
Model maxTokens is reported capability, not a product default: without an
explicit configuration the summarizer now requests the 4096 default, lowered
by model metadata when the model reports less, never raised by it. Explicit
values still win as-is.
* feat(vscode): remove YOLO mode setting, migrate old users to auto-approve all
The SDK extension's YOLO toggle was cosmetic: nothing in the approval
path read it, so runs were silently governed by the per-action
auto-approval settings underneath (cline/cline#13114). Instead of
keeping a parallel override system, remove the setting entirely and
make the auto-approve menu the single source of truth:
- drop yoloModeToggled (and the equally dead autoApproveAllToggled)
from state keys, settings handlers, state posts, telemetry, the
remote-config yoloModeAllowed transform, and the settings protos
(field numbers reserved)
- remove the Yolo Mode toggle from Settings -> Features (the whole
Experimental section, it was the only entry) and the
"Auto-approve: YOLO" AutoApproveBar takeover
- add a v3 storage migration that folds a previously-enabled YOLO /
auto-approve-all toggle into autoApprovalSettings by enabling every
action, so previously-unattended setups keep running unattended;
the dead keys are cleared from the file store
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* refactor(vscode): keep dead yolo keys in place instead of clearing them
Current builds never read the removed keys (the state loader only visits
known keys), so deleting them buys nothing - and the file store is shared
with older builds that still know them, so clearing would flip YOLO off
for a user who downgrades. Same downgrade-safety rule as the v1 export.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* chore(vscode): rename wasUnattended to shouldEnableAllActions in yolo migration
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* chore(vscode): drop dead toggleActModeForYoloMode and stale yoloModeAllowed comment
The method was a legacy-controller carryover nothing called, and it set
the mode without rebuilding the session, which is wrong for the SDK
architecture. The comment cited yoloModeAllowed as a live remote-config
example; it no longer maps to anything.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* chore(vscode): refresh checked-in proto descriptor_set.pb
The tracked descriptor set had not been regenerated since the repo
move and still advertised long-changed schemas (including the removed
yolo_mode_toggled fields) to gRPC reflection clients. Sync it with the
output of bun run protos.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* desktop: paste clipboard images into the composer as attachments (CLIENTS-78)
Pasting a screenshot into the composer did nothing: only drag-and-drop
and the paperclip file picker fed the attachment pipeline. Add an
onPaste handler on the composer textarea that extracts image files from
the clipboard, renames them to timestamped pasted-image-*.png files, and
routes them through the existing onAttachFiles flow. Text pastes are
untouched.
* desktop: only extract clipboard images in formats message serialization supports
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* desktop: context-aware welcome suggestions for non-code folders (CLIENTS-98)
* desktop: treat pending branch discovery as its own state for welcome cards
The welcome-card classifier read the "no-git" sentinel as a confirmed
non-repo, but page.tsx also used that value for the initial state and
while a workspace switch was awaiting branch discovery, so a git repo
could briefly show the plain-folder cards. Branch state is now null
while discovery is pending: the welcome screen shows no cards until the
folder is classified, and chat-mode cards (which never depend on git
state) still show immediately. Other branch consumers keep the string
contract via a "no-git" fallback.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* desktop: carry nullable branch state to all consumers
Propagate the pending-discovery null through ChatInputBar,
WorkspaceSelector, and the welcome workspace controls instead of
coercing to "no-git" at the page boundary, so only display leaves
fall back and the welcome classifier is the single consumer that
distinguishes pending from confirmed non-repo.
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* desktop: never let 'Add project…' fail silently; add manual folder path entry (CLIENTS-73)
- sidecar picker tries zenity then kdialog on Linux and throws a descriptive
error when neither exists, instead of returning null (indistinguishable
from user cancel); picked paths are trimmed of trailing separators
- picker failures now surface as visible error messages in both workspace
selectors, with a manual path-entry fallback (typed absolute or ~ paths
in the search box offer an 'Open folder' action)
- failed workspace switches (invalid/nonexistent paths) show an inline
error instead of silently doing nothing
- validate_workspace_directory expands ~ and returns the resolved path
* desktop: keep workspace menu search/error state through catalog refreshes
The welcome-screen workspace picker reset its search text and error
message whenever onRefreshWorkspaces changed identity, which happens on
every session-history poll. Typing a path or reading an inline error
raced against the timer: the menu would silently wipe mid-interaction.
Hold the refresh callback in a ref so the reset only runs when the menu
actually opens.
* desktop: format welcome-workspace-controls test
* desktop: distinguish picker launch failures from user cancellation
A zenity/kdialog rejection with a non-ENOENT spawn error (EACCES, EMFILE,
ENOMEM) or a crash signal was classified as a user cancel, which skipped
the kdialog fallback and suppressed the inline error - recreating the
silent no-op this branch is meant to eliminate. Only a clean exit code 1
from a dialog that actually opened now counts as cancellation; broken
backends fall through to the next candidate and surface a descriptive
error otherwise. Picker logic moved to sidecar/workspace-picker.ts with
an injectable exec so the classification is unit-tested.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Classify picker launch failures separately from user cancellation
zenity/kdialog failures like EACCES, EMFILE, or ENOMEM were treated as
user cancellation, suppressing the kdialog fallback and the inline
manual-entry error. Only a clean exit code 1 now counts as a cancel;
any other failure falls through to the next backend or throws the
picker-unavailable error. Picker logic moved to sidecar/folder-picker.ts
with an injectable exec so the classification is unit tested.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Revert "Classify picker launch failures separately from user cancellation"
This reverts commit 24d27a004d.
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(core): preserve queued prompts across user-initiated aborts
Pressing stop while prompts were queued silently destroyed them:
abort() called clearAborted(), which emptied the pending prompt queue
with no way to recover the typed input. The prompts vanished from the
UI, were never sent, and left no trace in session artifacts.
Aborting now only stops the in-flight turn. Queued prompts stay in the
queue and drain once the abort settles, matching the drain behavior
that already existed for self-aborted turns (loop detector / mistake
limit). The thrown-abort path (completeAbortedInteractiveTurn) now
schedules the same drain that runTurn schedules for turns resolving
with an aborted finish.
* fix(core): full-stop semantics and abort-window queue edits for surviving queues
Follow-up to the queued-prompt survival change: aborting a user turn keeps
the queue and auto-runs it, but two gaps remained.
1. No full stop: aborting a queue-initiated turn also kept draining, so
every Escape consumed one queued prompt and started a fresh provider
call - a session with queued messages could never be brought to rest.
Aborting a drained turn now discards the remaining queue: the first
Escape skips to your queued follow-ups, a second Escape stops the
queued work too.
2. Queue operations were still rejected while an abort settled: a prompt
typed right after Escape was silently dropped, and queued prompts were
briefly uneditable and undeletable even though they were about to
auto-run. enqueue/update/delete now work during the abort window;
scheduleDrain/drain still wait for the abort to settle.
* test(core): cover abort + host restart + seeded recovery durability
Adds an e2e regression guard for the reported "cancel a turn, lose the
conversation" failure: a cancelled turn, a daemon restart, a
client-side recovery seeded from disk, and a second restart before that
replacement ever runs a turn. Reverting the eager seeded-history
persistence makes the final read come back empty.
Materializing a seeded session at start also left its history row with
no prompt and no title, since there is no first prompt to derive one
from. Seed the title from the inherited transcript using the same
inference listSessionHistory hydration applies, so forks and recoveries
stay identifiable in unhydrated surfaces too.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(core): retitle seeded sessions from their first user prompt
Eagerly-materialized seeded sessions kept the interim transcript-
inferred title forever, a behavior change from pre-eager persistence
where a fork's history row was titled by the first post-fork prompt.
The interim title now only covers the window where no turn has run
(previously those rows were simply absent), and the first user prompt
after the seed backfills the row's prompt and retitles it — unless the
user renamed the session in the meantime, in which case only the prompt
column is backfilled. The resident manifest and session metadata are
updated in step so the end-of-turn usage-metadata merge cannot clobber
the title back through a stale in-memory fallback.
The e2e mock's updateSession now mirrors the real persistence-service
contract (row + manifest file), and the durability e2e covers both the
retitle and the rename guard; removing the retitle call fails the
'now add tests' assertion.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* refactor(core): collapse seeded-session titling to the old mechanism
The interim transcript-derived title, retitle flags, rename comparison,
and resident-manifest syncing existed only to title forks that never
run a turn - a new nicety, not parity. Dropping it collapses the whole
design back to what rows did before eager persistence: the persistence
service derives the title from the prompt when a row gains one, so the
host only needs to backfill the promptless row with the first user
prompt via updateSession. Renames win automatically because the service
preserves an existing title when no explicit title is passed.
Net production change vs main is a single 20-line backfill block in
executeTurn. The e2e mock's updateSession now models the service's
title semantics (explicit title wins, existing title preserved,
untitled rows derive from prompt), and the durability e2e asserts the
raw row stays untitled until first prompt while history hydration
infers a display title from the transcript.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
Pressing stop while prompts were queued silently destroyed them:
abort() called clearAborted(), which emptied the pending prompt queue
with no way to recover the typed input. The prompts vanished from the
UI, were never sent, and left no trace in session artifacts.
Aborting now only stops the in-flight turn. Queued prompts stay in the
queue and drain once the abort settles, matching the drain behavior
that already existed for self-aborted turns (loop detector / mistake
limit). The thrown-abort path (completeAbortedInteractiveTurn) now
schedules the same drain that runTurn schedules for turns resolving
with an aborted finish.
* fix(core): keep a hung MCP server from taking down session creation
A stdio MCP server that never finishes initializing used to hold its
connect open for the full DEFAULT_MCP_CONNECT_TIMEOUT_MS (doubled across
the newline/framed attempts). MCP tool discovery runs on the
session.create critical path, so that wait blew past the 30s hub command
timeout and the CLI tore the whole interactive session down instead of
just skipping the bad server.
- Bound MCP tool loading during session build with a startup budget that
is safely under the hub command timeout. Servers that connect in time
contribute their tools; slower/hung servers are skipped for the session
(their error still surfaces via the MCP manager) instead of failing
session creation. Budget is overridable via CLINE_MCP_STARTUP_BUDGET_MS
for tests.
- Add StdioMcpClient.close() (and optional McpServerClient.close) that
marks the client disposed so an in-flight connect() aborts its retry
loop instead of respawning the framed fallback.
- Dispose the manager by closing clients up front, outside the per-server
operation locks, so a server hung in initialize can no longer stall
teardown for the full connect budget.
Adds regression tests covering both the non-blocking build and prompt
disposal while a client is hung in connect().
* refactor(core): simplify hung-MCP-server fix to a startup budget
Replace the bespoke per-server race/tracking in loadConfiguredMcpTools
with a small withStartupBudget() wrapper around the existing
Promise.allSettled: a server that exceeds the budget becomes a normal
rejection that the existing loop already logs and skips. The connect
budget, MCP settings display (initialize timeout 30s), and the rest of
the loader are left untouched.
The client close()/manager.dispose() cleanup is kept minimal: it is what
lets teardown abort a still-in-flight connect instead of blocking on the
per-server lock (and clears the pending request timer).
* fix(mcp): cap the default initialize budget at 3s to protect session creation
Supersedes the startup-budget approach on this branch with the simple
constant fix.
MCP initialize runs on the session.create critical path, which the hub
caps at 30s, and connect() can spend the budget twice (newline then
Content-Length framing). The 30s default from #13067 meant a server that
never initializes held session.create for up to 60s, so the hub RPC
timed out and the CLI tore the whole session down and exited.
Return to the pre-#13067 shape with a bigger probe: 3s instead of 1.5s.
That still covers the ~2s starters the old probe killed (#13035) and
keeps the worst case at ~6s per server, far under the hub deadline.
Genuinely slow starters (JVM-based servers like Oracle SQLcl) now need
an explicit timeout in cline_mcp_settings.json, which continues to
override the default in either direction.
Tests: update the slow-start regression tests to the new policy (2s
connects by default, 4s connects with a configured timeout), refresh the
displayed initialize-timeout assertions, and add an invariant test that
keeps the doubled default well under HUB_DEFAULT_COMMAND_TIMEOUT_MS so
the budget cannot silently creep past the session deadline again.
* fix(desktop): stop opening a session from replacing the remembered model
The composer's ModelSelector mirrored every provider/model prop change
into the remembered last selection (localStorage), which seeds new
sessions via getInitialChatConfig(). Opening an existing session drives
those props to that session's config, so merely viewing an old session
silently replaced the user's explicitly picked default model.
The remembered selection is now written only from the explicit picker
handlers (provider select and model select). Passive prop changes, such
as opening a session, no longer touch it.
* fix(desktop): re-seed remembered provider/model on chat reset
reset() kept the previous config's provider/model and only cleared the
session ID, so a chat pane that had hydrated a historical session could
carry that session's model into the next chat. Re-seed provider/model
(and apiKey when the provider changes) from the remembered defaults --
the same source a freshly mounted thread uses -- so reset and remount
behave identically.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(desktop): stop treating leftover plugin install dirs as installed
isOfficialPluginInstalled() only checked that the marketplace install
directory existed. A failed or interrupted install can leave that
directory behind with no plugin inside, and the next install attempt
then short-circuited with a fake 'already installed' success: the
marketplace button flipped to Uninstall with no error while nothing
actually worked, and the installed-entries listing kept reporting the
broken entry as installed.
The check now requires a loadable plugin module inside the directory
(via discoverPluginModulePaths) before reporting the entry as
installed, so partial directories fall through to a real install
attempt whose outcome is surfaced to the UI.
* fix(desktop): reclaim leftover partial plugin install dirs with --force
* fix(desktop): canonicalize diff panel paths against the session cwd
Tool calls address the same file inconsistently across a session: one
edit uses a workspace-relative path (journal.txt), a later one the
absolute path (/tmp/ws/journal.txt). mergeToolDiffs keyed entries by the
raw string, so the same file was listed twice in the diff panel with
split +/- counts and inconsistent naming, most visibly after git was
initialized mid-session and the model switched to absolute paths.
Diff paths are now canonicalized against the session cwd before
merging: entries for the same file collapse into one, files inside the
cwd display as workspace-relative paths, and files outside it display
their resolved path. Without a cwd the previous raw-key behavior is
kept.
* style: collapse editorReplaceEvent signature per biome format
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(desktop): collapse dot segments and keep root cwd in diff path keys
* fix(desktop): compare Windows diff path keys case-insensitively
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(vscode): hide View Changes on completion rows until there are changes to show
The button previously always rendered on the latest completion row, faded
and disabled when the count check came back 0 - which covers both 'nothing
changed since your last message' and 'no checkpoint to compare against'
(non-git workspace, repo with no commits, comparison failure). A dead
button with a misleading tooltip in the non-git case is worse than no
button: now the row renders nothing until the host confirms there are
actual changes, and the button is always enabled when shown.
* fix(vscode): reset View Changes state when showViewChanges toggles
Greptile review: a stale positive hasChanges from a previous evaluation
could flash the button before the host confirms the new comparison when
showViewChanges flips false and back true on the same row. Reset to
'still checking' whenever the effect re-runs.
* fix(core): never run a foreign compiled plugin-sandbox bootstrap for a source host
When @cline/core runs from source (e.g. the desktop hub daemon in dev)
with CLINE_WRAPPER_PATH set, resolveBootstrap() picked the compiled
plugin-sandbox-bootstrap.js from a separately installed CLI platform
package (such as a published version sitting in the package-manager
cache) before falling back to the source bootstrap. That bootstrap
resolves modules against the other installation's layout, so every
plugin failed to load with "Cannot find module '@cline/core'" - and
the settings pipeline swallowed the failure, leaving Settings > Tools
showing "No plugin tools found" and plugins showing no contributions
even though the same plugins loaded fine in chat sessions.
Bootstrap selection now prefers, in order: a compiled bootstrap next to
this module (always matches the host build), the source bootstrap when
the host runs from source, and only then wrapper/executable-derived
bootstraps - which remain the path for compiled binaries where
import.meta points inside the bunfs bundle.
* chore(core): restore untouched settings-service formatting
* fix(core): keep session context durable across aborts and hub restarts
Users on slow self-hosted endpoints reported sessions losing their entire
conversation after cancelling a long-running request: the TUI still showed
the transcript, but the next turn greeted them like a brand-new session.
Root cause is a stack of two failures:
1. The hub daemon exits on any unhandled rejection that is not an
AgentRuntimeAbortError, so a floating abort-family rejection from a
cancelled provider stream kills every resident session.
2. When the CLI recovers the missing session it rebuilds from the persisted
messages file - but aborting a turn never flushed the transcript, and
lazy session persistence (SDK 0.0.70) kept seeded history (mode-switch
restarts, forks, previous recoveries) memory-only until the first
completed turn. Recovery then seeds an empty session: silent context wipe.
Fixes:
- completeAbortedInteractiveTurn now flushes the transcript to disk, so an
aborted exchange survives a hub restart.
- Sessions started with initialMessages persist them (and any compaction
sidecar) immediately; brand-new empty sessions stay lazy, so closing an
unused runtime still leaves no empty history entry.
- The hub daemon ignores abort-family unhandled rejections (DOMException
AbortError, Node ABORT_ERR) the same way it already ignores
AgentRuntimeAbortError, instead of exiting with every session resident.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(core): write seeded history atomically with session materialization
Greptile review flagged a residual crash window in the seeded-session
persistence: ensureSessionPersisted created the session row (with an empty
messages file) and only then called persistSessionMessages, so a crash
between the two left a discoverable session whose seeded history was gone.
Close the window by threading initialMessages/systemPrompt through
createRootSessionWithArtifacts: the messages artifact is now written with
the seeded transcript before the session row is committed, so every crash
point leaves either nothing discoverable or complete data. The follow-up
persistSessionMessages call at session start is gone; the seed travels
inside session materialization.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Revert "fix(core): write seeded history atomically with session materialization"
This reverts commit 5a7e0b37f1.
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
Turns drained from the pending-prompt queue resolve their errored
AgentResult inside PendingPromptsController.drain(), which discards it,
and the legacy 'error' agent event had no projection in the hub's
session-event projector — so a failed queued turn never produced any
terminal hub event. Interactive clients (e.g. the desktop app) hung on
'Thinking...' with no error shown.
The projector now publishes run.failed (with the error text and a core
session snapshot) for non-recoverable lead-agent error events, but only
when no RPC-driven turn is awaiting sessionHost.runTurn for that
session — the awaiting run.start handler already publishes the
authoritative terminal event, so this avoids double-reporting a turn
that resolves through both paths.
* fix(core/cli): drain queued prompts after self-aborted turns and surface the stop
When a run ends with finishReason "aborted" without a user abort request
(loop detector hard escalation or the consecutive-mistake safety stop),
runTurn skipped the pending-prompt drain, stranding user-queued messages
forever, and the CLI rendered nothing - the task appeared to silently
stop with queued messages never consumed (#13030).
- core: schedule the drain after every completed turn, including
aborted/error finishes. User-initiated aborts are unaffected because
abortSession() already clears the queue, and drain() stops after one
failed send so an erroring provider cannot spin the queue.
- cli: when a turn comes back aborted without the user having requested
an abort, append a "Task stopped before completion." status entry
instead of ending the turn silently.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(core): hold queued prompts on error finishes instead of consuming them
Addresses the Greptile P1 review on #13061: a drained prompt whose turn
resolved with finishReason "error" returned normally, so the
exception-only requeue path treated the send as successful - the failed
prompt was consumed and draining continued firing the remaining queue
into a failing provider.
- drain() now stops the chain when a drained send resolves with an
error finish. The errored entry itself is not requeued (its turn ran:
the prompt is in the conversation and the error is surfaced), but the
rest of the queue is held.
- runTurn() no longer schedules a drain after "error" finishes (the
skip is removed only for "aborted", which is the #13030 fix).
Held prompts still drain via the existing enqueue/update/delete
triggers or the next successful turn.
- Two new unit tests cover both layers.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* revert(cli): drop the 'Task stopped before completion' status line
Keep the change scoped to the queue-drain fix in @cline/core. The CLI
no longer prints a notice for non-user-initiated aborted finishes;
apps/cli is back to parity with main. When messages are queued, the
drain itself makes the stop visible (the queued message runs); richer
stop-reason surfacing can be a follow-up.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
The litellm builtin spec pinned protocol: "openai-responses", so every
request went to POST {baseUrl}/responses. Self-hosted LiteLLM proxies
commonly implement only /chat/completions, so all prompts failed with
404 Not Found on the SDK path (CLI, and now the Next extension bundle).
Drop the override so litellm inherits the openai-compatible family
default (openai-chat -> /chat/completions), matching every sibling
openai-compatible builtin and the Legacy extension behavior.
Fixes#13003, fixes#10781
* fix(cli): render MCP tool result text instead of escaped JSON in TUI
MCP tools return {content: [{type: "text", text}]} which
extractFullOutputText JSON-stringified, escaping newlines into one giant
line that word-wrapped across the whole terminal and never triggered the
line-based collapse. Extract the text parts with real newlines so the
existing collapse works.
Fixes#13038
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(cli): keep placeholders for non-text blocks in mixed MCP results
Addresses Greptile review on #13066: text-only filtering silently
dropped image/resource/audio blocks from mixed MCP content. Render them
as [type] placeholders instead.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(cli): surface non-text MCP block metadata in TUI output
Extract embedded resource text, and include resource/resource_link URIs
and image/audio mime types in placeholders so expanded mixed MCP
results keep identifying metadata.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
The stdio MCP client gave servers without a configured `timeout` only
1.5 seconds to answer initialize before killing the process, so
slow-starting servers (e.g. Oracle SQLcl's JVM-based `sql -mcp`) could
never load and were silently skipped at session start.
Raise the default connect budget to 30s, in line with the startup
budget other MCP clients allow. A configured `timeout` still overrides
it in either direction, dead commands still fail fast through the spawn
error/exit path, and the newline -> Content-Length framing fallback is
unchanged.
Fixes#13035
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* feat(desktop): route /team prompts through core runtime
Rewrite desktop `/team` commands as structured user command blocks before sending them to the core runtime. Validate task input and respect the globally disabled Teams tool setting.
Remove legacy agent spawn and team enablement flags from session configuration, and add coverage for prompt rewriting and disabled-tool behavior.
* fix(desktop): preserve team tool defaults
* fix(desktop): display queued /team prompts as their slash form
Queued prompts are stored in their runtime form, so a queued /team
command showed its raw <user_command> envelope in the prompt queue chip
and edit textarea. Fold queue items through formatDisplayUserInput for
display; saving an edit re-resolves the slash form through the sidecar,
so the round trip is lossless.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* chore(hub): align builtin tool catalog flags with the desktop sidecar
The desktop sidecar pins enableSpawnAgent/enableAgentTeams when listing
the builtin tool catalog; the hub's parallel listing did not, so the two
would drift if the preset defaults ever change. Pin the same flags in
the hub and cross-reference the two call sites.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* chore(desktop): drop inert enableSpawn/enableTeams config leftovers
buildCoreSessionConfig no longer reads these keys, so remove the dead
schema fields, default-config initializers, and chat-test payload
entries. The chat-session regression test still sends them on purpose
to prove legacy flags cannot override the runtime's tool presets.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(desktop): reject /team when the mode's tool preset disables teams
The /team guard only checked the global disabled-tools setting, but the
runtime resolves tool availability from the mode's preset, so a preset
without team tools (yolo) would still send the model a spawn-a-team
instruction it cannot act on. Resolve the teams catalog entry for the
session's mode and reject /team when it is unavailable, mirroring the
runtime's own availability logic.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(core): surface OAuth authorization for SSE MCP servers on 401
A 401 from an SSE MCP server never persisted authorizationRequired: the
fetch-boundary UnauthorizedError was consumed by EventSource and re-thrown
as a status-less SseError, so the instanceof check routed it to
markConnectionError and hosts never offered the OAuth connect action.
Give the SSE stream request a raw fetch so a 401 fails the connection with
the SDK's typed SseError(401), and recognize 401s across transports with a
single isMcpUnauthorizedError predicate at every detection site.
* style(core): apply biome formatting to MCP oauth changes
Toggling Plan/Act while a turn was streaming or waiting on a tool approval
aborted the turn but left the TurnStateTracker on its last live phase: the
aborted session's done event is fenced off as stale once the rebuild
unsubscribes it, so nothing ever settled the phase. The webview then kept
rendering that phase forever - an eternal Thinking spinner with the input
disabled (aborted while streaming), or dead Approve/Run Command buttons wired
to an approval that clearPending had already denied (aborted while awaiting
approval). Users experienced this as 'switched to act mode and nothing
happened / it never wrote the files'.
Mirror cancelTask: after aborting the turn for the mode change, append a
resume_task ask row and set the phase to resumable, so the footer offers
Resume Task with the input enabled in the new mode.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(llms): retry mid-stream network interruptions before any model output
* fix(llms): scale network retry backoff by network retry count, not shared attempt number
* desktop: native-feel polish and render-path performance fixes
- Suppress the WebView browser context menu on app chrome (keep it for
editable fields and active text selections)
- Make UI chrome unselectable app-wide; opt chat messages, markdown,
code, diffs, and error banners back into text selection
- Contain overscroll so inner scrollers don't rubber-band the window
- Lazy-load Settings/Sessions/Onboarding/Diff views out of the entry chunk
- Memoize ChatInputBar and AgentHeader; stabilize their props in the chat
pane so stream flushes only re-render the affected message bubble
- Stop refocusing the composer textarea on every keystroke (caret flicker)
- Cache slash commands across menu opens (stale-while-revalidate)
- Avoid rebuilding reversed message arrays and ask-question JSX per render
- Drop core info/debug console logging on the streaming hot path behind a
cline:debug-logs opt-in; remove leftover [webview:delete] debug logs
- SearchCombobox (provider/model picker): Escape closes and restores focus
- Remove unused @vercel/analytics, recharts, embla-carousel deps and the
unused chart/carousel UI components
* desktop: surface failed-turn errors instead of leaving the chat blank
On a failed run the runtime reports its error string in result.text.
The webview rendered that as an assistant bubble, which the canonical
history rehydration then wiped (the failed turn is never persisted),
so provider errors like a retired model id left the user staring at a
silently empty chat. Route failed-turn text to a persistent error-role
message added after rehydration instead.
* desktop: fade the welcome/conversation swap instead of hard-cutting
Sending the first message replaced the hero layout with the message
grid in a single commit, which read as a white flash. A 180ms enter
animation now plays when either side becomes visible; disabled under
prefers-reduced-motion.
* desktop: render new-chat panes instantly from the last catalog load
Clicking + remounts ChatThreadPane, which refused to render until the
provider catalog (a large fetch) and workspace list resolved again —
about a second of blank pane plus boot spinner on every new chat.
Seed remounts from a module-level snapshot of the last successful
load; the mount effect still refreshes both in the background.
* desktop: invalidate the provider-catalog snapshot with the cache
Seeding remounted chat panes from the last catalog load left a window
where a pane created right after a credential change could act on the
old keys. The snapshot now lives in the catalog module and is dropped
by invalidateProviderCatalogCache(), so credential edits force the
next remount to wait for fresh data.
* fix(cli): harden tool input/output formatters against malformed payloads
Tool inputs cross the model/tool boundary and may not match their
TypeScript annotations (e.g. run_commands with { command: null }).
truncate() called str.replace() on such values, crashing the TUI with
'.replace is not a function' and making persisted sessions containing
the payload non-resumable, since hydration replays the same input
through formatToolInput().
Normalize untrusted values at the formatting boundary: truncate() now
accepts unknown and safely stringifies null/undefined/objects (including
circular structures and throwing toJSON), formatStructuredCommand no
longer returns non-string commands verbatim, and fetch_web_content
request summaries tolerate malformed entries.
Fixes#13036
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(cli): keep valid empty-string args in structured command summaries
Greptile review: filtering normalized args by truthiness also dropped
genuine empty-string argv entries, so summaries could show a different
argument list than the one executed. Filter only nullish entries before
normalization instead.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(vscode): fall back to session cwd/Desktop for @-mention search in empty windows
* fix(vscode): use the shared chat workspace as the no-folder fallback root
ensureGitRepository cached a negative probe for the lifetime of the hook
instance, so a session started in a non-git folder never got checkpoints
even after the user ran git init. Cache only the positive answer and
re-probe otherwise; the probe runs at most once per user turn.
* fix(desktop): treat signed-out state as a typed result instead of a command error
* fix(desktop): sign out when the organization balance fetch reports the typed signed-out result
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* chore: retrigger checks after runner outage
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* refactor(ui): introduce Cline-owned semantic color system
* refactor(desktop): adopt shared semantic theme roles
* refactor(ui): set 15px root and recalibrate xs/sm type scale
Scale rem steps so xs/sm stay 12/13px visually, and slightly lift dark-mode neutral-4.
* refactor(ui): align SearchCombobox with package type and hover tokens
Use host-safe cline-ui utilities and keep option font inheritance from CSS.
* fix(ui): use standard stroke-2 utility on approval spinner
* refactor(desktop): modernize shared UI primitives for Tailwind v4
Replace legacy arbitrary/has selectors with current utility syntax.
* refactor(desktop): bump chat chrome typography to text-sm
Keep composer controls and pickers on the shared sm type step.
* refactor(desktop): use max-w-344 for page frame content width
* chore(desktop): disable Next.js dev indicators
* chore: ignore desktop-app Cursor settings
* docs(pr): add before/after screenshots for #12941
* chore: retrigger checks
---------
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
* desktop: fix silent turn failures, message duplication, and stuck composer; add first-run setup guidance
Findings from two full computer-use UX audits of the desktop app:
- Surface failed turns in the transcript: queued turns (incl. the first
prompt of a fresh session) only signal errors via chat_done, which the
UI previously ignored - sending a message with no credentials failed
in complete silence. Failed turns now show an error message enriched
with the latest core error log and a pointer to Settings -> Models.
- Fix duplicated user messages: a live send's optimistic user message
was materialized a second time by the runtime's queued-prompt-start
event.
- Fix composer stuck on 'Agent is working...': drop prompts from the
local queue snapshot when they start, emit a fresh queue snapshot from
the sidecar on pending_prompt_submitted, and double-check the server
queue on turn completion.
- Add a 'Connect a model' notice on the welcome screen when no provider
has credentials, with actions to reopen onboarding at the connect step
or jump to model settings; it reacts live to credential changes.
- Add 'Get an API key' links for popular providers in onboarding and
Settings -> Models (the catalog docUrl is never populated), and link
the Cline dashboard from the Cline API key form.
- Explain what Cline is on the onboarding welcome step.
- Make the stop button visible (was 8px with no padding) and support
Esc to stop; add Cmd/Ctrl+N (new session) and Cmd/Ctrl+, (settings).
- Remove leftover [webview:delete] console.error debug logging that
surfaced an error badge after deleting a session.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* desktop: remove remaining delete debug logging in session history hook
The sidebar right-click delete path had the same leftover [webview:delete]
console.error instrumentation, which made the Next dev-mode issues badge
appear after every deletion.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* desktop: fix Biome a11y error in WelcomeSetupNotice
biome's lint/a11y/useSemanticElements errors on role="status" divs;
use the semantic <output> element (implicit status role) instead. This
was failing the repo's 'bun run lint'.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* desktop: count structured-config and keyless providers as connected
The welcome setup notice previously only recognized apiKey/OAuth
credentials, so users running Bedrock/Vertex (structured configValues)
or a deliberately enabled keyless local endpoint (e.g. Ollama) were
nagged to connect a model they already use. isProviderConnected now
also counts an enabled provider whose required config fields are all
filled, or an enabled provider that has no API-key field at all.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* desktop: keep re-key eligible when chat_done lands in the same batch as its prompt start
When a turn fails fast, chat_queued_prompt_start and chat_done can be
dispatched in one React batch. Clearing the outstanding-optimistic-
bubble registry synchronously in the chat_done handler ran before the
re-key updater enqueued by the prompt-start event, so the optimistic
bubble was appended a second time instead of re-keyed. Clear the
registry inside a state updater so it executes in event order after
the re-key. Caught by the queued-turn-failure regression test.
* desktop: make the queued-prompt re-key updater idempotent under StrictMode
React StrictMode double-invokes state updaters in dev. The
chat_queued_prompt_start re-key updater consumed the optimistic
bubble's id from outstandingOptimisticUserIdsRef on its first run, so
the second run against the same prev found no eligible candidate and
appended the same user message a second time (and, without a promptId,
makeId() minted a different id per invocation). Hoist the message id
out of the updater and remember which optimistic bubble each queued
message id re-keyed so a re-run reaches the identical result. The memo
resets alongside the outstanding set (error state, reset, hydration).
Root-caused with runtime instrumentation: the duplicate only appeared
on turns that exercised the queue-drain re-key path, and hydration
later collapsed it to one message because the duplicate never existed
in persisted state.
* desktop: preserve failure messages across post-send canonical hydration
Persisted history never contains UI-only error bubbles, so the two
post-send read_session_messages replacements in sendPrompt wiped the
failure explanation appended from chat_done ~40ms after it rendered
(confirmed with runtime instrumentation). Re-append the active
session's error messages after the canonical history. Includes a
regression test reproducing the chat_done-error-then-RPC-resolution
race.
* desktop: don't let an optional API-key field veto a connected provider
Greptile P1 follow-up: Bedrock's catalog entry carries an optional
apiKey field ('Optional Bedrock bearer token') alongside IAM/profile
authentication, and keyless local endpoints can also surface one — so
treating the mere presence of an apiKey field as proof of disconnection
kept nagging configured users. An enabled provider (the user
deliberately persisted settings for it) now counts as connected unless
a required config field is unmet; auth may legitimately live outside
the catalog (IAM, env vars, local endpoints). Brand-new users have no
enabled providers, so the first-run notice still shows for them.
* desktop: tighten credential-error guidance and stop re-pinning stale failure bubbles
* desktop: invalidate the shared provider catalog after settings OAuth login
Greptile P1 follow-up: runOAuthProviderLogin only updated the settings
view's local provider state, so the shared catalog cache and its
invalidation subscribers (the composer selector and the welcome
screen's 'Connect a model' notice) kept reporting the provider as
disconnected until an unrelated invalidation or a pane remount. Notify
the shared cache on successful OAuth login, like the account view and
the API-key save path already do.
* desktop: clear the remembered core error on turn end, reset, and hydration
Greptile flagged that turn-start events are the only thing clearing
lastCoreErrorBySessionRef, and websocket events are not replayed: a
transport interruption that drops a turn's start event lets a later
detail-less failure resurrect an earlier turn's error. The remembered
error belongs to exactly one turn, so clear it whenever a turn ends
(chat_done, any reason) as well as on reset() and history hydration.
Regression test covers the dropped-start-event sequence.
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* chore(llms): regenerate model catalog from models.dev
* feat(llms): surface meta/muse-spark-1.2-contributor for the Cline provider
* test(llms): guard Vercel-only Cline model allowlist
* feat(vscode): explain when a free model promotion ends
Once a free promotion ends, the cline-free/ model is removed from the
catalog and the backend answers 'model not found' to requests against it.
The CLI has shown a dedicated 'Free model promotion ended' banner for this
since #12593; the extension instead rewrote the answer into generic
model-not-found guidance with no model-picker offramp.
Detect the case in the host where the active model id is known
(reshapeErrorForWebview, fed by a new MessageTranslatorState model-id
source), stamp the payload with a cline_free_promotion_ended code, and
render a dedicated card in the webview with a button into the model
picker. Classification is gated on the cline-free/ prefix so ordinary
model-not-found errors keep their generic path, and it runs before the
auth branch since the 404 status falls inside the generic auth range.
* fix(vscode): prefer the live task model over session-start metadata
A mid-task model-only switch updates the running session's model in place
(updateActiveSessionModel) and refreshes the task API shim, but never
touches the session's startConfig/manifest. Preferring the session-start
snapshot could therefore misclassify after such a switch: a genuine
retired-model 404 would miss the promotion-ended card, and the reverse
switch could show it for the wrong model. Provider switches restart the
session, so both sources agree there; the shim starts as "unknown"
(filtered out), so fresh sessions still resolve through start metadata.
* fix(desktop): dedupe chat_queued_prompt_start emitted for the same prompt
PendingPromptService.drain() emits a pending_prompts snapshot (head
removed) and a pending_prompt_submitted event back-to-back for the same
prompt. The sidecar translated both into chat_queued_prompt_start, so
the webview rendered the user's message twice until the chat was
re-hydrated from history. Track the last announced prompt id per live
session and emit the start chunk once.
* fix(desktop): re-key optimistic user bubble when the runtime queues the prompt
The send path renders an optimistic user bubble for prompts dispatched
while the session is idle, keyed by a random id. When the runtime
routes that prompt through its pending queue (e.g. during session
startup), the queued-prompt-start event appended a second bubble under
queued_user_<promptId> — the same message rendered twice until the
chat was re-hydrated from history. Re-key the trailing optimistic
bubble to the event's id instead of appending.
* fix(desktop): re-key only outstanding optimistic bubbles on queued prompt start
Review follow-up: matching by content alone could swallow a new queued
prompt that repeats the text of a message left at the transcript tail
by an earlier cancelled/failed turn. Track in-flight optimistic bubble
ids explicitly (registered on optimistic append; cleared on re-key,
turn end, error, and history hydration) and only re-key those.
* fix(core): don't count plan-mode guard-blocked commands as model mistakes
The plan-mode command guard (#12906) rejects file-editing run_commands
calls with a tool error. The orchestrator counted that error as a failed
tool call, so a turn whose only tool call was guard-blocked fed the
MistakeTracker, which emits a recoverable "error" AgentEvent
("1 tool call(s) failed: [run_commands] ...").
Hosts render that event as a failed turn. In the VS Code extension the
turn ended in the "error" phase (Retry / Start New Task footer), the
final plan text was never retagged to plan_completion_result, and
toggling to Act therefore rebuilt the session without the auto-continue
send - the toggle appeared to do nothing and the presented plan was
never acted on. In the CLI TUI the same event flipped the footer to
idle mid-turn.
A guard rejection is deliberate session policy, not a model mistake:
the run continues and the model is expected to fold the change into
its plan. Tag the guard error with a stable marker sentence, expose
isPlanModeBlockedCommandError, and skip the failed-tool bookkeeping for
matching results so no mistake is recorded and no error event is
emitted. Repeated blocked attempts are still bounded by loop detection
and maxIterations.
* docs(core): flag plan-mode guard error string matching for typed skip channel
FIXME on isPlanModeBlockedCommandError: recognizing guard rejections by
sniffing the error text is brittle. The intended replacement is a typed
skipSource/skipCode on the tool-finished runtime event so the
orchestrator (and the VS Code approval-denial suppression) can identify
skipped tools structurally instead of via string matching.
* Revert core mistake-counting change for plan-guard blocks
A model attempting a file-editing command in plan mode is disobeying
its instructions - that IS a model mistake, and the MistakeTracker
should keep counting it (it is the brake that stops weak models from
flailing at blocked commands indefinitely). The real bug is host-side:
a recoverable mid-turn mistake must not kill a turn that afterwards
completes with a presented plan. The follow-up commit fixes that in
the hosts instead.
* fix(vscode,cli): treat recoverable agent errors as in-run notices, not turn outcomes
The MistakeTracker emits a recoverable error event for every recorded
mistake while the run continues - e.g. a plan-mode guard-blocked
run_commands call as the turn's only tool call. Both hosts treated any
error event as terminal:
- The VS Code translator cleared the pending completion retag, set
errorSeen (turn phase "error": Retry / Start New Task footer), marked
the turn complete, and rendered the error recovery UI. A plan turn
that recovered from the mistake and completed cleanly therefore never
produced plan_completion_result, so togglePlanActMode's planPresented
check failed and switching to act mode rebuilt the session without
the auto-continue send - the toggle appeared to do nothing.
- The CLI TUI flipped isRunning/isStreaming to idle mid-turn, so the
footer lied about the still-running turn.
Recoverable errors are informational: the turn's outcome is decided by
how it actually ends (done/error). VS Code now logs them and keeps them
out of the chat (the tool failure is already shown inline on its tool
row, and provider-failure telemetry already ignores recoverable events
for the same reason); the CLI keeps its running state and surfaces them
only in verbose mode, as it already did for display. Genuine run
failures carry recoverable: false and keep the existing error UI.
Anthropic (and several providers OpenRouter fans out to) rejects any tool
whose input_schema has oneOf, allOf, or anyOf at the top level, failing the
whole request with:
tools.N.custom.input_schema: input_schema does not support oneOf, allOf,
or anyOf at the top level
MCP servers commonly advertise tools whose input schema is a union of object
shapes (e.g. generated from a Zod union), so one such tool bricked every
turn of the session. Merge union branch properties into a single object
schema at the provider boundary; tools still validate their real input
shapes in execute().
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Add plan-mode command blocklist to run_commands
Plan mode kept run_commands available (needed for read-only
investigation) but relied on prompting alone to prevent file edits,
and weaker models routinely ignore that. Add a hard guard in
@cline/core's createShellTool that inspects each command before
execution and rejects file-editing constructs with a plan-mode tool
error instead of running them.
The guard is a quote/heredoc-aware scan that blocks file-manipulation
commands (rm/mv/cp/tee/touch/...), in-place editors (sed -i, perl -i,
gawk -i inplace, sort -o), output redirection to files (allowing /dev
sinks and /tmp for the documented output-capture pattern), mutating
git subcommands, package-manager installs, find -delete/-exec, and
nested command strings (sh -c, eval, sudo, xargs, ...). Windows and
PowerShell equivalents are covered too.
Enabled via a new blockFileEditingCommands flag on DefaultToolsConfig,
set by the plan tool preset (CLI and core runtime) and plumbed through
the VS Code extension's custom run_commands tool from the session mode.
The tool description and PLAN_MODE_INSTRUCTIONS now state the hard
block so models are forewarned.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Simplify plan-mode command guard to a plain blacklist
Replace the char-by-char shell tokenizer (heredoc queues, process
substitution, recursion into sh -c/eval, find -exec analysis) with a
simple scan: mask quoted text/heredoc bodies/escapes/comments so they
cannot false-positive, split on shell separators, and compare the
leading command word of each part against flat blacklists (commands,
mutating subcommands, in-place edit flags), plus one redirect check.
Quoted nested commands (bash -c 'rm x') are a documented false
negative. Also drop the guard from the package's public exports; it
is internal to createShellTool.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Move plan-mode command guard into a built-in beforeTool hook
Review feedback (abeatrix): command blocking is session policy, not
shell-executor configuration. Replace the blockFileEditingCommands
flag threaded through preset -> tool config -> VS Code host with a
core extension registered by the runtime builder for plan-mode
sessions. The beforeTool hook intercepts every run_commands tool in
the runtime - the SDK builtin, host replacements like the VS Code
terminal tool, and delegated sub-agents - and rejects file-editing
calls with the plan-mode error before tool policy and user approval,
so users are no longer prompted to approve a command that would only
fail. All VS Code wiring for the guard is removed.
Also adds block telemetry (sdk.plan_mode_command_blocked with the
blocked construct, never raw command content), per review.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Address review feedback on the plan-mode command blacklist
False positives (mkondratek):
- perl -Ilib / uppercase value-taking flags no longer match the
in-place check; the flag cluster must end at a lowercase i
(sed -Ei still blocked)
- awk inplace detection is tied to the -i/--include flag instead of
matching the substring anywhere (filenames like inplace-notes.txt
no longer trip it)
- read-only git forms allowed: stash list/show, worktree list,
submodule status/summary, and any git subcommand with --help/-h
- arithmetic expansion (1) is masked before the redirect scan
Hardening and coverage (mkondratek, dominiccooney):
- temp-path redirect allowance rejects .. traversal (/tmp/../...)
- Windows gets a temp escape hatch: %TEMP%/%TMP%/$env:TEMP redirect
targets are allowed and the block error mentions it
- curl -o/-O/--output/--remote-name and wget downloads blocked
(--spider and -qO- stdout forms stay allowed)
- python -m pip resolves to the pip subcommand check
- unambiguous PowerShell aliases (mi, ri, cpi, rni, ac, clc) plus a
case-variant test
- more package managers: winget, nuget, gem, composer, dotnet add,
go install/get; bare classic yarn blocked again
- find -exec/-execdir/-ok chains and xargs -I {} placeholders are
checked for mutating commands
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* chore(llms): regenerate model catalog from models.dev to pick up reasoning options
The baked fallback catalog was last regenerated before toModelInfo started
mapping models.dev reasoning_options into ModelInfo.reasoningOptions, so it
carried no reasoning metadata. Whenever the live models.dev fetch fails or a
model resolves from the baked catalog, adaptive-era Claude models (4.6+/5.x)
fell through the missing-reasoningOptions path to Anthropic manual thinking
and the API rejected the request with 'thinking.type.enabled is not
supported'.
This regen also picks up upstream models.dev drift; test expectations that
hardcoded stale catalog values (GLM 5.2 context window, OpenRouter GLM 4.7
reasoning controls, Vercel AI Gateway Qwen 3.6 Plus budget controls) are
updated to the current published values.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(llms): infer adaptive thinking for adaptive-era Claude ids when catalog reasoning options are missing
Claude 4.6+ and 5.x models reject the manual thinking wire shape
(thinking.type 'enabled') on the Anthropic API. When a model resolves
without reasoningOptions metadata (offline baked catalog before the regen,
or user-typed unlisted ids such as claude-opus-4-6:1m), the reasoning
policy previously fell through to anthropic-manual and every
reasoning-enabled request failed with a hard API error.
Add isClaudeAdaptiveEraModelId as a narrowly scoped id fallback (name-first
Claude ids with version 4.6+ or 5.x, plus the Fable line) and use it in the
missing-reasoningOptions branch of resolveAnthropicReasoningRequestPolicy.
Genuinely old or unknown Claude-compatible ids keep the manual default,
which remains the safe shape for third-party Claude-compatible endpoints.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(llms): prefer adaptive thinking over manual when a model advertises an effort control
A numeric reasoning.budgetTokens (e.g. a thinkingBudgetTokens setting
migrated from the legacy extension) used to force the anthropic-manual
policy whenever the model advertised a budget_tokens control. Claude 4.6+
models advertise both effort and budget_tokens on models.dev but reject
thinking.type 'enabled' on the Anthropic API, so those requests failed.
Effort now wins: adaptive is selected and the numeric budget is ignored.
Budget-only models (Sonnet 4.5 and older) keep honoring explicit budgets
via the manual shape.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* test(llms): guard baked catalog reasoning options for adaptive-era Claude models
Resolve adaptive-era Claude models through the generated (offline fallback)
catalog and assert their entries carry effort reasoning options that the
Anthropic reasoning policy resolves to adaptive thinking.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* test(core): update GLM 5.2 context window to current models.dev value
The catalog regen picked up upstream drift: models.dev now publishes a
1,000,000-token context window for zai/glm-5.2 (was 1,040,000).
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* revert(llms): drop the Claude id-based adaptive-thinking fallback
Keep the fix surface minimal: the catalog regen covers every model
models.dev lists (the overwhelming share of the production failures), and
the effort-over-budget policy covers listed models that advertise both
controls. Unlisted id variants (e.g. claude-opus-4-6:1m) keep the
pre-existing manual fallback rather than introducing id-version parsing in
model-facts.ts; if they appear in models.dev the catalog picks them up
automatically.
This reverts commit 6d725b1ecd7c896a08c3c58dbb37afeaf34bb31e, keeping the
regenerated catalog and the effort-precedence change.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(llms): default unknown Claude ids to adaptive thinking when catalog options are missing
Reintroduce the id-based fallback with the forward-compatible policy the
ecosystem converged on (vercel/ai#17804 for @ai-sdk/anthropic's capability
lookup; opencode's transform.ts after repeated allowlist misses for
opus-4.7, sonnet-5, and opus-5): when catalog reasoningOptions metadata is
unavailable, treat unrecognized Claude ids as newer than the known model
list and use adaptive thinking, since new Claude releases reject the manual
wire shape. Known legacy families (Instant, 2.x, 3.x, and name-first 4.0-4.5)
keep manual, as do non-Claude Anthropic-compatible ids and unknown Claude
ids carrying an explicit numeric budget (a custom-endpoint signal).
Unlike the earlier reverted allowlist (which defaulted unknown ids to
manual), this fails open for future models: claude-opus-4-6:1m-style
variants and next year's Claude work without a code change.
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(llms): retry empty model turns on all providers, not just Ollama
Production telemetry shows 'Model returned empty response' hard failures
on hosted backends (openrouter, cline, openai-compatible endpoints), not
just local Ollama — 46 tasks / 120 events in 24h on the SDK extension vs
~0 on legacy, which has its own empty-response fallback.
The retry-empty-response middleware already existed but was wired only
into the Ollama vendor. Move the wrap to the central AI SDK composition
point (createAiSdkProvider in ai-sdk.ts), where every vendor's model is
constructed, so all providers get it: retry only when a turn produced
genuinely nothing (no text, no reasoning, no tool call), tool-call-only
turns are never retried, non-empty turns stream through live, and error/
token-limit finishes pass through unchanged. Vendors can opt out or tune
attempts via ProviderFactoryResult.retryEmptyResponses. The agent
runtime's loud failure after persistently empty turns is unchanged.
* docs(sdk): drop changelog edit — release commits own the changelog
v0.0.69 is already published; its section must not be edited
retroactively. The next release commit will describe this change.
* ci: re-trigger checks (flaky Windows runner test timeouts)
* ci: re-trigger checks (flaky Windows runner test timeouts)
* fix(llms): classify stream parts exhaustively, buffer retry attempts, aggregate usage
Review follow-up (dominiccooney): the retry predicate and the response
parser were two independent, incomplete interpretations of the
LanguageModelV4StreamPart union, and rejected attempts leaked structural
parts and dropped billable usage.
- stream-part-classification.ts is now the single exhaustive boundary:
every part is converted content, explicitly unsupported output,
structural metadata, stream-start, finish, or error, with a never
check so new AI SDK part types fail compilation. Retry eligibility
derives from it: only turns with no output at all are retried;
unsupported-but-real output (custom, reasoning-file, source,
provider-executed tool-result) is never retried.
- Generated file parts are converted end to end: emitAiSdkEvents emits
a new file AgentModelEvent and the agent runtime assembles it onto
the assistant message (image part for image/*, file part otherwise),
so a file-only turn is no longer an empty message. The legacy
ApiStream bridge explicitly skips file events (no chunk type).
- Each retry attempt is buffered until it proves non-empty (first
output or error part), so discarded attempts leak nothing — one
retried request produces one clean stream with exactly one
stream-start.
- finish.usage from discarded attempts is aggregated field-by-field
(cache and reasoning detail included) into the emitted finish, so a
three-request turn reports three requests' worth of tokens.
* fix(vscode): show cwd-relative tool paths in the chat view
The SDK message translator copied the model's absolute file paths straight
into the ClineSayTool messages, so chat cards like "Cline wants to read this
file" showed full absolute paths. Relativize them against the task's cwd for
display (classic getReadablePath behavior: relative inside the cwd, basename
for the cwd itself, absolute when outside), including apply_patch's
"*** Update File:" markers which DiffEditRow parses for its headers.
Also restores the readFile card's click-to-open target by setting content to
the absolute path, matching the classic extension.
* refactor: apply display-path relativization as a single ClineSayTool transform
Instead of threading cwd through every case of sdkToolToClineSayTool, leave
the tool mapping untouched and apply one toDisplaySayTool transform (with a
filesystem-path tool whitelist) at the points where tool cards are emitted.
Same behavior, much smaller footprint; MCP/unknown tools keep their exact
prior behavior.
* fix: keep absolute readFile open-target untouched on Windows; match '..' as whole segment
path.resolve(cwd, absPath) rewrites a drive-less absolute path onto the
current drive on Windows, breaking the readFile card's click-to-open target
(and the tests asserting it). Guard with path.isAbsolute instead.
Also match '..' only as a whole path segment in toDisplayPath so an in-cwd
entry literally named '..config' is not misclassified as outside the cwd
(greptile P1).
* fix: keep Desktop-fallback paths absolute; relativize '*** Move to:' destinations
When VS Code has no workspace open, getWorkspaceRoot() falls back to the
Desktop; classic getReadablePath deliberately keeps full absolute paths in
that case so the user can see where operations occur. Restore that guard in
toDisplayPath.
Also enroll PATCH_MARKERS.MOVE in relativizePatchPaths so a rename renders
both source and destination relative (covers the split-patch path too).
* Fix Bedrock prompt caching: emit Converse cachePoint markers instead of anthropic cache_control
The Bedrock provider manifest routed prompt caching through the
anthropic-cache-control format, so requests carried cache_control
provider options that @ai-sdk/amazon-bedrock silently drops - its
Converse message converter only reads providerOptions.bedrock.cachePoint.
Bedrock never received a cache checkpoint, cacheRead/cacheWrite were
always 0, and a stray top-level cache_control field leaked into the
Converse request body.
Adds a bedrock-cache-point prompt-cache format that attaches a
message-level cachePoint marker to the last user message, which the
converter appends as a cachePoint content block, caching the whole
prefix up to it.
Fixes#12913
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Format gateway.test.ts assertion
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(desktop): show skills in slash command menu
* fix(core): normalize runtime slash command names
* fix(core): disambiguate colliding slash commands
* fix(core): preserve same-kind slash commands
* fix(core): stabilize colliding slash command aliases
* fix(core): avoid slow runtime command regex
* fix(core): remove quadratic hyphen trim
* fix(core): prefer skills over workflows on slash command collisions
Workflows are effectively deprecated in favor of skills, so when a
workflow's normalized name collides with a skill the skill now owns the
token and the workflow is dropped. This removes the collision
qualification machinery (-skill/-workflow/-hash aliases), which silently
renamed established CLI and VS Code command tokens, and removes the
duplicate-token throw that sat in the CLI send path, the hub snapshot
capability, and the desktop list_user_instruction_configs command.
Same-kind collisions resolve to the first entry of the deterministic
(name, id) sort, stable across discovery order.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(vscode): resolve typed workflow filenames through record ids
Name normalization broke the legacy /my-workflow.md fallback for
workflows renamed via frontmatter: the configured record name (e.g.
"Ship It") no longer compares equal to the normalized command token
("ship-it"), so a typed filename stopped expanding. Match the discovered
record to its runtime command by the stable record id instead, keeping
the canonical-name comparison as a fallback for callers that pass
records without ids.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(core): normalize snapshot names in hub slash command proxy
The hub-side proxy normalizes the typed token but compared it against
snapshot command names verbatim. Snapshots served by older clients carry
raw configured names (e.g. "Ship It"), which could previously exact-match
typed input and would now never match. Normalize both sides of the
comparison so mixed-version hub setups keep resolving.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(desktop): expand slash commands in the sidecar send path
Selecting a skill or workflow from the desktop slash menu inserted the
token but the sidecar dispatched it verbatim, so the model received
literal text like '/publish-ui write docs' instead of the configured
instructions. handleSend now expands a leading runtime slash command via
the core user-instruction service before dispatch (mirroring the CLI's
buildUserInputMessage), keeping the raw token as the session's display
prompt. Built-in webview commands (/fork, /team) and unknown tokens pass
through unchanged, and discovery failures fall back to the raw prompt.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(desktop): expand slash commands when editing queued prompts
Editing a pending prompt stored the raw slash token, which the runtime
later delivered to the model unexpanded — only the initial send path
went through expandRuntimeSlashCommand. handleUpdatePendingPrompt now
expands a leading skill/workflow token before persisting the update,
matching the enqueue behavior in handleSend.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(core): preserve Unicode letters in slash command tokens
Normalization stripped all non-ASCII characters, so a skill named 发布
got an unrelated generated token while typing /发布 could never resolve —
a regression from pre-normalization behavior where the exact name
matched. Keep Unicode letters and numbers in normalized tokens and only
collapse whitespace and symbol runs into hyphens.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
* fix(shared): emit AI SDK 7 file parts for images
formatMessagesForAiSdk still built the retired shapes: user images as
{type:'image'} message parts and tool-result images as {type:'image-data'}
content parts. AI SDK 7 auto-migrates both at runtime, but logs a
DeprecationWarning through process.emitWarning on every image-bearing
request, and the shims are slated for removal in the next major.
Emit the canonical shapes instead: {type:'file', data, mediaType} for
user images and {type:'file', data:{type:'data', data}, mediaType} for
tool-result media. mediaType is required on file parts, so URL-backed
images without a known type use the bare 'image' top-level segment,
which AI SDK 7 resolves per provider.
* fix(llms): allow the AI SDK 7 major of ai-sdk-provider-claude-code peer
The AI SDK 7 upgrade moved the ai-sdk-provider-claude-code
devDependency to ^4 but left the peer range at ^3.4.3, so consumers
resolving the peer would install the AI SDK 6 (Provider V3) major.
Align the peer range with the version the package is built against.
* fix(llms): route Bedrock foundation models through geo inference profiles
AWS Bedrock offers no on-demand throughput for newer foundation models;
they must be invoked through an inference profile. The SDK Bedrock vendor
passed model ids through unmodified, so every request with a bare modern
model id (e.g. anthropic.claude-sonnet-4-6) failed with "Invocation of
model ID ... with on-demand throughput isn't supported".
Resolve the wire-level model id in the Bedrock vendor: honor the existing
useCrossRegionInference / useGlobalInference settings (already plumbed
through provider config but previously ignored), and auto-prefix bare ids
of models known to have no on-demand throughput so they work without the
toggle. Ids that are already profile-prefixed, ARNs, and custom-model
configurations are never rewritten; unknown regions fall back to the raw
id. Country profiles (jp./au.) are preferred over apac. where the model
catalog shows AWS ships them.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(llms): future-proof Bedrock profile-required model patterns
Match Anthropic tier-first naming generically (excluding the frozen
legacy claude-3-*/claude-v2/claude-instant naming schemes) instead of
enumerating tier names, so future profile-only Claude tiers work without
pattern-list updates. Also cover the profile-only Amazon Nova 2 series.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(llms): gate Bedrock geo profiles on catalog availability
Address review feedback on inference-profile resolution:
- Never manufacture apac. (or other unconfirmed) profile ids: prefer a
catalog-confirmed variant among the region's candidates (jp./au./apac.),
and otherwise keep the raw id so AWS returns the actionable on-demand
error instead of "provided model identifier is invalid". Profile-only
models still fall back to us./us-gov./eu. prefixes, where AWS reliably
ships geo profiles for such models.
- Drop the customModelBaseId short-circuit: legacy migration copies the
base id without the custom-selected flag, so its presence must not
disable profile routing for a normal catalog model. Custom/provisioned
ids stay raw on the cross-region path because no catalog variant can be
confirmed for them, and ARN-based custom models were already passed
through.
Adds a per-region wire-id table test, an injected-catalog apac test, and
a stale-customModelBaseId regression test.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(llms): require catalog confirmation for every Bedrock geo profile
Remove the us./us-gov./eu. pattern fallback: AWS documents
inference-profile availability per model and geography, so no geographic
prefix is assumed valid without a catalog-confirmed variant (the catalog
had bare amazon.nova-lite/micro/pro ids with no geo variants, which the
fallback would have rewritten to unconfirmed eu./us. ids). The pattern
list now only gates eligibility for automatic routing; the catalog
always picks the actual prefix, and the raw id is preserved when no
variant is confirmed.
Adds boundary tests asserting pattern-matched models without confirmed
variants stay raw (with and without cross-region inference), plus
injected-catalog positive tests for us-gov. and future tier-first ids.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Fix installed plugins all displaying as "index" in the desktop app
Hoist getPluginDisplayName (nearest-ancestor package.json name with
basename fallback) into @cline/shared storage paths, re-export it via
@cline/core, and replace the duplicated copies in cline-hub, the CLI
TUI, and VS Code marketplace helpers. Fix the desktop sidecar and
'cline config plugins', which still named plugins by entry-file
basename, so package-backed installs showed up as "index".
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Use shared getPluginDisplayName in desktop sidecar after #12933
Merging main brought in PR #12933, which fixed the desktop plugin
naming with another local copy of the helper. Drop that copy in favor
of the shared @cline/shared implementation this branch introduces, and
remove the node:path imports it needed.
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
The webview posts an optimistic say:'task' message carrying the user's
images/files, and only clears it once an identical authoritative message
arrives from the extension. emitInitialTaskMessage omitted attachments,
so the optimistic copy was never confirmed and withPendingUserMessage
kept re-injecting the old task into the transcript even after New Task
cleared it - leaving the chat permanently stuck on the previous task.
- Include images/files on the authoritative initial task message so the
optimistic pending copy is confirmed and cleared as designed.
- Defensively drop any unconfirmed optimistic message in startNewTask so
an explicit New Task click always yields a clean slate.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(cli): track external git branch changes in the TUI status bar
The branch shown below the prompt was read once at startup and only
refreshed after an agent turn, so checkouts made from another terminal
or an editor left the TUI showing a stale branch (#12911).
Watch the repo's git dir for HEAD changes (git replaces HEAD via
rename, so a directory watch is used) and refresh the status bar
immediately, with a slow 5s poll as a fallback for filesystems where
fs.watch is unreliable. State updates are skipped when nothing changed
to avoid needless re-renders.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* refactor(cli): replace subprocess polling with stat-based HEAD backstop
Drop the unconditional 5s git-subprocess poll from useRepoStatus. The
fs.watch directory watcher stays for instant updates where the runtime
delivers HEAD events, but Bun on Linux drops them, so add fs.watchFile
on the HEAD file as the backstop: one in-process stat() every 2s that
only triggers git subprocesses when HEAD actually changed. Verified in
the Bun-run TUI that external checkouts show up within ~2s.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* refactor(cli): simplify HEAD watching to a single fs.watchFile
Drop the fs.watch directory watcher (Bun on Linux never delivers its
HEAD events, making it dead weight on the runtime the CLI ships on) and
the debounce it required. watchGitHead now just stat-watches the single
.git/HEAD file via fs.watchFile, which survives git's rename-based HEAD
updates and works on network mounts.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* refactor(cli): use a plain 5s poll for repo status
Remove the HEAD watcher entirely per review preference for minimal
code: root.tsx now just polls readRepoStatus every 5 seconds, skipping
state updates (via isSameRepoStatus) when nothing changed so idle ticks
don't re-render the app.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(cli): skip repo status poll ticks while a read is in flight
Bounds concurrent git subprocesses when a read exceeds the 5s interval
(slow git on huge repos) and prevents an older completion from
overwriting newer status. Addresses Greptile review feedback.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
The copy/fork (and user copy/edit/restore) action row was pulled up 8px
(-translate-y-2), which made the icons collide with the descenders of the
message's last line of text. Reduce the raise to 4px (-translate-y-1) so the
actions sit with a small, deliberate gap under the message content while
still hugging the message closely enough to read as attached to it.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(telemetry): dedupe sdk.error across layers and rate-limit repeated failures
Every provider failure was emitted twice — once by the model layer
(provider.stream, handled: true) and again verbatim by the agent loop
(agent.run, handled: false) — and unattended retry loops emitted the
same failure every iteration, unbounded. 24h of CLI data: 300K events
from 1.7K users, top 10 machines at 70% of volume.
Two changes:
- The agent loop no longer re-reports model stream failures. run-failed
events carry errorClass exactly when the run failed on a model stream
error, and the model layer already reports those at its own error
boundary — so the run loop reports only failures that originate in
the loop itself (empty response, max iterations, ...).
- captureSdkError caps identical failures per process: 5 per hour per
(event, component, operation, error_type, normalized message), with
digit runs collapsed so retry counters coalesce. Suppressed emissions
surface as suppressed_count on the next emission after the window
rolls over. In-memory only; the cap never blocks reporting.
Event name, attributes, and all call sites are unchanged;
suppressed_count is the only additive field.
* fix(telemetry): make sdk.error dedup ownership explicit and key limiter on status/code
Review follow-ups (#12931):
- Reporting ownership is now an explicit signal instead of being inferred
from errorClass. captureSdkError returns whether the failure was
recorded, the model layer forwards that as errorReported on the finish
event, and the run loop skips only failures marked reported. Custom
AgentModel implementations that never call captureSdkError leave the
bit unset, so their failures still produce exactly one sdk.error from
the run loop (regression test added).
- The rate-limit key now includes the structured error_status and
error_code that normalizeSdkError already extracts, so an HTTP 429
hot loop cannot consume an HTTP 401's budget even though their
messages differ only by digits (tests added for both fields).
- resetSdkErrorRateLimiterForTests is tagged @internal; it stays
re-exported because package test suites can only reach it through the
package entry point.
The publish job ran under the shared `Publish` GitHub environment, whose
required reviewers turned every @cline/ui release into a two-person
ceremony. Nothing in the job reads secrets from that environment — it
authenticates to npm purely over OIDC trusted publishing — so the
environment bought us an approval prompt and nothing else. sdk-publish
and cli-publish already publish unattended the same way.
The npm trusted publisher for @cline/ui was registered with
`environment: Publish`, which pins the OIDC token's environment claim, so
it has been re-registered without it (same repo, workflow file, and
permissions). That change is already live; landing this without it would
have broken publishing.
Access is still gated by workflow_dispatch (write access required), the
`refs/heads/main` ref check, and the typed `publish` confirmation.
* Remove model-initiated plan-to-act switching from the VS Code extension
Match the legacy extension: the model can no longer call switch_to_act_mode
to move itself from plan mode to act mode. The user must flip the Plan/Act
toggle manually. The CLI keeps the tool and its prompt unchanged.
- Stop registering the switch_to_act_mode extra tool in plan-mode sessions
and drop the pending-mode-change queue, beforeModel stop hook, and idle
apply path that existed only for the tool-initiated switch.
- Add a planModeSwitchTool option to buildClineSystemPrompt (default true,
CLI output unchanged) and a PLAN_MODE_INSTRUCTIONS_MANUAL_SWITCH variant
that directs the model to ask the user to toggle to Act mode instead of
calling a tool it does not have; the extension passes false.
- The user-driven toggle path (togglePlanActModeProto), including
auto-continue when a completed plan is presented, is unchanged.
* Use a generic completing-tool name in translator retag test
Review feedback: submit_and_exit is a yolo-mode tool and does not exist
in plan/act sessions. The test exercises tool-agnostic translator
behavior, so use a neutral example name and clarify the comment.
* fix(cli): claim connector instance before socket connect
Prevent racing foreground or detached connector launches from
both opening socket-mode with the same bot token by exclusively
claiming the state file via tryClaimConnectorStateFile before
connecting, and exit with CONNECT_ALREADY_RUNNING_EXIT_CODE when
another live instance already holds the claim.
* serializes stale-generation replacement without a removable mutex
* lint
* fix
* base
* fix(cli): keep pre-claim Slack state files manageable
State files written by CLI versions that predate connector claiming have
no claimId, so requiring it in readConnectorState made a live legacy
connector invisible: stop deleted its state without stopping it, which
let the next connect open a second socket-mode connection with the same
bot token. Treat claimId as optional metadata; claiming itself never
relied on the validator.
---------
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* chore(telemetry): remove duplicate capture defs owned by @cline/core
The cline-core bundle layer (apps/vscode, which also compiles into the
cline-core.js that JetBrains runs) still carries event constants and
public capture methods inherited from the legacy architecture. Where
@cline/core now emits the same event, the bundle-layer twin is a second
capture path on the same telemetry service — which is how the
task.provider_api_error double-emission happened (cline/cline#12820,
follow-up to the removals in cline/cline#12818).
Removes 20 such methods and the EVENTS constants only they referenced:
- 15 whose signal @cline/core emits today (task lifecycle, tokens, tool
and skill usage, auth start/success/failure, opt-out, workspace init,
and summarize_task, which core replaced with task.compaction_*).
- 3 obsolete on both architecture lines: captureModelSelected (its
model_selected signal survives as an action on
captureOnboardingProgress), captureRulesMenuOpened, captureHostEvent.
- 2 whose trigger moved into core/sdk, so this layer can no longer
observe them and re-wiring here would be wrong:
captureWorkspacePathResolved (core already owns workspace.path_resolved)
and captureGeminiApiPerformance (providers live in core; generic
provider-timing events supersede it).
Deliberately NOT removed: capture methods with no caller here but a live
caller on legacy-extension. Those emit signals originating in this bundle
(webview UI, VS Code storage, host terminal, checkpoints, focus chain,
legacy-task migration), so core cannot emit them and the missing piece is
a call site on this line, not a redundant definition. They are the
SDK-parity backlog and are flagged as such in the file.
Verified against the JetBrains plugin repo: it references none of these
methods or event names, and no proto surface changes.
Also drops the unused TokenUsage interface, the taskTurnCounts and
taskToolCallCounts maps (only deleted methods wrote to them), and EVENTS
constants that were already orphaned before this change.
Tests that only exercised a removed method are gone; tests that used one
merely as a vehicle for provider/metadata assertions now use a surviving
method, so that coverage is preserved.
* fix(telemetry): keep agent identity on events dispatched after session teardown
A small share of task.tool_used events (~285 of 229k over 48h on
extension_variant=next) arrive without any agentId/agentKind/isSubagent
attributes. Root cause: AgentEventBridge.dispatchAgentEvent resolves
identity solely from the live-session map (AgentEvent metadata never
carries agentId in practice), and session teardown deletes the map entry
before the agent's run fully drains — dispose/stopSession paths can skip
or fail agent.shutdown() without aborting first. Late events from the
still-draining run then hit the session-map miss branch, which passed no
identity at all, so buildTelemetryAgentIdentity returned undefined and
the event was emitted bare.
Fix: snapshot the identity stamped on each session's events while the
session is registered (bounded FIFO map) and reuse it on a session-map
miss. Purely additive — no event is added, removed, or renamed; the
live-session and sub-agent paths emit byte-identical properties.
* feat(sdk): add session initiation mode and lazy session persistence
- Introduce top-level `StartSessionInput.mode` (`user`, `automation`, `subagent`, `team`) alongside `source`, so persisted history records both the client surface and how the session began; missing mode defaults to `user`.
- Make root-session persistence lazy: starting a runtime allocates the session ID in memory without creating a database row, manifest, or messages artifact. The first accepted user turn persists that same ID, so closing a runtime before any user turn leaves no empty history entry, and persistence never allocates a replacement ID for unknown sessions.
- Require automation runtime adapters to explicitly persist `mode: "automation"` for every run.
- Document the provenance model in `sdk/ARCHITECTURE.md`, update the VS Code session factory comment, and add tests for the automation runtime handlers.
* fix(sdk): persist automation trigger source as session provenance
The runtime adapters stopped writing the cron request source into the
session row when source became the client surface, which silently
dropped the spec-defined trigger label. Record it as
sessionHistoryOrigin.trigger instead, surface it in the messages-file
origin, and sort the new history-origin import in HubRuntimeHost.
---------
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
* fix(ollama): use native AI SDK provider
* fix(ollama): patch ollama-ai-provider-v2 wire contracts and lock them with real-provider tests
The pinned ollama-ai-provider-v2@4.0.1 breaks four native Ollama wire
contracts (review findings on #12892). Patch the package via Bun
patchedDependencies:
- omit think from the request when no reasoning setting resolves,
instead of forcing think: false (lets the server default apply)
- surface mid-stream {"error": ...} objects as error stream parts with
an error finish reason, instead of dropping them before a clean finish
- serialize attachment-only user turns as string content (""), not []
- include the documented tool_name field on tool result messages
Add ollama.wire.test.ts exercising doStream through the vendor module
against the real (patched) package with a stubbed fetch, asserting on
the actual /api/chat request bodies and parsed stream so regressions in
the dependency's request converter or stream parser are caught.
* fix ollama model list refresh
---------
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
* fix: prevent duplicate connector launches during doctor/connect
Mark connectors as starting before the hub daemon spawns so autostart
skips in-flight instances, and improve doctor process filtering with
container-aware namespace/cgroup checks plus detached log rotation.
* Update apps/cli/src/connectors/common.ts
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* feat(hub): supervise connector processes
* feat(connectors): enable tools by default, and stop replaying the Slack greeting
Tools were on by default only for Telegram (via --no-tools); Slack, Discord,
Linear, Google Chat and WhatsApp all required an explicit --enable-tools. All
six now default to tools on and opt out with --no-tools.
--enable-tools still parses everywhere, including Telegram which never accepted
it, so deployed scripts, systemd units and persisted autostart arguments keep
working. Passing both resolves to the safer answer: --no-tools wins. This also
affects hub/webview starts, which never emitted a tools flag and so ran those
five connectors with tools off.
Slack no longer posts the "Connected to Cline." first-contact message. It was
gated on per-thread welcomeSentAt, so a connector restart or a cleared history
made the next user message look like first contact and replayed the greeting.
The host mechanism is unchanged and the other adapters still greet.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(connectors): recover a thread whose session is wedged mid-run
A connector thread keeps a long-lived mapping to a hub session. When that
session's runtime still had a run in flight and no abort had been requested,
every message in the thread came back as "SessionRuntime.shutdown called while a
run is in progress" instead of an answer, and stayed that way until someone
cleared the binding by hand. Observed on the Cline Mom Slack bot after a stack
restart.
The connector host already recovers from a session the hub no longer knows
about: it forgets the mapping and replays the turn once against a fresh session.
This widens the trigger from "session not found" to "session cannot serve
another turn" via isUnusableSessionError, so a wedged runtime takes the same
path.
The shutdown error now carries a stable code (SessionRunInProgressError,
session_run_in_progress) so callers can recognise it structurally. The predicate
also matches on message, because an error reaching a connector has crossed the
hub's JSON boundary and arrives as a bare message - and because a host commonly
runs a hub and CLI of different versions. Ordinary run failures still propagate
untouched: replacing the session on those would hide real errors and drop the
conversation.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(connectors): serialise turns that share a session
Answering "what happens if I message the bot in another thread while it is still
replying": channel threads were already independent, but DMs were not.
findBindingForThread deliberately reuses one binding — and therefore one runtime
session — for every message in a DM channel, so a DM stays one continuous
conversation. The turn queue, though, was keyed by thread id, and a DM thread id
carries the message timestamp. Two messages in flight in the same DM therefore
got two independent queues and ran concurrently against a single session, which
fails with "shutdown called while a run is in progress" or interleaves two
conversations in one session history.
The queue key now follows the same identity rule as the binding lookup, via
resolveThreadTurnQueueKey next to findBindingForThread so the two cannot drift.
DM messages queue behind each other on the shared session; channel threads keep
their own key and still run in parallel. Applied to all six adapters, which all
had the same mismatch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(core): abort an in-flight run before tearing its session down
Where the Slack bot's "plugin-sandbox process exited (code=null, signal=SIGTERM)"
came from, and its "shutdown called while a run is in progress" sibling: both are
one event, a session released while a run was still going.
stopSession aborts the agent first "so shutdown can proceed", but callers that
reach shutdownSession or releaseSessionRuntime another way did not - hub
dispose() on a restart being the one that hurt. Without an abort the runtime
refuses to shut down, that error is rethrown from the cleanup, and the plugin
sandbox is SIGTERMed while tool calls are still pending, so those calls reject
with "plugin-sandbox process exited". A connector turn awaiting the run reports
whichever surfaced first instead of answering.
Both paths now abort and let the run drain before shutting the agent, runtime and
sandbox down, guarded on session.aborting so callers that already aborted do not
abort twice.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(connectors): stop announcing "Steering current task."
Every follow-up sent while the bot was replying added an acknowledgement line to
the thread, and the wording overstated what happens: the host treats delivery
"steer" the same as "queue", enqueuing the prompt for the session rather than
injecting it into the loop already running. The follow-up is now handed over
silently.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(core): retire a dead supervised entry before replacing it
A start arriving while an instance sat in backoff left the old entry's
restart timer live. The timer closes over the old entry object, so when
it fired it spawned a second process for the same (channel, instanceId)
- untracked by the supervisor's map, so invisible to list() and
unreachable by stop() - two connectors holding one bot token, which is
the exact failure supervision exists to prevent. Its exit handler then
kept reaping the live instance's state and rescheduling restarts.
The same window exists before the timer is even scheduled: the
exit-cleanup chain runs first, and a replacement made mid-chain would be
followed by a restart scheduled for the retired entry.
start() now retires a dead existing entry explicitly - cancel its timer,
mark it stopped, drop its exit listener. Both the timer callback and the
cleanup chain already stand down on "stopped", so one mark covers both
phases.
* fix(core): serialise supervisor start/stop and wait for stopped processes to die
Found by exercising a hub restart against a live webhook connector: the
new hub's boot reconnect restarts the adopted survivor - which suspends
inside stop() on the CLI cleanup - while the user's `cline connect`
arrives as connector.start. With no per-instance serialisation the two
starts interleaved across that suspension and both spawned. The map
tracked one process while the other lived on untracked, holding the
connector's webhook port; the tracked chain crash-looped on EADDRINUSE
through all five attempts and ended state=failed, while the ghost kept
running with no way to reach it through list() or stop().
Two changes:
- start/stop (and the backoff-restart spawn) now run under a per-
instance-key promise queue, so one instance has exactly one lifecycle
operation in flight. The exit-cleanup chain also stands down when its
entry is no longer the one in the map.
- stop() waits for the process to actually die after SIGTERM (bounded,
then SIGKILL) instead of returning while it still holds its listen
port - the race that turned the double-spawn into a crash loop, and
that could burn a backoff cycle on any webhook connector restart.
process.kill is now injectable (killProcess), which also stops the test
suite from signalling arbitrary real pids like 600 on the host.
Verified live: the same kill-hub-then-reconnect sequence now converges
to one tracked running process, with the concurrent user start
correctly answered "already running under the hub".
---------
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
Tauri's universal-apple-darwin target lipos the Rust binary but expects
sidecars to already be fat binaries, so build-sidecar-bin.ts now compiles
both Bun slices and merges them when the target triple is universal.
The publish workflow builds one universal bundle instead of a two-leg
matrix, verifies every Mach-O in the bundle carries both slices, and the
updater manifest points both darwin-aarch64 and darwin-x86_64 at the same
universal artifact so existing per-arch installs migrate automatically.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Add user-selectable color themes to the CLI TUI
Adds a theme system to the interactive TUI (cline -i):
- New tuiTheme global setting persisted in global-settings.json
- Built-in themes: Auto (terminal-adaptive, default), Cline Dark,
Cline Light, Tokyo Night, Gruvbox Dark, Nord, Dracula, Catppuccin
Mocha, One Dark, Solarized Dark, Solarized Light
- /theme command, command palette entry, and a Theme row in
/settings General tab, all opening a live-preview theme picker
- Named themes paint their background, default foreground, accents,
syntax highlighting, and derived diff colors across the TUI
- CLINE_THEME env var overrides the persisted theme at startup
Closes#12872
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Widen theme picker dialog and label column
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Format theme picker
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Give each theme a descriptive picker blurb
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Theme all main-surface components instead of static palette colors
The ask-question / tool-approval element, toasts, queued prompts,
autocomplete dropdown, chat error cards, searchable lists, and the
onboarding screens hardcoded the brand palette (act blue, selection
highlight, black-on-selection text) and fixed dark grays, so they
ignored the active theme.
- ResolvedTheme gains selection/textOnSelection; the selected-row text
flips between black and white by WCAG contrast against the accent
- Inline ask-question / tool-approval, Toast, QueuedPrompts,
AutocompleteDropdown, SearchableList, and chat error cards now use
theme accents and the themed selection pair
- Onboarding screens derive subtle borders/details from the theme
background instead of #333333/#555555, and use themed accents
- Dialog surfaces (settings, pickers, history) intentionally keep their
static dark surface styling
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix: correct Linux keybinding label in Plan/Act mode tooltip
On Linux, event.metaKey maps to the Super (Win) key, not Alt.
detectMetaKeyChar was returning "Alt" for Linux, causing the Plan/Act
mode toggle tooltip to display "Alt+Shift+A" instead of "Super+Shift+A".
Fixes#11026
* fix: update platformUtils.spec.ts Linux test expectation to Super
---------
Co-authored-by: Dominic Cooney <dominic.cooney@cline.bot>
* ci(desktop): drop the Rust build cache from the code-signing job
The `build` job is the only one that can read the Apple Developer ID
certificate and the Tauri updater signing key, and it restored a
swatinem/rust-cache archive before running them. A restored cache archive
is attacker-controlled the moment the Actions cache is poisoned, which is
the pivot used against this repo's nightly workflow in Feb 2026 and the
reason actions/cache was stripped from the credential-bearing publish
jobs at the time. This workflow was added months later and reintroduced
the pattern. The updater key is the worst thing here to leak: it signs
every auto-update the installed desktop app accepts.
The cache was also not buying anything. Across the eight runs of this
workflow, seven logged "No cache found" on both matrix legs; only the run
32 minutes after another one hit, saving 1-3 minutes. A release cadence
measured in days does not outlive the entry under the repo's 10 GB LRU
eviction, so the steady state was a cold build regardless. Cold builds
took 5-7 minutes against a 90-minute timeout.
No behaviour change otherwise: the step had no id and no outputs, so
nothing referenced it.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* ci(desktop): trim the cache-removal comment to the constraint
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
The @chat-adapter/telegram library intercepts any message whose leading
entity is a bot_command and routes it to slash-command handlers instead
of the mention/subscribed-message handlers. The Telegram connector
registered no onSlashCommand handler (unlike Discord and Slack), so
commands like /clear were consumed by the library and silently dropped.
Register a slash-command handler that rebuilds the originating chat
thread and forwards the original message text (preserving @bot
addressing for group chats) into the same turn pipeline as regular
messages, so connector commands reach the chat command host.
Fixes#12871
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(vscode): surface a clear error when a provider has no API key
A key-based provider with no API key sends the request without an Authorization header, and the provider's raw 401 reached the chat panel unclassified because reshapeErrorForWebview falls through to the raw message and ClineError's auth regexes do not match it. Rewrite the missing-Authorization-header case into actionable guidance naming the provider, alongside the existing model-not-found matcher. Matching is limited to the no-header signature so a present-but-wrong key is never relabelled as missing, and no preflight is added because authMethod misclassifies local providers and 175 of 179 builtins resolve keys from the environment.
* fix: don't name a fallback provider in the missing-key message
reshapeErrorForWebview defaults providerId to "cline" for its
ClineError-JSON branches, but state.activeProviderId() can be undefined —
the missing-credential message would then blame the cline provider for a
key it doesn't take. Keep the "cline" fallback for the JSON branches and
pass the raw id to the credential matcher, which now only names a provider
it was actually given.
---------
Co-authored-by: Mikołaj Kondratek <19799111+mkondratek@users.noreply.github.com>
* refactor(llms): classify typed AI SDK errors before the structural walk
Add a typed pre-pass to classifyProviderError that recognizes real AI SDK
error instances via their symbol-based isInstance() guards: RetryError
unwraps to its last attempt, APICallError is judged on message/responseBody/
data with its typed statusCode as the sole authoritative status,
TypeValidationError on the payload in value, and any other AISDKError
recurses into its cause. The detection rules (overflow patterns, provider
codes, rate-limit vetoes, invalid-request status gate) are extracted into a
shared verdict function used unchanged by both the typed pass and the
existing structural walk, which remains the fallback for gateway-forwarded
plain-JSON payloads that only name an AI SDK error (ENG-2394).
* fix(llms): classify a RetryError by its final attempt even when untyped
When a RetryError's last attempt was not a typed AI SDK error, the typed
pre-pass fell back to structurally walking the whole wrapper, letting
signals from earlier (retried-away) attempts veto or fake the final
attempt's verdict — e.g. a retryable 429 on attempt one vetoing a plain
overflow rejection on the final attempt. Walk the final attempt alone
instead; a RetryError with no recorded attempts still falls back to the
plain structural walk.
* fix(llms): gate typed APICallError verdicts on the authoritative statusCode
verdictFromSignals checks the explicit context_length_exceeded code before
the rate-limit veto and the invalid-request status gate, so a typed
APICallError with statusCode 429 or 500 whose body echoed that code was
still classified as an overflow, contradicting the branch's contract that
the typed statusCode is the sole authoritative status. Gate the whole
payload verdict on the typed statusCode first (absent a statusCode the
payload still decides), and cover the explicit-code case at 429/500/400
with real instances.
* feat(sdk): detect and recover from context-window overflow errors
Port of the legacy arch's context-window-exceeded handling to the SDK
arch (the SDK arch previously surfaced these as raw unclassified stream
errors with no recovery; see ENG-2394 root-cause investigation).
- llms: new classifyProviderError() walks the raw provider error
structure (AI SDK wrappers, gateway value.error_message, responseBody,
cause chains) and classifies it before extractErrorMessage flattens
it. Reuses the legacy detectors' message patterns with rate-limit
vetoes and an invalid-request status gate.
- shared: ProviderErrorClass union; errorClass on the model finish
event, run-failed event, runtime snapshot, and prepare-turn contexts.
- core: prepare-turn overflowRecovery flag forces a compaction that
bypasses the token-estimate trigger (the estimate just proved wrong)
and runs the deterministic basic strategy directly, so recovery never
depends on another successful LLM request. New overflow_recovery
compaction mode in status notices and compaction telemetry.
- agents: on a classified overflow the runtime force-compacts and
retries once per run, emitting a status notice. Terminal states fail
with actionable messages instead of raw provider dumps: nothing to
compact (first-prompt overflow), no prepare-turn pipeline, or a retry
that still overflows. The doomed request is not re-sent when forced
compaction cannot shrink the transcript.
- telemetry: task.provider_api_error gains errorClass and
task.provider_stream_failed gains error_class, populated from the
same classification, so context-overflow failures become countable.
* fix(core): keep overflow-recovery compaction deterministic with custom compactors
A session-supplied compaction.compact previously took precedence over
the overflow_recovery basic-strategy branch, so an LLM-backed custom
compactor could hit the same context overflow mid-recovery. The custom
compactor still gets first shot (it sees mode overflow_recovery and
owns its transcript invariants), but if it throws or declines, basic
compaction now runs so recovery never depends on another successful
LLM request. Cancellation still propagates.
* fix(core): fall back to basic compaction when a custom compactor does not shrink during overflow recovery
A custom compactor that returns unchanged or larger messages would
previously satisfy the recovery branch, and the runtime would then
reject the retry as non-shrinking and fail terminally even though
basic compaction could still prune the transcript. Recovery now
treats a non-shrinking custom result like a decline and runs basic
compaction.
* fix(core): hold custom overflow-recovery compaction to the recovery token target
A custom compactor result that was only marginally smaller than the
input passed the shrink check, skipped the basic fallback, and spent
the run's single recovery retry on a request that still could not fit.
The custom result is now accepted only when it is strictly smaller AND
within the recovery token target basic compaction aims for; otherwise
basic compaction runs.
* fix(core): reject empty custom compaction results during overflow recovery
An empty transcript from a custom compactor passed both the shrink and
token-target checks (trivially smaller, zero tokens) and suppressed the
basic fallback, so the retry would have been sent without the request
it was supposed to re-send. The acceptance bar now covers the full
input space in one predicate: non-empty AND strictly smaller AND within
the recovery token target.
* feat(core): expose the turn abort signal to custom compactors
CoreCompactionContext now carries the prepare-turn abort signal, so a
custom compact implementation that calls a model or external service
can observe cancellation instead of blocking the turn (including the
overflow-recovery path) on a stalled request. Builtin strategies
already received the signal via providerConfig; this closes the gap
for custom compactors across auto, manual, and recovery modes.
* fix(sdk): classify provider errors from registered ApiHandler models
Registered handlers (VS Code LM and any other host-supplied provider)
reach the runtime through createAgentModelFromApiHandler, which flattens
failures to a message string — so context-window rejections on that path
were never classified and never entered overflow recovery.
- The adapter now classifies at its own error boundary, where the raw
error is still structured (status codes, response bodies), for both
thrown errors and failed done chunks. Aborts stay unclassified.
- The runtime falls back to classifying the finish message when a model
supplies no class, so custom AgentModel implementations are covered
too.
- Hold the custom-compactor acceptance check to token estimates on both
sides instead of mixing serialized length with a token target, and
document why the runtime's shrink backstop keeps a serialized-size
proxy (the shared estimator is linear in characters, so the verdict is
identical) with a TODO to surface real estimates from prepareTurn.
- Drop the now-unused errorClass parameter from captureProviderApiError:
#12820 removed core's capture site, so host adapters own that event.
* test(core): reuse the handler harness for the overflow classification case
The hand-rolled throwing generator had no yield, which biome's
correctness/useYield rejects as an error (the repo's lint gate runs on
sdk/ and apps/, and biome does not honor the eslint require-yield
directive the existing harness carries). fakeHandler now accepts the
error to throw, so the new case reuses it instead.
* fix(mcp): refresh lists on list_changed notifications instead of toasting
Servers emit notifications/tools/list_changed in bursts (a toolset change
or shutdown can produce a dozen at once), and the fallback notification
handler surfaced every one of them as a host toast, flooding the user
with identical messages (ENG-2298, found testing the JetBrains IDE MCP
server integration).
Handle tools/resources/prompts list_changed notifications by refreshing
the corresponding cached lists, debounced 300ms per server and list
kind, then pushing the update through notifyWebviewOfServerChanges() so
the webview and the SDK session tool-list check pick it up. Downgrade
remaining unhandled notification types to logger output.
* fix(mcp): guard list_changed refreshes against races and failed fetches
Address review: serialize per-key refreshes by chaining onto any
in-flight one, so overlapping fetches can't complete out of order and
publish a stale list. Make the fetch helpers return undefined on
failure (instead of an empty list) so the refresh path can keep the
previous cached list and skip the webview notification, rather than
erasing valid entries on a transient error; connect-time call sites
keep their old empty-list fallback.
* fix(mcp): drop in-flight list refresh when the connection was replaced
Address review: refreshChangedList captured the connection object before
awaiting the list fetches, so a reconnect mid-fetch wrote the result to
the removed connection while the replacement kept its own state. Re-check
connection identity after the fetches and drop the result when it
changed — the replacement fetched fresh lists at connect time, after the
change that produced the notification, so the in-flight result is older.
* fix(mcp): retry failed list refreshes and publish state after reconnect
Address review. A list_changed notification consumes the server's change
signal, so a transiently failed refresh left the cached list stale until
the next notification; retry with exponential backoff (1s/2s/4s, max 3)
per server and list kind, with a fresh notification superseding any
pending retry. Also publish server state after a successful streamable
HTTP reconnect: connectToServer() loads fresh lists but never sent them,
leaving the webview on 'connecting' with pre-reconnect capabilities.
* fix(mcp): don't restart a live connection when post-reconnect publish fails
Address review: the post-reconnect notifyWebviewOfServerChanges() sat
inside the connect retry loop's try block, so a publication failure
(e.g. a settings file read error) was treated as a transport failure
and re-ran connectToServer() against the already-live connection,
leaking its client/transport. Publication now happens outside the
connect try/catch and only logs on failure.
* fix(mcp): drop superseded in-flight list refreshes instead of publishing
Address review: a newer list_changed notification queued its refresh
behind one already in flight without invalidating it, so the older run
could briefly publish an obsolete list (and churn the SDK session)
before the newer refresh corrected it. Each schedule now starts a new
generation per server+kind; a run whose generation is no longer current
skips fetching (when caught early), drops its result before publishing,
and doesn't schedule retries — the superseding refresh covers it.
* fix(mcp): harden list refresh and reconnect publication paths
Address review (post-reconnect publish failure leaving consumers stuck
on 'connecting' with stale lists) plus an adversarial pass over the
whole change to close the remaining gaps in one batch:
- Retry publications bounded (publishServerChanges) after a successful
reconnect AND in both terminal disconnected paths, which are equally
terminal; never throw from handleError, whose promise transport.onerror
discards. Guard the stdio/SSE onerror publishes the same way.
- Treat an undeclared capability or a method-not-found answer as an
authoritatively empty list instead of a retryable failure, so servers
without e.g. resources/templates/list don't burn the full retry ladder
on every list_changed notification.
- Retry when the fetch succeeded but the webview publish failed: the
cache is updated but consumers haven't seen it.
- Cap debounce deferral at 2s so a sustained sub-300ms notification
stream can't starve the refresh indefinitely.
- Cancel pending refresh timers in deleteConnection; return 'skipped'
(not 'failed') when a fetch failure coincides with connection
teardown or supersession, so no retry fires against a replacement
connection that already fetched fresh lists.
- Clear the pre-existing toolListChangeDebounceTimer in dispose().
* fix(mcp): supersede in-flight refreshes during connection teardown
Address review: deleteConnection removes the connection from
this.connections only after awaiting transport/client close, so a list
refresh completing inside that window passed its identity check and
published state for a connection being torn down. Bump the per-key
generation at the start of deleteConnection so any in-flight refresh is
superseded and drops its result; bumping (never resetting) keeps
generations monotonic across reconnects.
* fix(mcp): close reconnect-retry and teardown-publication races
Address review (cline-cloud):
1. The streamable HTTP reconnect loop revalidated only after the first
backoff. Later retries could resurrect a server removed or disabled
from settings during a delay, or displace a replacement connection
another path had installed — connectToServer() drops a same-name
connection without closing it, leaking its transport. The loop now
revalidates before every attempt: it aborts when a live replacement
exists (our own original connection and the 'disconnected' husk left
by our own failed attempt don't count) or when fresh-read settings no
longer define the server as enabled (isStillWanted callback; a
settings read failure keeps the chain alive).
2. deleteConnection removed the connection from this.connections only
after awaiting transport/client close, so a publication passing its
suspension points inside that window could still serialize and
publish the dying connection's state. The connection is now removed
from published state before the close handshake is awaited.
The exhausted-retries test's partial-connection mock now carries status
'disconnected', matching what connectToServer's error path actually
leaves behind — that status is what distinguishes our own husk from a
live replacement.
* fix(mcp): don't displace an OAuth-required replacement during reconnect retries
Address review: the retry loop's replacement guard treated every
'disconnected' connection as our own failed-connect husk. An
OAuth-required connection is also 'disconnected' but retains its
client, transport, and authProvider for authentication — a retry that
displaced it would orphan that session and clobber the pending auth
state. Distinguish by client presence: the husk's creation sites set
client: null, so a 'disconnected' connection holding a client is a
replacement and aborts the retry chain.
* fix(mcp): distinguish OAuth replacements by flag, not client presence
Address review: an ordinary post-registration connect failure leaves a
'disconnected' connection with its (already-closed) client still
attached, so the client-presence check classified it as an OAuth-style
replacement — aborting the reconnect chain after a single failure,
including our own retries. Discriminate on server.oauthRequired
instead: only the OAuth-required connection retains live
client/transport/authProvider state worth protecting; ordinary failed
connections closed their client before being marked disconnected, so
retrying past them displaces nothing live. The exhausted-retries test
mock now carries the real husk shape (closed client attached) to pin
this regression.
* test(mcp): cover retry succeeding after a failed attempt's registered connection
Requested in review: the guard must recognize the 'disconnected'
connection a failed non-OAuth connectToServer() leaves behind (closed
client still attached) as our own attempt, and the following retry must
proceed and succeed.
* fix(mcp): don't retry reconnects with a config settings no longer define
Address review: a retry reconnects with the config captured at
connection creation, so if the user changed the server's config during
a backoff delay (and the watcher's reconnect with the new config
failed, leaving a disconnected husk our guard rightly retries past),
the retry would resurrect the obsolete URL/headers/command — and a
successful stale connection would contradict settings until the next
file touch. isStillWanted now also compares the captured config against
current settings via configsRequireRestart (connection-relevant fields
only), aborting the chain when they differ: the settings watcher owns
reconnection after a config change.
* feat(desktop): show token usage in input toolbar
Load per-model context window sizes from the provider catalog and
pass the active model's limit down to ChatInputBar. Render a token
ring that visualizes current token usage against the model's context
window, and hydrate token usage plus cumulative cost from messages
and chat_usage events so the indicator stays accurate across turns
and session reloads. Add tests covering the ring rendering and usage
hydration.
* bigger ring
* move submit button to input box
* fix
* add cost tracker
* fix
* fix queued turn cost tracking
* ci(desktop): gate desktop publish secrets behind PublishDesktop environment
The Apple signing/notarization and Tauri updater secrets were repository
secrets, readable by any workflow in the repo and by anyone with push
access via a branch carrying a modified workflow. Move them behind the
PublishDesktop environment, which requires reviewer approval and
restricts deployments to main.
The build job now declares the environment, so those secrets are readable
only there and only after an approval. Add a preflight check because a
missing secret fails dangerously rather than loudly: Tauri silently skips
code signing when APPLE_CERTIFICATE is empty and skips notarization when
APPLE_API_KEY is empty, so a misconfigured environment would still
publish an unsigned, un-notarized bundle. Only a missing updater key was
already caught, by the .sig check in Collect artifacts.
validate stays ungated so a bad tag fails in seconds rather than after an
approval, matching the ungated-build/gated-publish split in
ext-vscode-ab-package. The shared Slack and telemetry secrets stay where
they are; scoping them to this environment would silently empty them in
the CLI, SDK, and extension publish workflows.
* ci(desktop): verify signing secrets are not repository-scoped
The preflight added in the previous commit checks that the signing
secrets are non-empty, which proves presence but not scope, and then
reported that they had resolved from PublishDesktop. An environment-gated
job resolves repository and organization secrets too — environment values
merely take precedence — so a credential left at repository level would
pass that check while the message claimed the migration had worked. This
workflow already demonstrates it: the gated build job reads the shared
Slack and telemetry secrets, none of which are on the environment.
Add the complementary check to validate, which declares no environment: a
signing secret that resolves there can only be repository- or
organization-scoped, so it fails the run and names the offenders. Neither
check establishes provenance alone; together they do. validate is
ungated, so a misplaced secret now fails before the approval rather than
after it.
Also drop the provenance claim from the build message and correct the
skill doc, which stated that a repository-level secret would be invisible
to the gated job.
Reported by greptile on #12854.
Local backends (Ollama especially) intermittently return a turn that
finishes normally but carries no text, reasoning, or tool call. In the
SDK runtime an empty assistant turn is a hard failure ("Model returned
empty response"), so one flaky generation kills the whole task.
Adds a LanguageModelV3 middleware that retries the stream only when a
turn produced genuinely nothing, wired as the outermost middleware on
the Ollama vendor. A tool-call-only turn counts as content and is never
retried; non-empty turns stream through live with no added latency; and
turns that error or hit the token limit are passed through unchanged.
This is the streaming-safe slice of ai-sdk-ollama's reliability story:
its own reliability layer lives in doGenerate and owns the tool loop
(executes tools and force-synthesizes text), which is incompatible with
Cline running its own loop over doStream.
* fix(migration): fall back to the default Cline model for unknown legacy model ids
Some migrated users ended up making Cline provider requests with a model
id the new extension doesn't have because the legacy migration carried
their stored model id over verbatim and never applied a default.
Two small fixes in the provider settings migration:
- Drop a legacy Cline model id the catalog doesn't know so the entry
falls back to the default model instead of carrying the unknown id
into inference requests.
- getDefaultModelForProvider only accepted defaults present in the
generated model block; Cline's generated block holds a few free models
while its declared default (anthropic/claude-sonnet-5) lives in the
collection catalog, so the fallback previously landed on an arbitrary
free model instead of the default.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(migration): validate legacy Cline models against the full runtime catalog
The known-model check used the curated Cline collection plus the tiny
generated cline block, but the runtime Cline catalog is OpenRouter-backed
and also resolves Vercel AI Gateway alias ids. Legacy users on
runtime-served ids outside the curated collection (e.g. the z-ai/glm-5
family) would have been wrongly defaulted. Suffixed variant ids like
...:1m still fall back to the default Cline model.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(migration): validate Cline models against the canonicalized runtime catalog
Greptile review: the raw generated-catalog checks accepted alias ids
(e.g. OpenRouter's z-ai/...) that buildClineModels canonicalizes away
(to zai/...), persisting models absent from the exposed runtime catalog.
Validate against the collection model list (which the runtime catalog
mirrors exactly) and fold alias spellings onto their canonical ids via
the shared VERCEL_OPENROUTER_MODEL_ID_ALIAS_RULES, so legacy z-ai users
keep their model under the canonical id instead of being defaulted.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(llms): retry Ollama response-start timeouts through the AI SDK retry loop
The pre-SDK handler wrapped Ollama chat calls in withRetry({ retryAllErrors:
true }), which silently rode out model cold loads: Ollama holds /api/chat open
while loading and only sends response headers once the model is ready, so the
first attempt of a large model routinely times out at 30s and a later retry
lands on the loaded model. The SDK path lost that behavior twice over: the
response-start timeout rejected with a plain Error (the AI SDK only retries
APICallError with isRetryable), and ai-sdk-ollama wraps every doStream failure
in its own OllamaError, hiding even a correctly-typed error from the retry
predicate. Net effect: one attempt, a surfaced timeout error, and no automatic
recovery - a regression vs the legacy extension for local models that load
slower than the timeout (cline/cline#12829).
Fix: withOllamaResponseTimeout now rejects with APICallError(isRetryable:
true) when its own timer fired (upstream aborts still propagate untouched),
and a restoreOllamaApiCallErrorMiddleware unwraps the buried APICallError from
OllamaError cause chains so streamText's built-in retry (2 retries with
backoff, ~96s of cold-load coverage) engages.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* style: biome format
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* rework: raise Ollama response-start default to 5 minutes instead of retrying
Replaces the APICallError/retry-middleware approach: the 30s guillotine was
the actual root problem (Ollama sends response headers only after the model
cold-loads; killing a healthy request forces error/retry churn), so give the
response-start budget the same order of generosity other AI SDK-based agents
use (opencode: no default header timeout for custom providers, 5 minutes for
its only default) and delete the retry machinery. Unreachable servers still
fail instantly at the connection level, users can still cancel from the UI,
and an explicit requestTimeoutMs is still honored.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* revert Ollama timeout description copy, keep the new default values
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix: forward Ollama request timeout and context window to standalone handlers
Greptile review catch on #12839: buildSdkProviderConfig never carried
requestTimeoutMs, so handlers built via buildApiHandler (commit message
generation) ignored an explicit user timeout — pre-existing, but material now
that the fallback default is 5 minutes. Reuse the session factory's
resolveOllamaProviderConfig so the standalone path honors the configured
timeout and the user's context window (num_ctx) instead of Ollama's 4096
default, keeping the two paths on one source of truth.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
OpenAI Compatible and LiteLLM gated their base URL onChange (and API key
writes via canWrite) on the async provider config having loaded. Text typed
in that window hit a no-op onChange after the debounce cleared the
pending-edit flag, so the late initialValue resync wiped it and nothing was
saved. write() never needed loaded config, and useProviderConfig's request
sequencing already drops the stale initial read, so the guards are removed.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
Sending a message (Enter or send button) cleared the isTextAreaFocused
flag without blurring the textarea. Since the DOM element stayed focused,
onFocus never re-fired (programmatic .focus() on an already-focused
element is a no-op), so the mode-colored outline stayed hidden until a
real blur/refocus cycle - which is why toggling Plan/Act mode brought it
back. Stop clearing the flag on send; blur is already handled by the
onBlur handler.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
Editing a previous message while a tool approval prompt was pending left the
old session's approval promise parked forever: the superseded run stayed
suspended awaiting an answer that could never come, and the stale resolver
kept intercepting later ask responses. Clear pending interactions before
starting the replacement session, exactly like cancelTask / clearTask /
task-switch / mode-change already do.
Ref #12827
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
The runtime baseUrlMap in resolveBaseUrl lacked the asksage ->
asksageApiUrl mapping (present in store.ts, effective-config.ts, and the
legacy migration), so a custom AskSage API URL saved in legacy state was
never read and requests fell through to the builtin default
https://api.asksage.ai/server.
Also write the URL through the SDK provider-config store in
AskSageProvider.tsx (mirroring AnthropicProvider) so providers.json
stays in sync for CLI/desktop hosts; the store mirrors baseUrl back to
the legacy asksageApiUrl state key, keeping the /get-models fetch and
legacy readers working.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
Convert the Qwen and Moonshot regional API line dropdowns from
legacy-state-only writes to useProviderConfig().write({ apiLine }),
matching the Z AI pattern. The host store mirrors the write back to the
legacy qwenApiLine/moonshotApiLine state keys, so a single write keeps
providers.json (read by the CLI and desktop app) and the legacy
StateManager (read by the VS Code session factory) in sync.
Adds store tests pinning the dual-write mirroring and webview component
tests for the dropdowns' write and display behavior.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(vscode): include untracked files in commit message generation
getGitDiff only ran git diff --staged and git diff HEAD, neither of which reports untracked files, so an add-only working tree failed with 'No changes in workspace for commit message'. Gather untracked files and diff each against /dev/null via execFile (argv, no shell) so add-only trees work and special-char filenames are safe.
Closes#12060
* fix(vscode): include untracked files alongside tracked changes
Address review: append untracked-file diffs in the non-staged path instead of gating on an empty diff, so a mix of edited tracked files and new untracked files includes both. Re-throw git exit codes other than 1 (files differ) so real errors aren't swallowed. Use a named, non-runnable label for the output header. Adds a mixed tracked+untracked test.
Refs #12060
---------
Co-authored-by: Minhkunn <minh.12072k6@gmail.com>
#12831 removed the import while rewriting checkpoint restore, and #12830
landed on top of it adding a usage that assumed the import was still
there. apps/cli typecheck has failed on main since, which fails the
sdk-test Quality Checks job on every PR touching sdk/**.
Co-authored-by: Cline Agent <cline-agent@users.noreply.github.com>
After a checkpoint restore (/undo or Esc Esc), the rewound user message is
dropped into the input box to edit and re-send. It was prefilled from the
raw stored text, which the runtime wraps in a <user_input mode="...">
envelope, so the input showed '<user_input mode="act">...</user_input>'
instead of what the user typed. Prefill the display form via
formatDisplayUserInput (already used for the picker preview and imported in
this file), which strips the envelope and preserves slash-command display
form.
* fix(core): create checkpoints reliably across hosts, restarts, and compaction
#12691 moved checkpoint run-boundary detection into a beforeRun hook that
recorded snapshot.messages.length, assuming the run's user prompt is
appended afterwards. SessionRuntime (VS Code + CLI) instead seeds the
prompt into initialMessages and calls run(""), so the beforeRun delta is
always empty and no checkpoints were ever created in either surface.
Gate checkpoint creation on two signals instead of the fragile in-memory
delta alone:
- introducedUserRun: the beforeRun delta contains a new user turn. Covers
hosts that pass the prompt as run input and refreshes the entry on
edit-and-regenerate.
- alreadyCheckpointed: the run count already exists in the DURABLE session
checkpoint history. Covers the seeded-prompt path and, unlike an
in-memory counter, still holds after a process restart.
Skip only when neither applies (a continuation/resumption re-running an
already-checkpointed run), so a reopened session can't overwrite a good
pre-run snapshot with the mutated workspace. Run numbering uses the
span-aware countUserRunMessages so it survives compaction folding turns
into one summary message.
Adds regression tests for the seeded-prompt creation, the reopen-without-
new-turn overwrite case, and the first-turn-after-compaction case.
* fix(cli): number /undo checkpoints span-aware so restore can map them
The interactive /undo picker counted every role="user" message when
assigning run numbers to checkpoints. Tool-result messages also carry
role "user", so any turn that used tools got an inflated run number; the
picker then handed that number to the core, whose span-aware
findUserRunMessage could not map it and aborted with 'Could not find user
message for run N'. Restore was effectively unusable whenever the agent
called a tool.
Count runs with the core's getUserRunSpan (tool results contribute 0, a
compaction summary spans the turns it folded) so the picker's run numbers
match what the core records and resolves. Extracted the item-building into
a pure buildCheckpointPickerItems helper with unit coverage for the
tool-result and compaction cases.
* fix(core): capture untracked files in checkpoints as a third parent
Checkpoint creation used plain `git stash create`, which cannot include
untracked files (no -u support). Restore therefore had no way to bring back
a file Cline created during a task, so a full rewind was impossible.
Synthesize a stash-shaped snapshot commit that also records untracked,
non-ignored files as a third parent - exactly like
`git stash create --include-untracked` - without touching the working tree,
the real index, or the stash list: list `ls-files --others
--exclude-standard`, stage into a temp GIT_INDEX_FILE, write-tree +
commit-tree to get the untracked parent, then rebuild the stash commit with
that extra parent. When the tracked worktree is clean but untracked files
exist, synthesize the stash from HEAD so they are still captured instead of
falling back to a bare HEAD-commit checkpoint. Fully clean worktrees still
use the HEAD-commit fallback.
* fix(core): full workspace rewind on restore for snapshot checkpoints
Restore now rewinds untracked files generation-aware:
- If the checkpoint carries an untracked third parent (a snapshot from
createWorktreeStashCommit), do a full rewind: reset tracked to the base,
`git clean -fd` to drop files created after the checkpoint (and clear the
worktree so `stash apply` cannot hit an "already exists" conflict), then
`git stash apply`, which restores each captured untracked file to its
checkpoint-time content from the third parent. `git clean -fd` (no -x)
leaves .gitignored paths - build output, node_modules, .env - alone. This
is safe because everything removed is either recreated from ^3 or postdates
the checkpoint, and the pre-restore recovery snapshot (stash push
--include-untracked) can roll the whole operation back.
- If the checkpoint has no third parent (legacy 2-parent stashes and
HEAD-commit fallbacks from before capture existed), keep the conservative
behavior: never touch untracked files, since nothing can reconstruct them.
This makes 'Reset Code' / '/undo' a true rewind: a file Cline created in an
early turn and ruined later comes back to the early-turn version.
* fix(telemetry): single classified emitter for provider API errors, gated on terminal failures
* chore: remove explanatory comment block from agent-events.ts
* feat(telemetry): stamp terminal=true on SDK provider failure events
* chore: remove dead notice api_error capture (no producer emits that reason)
* refactor(telemetry): rename provider-failure 'terminal' flag to 'fatal' (terminal is the shell in Cline)
* refactor(telemetry): drop the fatal flag - only user-surfaced failures are reported on both bundles
* fix(cli): don't let the ClinePass promo dialog trap users whose terminal drops Esc
The promo dialog could only be dismissed with Escape, and Esc is the
least reliable key across terminals: it arrives as a bare \x1b that
needs timeout disambiguation, and Bun's Windows console input layer is
known to swallow it (Windows PowerShell users reported being unable to
dismiss the dialog at all). Worse, the 'shown' marker was only written
when the dialog closed, so a user who force-quit saw the promo again on
every launch.
- Any key other than Enter now dismisses the dialog (Enter still opens
the subscription page)
- The shown marker is persisted when the dialog is displayed, not when
it is dismissed, so a force-quit never loops the promo
- Add a tuistory e2e test covering marker timing and any-key dismissal
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(cli): let any key cancel the OAuth waiting screen
Like the ClinePass promo, the OAuth wait screen was dismissible only
with Esc (plus K for the API-key fallback when offered) while blocking
on a browser flow that may never complete — a trap on terminals that
drop Esc. Any key other than K now cancels the pending auth attempt;
K still switches to manual API key entry when available.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(cli): don't let a modifier keypress dismiss link-bearing dialogs
The ClinePass promo and OAuth wait screens both render a URL the user
opens by holding Cmd/Ctrl and clicking. With 'any key closes', that
modifier keystroke could tear the dialog out from under the click. Add
a shared isAnyKeyDismiss() guard so only unmodified keys dismiss; keys
held with ctrl/meta/super/hyper (and bare modifier presses) are ignored.
Enter still opens the promo and K still opens manual API key entry.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* revert(cli): persist promo shown-marker on dismiss again
Now that any key dismisses the promo, users can reliably close it, so
there's no need to write the shown-marker eagerly on display. Restore
persisting it in the dialog's finally() and update the e2e assertion.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
Follow-up to #12782, which replaced the open package with openUrlInBrowser
but missed two call sites and dropped some platform handling the package
provided:
- Migrate the two remaining open users (skills marketplace open in
tui/root.tsx and ACP OAuth in acp/auth.ts) to openUrlInBrowser; the
listenerless-child crash fixed by #12782 was still reachable there.
- Treat containers running on a WSL2 kernel (Docker Desktop for Windows,
devcontainers) as plain Linux: /proc/version says microsoft but there is
no Windows interop, so use xdg-open instead of powershell.exe (matches
the is-inside-container check open@10 performed).
- Try opener candidates in order: on WSL, powershell.exe on PATH, then the
absolute /mnt/c/... path (covers appendWindowsPath=false), then xdg-open
(sandboxed WSL with WSLg); on win32, the %SystemRoot% absolute PowerShell
path first (what open@10 used), then PATH lookup.
- Convert Linux file paths to \\wsl$ UNC paths via wslpath before handing
them to Start-Process, so 'cline doctor log' works on WSL.
- Remove the now-unused open dependency from apps/cli.
captureDiffEditFailure and captureWorkspaceInitError have no callers: SDK core
is the sole emitter of task.diff_edit_failed and workspace.init_error. Keeping
callable host-side capture APIs for core-owned events is how the
task.provider_api_error double-emission happened — a future host caller would
silently double-count these events with no type error or failing test. Also
drops the two event-name constants, which were only referenced by the removed
methods.
* feat(core): resolve display-ready names in fetchClineRecommendedModels
* refactor(cli,vscode): render recommended-model names from the enriched feed
* fix(core): resolve catalog names through vercel/openrouter id aliases
* fix(core): share one timeout budget across the feed and catalog lookups
Greptile flagged that resolveDisplayNames started a fresh timeoutMs window
after the recommendation request finished, so a slow endpoint plus a cold
or hung catalog could keep the picker loading for ~2x the timeout. The
catalog race now gets only the budget remaining from a single deadline;
an already-cached catalog still applies on an exhausted budget because
its promise resolves ahead of the zero-delay timer.
* fix(cli): strip the Slack bot mention from incoming connector messages
Slack delivers an at-mention of the app as `<@U0B8E8H3U1F> hi`, and the chat
SDK deliberately leaves the bot's own mention unresolved so mention detection
keeps working - flattening it to `@U0B8E8H3U1F hi`. The connector forwarded
that verbatim, so the agent saw the raw bot id at the front of every
mention-triggered turn.
Strip the leading self-mention in onNewMention/onSubscribedMessage before the
approval-reply check and handleTurn, resolving the bot id from the adapter
(request-scoped in multi-workspace mode) with a fallback to the event envelope
authorizations. Mentions of other users and inline mentions are preserved, and
a bare mention is left as-is so the turn is not dropped as empty input.
* fix(cli): only strip a complete Slack bot mention, not an id prefix
The `<@ID>` and `<@ID|name>` alternatives in stripSlackBotMention are
terminated by `>`, but the SDK-flattened bare `@ID` alternative had no
trailing boundary, so it also matched the start of a longer id. With bot id
`U123`, a message addressed to a different user - `@U1234 help` - was
rewritten to `4 help`, corrupting both the approval-reply check and the text
handed to the agent.
Require the flattened alternative to be followed by a non-id character with a
`(?![A-Za-z0-9])` lookahead, so it only matches a complete Slack id. A plain
`\b` cannot express this, because Slack ids end in word characters and `\b`
still matches between `U123` and `4`.
Existing behaviour is unchanged: angle-bracket and flattened self-mentions are
still stripped, repeated leading mentions still collapse, trailing `[\s,:]`
separators are still consumed, other users' and inline mentions are preserved,
and a bare mention is still left untouched so the turn is not dropped as empty.
Adds regression tests for the prefix collision, which fail against the previous
regex and pass with this one.
---------
Co-authored-by: cline-test-bot <cline-test-bot@users.noreply.github.com>
* fix: surface upstream provider error from gateway-forwarded stream failures
Vercel AI Gateway streams upstream rejections (e.g. Alibaba Qwen context-
length errors) wrapped in its own parse failure: the top-level message is
just 'Stream error occurred' and the cause is an internal ZodError, while
the real rejection is JSON-encoded in value.error_message. Unwrap it so
users see 'This model's maximum context length is 40960 tokens...' instead
of a raw Zod issue dump.
Also fall back to JSON.stringify for opaque object errors so the UI never
renders '[object Object]'.
* refactor(llms): use shared safe-JSON helpers and a named type guard in extractErrorMessage
* fix(llms): resolve OpenRouter display names for all Cline free models
* fix(vscode): resolve featured model card display names from the provider catalog
* fix(vscode): fall back to endpoint-provided names on featured model cards
* Add tuistory-based TUI e2e harness for the CLI
Evaluates https://github.com/remorses/tuistory as a Playwright-style
driver for the interactive TUI. Adds:
- tuistory devDependency in apps/cli
- test:e2e:tuistory script + vitest.tuistory.e2e.config.ts
- src/cli.tuistory.e2e.test.ts: ports the script(1)-based interactive
smoke tests to reactive waitForText/screen-state assertions against a
real PTY + Ghostty terminal emulator (5 tests, ~11s, no fixed sleeps)
- DEVELOPMENT.md docs for the vitest suite and the tuistory session CLI
agents can use to manually drive the TUI headlessly
* Add tuistory agent skill (.cline/skills, symlinked to .claude/.agents)
Teaches coding agents to drive the Cline TUI headlessly via tuistory
sessions (launch with isolated env, reactive wait, snapshot/screenshot,
observe-act-observe loop) and to write launchTerminal()-based e2e tests,
closing the loop for cloud agents testing apps/cli.
* fix(cli): don't crash on browser-open failure when no opener binary exists
open() with { wait: false } resolves to the detached child process before
the opener binary is known to exist. On hosts without one (e.g. xdg-open
on headless Linux), the failure arrives as an async 'error' event on the
listenerless child, escalating to an uncaughtException that kills the CLI
— bypassing every try/catch and .catch() at the call sites. Hitting
"Sign in with Cline" from the welcome screen reliably crashed the TUI in
containers.
Route all browser opens through a shared openUrlInBrowser() helper that
attaches the error listener and reports failure via its returned promise,
so flows fall back to their existing "visit the URL below" messaging.
* fix(cli): attach opener error listeners in the same tick as spawn
Greptile's review caught that the helper attached its listeners only after
awaiting open()'s promise. Empirically that window is safe under Node 22
(the listener wins) but real under Bun — the runtime the compiled CLI
ships on — where the missing-binary ENOENT 'error' event fires before the
microtask queue drains, reproducing the exact crash this helper exists to
prevent.
macOS and non-WSL Linux now spawn their opener (open / xdg-open) directly
with listeners attached in the same synchronous tick, which both runtimes
guarantee can never miss the event. Windows and WSL keep delegating to the
open package for its shell quoting and interop routing; their openers
(cmd/powershell) always exist, so the post-await path cannot hit ENOENT.
The regression test now emits the error on nextTick — before microtasks —
which fails against the previous implementation.
* fix(cli): drop the open package — same-tick opener spawn on every platform
The win32/WSL delegate path still attached listeners after awaiting
open()'s promise, leaving a narrow uncaught-error window under Bun for
emittable spawn failures (e.g. AV-blocked EPERM). Spawn the opener
directly everywhere instead: open on macOS, xdg-open on Linux, and
powershell -EncodedCommand on Windows/WSL — the base64-encoded
Start-Process command sidesteps cmd/PowerShell quoting of URLs entirely,
so nothing is ever shell-interpolated.
On CLI exit, renderer.destroy() runs root.destroyRecursively() before
React flushes the DialogProvider's passive unmount cleanup, so the
dialog container is already detached when the cleanup calls
renderer.root.remove(container), triggering OpenTUI's 'Renderable with
id dialog-container is not a child of __root__, skipping remove'
warning. Drop the explicit remove from the patched @opentui-ui/dialog
provider cleanup (react + solid): Renderable.destroy() already detaches
from its parent when attached and no-ops when already destroyed.
basename("/") is an empty string, which WorkspaceInfoSchema rejects
(hint is z.string().min(1).optional()), so upsertWorkspaceInfo threw a
ZodError for any session rooted at the filesystem root — e.g. the
desktop app launched from the Dock with cwd "/" — and commands never
ran. Omit the hint instead of storing an empty string.
Launch the hub through CLINE_WRAPPER_PATH after Unix self-updates so npm 12 does not reuse a deleted cached executable. Preserve the in-process fallback for Windows and development builds, and add coverage for success and failure paths.
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
The uid/mcpServerKeys registry existed to encode server names into
native tool-call function names and decode them back at dispatch.
That encode/decode path was removed with the extension host
(c4c126bee): tool names are now built by the SDK's deterministic
defaultMcpToolNameTransform and execution closes over the server
name directly, so getMcpServerByKey has no callers and the keys are
write-only state. Delete the registry, the uid field, and the
deleteServerKey callback plumbing.
Co-authored-by: Cline Agent <cline-agent@users.noreply.github.com>
BannerService tests still used 10ms sleeps for background fetch completion.
On slow CI runners that races mocha timeouts. drainForTesting() already
exists and awaits the in-flight fetch promise deterministically.
Rebased onto monorepo main (apps/vscode path).
Signed-off-by: Sebastien Tardif <sebtardif@ncf.ca>
* ci(vscode): gate the combined A/B package workflow on both bundles' test suites
* docs(skills): add publish-extension skill for VS Code extension releases
* ci(vscode): pin tested revision for next bundle and refuse publishing untested next-refs
* ci(vscode): pin legacy bundle to the revision its test gate ran against
* fix(vscode): show migrated model in settings instead of hardcoded default
After the SDK provider migration a user who never explicitly picked a model
(i.e. took the legacy default) ends up with the model recorded in
providers.json but not in the mode-specific globalState fields the settings
picker reads. The OpenRouter picker and its info card then fell back to the
hardcoded openRouterDefaultModelId (claude-sonnet-4.5) and its pricing, while
the extension actually ran the migrated model (claude-sonnet-5).
- resolveModelInfo: when no model id is requested, honor the provider store's
committed selection (which reads providers.json when the state field is
empty) before substituting a catalog default.
- OpenRouterModelPicker: source the displayed model id/info from the
authoritative resolver as the fallback when the mode fields are empty,
instead of the hardcoded constant. Committed-field users are unaffected.
* fix(vscode): guard picker model info against resolver default substitution
Review hardening: the resolver substitutes its provider default for ids it
cannot resolve, so only trust its info when it answered for the id actually
displayed. Prefer the live catalog entry for the displayed id (synchronous
once fetched, which also removes the transient placeholder while the resolver
is in flight), and never render another model's metadata under the displayed
model's name. Also document why the act-then-plan readSelection order in the
empty-id branch cannot misattribute a mode-specific selection.
* Restore task export to markdown and show download button in all builds
* Render untyped tool outputs and object tool inputs as JSON in task export
* Open the task's SDK session folder from the export button and show it in all builds
* Keep the task header session-folder button dev-only
* Fix queued prompt row alignment and auto-scroll on queue
Center the dot, badges, and cancel button on the first text line of each queued prompt row (the X previously sat ~3px below the text), and re-pin the chat view to the bottom when a prompt is queued so the queue banner doesn't cover the end of the conversation.
* Don't treat task switches as queue growth for auto-scroll
Guard the queued-prompt auto-scroll effect on the displayed task's ts: switching to a task that already has queued prompts grows the count without a send from this webview, and should not hijack the newly opened conversation's scroll position.
* feat(desktop): support message editing & checkpoints
Fork sessions before a selected user run, trim checkpoint history, and restore prior messages so prompts can be edited safely. Update the chat UI and tool activity panels to support the editing flow and preserve horizontal scrolling for long content.
* fix(desktop): restore checkpoints when editing messages
* fix(core): infer kindless checkpoint types
* fix(core): preserve checkpoint run numbering
* fix(desktop): make message edit restores transactional
* fix(desktop): make checkpoint restores workspace-atomic
The confirmation that appears when clicking the compact button in the
task header was a bare unstyled row with a stray bottom margin (my-2)
that stacked on the header card's own bottom padding, leaving a dead
gap under the buttons. It is now a distinct bordered card (editor
background against the header's toolbar surface) with a title, a short
description of what compacting does, and right-aligned Cancel/Compact
buttons, with symmetric spacing above and below.
Also drops the ContextWindow wrapper's bottom margin (my-1.5 -> mt-1.5)
so the row's bottom spacing matches the header padding, and adds a
ContextWindow Storybook story that mirrors the expanded TaskHeader
surface so the confirmation can be previewed in isolation.
Selecting ClinePass in settings awaited a network round-trip (PUT
/active-account + possible token refresh) before postStateToWebview,
so the settings panel stayed on the previous provider until the
request finished. Make the personal-account switch fire-and-forget:
it was already best-effort, and auth state changes propagate to the
webview separately once it completes.
Also convert the helper's test to bun:test so it actually runs (the
mocha version was excluded by both the bun unit runner and the
vscode-test glob) and fix its stale null-vs-undefined assertion from
the SDK migration.
The xAI, Z AI, and Moonshot settings components were never wired to the
catalog's reasoning capability: xAI only offered a legacy low/high
checkbox hardcoded to grok-3-mini model ids, and Z AI / Moonshot had no
reasoning control at all, even though models.dev marks grok-4.5, glm-5,
kimi-k2-thinking, etc. as reasoning models. Every catalog-driven
provider (GenericProviderSettings, OpenRouter/Vercel/Requesty pickers)
already gates ReasoningEffortSelector on supportsReasoning.
Render the shared ReasoningEffortSelector in these three components when
the selected model's catalog info advertises reasoning, persisting the
choice to the provider config the same way GenericProviderSettings does.
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
* fix(vscode): stop queued-prompt turns from getting stuck on Thinking
When the SDK drains a queued prompt at the end of a turn, the new turn's
pending_prompt_submitted bookkeeping (isRunning=true, phase=streaming) always
runs before the previous turn's send promise unwinds in fireAndForgetSend.
That .then then unconditionally called setRunning(false), so the queued turn
ran with isRunning=false and its own turn-complete was mistaken for a
cancelled-turn straggler - the phase never left "streaming" and the chat
showed an endless Thinking indicator.
Track a monotonic turn epoch on SdkSessionLifecycle: immediate sends and
drained queue prompts bump it, and both the send-settled callbacks and the
event coordinator's turn-end handling skip their bookkeeping when a newer
turn has started since (covers the symmetric interleaving where the done
handler resumes after the drain and would clobber the queued turn's
streaming phase).
* Simplify: preserve only an actual cancel phase in the turn-complete straggler guard
Replaces the turn-epoch machinery with the minimal fix: the straggler
guard's intent is to preserve the cancel-set "resumable" phase, so key it
on the phase itself instead of the isRunning proxy. When the SDK drains a
queued prompt at turn end, the previous turn's send promise settles after
the queued turn already started and flips isRunning back to false
mid-turn; with the old guard the queued turn's real completion was then
mistaken for a cancel straggler and the phase stayed stuck on
"streaming" (endless Thinking). Checking for "resumable" lets that
completion resolve the terminal phase normally while cancel behavior is
unchanged.
The anti-flash grace period (added to stop the loader flashing at turn
end) also fired mid-turn, causing a visible hide/show/hide flicker right
before a tool row appeared:
- When a reasoning tail finalized while the turn kept streaming, the
reasoning shimmer collapsed, the loader stayed hidden for the 500ms
grace, popped in, then hid again when the tool row landed. Reasoning
never ends a turn, so skip the grace for reasoning tails and hand the
shimmer straight to the loader.
- When the loader was already visible below a streaming tool group, the
group tail finalizing blinked it off for the grace period. The grace
now only delays hidden -> visible transitions, never hides an
already-visible loader.
* Show user message immediately when sending to a history-resumed task
Sending a message to a task opened from history routed through the
resume_task/resume_completed_task askResponse branch, which forced the
Thinking loader but never set the optimistic user_feedback bubble. The
extension only echoes the user's message after the full SDK session
resume completes, so the chat showed a Thinking indicator with no user
message until the (slow) resume finished.
Pass showPendingMessage on the resume branch like the other
non-streaming follow-up paths, so the user's message appears in the
chat immediately. The optimistic bubble reconciles with the extension's
say:user_feedback echo once the resume completes (identical raw text).
* Add changeset
* fix(core): add a plugin telemetry bridge
* fix(core): address plugin telemetry bridge review feedback
- Sanitization fallback now covers the whole executeTool IPC payload:
`input` can be rewritten by beforeTool hooks or programmatic callers,
so a non-serializable input degrades gracefully like the context does.
- The sandbox only offers ctx.telemetry when the host actually has a
telemetry service (new PluginSandboxOptions.telemetryAvailable, derived
from options.telemetry in the config loader), so feature-detecting
ctx.telemetry means "someone is listening" in both execution modes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(core): plugin telemetry review round 2 — timer leak and setup-time fallback
- SubprocessSandbox.call: a synchronous child.send() throw (cyclic payload)
left the pending timeout timer armed; it later fired and shut the sandbox
down, killing unrelated in-flight calls. Cancel the pending entry and
reject with the original error so serialization failures stay classifiable.
- plugin_telemetry events emitted during plugin setup() arrive before the
session is registered, so the session-config lookup missed and setup-time
telemetry was silently dropped. Route through a fallback telemetry service
(extensionContext/local config/host default), mirroring handlePluginLog's
fallback logger.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix(core): classify BigInt IPC serialization errors for the sandbox fallback
Bun ("cannot serialize BigInt") and Node ("Do not know how to serialize a
BigInt") raise messages that did not match the cyclic/circular predicate, so
a bigint smuggled into tool input or context by a hook or programmatic caller
rethrew instead of retrying with the JSON-safe clone — which already drops
bigint leaves.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Auto-compaction state was silently rejected on every save ("Skipped
stale session compaction state"), forcing a full re-compaction — an
extra summarizer LLM call — on every turn past the trigger, and a
resume-time identity churn could leave a dead sidecar permanently
blocking replacements.
Three changes:
1. Stop hashing volatile transport identity. The source-prefix hash no
longer includes message id/ts, which the codec regenerates on every
wire/storage round-trip (a store's just-appended user turn has none
yet; consolidated parallel tool results are re-split with minted ids
on resume). The fingerprint now covers role, content, and durable
metadata. Hash seed bumped to v2; v1 sidecars fail projection once
and are replaced by the next compaction.
2. Validate persists against the exact source messages the state was
computed over. createCompactionStateAwarePrepareTurn passes
context.messages to saveState, and the local runtime host threads
them into persistActiveSessionCompactionState instead of falling
back to the conversation store's mid-turn shape.
3. Scope the count-based stale-write guard to states that still
project. An unprojectable current state no longer blocks a
newer-timestamped replacement, so invalidated sidecars self-heal
instead of deadlocking the session.
All three regression tests fail on main and pass with this change.
* fix(vscode): mark onboarding complete only after OAuth succeeds
The onboarding webview marked welcomeViewCompleted immediately after the
sign-in URL opened (accountLoginClicked resolves at URL-open time), so
Free/Frontier/ClinePass signups landed in chat signed out when the user
abandoned or failed browser auth, and the flag persisted across reloads.
Restore the classic extension behavior: the host (SdkAuthService) now
sets welcomeViewCompleted after the OAuth token exchange succeeds, in
createAuthRequest, handleAuthCallback, and the E2E mock login. The
webview persists the model selection up front, stays on the 'Almost
there!' step until auth completes, and fires the 'completed' funnel
event via a pending-intent module once clineUser arrives (mirroring the
pendingClinePassSubscribe pattern). This also fixes the legacy
WelcomeView fallback, whose 'Get Started for Free' never completed
onboarding after login.
* refactor(vscode): slim the onboarding-completion fix to its essentials
Drop the pending-telemetry module and App hook (the 'completed' funnel
event keeps its existing main-branch semantics, firing when the flow is
initiated, so no telemetry change in this PR), restore finishOnboarding
to its original shape with just a markCompleted parameter, and reduce
the host helper to a single setGlobalState call.
* Post streaming turn state to webview before session startup
The webview only learns the turn phase through full state posts, and the
first post after initTask happened only after startNewSession settled —
so the chat mounted with a stale idle TurnState and the thinking
indicator popped in noticeably late. Ship a state post right after the
initial task message is emitted, in parallel with session startup.
* Show thinking indicator optimistically on new-task submit
Capture the TurnState seq at the moment the newTask RPC is sent and
force the in-list Thinking loader row until a fresher TurnState arrives
(any phase), so the indicator renders together with the task message
instead of waiting for the streaming TurnState to round-trip. Rolled
back if the RPC fails; legacy (no turnState) hands off to the existing
tail heuristic once the task message lands.
* Paint the initial Thinking loader without waiting for Virtuoso
Frame-by-frame measurement showed the loader decision was true on the
chat view's first paint, but the synthetic in-list row still appeared
~150-200ms later: a cold-mounting virtualized list needs several frames
to measure and paint its first item. When the list has no visible rows
yet (new task just submitted), render the waiting row as a plain
element over the (empty) list instead; once any real row exists the
warm list takes over with the in-list row as before.
* Show thinking indicator immediately for follow-up messages too
Follow-ups had the same delay as new tasks: SdkController.askResponse
moves the phase to streaming but never posted state, so the webview
kept the stale terminal phase (hiding the loader) until the new turn's
first session event posted state. Post right after the phase change,
and generalize the webview's optimistic marker from new-task-only to
any turn-starting send (askResponse outside a streaming phase), with a
guard that never shows the loader while a content row is actively
streaming. Renames pendingNewTaskSeq to pendingTurnStartSeq.
* fix(vscode): render thinking loader synchronously
* fix(vscode): clear chat input immediately when /compact is submitted
* fix(vscode): let the compaction divider label wrap at narrow widths
* fix(vscode): update context-window header even when compaction grows the context
* chore: add changeset for /compact UX fixes
* docs(vscode): align getLastApiReqTotalTokens return doc with unclamped rescale
The desktop chat integration test renders components from @cline/ui, but
it also pulls @cline/shared/browser through the desktop app's own
message-content module. That subpath resolves to dist output no step in
this job produced, so the suite failed to collect.
Build @cline/shared before the test, and install the full workspace: the
two-package filter did not provide enough of the tree for that build.
The ui-publish workflow installs only the @cline/ui and @cline/code
workspaces, so the root devDependencies that previously supplied the
'bun' and 'node' type roots were absent and tsc failed with TS2688.
Declare them on the package that requires them in its tsconfig types.
Also refreshes the stale @cline/code version recorded in bun.lock.
* feat(vscode): enable Auto Compact by default
The SDK-based extension has no fallback context management: with auto
compact off, hitting the model's context window fails the request with a
provider error and retrying keeps failing (the legacy extension truncated
the oldest half of the conversation in this situation). The CLI already
defaults compaction on (agentic); align the extension with it.
* chore: add changeset for Auto Compact default-on
* fix(models): tolerate null contextWindow/maxTokens in SDK catalog shapes
Live LiteLLM proxies report unknown model limits as explicit nulls in
/model/info (e.g. max_tokens: null). adaptSdkModelInfo only tolerated
undefined, so a single such model failed the entire catalog refresh and
left the model picker empty. Treat null like a missing value (matching
the existing pricing handling) and fall back to the safe defaults.
* Update apps/vscode/src/sdk/model-catalog/shape-adapter.ts
Co-authored-by: Tomás Barreiro <52393857+BarreiroT@users.noreply.github.com>
* Update apps/vscode/src/sdk/model-catalog/shape-adapter.ts
Co-authored-by: Tomás Barreiro <52393857+BarreiroT@users.noreply.github.com>
* fix(models): restore missing limit fallbacks
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Tomás Barreiro <52393857+BarreiroT@users.noreply.github.com>
* fix(cli): open history in the existing TUI
* refactor(cli): clarify history TUI startup target
* feat(cli): add history actions to TUI
* fix(cli): avoid empty session when resuming history
* fix(cli): fail history delete without session id
* fix(cli): dispatch resume hook from history picker
* fix(vscode): show per-file diff for multi-file apply_patch
apply_patch edits to multiple files rendered the entire multi-file patch in every per-file diff row. Split the patch into one tool message per file at content_end (mirroring the read_files split) so each row shows only that file's changes.
Closes#9904
* fix(vscode): address review on multi-file apply_patch split
Import the canonical PATCH_MARKERS from @cline/core instead of the local AP_MARKERS duplicate and export it through the core barrel. The cross-world import barrier the old comment claimed does not exist - apps/vscode already imports runtime values from @cline/core.
Route the apply_patch branch in sdkToolToClineSayTool through getApplyPatchString so the streaming and finalized rows derive their content from one source.
Handle the bare-string apply_patch input. ApplyPatchInputUnionSchema accepts { input: string } | string; a bare two-file patch made getApplyPatchString return undefined, so both content_start and content_end produced one empty-path row instead of the per-file split. Return the raw string when the field lookup finds nothing, with a start/end reconciliation test.
Refs #9904
---------
Co-authored-by: Minhkunn <minh.12072k6@gmail.com>
* fix(connectors): recover Slack thread mapping when session is gone
A connector thread binding can outlive its runtime session (hub restart,
session abort, retention cleanup). When that happened the thread stayed
pinned to a dead session id and every subsequent turn failed with
`session_not_found`, so the bot replied "Slack bridge error: session not
found" forever with no way to recover short of editing threads.json.
Drop the stale binding and replay the turn once against a brand new
session. Both the normal turn path and the steering path are covered.
Adds forgetThreadSession() to session-runtime and 3 regression tests.
* fix(connectors): serialize stale session recovery
---------
Co-authored-by: cline-test-bot <cline-test-bot@users.noreply.github.com>
Co-authored-by: abeatrix <beatrix@cline.bot>
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
* Consolidate per-provider model refresh handlers into the SDK
The VS Code extension resolved model catalogs from two sources: the SDK
catalog (models.dev-backed) used by resolveModelInfo/task header, and
host-side refresh handlers (refreshOpenRouterModels & co.) used by the
settings pickers. This dual-source split produced inconsistencies like
ENG-2345.
SDK (@cline/core):
- New rich live model sources (live-model-sources.ts) ported from the
extension handlers: OpenRouter (pricing incl. cache read/write,
descriptions, image support, thinking config, tiers/global-endpoint
metadata, curated overrides, stealth models), Vercel AI Gateway, and
Hugging Face. Keyed by generated catalog key so cline shares
OpenRouter's live data.
- mergeKnownModels layers rich live entries field-wise on top of the
curated catalog (live fields win, curated fields fill gaps) instead of
the modelsSourceUrl replace semantics.
- New Groq and Requesty private fetchers (API-key gated); Baseten
private fetcher now parses live pricing and reasoning support and is
enriched from the curated catalog.
Extension (apps/vscode):
- refreshOpenRouterModels/Groq/Baseten/VercelAiGateway/HuggingFace/
Hicap/Requesty are now thin delegates over the SDK provider catalog;
all bespoke fetch/parse/disk-cache code is deleted.
- shape-adapter maps the SDK's thinkingConfig, temperature,
global-endpoint capability, and metadata tiers onto the extension
ModelInfo.
- Removed the now-unused StateManager models cache, per-provider disk
cache files, and the dead readOpenRouterModels stub.
Fixes ENG-2381.
* Simplify: rely on the SDK's models.dev catalog, no rich live sources
Drop the ported per-provider live fetchers and curated overrides
(live-model-sources.ts) and all SDK merge changes. The extension now does
exactly what the CLI does: refresh handlers resolve through
resolveProviderConfig, which serves the models.dev-backed catalog
(bundled + runtime live refresh) plus the SDK's pre-existing
authenticated fetchers (Baseten/Hicap/LiteLLM/Poolside). No hardcoded
model info or per-model pricing workarounds remain anywhere.
Also reverts the shape-adapter additions since no SDK catalog source
populates thinkingConfig/temperature/metadata tiers today.
* Replace thinking-budget sliders with catalog-driven reasoning effort selection
Match the CLI's UX: every reasoning-capable model (SDK catalog
'reasoning' capability -> supportsReasoning) gets the Reasoning Effort
selector (none/low/medium/high/xhigh); the legacy 'Enable thinking' +
budget-tokens slider is removed everywhere, along with the hardcoded
per-provider thinking-model id lists (Anthropic, Claude Code, Bedrock,
Qwen) and claude/grok model-id heuristics in the OpenRouter, Vercel,
and Requesty pickers.
Effort changes now dual-write the provider-config reasoning settings
({enabled, effort}) that the session factory actually consumes - the
budget slider wrote legacy plan/act thinkingBudgetTokens state that
sessions already ignored. The utility request path
(buildSdkProviderConfig) drops its budget preference and forwards
effort only; the SDK translates effort into each provider's wire
format (including budget-token mapping where required).
* Gate picker reasoning-effort UI on live catalog entries
The OpenRouter/Vercel/Requesty pickers read the committed legacy
model-info snapshot, which provider-config writes can clear when a
resolution lands on a fallback source - selecting an effort made the
selector disappear. Gate on the live catalog map (with snapshot
fallback) instead; Requesty gates on the catalog only, since its
safe-default fallback over-reports reasoning support.
* Address review: honor legacy thinking budgets, dedupe refresh handlers
- Persisted thinking budgets are honored again (greptile P1 / review
request): normalizeProviderReasoningSettings maps a stored
reasoning.budgetTokens (written by older versions or the SDK's
legacy-state migration) onto the effort scale and treats it as
thinking-on, and buildSdkProviderConfig derives an effort from the
legacy plan/act budget fields when no explicit effort exists. An
explicit 'none' still wins. Shared mapping lives in
reasoningEffortFromThinkingBudget with low/medium/high buckets.
- Extract resolveProviderModelsRecord into providerCatalogShared and
collapse the seven refresh handlers onto it.
- Document the explicit OCA decision: its reasoning control is the
API-driven effort dropdown; the removed budget slider wrote state no
OCA request path consumed.
* Harden OpenRouter picker reasoning gate against placeholder metadata
Gate on the raw committed model-info snapshot instead of the hook's
default-info fallback, so a selected id that is absent from the catalog
can never inherit reasoning support from placeholder metadata (the
fallback carries no supportsReasoning today, but reading the raw field
removes the latent dependency).
Cancelling a task previously only detached Cline's listeners from an
in-flight foreground command (process.continue()); the spawned process
kept running in the user's terminal after cancellation.
Send Ctrl+C (ETX) to the terminal before detaching so the shell delivers
SIGINT to the foreground process group, actually stopping the command.
The terminal is left open for reuse, and cancellation still succeeds even
if the interrupt write throws (e.g. terminal already disposed).
The legacy-provider migration seeded the openai-compatible models.json
entry with hardcoded defaults (contextWindow 128k, no pricing/temperature/
maxTokens/R1 flag). Because later override migrations skip models that
already exist in models.json, the user's legacy planMode/actModeOpenAiModelInfo
overrides were silently discarded on first upgrade: context window,
max output tokens, input/output prices, temperature, supportsImages=false,
and isR1FormatRequired all reset to defaults.
Seed the entry from the mode-appropriate legacy model-info snapshot
instead, treating legacy sentinels (maxTokens -1, temperature 0,
prices 0) as unset.
Matches the legacy extension's OpenRouter default (openRouterDefaultModelId),
so users migrating from the legacy build without an explicitly selected model
keep the same default model instead of being silently moved to
anthropic/claude-sonnet-4.6.
A legacy single-file .clinerules at the workspace root made the config
watcher's scans of .clinerules/skills and .clinerules/workflows throw
ENOTDIR, which aborted the entire user-instruction refresh: workspace
rules, global rules, and the Skills view all silently failed to load.
Treat ENOTDIR like ENOENT in isIgnorableDirectoryError so a file in a
directory position simply yields no candidates. The .clinerules file
itself is still picked up by the file branch of discoverRulesLikeFiles.
* fix(vscode): hide /newrule and /deep-planning until their prompt expansions are ported to the SDK runtime
* feat(vscode): port the /newtask context handoff to the SDK runtime
Expand /newtask into explicit new_task-tool instructions in
SdkController.resolveSlashCommands (ported from legacy
newTaskToolResponse), register a custom new_task AgentTool that captures
the model-generated context summary and completes the run, and emit the
ask:"new_task" message on turn completion so the existing webview
"Start New Task with Context" button (which preloads a fresh task with
the ask text) becomes reachable again. Set the turn phase to
awaiting_followup when emitting the ask, since the completesRun
termination path skips the translator's usual end-of-turn status
handling.
* fix(vscode): hide /reportbug until its prompt expansion is ported to the SDK runtime
Also drop the feature tip promoting /reportbug so the UI doesn't
advertise a command that no longer autocompletes.
* Revert "feat(vscode): port the /newtask context handoff to the SDK runtime"
This reverts commit d9ad153aec.
* feat(vscode): make /newtask an alias of /compact
Condensing achieves /newtask's goal (continue working with a fresh,
summarized context window) without the legacy new_task tool, so the
webview intercepts /newtask alongside /compact and /smol and runs the
condense RPC. Menu description updated to match.
* feat(vscode): port the /deep-planning prompt expansion to the SDK runtime
Expand /deep-planning into the legacy generic-variant instructions
(silent investigation, targeted questions, implementation_plan.md) in
SdkController.resolveSlashCommands, ahead of workflow/skill expansion.
Legacy's STEP 4 created an implementation task via the new_task tool,
which doesn't exist on the SDK runtime; the ported prompt instead has
the agent present the plan and wait for explicit user confirmation.
Re-adds /deep-planning to the slash menu.
* refactor(vscode): simplify the /deep-planning expansion
Drop the custom regex/expander and shell-specific research-command
blocks: the builtin is now a plain AvailableRuntimeCommand appended to
the discovered workflow/skill commands, so the existing
expandSlashCommands machinery handles matching and replacement. The
prompt keeps the four-step protocol and implementation_plan.md
structure with a generic investigation paragraph instead of embedded
OS-specific commands.
* fix(mcp): honor per-server timeout (seconds) across all clients
The per-server timeout field in cline_mcp_settings.json was only read
by the VSCode extension's tools/call path. Everywhere else used
hardcoded constants: the SDK client timed out all requests at 5s and
initialize at 1.5s, and the extension's metadata requests (tools/list,
resources/*, prompts/*) timed out at 5s. Slow servers failed despite a
configured timeout (#7635, #12344).
Resolve the timeout once per client and apply it to every request:
- @cline/shared exports the default (60s) and bounds (1s-3600s) plus a
resolver that clamps out-of-range values, so a milliseconds/seconds
mix-up can no longer become hours.
- The SDK config loader parses timeout into
McpServerRegistration.timeoutSeconds; StdioMcpClient and
SdkUrlMcpClient apply it to initialize, tools/list, and tools/call.
Unconfigured servers keep the fast 1.5s initialize probe so startup
is no slower than before; a configured timeout raises that budget
for slow-starting servers.
- The extension routes every request (including metadata) through one
resolver and drops the hardcoded 5s DEFAULT_REQUEST_TIMEOUT_MS.
- createMcpTools derives the agent tool timeoutMs from the same value,
keeping the wrapper and request timeouts in agreement.
- Timeout errors now name the bound and the field to increase; the
VSCode server row and the CLI server list show the effective timeout
and how to change it.
* fix(mcp): harden timeout lifecycle handling
* fix(mcp): address timeout review feedback
* fix(mcp): bound initialization and reconnect
* fix(mcp): keep timeout snapshots consistent
* fix(mcp): use standard stdio framing
* fix(mcp): bound legacy stdio fallback
* fix(mcp): honor timeout in framed fallback
* test(vscode): use SDK Vitest runner
* fix(mcp): fetch server capabilities in parallel
The four post-connect metadata requests (tools/list, resources/list,
resources/templates/list, prompts/list) ran sequentially, so a server
that hangs after initialize blocked connectToServer for four timeout
bounds. The MCP client correlates concurrent requests by JSON-RPC id
and the stdio transport writes each message atomically, so the fetches
now run in parallel and the worst case is one bound.
Also delete McpHub.readResource and McpHub.getPrompt and their response
types: nothing calls them since the SDK migration removed the
access_mcp_resource tool and prompt expansion.
* fix(mcp): keep failed servers and both framing errors visible
When both stdio framing attempts fail differently during initialize,
name each attempt's error instead of discarding the Content-Length
fallback's diagnostics. When they fail identically (both timed out),
rethrow the newline error unchanged so the timeout hint is the whole
message.
When connectToServer fails before the connection is registered (e.g.
the transport fails to start), register a disconnected entry carrying
the error so the server stays visible in the list instead of silently
disappearing, and notify the webview so the row leaves the connecting
state.
* fix(mcp): reject tool calls on connections without a client
A failed (re)connect registers a disconnected entry with a null client
so the server stays visible in the list. A tool wrapper captured by an
active session can still target that server; callTool now rejects it
with a controlled error naming the server and its last connection
error, instead of dereferencing the null client and throwing a
TypeError.
parseKeyPairsIntoRecord wrapped the whole forEach in one try/catch, so a single entry that broke decodeURIComponent (e.g. a stray % in OTEL_EXPORTER_OTLP_HEADERS) aborted the loop and silently dropped every remaining header. Move the try/catch inside the loop to skip only the malformed entry. Adds regression tests.
* fix(desktop): disable timeout for chat send commands
Add per-invocation timeout options to the desktop client and disable the deadline for long-running chat send requests. Extract the shared command response type and verify send commands use the timeout override.
* fix(desktop): clean up failed websocket sends
useDebouncedInput scheduled its debounced onChange on mount and on every
external initialValue resync, not just user edits. Settings fields mount
with a placeholder value while their backing provider config is still
loading asynchronously, so the mount-fire echoed that placeholder back
to the backend ~100ms later.
For DebouncedTextField-backed secret fields (e.g. the OpenRouter API key,
which renders a masked value derived from the async readProviderConfig
response), losing that race meant writing apiKey: "" — silently deleting
the stored key from both providers.json and the legacy secrets store,
and leaving the field rendering empty despite a previously persisted key.
Non-secret fields similarly re-saved stale placeholder values on every
mount.
Gate the debounced save on an actual user edit: only values set through
the returned setter fire onChange; mount and external resyncs never do.
* fix(openai-compatible): keep user model metadata when only the model id changes
Changing the OpenAI Compatible model id committed the new id without
overrides, so an id unknown to the catalog resolved to safe defaults
(inputPrice/outputPrice 0, supportsPromptCache false) and paid requests
billed as $0.0000. The legacy extension kept this user-authored metadata
in a single id-independent blob, so custom prices survived id edits.
Recommit the currently displayed overrides under the new id when the
model id changes, and let edits made while that commit is round-tripping
target the pending id instead of the stale read-back id.
Fixes ENG-2341
* fix(openai-compatible): scope pending selection state per mode
Review follow-up: the pending-override accumulator and pending-commit
counter were shared across Plan and Act. Changing the Act model id,
switching to Plan while that commit was round-tripping, then editing an
override committed the Plan edit under the pending Act model id (and the
shared pending count blocked Plan's reseed at the mode boundary).
Record the mode alongside the pending selection and only trust it for
edits in the same mode, keep per-mode pending counts so a mode switch
reseeds from that mode's committed state, and cover the deferred-commit
mode-switch scenario with a component test.
* fix(openai-compatible): give each mode its own pending-selection accumulator
Review follow-up: tagging the single shared accumulator with a mode still
lost state on a mode round trip. With an Act commit pending, visiting
Plan reseeded the shared slot to Plan; returning to Act could not reseed
(Act's read-back was still in flight), so the next Act edit merged onto
an empty set and silently dropped the pending prices/context/capabilities.
Keep one accumulator slot per mode so a round trip through the other
mode never disturbs a mode's pending state, and cover the scenario with
a deferred-commit round-trip test.
The OpenRouter model picker (refreshOpenRouterModels) still applied the
legacy 200k context-window restriction to Anthropic Claude models, while
the task header and auto-compaction resolve model info through the SDK
catalog, which reports the full 1m extended context window. The same
model showed Context: 200K in the picker and 1.0m in the task header.
Per the current product direction the 200k restriction (and its :1m
opt-in variants) is dropped entirely — everyone gets the 1m context
window. Remove the artificial clamps from refreshOpenRouterModels
(keeping the prompt-cache pricing overrides) and update the
openRouterDefaultModelInfo fallback to match, so the picker, the task
header, and compaction thresholds all agree on 1m.
Closes ENG-2345.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(core): migrate legacy API keys for all secret-backed providers
collectCandidateProviderIds only nominated 11 provider ids while
buildLegacyProviderSettings can copy keys for 34, so stored keys for the
other 25 providers (deepseek, mistral, xai, groq, ...) were silently
dropped during migration unless the provider was the active plan/act
provider. Add the missing candidate checks so any stored key makes its
provider a migration candidate.
Also pick the legacy mode per candidate: a split plan/act config applied
the single globalState.mode to every provider, so the non-current mode's
configured model was replaced by the catalog default.
Migration re-runs on manager construction and never overwrites existing
entries, so users who already ran the buggy migration get dropped keys
backfilled from the still-present legacy secrets.json on next launch.
Fixes ENG-2337
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(core): normalize legacy provider-id aliases during migration
Address review: the mode selection and model fallback compared raw
legacy provider ids, so a declared alias (togetherai -> together,
sap-ai-core -> sapaicore) in globalState would miss its canonical
secret-derived candidate, read the wrong mode, and could write duplicate
alias/canonical entries. Route candidate collection, mode comparison,
and the generic model fallback through the existing normalizeProviderId
boundary. resolveMigratedProviderId now delegates to normalizeProviderId
(identical for the openai -> openai-compatible case it already handled).
Legacy ApiProvider never actually stored alias forms, so this is
hardening for hand-edited state rather than a live regression.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(vscode): reconcile the two provider state stores (ENG-2332)
- createStorageContext now honors CLINE_DATA_DIR with the same priority as
the SDK's resolveClineDataDir and the legacy reader's resolveDataDir
(explicit option > CLINE_DATA_DIR > CLINE_DIR/data > ~/.cline/data), so
globalState.json/secrets.json live in the same data dir as providers.json
and legacy task state instead of silently splitting across directories.
- Add setLastUsedProvider and call it on active provider switches
(SdkProviderChangeCoordinator) and when a session resolves its provider
from StateManager (buildSessionConfig), so providers.json's
lastUsedProvider no longer goes stale across provider switches.
- Trim env vars in legacy-state-reader's resolveDataDir to match the SDK's
resolution exactly.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* refactor: trim ENG-2332 fix to the minimal change set
Revert the cosmetic legacy-state-reader trim, restore the original CLINE_DIR
line in createStorageContext, and tighten comments. No behavior change to
the two core fixes.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* refactor: drop lastUsedProvider sync, keep only the data-dir alignment fix
Scope ENG-2332 to the root-cause fix: createStorageContext honoring
CLINE_DATA_DIR like the SDK resolvers. The providers.json lastUsedProvider
staleness is deferred to a follow-up.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix: trim CLINE_DATA_DIR in resolveDataDir to match createStorageContext
Addresses Greptile P1: a whitespace-padded CLINE_DATA_DIR was trimmed by
createStorageContext but used verbatim by the legacy reader, which could
resolve the two stores to different directories again.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* chore: retrigger CI (windows e2e flake in chat.test.ts)
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* refactor: share one data-dir resolver between storage context and legacy reader
Per review feedback: extract resolveDataDirFromEnv in storage-context.ts and
have legacy-state-reader's resolveDataDir delegate to it, so the two stores
structurally cannot drift apart again.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix: trim CLINE_DIR in the shared data-dir resolver to match the SDK
The SDK's resolveClineDir trims CLINE_DIR; a whitespace-padded value would
otherwise still resolve VS Code state and providers.json to different
directories.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* feat(vscode): show working-directory badge when a task runs outside the open workspace
Tasks resumed from the CLI or another workspace keep their original cwd,
so Cline reads, edits, and runs commands in a directory that is not the
one visible in the window - previously with no indication anywhere.
- Add TaskWorkingDirectoryBadge: a persistent warning chip in the task
header (folder icon + cwd basename, full path + explanation in the
tooltip) shown only when the task cwd is neither an open workspace
root nor inside one. Hidden when roots or cwd are unknown to avoid
false positives.
- Fix SdkController.getStateToPostToWebview to pass its workspace
manager into the shared state builder; the SDK path previously always
sent workspaceRoots: [] to the webview.
- Unit tests for the outside-workspace predicate (case, separators,
multi-root, prefix collisions) and badge render states.
* fix(vscode): platform-aware path comparison in working-directory badge
Address PR #12637 review findings:
- Case folding is now platform-aware (win32/darwin insensitive, linux
and unknown strict), so case-only path differences on Linux are no
longer hidden; mirrors arePathsEqual in src/utils/path.ts.
- Backslashes are treated as separators only on win32; on POSIX a
backslash is an ordinary filename character.
- Containment prefix no longer doubles the separator when a workspace
root already ends with one, fixing false warnings for '/' and drive
roots.
- Tests cover case-only pairs under win32/darwin/linux/unknown,
POSIX-backslash filenames, '/' and 'C:\' workspace roots.
* fix(vscode): make darwin path comparison strict in working-directory badge
Follow-up to PR #12637 review: darwin volumes can be case-sensitive, and
the host's canonical arePathsEqual (src/utils/path.ts) already treats
only win32 as case-insensitive. Align the badge predicate with that
convention: case folding and backslash separators apply on win32 only;
darwin, linux, and unknown compare strictly. For a warning badge a rare
spurious warning beats silently hiding a real mismatch.
* fix(vscode): restore legacy workflow invocation and management UI
- Expand /workflow slash commands typed with the legacy .md filename
spelling (what the autocomplete menu inserts) and mid-message, and
honor the user's workflow enable/disable toggles, instead of only
expanding a leading extension-less /name via the SDK resolver.
- Restore the Workflows tab in the rules modal (view, toggle, create,
edit, delete; enterprise section) that was dropped in the SDK-backed
extension while all its gRPC handlers remained wired.
* chore: add changeset for workflow fixes
* fix(vscode): refresh workflow toggles on webview launch
The slash command menu is driven by workflowToggles state, but nothing
refreshed it at startup in the SDK-backed extension (only opening the
rules modal or creating a rule file did), so workflows never appeared in
the chat autocomplete until the user opened the modal. Legacy refreshed
toggles on task init.
* feat(vscode): move Workflows tab last and add deprecation warning
Workflows tab now appears after Rules/Hooks/Skills, and its view leads
with a warning banner: workflows are being deprecated in favor of
skills, with a docs link.
* chore: update changeset for workflow deprecation notice
* fix(vscode): address review findings on workflow expansion
- Honor remoteWorkflowToggles (and locked alwaysEnabled remote
workflows) when building the disabled set, so disabled enterprise
workflows no longer expand.
- Treat a workflow as disabled only when no scope has it enabled, so a
disabled workspace file no longer shadows a same-named enabled global
one (legacy expanded the enabled scope).
- Strip all workflow extensions the SDK discovers (.md/.markdown/.txt)
when matching typed commands, not just .md.
- Re-read toggle state after the async directory scan in
refreshWorkflowToggles so a toggle flipped mid-scan is not overwritten
by the stale snapshot.
* fix(vscode): map workflow toggles to records so frontmatter names are governed
Compute the disabled set from the discovered workflow records
(listRecords) instead of toggle-path basenames alone: a file's toggle is
matched by its basename and disables the record's actual command name,
so a frontmatter 'name' that differs from the filename is still governed
by the Workflows toggle. Remote-config-materialized records are governed
by the name-keyed remote toggles (locked alwaysEnabled remain on).
* fix(vscode): harden workflow toggle-name mapping for expansion
- A command name shared by several records now counts as enabled when
any record is enabled, so a disabled local workflow can no longer
suppress an enabled or locked (alwaysEnabled) enterprise workflow.
- Remote toggles/locks are matched via a sanitizeSegment-compatible key,
so config names that get rewritten during materialization (e.g. 'Org
Standards' -> org-standards.md) still govern expansion.
- Typed filenames (e.g. /my-workflow.md from autocomplete) now resolve
to workflows whose frontmatter renames the command, via the record's
file basename.
* fix(vscode): govern each workflow command by its own record's toggle
Key the disabled set by exact command name and decide each record
independently instead of OR-aggregating by canonical name: distinct
commands whose names only differ by case or extension (e.g. a local
'Release' and a remote 'release') no longer influence each other, so an
enabled local workflow cannot keep a disabled enterprise workflow
expandable, and a disabled one cannot suppress a locked enterprise
workflow.
* fix(vscode): exact remote-name sanitization and keep mid-scan toggle additions
- Port @cline/shared's sanitizeSegment verbatim (incl. the 80-char cap)
for remote workflow name comparison, so long enterprise workflow names
cannot bypass a disabled toggle after filename truncation.
- The post-scan toggle merge now also keeps entries added while the scan
was running (e.g. a workflow created via the modal), instead of
pruning them with the deleted files.
* fix(vscode): handle mid-scan deletions and sanitized remote-name collisions
- The post-scan toggle merge now also drops entries that were removed
from state while the scan ran, so a workflow deleted mid-refresh is
not restored by the stale scan result.
- Remote toggle names that sanitize to the same materialized name merge
as enabled-if-any-enabled instead of last-write-wins.
* fix(vscode): serialize workflow toggle refreshes
Queue refreshWorkflowToggles runs on a promise chain so overlapping
refreshes (webview launch, modal open, file create/delete) cannot
interleave scans and writes. Combined with the post-scan merge for
direct toggle flips, this closes the remaining stale-refresh races.
* fix(vscode): key remote workflow toggles off the materialized filename
The materializer names remote workflow files from the config name, so
derive the remote toggle key from the file basename instead of the
parsed command name; a frontmatter alias can no longer bypass a
disabled remote toggle.
* fix(vscode): make interrupted tasks findable in History and restore Resume button
Interrupted/cancelled sessions were presented as gone (ENG-2336):
- History fuzzy search used location-based Fuse scoring (ignoreLocation:
false, threshold 0.6), so any match more than ~60 characters into the
task title scored above the threshold and the task silently vanished
from search results even though it was in the list. Search now matches
anywhere in the title.
- Opening a task from History never updated the authoritative TurnState,
so the footer kept the previous context's phase (usually idle) and the
Resume Task button never appeared for interrupted/failed sessions.
showTaskWithId now derives the phase from the reopened conversation:
resumable for interrupted tasks, completed for completed ones.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(vscode): decide Resume vs Start New Task from persisted session status
SDK conversations do not record a completion tool call in the transcript
(a completed turn and one interrupted mid-stream both end with plain
assistant text), and history rendering appends a synthetic trailing
ask:"completion_result" either way, so the message tail always looked
"completed". Reopening a task from History now reads the persisted
session status: "completed" gets the Start New Task affordance, while
cancelled/failed (interrupted) sessions get Resume Task. When reopening
the currently-active task, the stop is awaited first so the status read
reflects how the last turn actually ended.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(vscode): fence concurrent history opens and default unknown status to Resume
Address review feedback:
- showTaskWithId now takes a generation fence: a request that loses the
race to a newer showTaskWithId or clearTask abandons installation after
its awaited reads, so a slow older request can never clobber the user's
latest selection (task proxy, messages, or turn phase). clearTask bumps
the generation too so New Task wins over an in-flight history open.
- The resume affordance no longer falls back to the message tail when the
persisted session status is unavailable: the tail always ends with the
synthetic ask:"completion_result" that history rendering appends, which
misclassified interrupted tasks as completed on a failed status read.
Only an explicit "completed" status gets Start New Task; anything else
(including unknown) gets Resume Task, the safe direction.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(vscode): allocate history-open generation before the lookup and fence before session stop
Address review feedback: the latest-selection-wins fence started too late.
SdkController.showTaskWithId awaited findHistoryItem() before entering the
coordinator, so a stalled preflight for an older selection could re-enter
with a NEWER generation than a later selection and replace it — and since
the first fence check sat after endActiveSession, a superseded request
could also stop a session the newer selection had just installed.
The history lookup now lives inside the coordinator (skipHistoryLookup is
gone), the generation is allocated synchronously before all asynchronous
work, and a fence check runs before endActiveSession so a superseded open
never stops the newer selection's session. The coordinator returns the
HistoryItem so SdkController keeps its TaskResponse contract. Regression
test covers the exact reported sequence: stalled lookup for task A, task B
selected and loaded, A resolves last — B stays installed and A stops
nothing.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Show completion feedback box for inferred turn-final responses in SDK path
The SDK agent usually ends a turn with a plain text response instead of an
attempt_completion / plan_mode_respond tool call, so the legacy green 'Task
Completed' box (act) and 'Plan Created' box (plan) never rendered in the new
extension — making a finished turn look stuck or frozen.
Now, when a turn ends cleanly (done reason 'completed', no completion tool
used) and its last content is a text response, that text row is retagged in
place to say:'completion_result' (act, green box) or the new
say:'plan_completion_result' (plan, yellow-accented 'Plan Created' box).
- Track the turn-final text candidate in MessageTranslatorState; cleared on
tool activity, errors, aborts, and new user turns
- Replay the same inference during history rehydration, recovering each
turn's plan/act mode from the persisted <user_input mode="..."> wrapper
- Add plan_completion_result ClineSay type (+ proto enum) rendered via
PlanCompletionOutputRow, restyled with the plan-yellow accent to match
the plan/act toggle and the CLI's plan color
- Turn phase semantics unchanged: footer buttons still come from TurnState
* Remove attempt_completion tool and strip completion box headers
- Drop the attempt_completion extra tool (and its shell-command executor)
from VS Code SDK sessions; the SDK's built-in submit_and_exit is already
disabled for act/plan presets, so the agent now always ends its turn with
a plain text response and the turn-end inference styles it.
- Translator keeps recognizing attempt_completion/submit_and_exit for
replaying persisted transcripts from older sessions.
- Remove the 'Task Completed' header, check icon, and copy button from the
green completion box, and the 'Plan Created' header, notepad icon, and
copy button from the yellow plan box. The final text of a turn may be a
question rather than an actual completion or plan, so the boxes are now
quiet color cues that make no claim.
* Skip completion retag for terminal text of failed/cancelled sessions
The trailing text of a session whose last run failed or was cancelled is a
dangling partial response, not a completion. Gate the history converter's
final synthesized turn end on the session record's status so reopening a
broken task keeps its terminal text as a plain row instead of an inferred
completion box. Mid-transcript turns are unaffected: the user continued
after them and history carries no per-turn outcome.
* Require clean at-rest session status before retagging terminal text
Tighten the negative failed/cancelled check into an allowlist: the history
converter now only retags the transcript's terminal text when the session
record is 'completed' (formally stopped clean run) or 'idle' (the normal
at-rest state between interactive turns). 'running'/'pending' at rest means
the process died mid-turn, so its dangling partial response stays plain.
* Restrict history completion retag to the transcript's final turn
Persisted SDK transcripts carry no per-turn outcome, so a mid-conversation
turn the user cancelled mid-response (then followed up on) is
indistinguishable from one that ended cleanly. Retagging those presented
interrupted responses as deliberate turn ends. History rehydration now only
retags the final turn's terminal text, gated on the session record's
at-rest status; earlier turns always render as plain text. Live sessions
are unaffected — their per-turn boxes come from real done events.
* Trust only status 'completed' for the history completion retag
'idle' is written by markTurnIdle for every interactive finish reason,
including aborted turns, so an at-rest idle record cannot prove the last
turn ended cleanly. Terminal statuses are reliably written when sessions
are released (task switch, clear, dispose), so requiring 'completed' keeps
the box on normal reopened tasks while never styling an interrupted
response as a deliberate turn end.
* Treat missing session records as unknown outcome in history retag
A transcript with no session record has no recorded outcome, so its
terminal text stays a plain row instead of getting completion styling.
migrateWelcomeViewCompleted derived the flag solely from VS Code's
per-profile stores, which are empty for users upgrading from the live
4.x extension (file-backed config under ~/.cline/data). The flag landed
as false and fully configured users were pushed back through onboarding.
Purely additive: the existing VS Code checks are untouched; the same
signals (completed flag, provider secrets, keyless provider configs) are
now also read from the file-backed globalState.json/secrets.json and
OR-ed into the result.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
Models sometimes emit numeric tool arguments as JSON strings. `insert_line`
and the `read_files` line bounds were plain `z.number()`, so an
`insert_line: "3"` rejected the whole tool call before it ran:
1 tool call(s) failed: [editor] {"error":"✖ Invalid input: expected number,
received string\n → at insert_line"}
The model is handed that error and burns a round trip re-deriving the argument.
`z.coerce` leaves the JSON Schema advertised to the model untouched (still
`integer`), and `.int()` / `.positive()` still reject "abc", "3.5" and 3.5.
* fix(webview): stop unbounded polling of local model endpoints (ENG-2344)
The Ollama provider form polled /api/tags every 2s from two places at once
(OllamaProvider and a dead duplicate poll in ApiOptions whose result was
never read), producing ~1 req/s for as long as the settings pane was open.
Since the base URL is user-configurable, this could hammer a remote or
metered endpoint. VSCodeLmProvider and LMStudioProvider had the same
interval pattern.
- Remove all useInterval model polling; fetch on mount and when the
base URL changes instead
- Refresh the Ollama model list when the picker field gains focus so a
server started after the pane opened is still discovered
- Delete the dead _ollamaModels poll in ApiOptions
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(webview): add on-demand model refresh for LM Studio and VS Code LM
Greptile review follow-up: removing the polling intervals left these two
pickers pinned to their mount-time snapshot. Mirror the Ollama picker's
interaction-driven refresh:
- LM Studio: refetch models when the model dropdown or the manual model
id field gains focus
- VS Code LM: refetch when the dropdown gains focus, and add an explicit
'Refresh the model list' link to the empty state
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
An `editor` tool call with `insert_line` (e.g. a prepend) targets an
existing file — the SDK editor executor requires the file to already exist
for inserts — but sdkToolToClineSayTool only treated `old_text`/`replace_in_file`
as edits, so inserts were classified as newFileCreated and the approval card
read "Cline wants to create a new file:" for an existing file.
Treat insert_line as an edit so the card reads "Cline wants to edit this file:".
The webview-ui-toolkit VSCodeDropdown fires a spurious change event with
the wrong option (index 2, Portuguese - Brasil) while its slotted options
initialize after a window reload, and the handler persisted that value
unconditionally. Any saved language not at the top of the list could be
silently rewritten to Portuguese just by opening the General settings tab.
Replace the toolkit dropdown with the ui/select component already used by
the other settings dropdowns (Auto Compact Strategy, MCP Display Mode),
which only emits onValueChange for real user selections, and render the
options from the shared languageOptions list instead of a hardcoded copy.
At turn end the final message is finalized (partial: false) via the fast
partial-message stream a moment before the done event flips turnState out
of "streaming" via a full state post. During that gap the in-list
"Thinking..." loader row appeared and immediately disappeared, flashing
on every turn completion.
- Extract the loader show/hide logic from MessagesArea into a testable
useThinkingLoaderRow hook.
- Debounce the loader when its trigger is the tail message finishing
streaming: mid-turn a real wait outlives the grace period, while the
turn-end phase change cancels it before it ever shows.
- Add the legacy path's say("completion_result") anti-flicker guard to
the turnState path so attempt_completion turns never flash regardless
of timing.
* fix(vscode): restore Retry/Start New Task buttons after API failure
A provider stream error emits ask:'api_req_failed', but the session-event
coordinator resolved the turn-end phase to 'awaiting_followup', clobbering
the error state — so the footer never showed the error-recovery buttons and
the error surface offered no way to recover (ENG-2339).
Record the error outcome in MessageTranslatorState when the error event is
translated, and resolve turn end to the 'error' phase so the existing
api_req_failed button config (Retry / Start New Task) is reachable again,
matching legacy behavior.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* fix(vscode): also record error outcome for done(reason:'error') terminations
A turn can terminate with done(reason:'error') without a separate 'error'
event; record the error outcome there too so turn end still resolves to the
'error' phase and the Retry / Start New Task buttons appear.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Fix ignored China/international API line toggles for Qwen, Moonshot, Z AI (ENG-2340)
The regional apiLine setting was persisted through both storage layers but
never consulted when resolving the request endpoint, silently sending
regional users to the wrong host.
- @cline/llms: record china/international base URLs on the builtin specs
for qwen, qwen-code, moonshot, zai, zai-coding-plan, and minimax; expose
resolveProviderApiLineBaseUrl; resolve options.apiLine against the
registered apiLineBaseUrls in GatewayRegistry.createProvider (explicit
base URLs still win).
- @cline/core: toProviderConfig now resolves the base URL from apiLine
between the explicit setting and the static provider default.
- VS Code: buildSessionConfig resolves the API line from legacy state
(qwenApiLine/moonshotApiLine/zaiApiLine/minimaxApiLine) with a
providers.json fallback and forwards it on the provider config so the
gateway can route regionally.
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Share the base provider's legacy API line with qwen-code and zai-coding-plan
The coding variants have regional endpoints in the SDK but no legacy
state field of their own, so a China-line user selecting them from the
VS Code UI would silently fall back to the international default. The
variant's own providers.json apiLine still wins over the shared legacy
field.
---------
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
Session rebuilds seed the replacement session from readMessages, but the
persisted transcript only catches up at assistant-message/turn boundaries
and abort() does not flush. Toggling plan/act mode while a task's first
turn is mid-flight (e.g. a command approval pending) therefore rebuilt the
session with no history at all and the new mode's model lost the task.
Add RuntimeHost.readLiveSessionMessages (optional) which prefers the
resident session's agent.getMessages() and falls back to the persisted
transcript, expose it as ClineCore.readLiveMessages, and use it in the
VS Code history loader that feeds session rebuilds. readSessionMessages
keeps its persisted-transcript semantics for existing callers (compaction
validation, session snapshots, history).
Co-authored-by: Saoud Rizwan <saoudrizwan@users.noreply.github.com>
* Add a built-in cline-settings skill and broaden the legacy resume
warning
Models diagnosing configuration problems have no authoritative source
for where Cline stores settings: the SDK migration removed the old MCP
documentation tool, and resumed legacy conversations can carry stale
paths and instructions from older runtimes (CLINE-2570).
Add a core-owned virtual skill, cline-settings, whose instructions are
generated at invocation from the shared storage path resolvers. It is
listed and invoked through the existing skills registry on both the
local and Hub session paths, is reserved against shadowing by
file-backed skills (case-insensitive), honors session skill allowlists
(an explicit empty allowlist disables all skills including built-ins),
and never appears in editable listRecords.
Broaden LEGACY_RESUME_MODEL_WARNING to cover stale configuration
paths, file formats, and product instructions, not just tool names.
Anchor the persisted history boundary on a stable marker; recognize
and upgrade the historical warning in place so previously resumed
tasks get the new wording without duplicate warnings, and preserve
resumed user text that shares a message with the warning.
* Fix Windows MCP stdio spawn for paths with spaces; add settings-skill
rule
The runtime-builder MCP test failed on Windows because the stdio
client spawns with shell: true there, and cmd.exe split the unquoted
executable path at the space in "C:\Program Files\nodejs\node.exe".
Quote the command and arguments for cmd.exe so any server whose
command or arguments contain spaces can start. Also raise the connect
timeout to match the request timeout: connect covers process spawn
plus the first initialize round-trip, and 1.5s is tight for cold
starts on loaded machines.
Add a brief .clinerule noting that settings/storage-path changes may
require updating the cline-settings built-in skill.
* Quote empty MCP arguments for cmd.exe
An empty-string argument passed through unquoted disappears when
cmd.exe re-parses the concatenated command line, silently shifting the
server's argument list. Quote empty values so they survive as "".
---------
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
* fix(desktop): order sessions by last activity and unify status dot colors
* feat(desktop): add pagination to session history view
Display sessions in ten-item pages and fetch older history only after reaching the final local page. Add coverage for pagination, session opening, and compact token formatting.
* page numbers
* Support adding a session to favorite list
* apply feedback
* fix
* feat(chat): refine tool and reasoning message disclosures
Add tool-specific icons with a fallback, display elapsed reasoning time, and restyle reasoning disclosures. Position hidden message actions outside the layout and update attachment sizing to valid Tailwind utilities.
* fix(desktop): improve chat message actions and scrolling
Refine action positioning, sizing, timestamps, and visibility for chat messages. Remove nested overflow constraints so scrolling remains controlled by the conversation viewport, and tighten tool disclosure spacing.
* tools icon mapping
* fix(desktop): align chat timestamps and tool icons
2026-07-28 15:30:33 -07:00
994 changed files with 129963 additions and 28094 deletions
fix: restore workflow support regressions — expand `/workflow.md` slash commands (the legacy filename spelling the autocomplete inserts) and mid-message commands, honor workflow enable/disable toggles during expansion, refresh the slash menu's workflow list on webview launch, and bring back the Workflows management tab in the rules modal (now last in the tab list, with a deprecation notice pointing to Skills)
Enable Auto Compact by default so long chats automatically compress conversation history instead of failing at the model context limit. It can be disabled in Settings → Features → "Auto Compact".
Fix /compact UX: clear the chat input as soon as the command is submitted, wrap the compaction divider row at narrow sidebar widths, and update the context-window header even when compacting a small conversation grows the estimated context
Hide the "View Changes" button on completion rows until there are actually changes to show, instead of rendering it faded and disabled. Turns that changed nothing, non-git workspaces, and repos without commits no longer show a dead button with a misleading tooltip.
Show the user's message in chat immediately when sending to a task opened from history, instead of only a thinking indicator until the session resume finishes
@@ -9,7 +9,7 @@ Use this skill when the user asks to release the desktop app, publish Cline Code
> Working directory: run every command below from the repository root.
Desktop releases are macOS-only today (signed + notarized DMG for Apple Silicon and Intel) and are built entirely in GitHub Actions — there is no local publish path. Installed apps discover new releases automatically through the Tauri updater, so publishing a release is what ships the update to every existing user.
Desktop releases are macOS-only today (a single signed + notarized universal DMG that runs natively on both Apple Silicon and Intel) and are built entirely in GitHub Actions — there is no local publish path. Installed apps discover new releases automatically through the Tauri updater, so publishing a release is what ships the update to every existing user.
## Release contract
@@ -17,7 +17,7 @@ Desktop releases are macOS-only today (signed + notarized DMG for Apple Silicon
- Release tag: `desktop-vX.Y.Z`, where `X.Y.Z` matches both version files.
- Release prep includes approved release notes, the version bumps, and an `apps/examples/desktop-app/CHANGELOG.md` update.
- Publish path: `.github/workflows/desktop-publish.yml` (workflow_dispatch, requires the tag to exist, point at the checked-out commit, and be reachable from `origin/main`).
- The workflow creates the `desktop-vX.Y.Z` GitHub release (DMGs + updater artifacts + `latest.json`) and refreshes the rolling `desktop-latest` release, which is the static auto-update feed every installed app polls. Never delete the `desktop-latest` release or tag.
- The workflow creates the `desktop-vX.Y.Z` GitHub release (universal DMG + updater artifact + `latest.json`) and refreshes the rolling `desktop-latest` release, which is the static auto-update feed every installed app polls. Never delete the `desktop-latest` release or tag.
- The changelog's top `## X.Y.Z` section is extracted verbatim into the GitHub release body, the Slack announcement, and the updater manifest notes.
gh run list --workflow=desktop-publish.yml --limit=1 --json url,status,conclusion,createdAt --jq '.[0]'
```
The workflow builds both architectures in parallel (aarch64 native, x86_64 cross-compiled), signs with the Developer ID certificate, notarizes with the App Store Connect API key, signs updater artifacts with the Tauri updater key, creates the GitHub release, refreshes`desktop-latest/latest.json`, and posts to Slack. Notarization typically adds 2–10 minutes.
**The run pauses for approval.**`validate` runs immediately, then the `build`
job waits on the `PublishDesktop` environment until a required reviewer approves
it — the run sits in `waiting`, which is expected, not a hang. Approve it in the
run's web UI ("Review deployments"), or:
If the workflow fails on missing credentials, see "Repo secrets (one-time setup)" below.
```sh
gh api repos/cline/cline/actions/runs/<run-id>/pending_deployments \
--method POST -f state=approved -f comment="desktop vX.Y.Z"\
Nothing after `validate` runs — and no signing key is readable — until then.
The workflow builds one universal macOS bundle (`tauri build --target universal-apple-darwin` lipos the aarch64 + x86_64 Rust binaries; the Bun sidecar is lipo'd by `build-sidecar-bin.ts`), verifies every Mach-O in the bundle carries both slices, signs with the Developer ID certificate, notarizes with the App Store Connect API key, signs the updater artifact with the Tauri updater key, creates the GitHub release, refreshes `desktop-latest/latest.json`, and posts to Slack. Notarization typically adds 2–10 minutes.
If the workflow fails on missing credentials, see "Publish secrets (one-time setup)" below.
9. Verify the update feed after the run succeeds.
@@ -100,17 +113,30 @@ If the workflow fails on missing credentials, see "Repo secrets (one-time setup)
curl -sL https://github.com/cline/cline/releases/download/desktop-latest/latest.json | head -30
```
The `version` field must be the new release and both `darwin-aarch64` and `darwin-x86_64`URLs must point at the new `desktop-vX.Y.Z`assets. Installed apps pick the update up on next launch or within 2 hours.
The `version` field must be the new release and both `darwin-aarch64` and `darwin-x86_64`entries must point at the same new `desktop-vX.Y.Z`universal `.app.tar.gz` asset (each slice of the fat binary requests its own arch key at runtime, so both keys serve the one artifact). Installed apps — including older per-arch installs — pick the update up on next launch or within 2 hours.
10. Final response.
Report: version, tag, changelog updated, commit hash, what was pushed, workflow URL, and the feed verification result.
## Repo secrets (one-time setup)
## Publish secrets (one-time setup)
The workflow needs these repository secrets. The Apple ones come from the same
Apple Developer account used for manual signing (see the app README's "macOS
signing & notarization" section for how to obtain them):
These live on the **`PublishDesktop` environment**, not at repository level, so
only the `build` job can read them and only after an approval. Set them under
Settings → Environments → PublishDesktop → Environment secrets. The environment
also restricts deployments to `main` and requires a reviewer.
Adding one of these as a *repository* secret is the common mistake. The build
would still succeed — an environment-gated job resolves repository secrets too,
with environment values simply taking precedence — so the credential would sit
repo-wide while everything looked fine. `validate` therefore fails the run if any
of them resolves in a job with no environment. If you hit that, delete the
repository-level copy rather than duplicating it.
If a secret is missing everywhere, the preflight in `build` fails the run naming
the missing entries. The Apple values come from the same Apple Developer account
used for manual signing (see the app README's "macOS signing & notarization"
section for how to obtain them):
| Secret | Value |
| --- | --- |
@@ -124,4 +150,8 @@ signing & notarization" section for how to obtain them):
| `TAURI_SIGNING_PRIVATE_KEY_PASSWORD` | Password for that key |
The Slack + telemetry secrets (`SLACK_RELEASE_BOT_TOKEN`, `TELEMETRY_SERVICE_API_KEY`,
OTEL settings) are shared with the CLI publish workflow and already configured.
`ERROR_SERVICE_API_KEY`, OTEL settings) are shared with the CLI, SDK, and extension
publish workflows and already configured. **Do not move these into
`PublishDesktop`** — scoping them to this environment empties them in every other
publish workflow, silently, with no error beyond missing telemetry and a failed
description:Use when releasing the Cline VS Code extension — stable (currently the combined legacy+next A/B VSIX via ext-vscode-ab-package), nightly (ext-vscode-publish-nightly), or a legacy-branch hotfix (ext-vscode-publish-legacy). Guides version selection, changelog, PostHog rollout-flag coordination, workflow dispatch, environment approvals, tagging, and post-publish verification, plus the eventual cutover to publishing the SDK extension standalone.
---
# VS Code Extension Release
Use this skill when the user asks to release, publish, or ship the VS Code extension — stable, nightly, or a legacy hotfix — or to dial the rollout, or to cut over to the SDK extension permanently.
> Working directory: repo root. All workflows are dispatched from `main` (GitHub requires the workflow file on the default branch; each workflow checks out the refs it actually builds).
## The current era: combined A/B rollout
We are mid-migration from the legacy (npm, pre-SDK) extension to the next (SDK-based, bun) extension. Until the cutover is complete, **the stable and nightly listings ship a combined VSIX**: a small loader + two complete extensions (`next/` built from `main`, `legacy/` built from the `legacy-extension` branch). The loader picks one per window based on the PostHog flag `ext-sdk-bundle-rollout`. Deep-dive docs: `apps/vscode-rollout/README.md` (authoritative) and PR #12253 (design + runbook comments).
Endgame (see "Cutover" at the bottom): once the next bundle is trusted at 100%, stable goes back to a plain build of `main` via `ext-vscode-publish-stable.yml` and all the legacy/rollout machinery is retired.
### The listings and the workflows
| Channel | Marketplace ID | Workflow | Trigger | Version |
| Nightly (combined) | `saoudrizwan.cline-nightly` | `ext-vscode-publish-nightly.yml` | cron 12:00 UTC + dispatch | auto `<major>.<minor>.<unix-ts>` from main's `apps/vscode/package.json` |
| Legacy hotfix (standalone) | `saoudrizwan.claude-dev` | `ext-vscode-publish-legacy.yml` | dispatch | from `apps/vscode/package.json` on `legacy-extension` |
| Stable standalone (post-cutover) | `saoudrizwan.claude-dev` | `ext-vscode-publish-stable.yml` | dispatch | from `apps/vscode/package.json` on `main` |
All three publish paths gate on tests before publishing: nightly and ab-package run the reusable bun suite (`ext-vscode-test.yml`, tests `main`) — ab-package additionally runs the legacy branch's npm suite — and the legacy workflow inlines the npm suite. Environment gates: stable paths use `publish` → `Publish` environment (required reviewers approve in the Actions UI); nightly uses `PublishNightly` (branch policy only, no reviewers — a reviewer requirement would block the cron).
## Golden rules (read before any release)
1.**One listing, one version line.**`claude-dev` is published from multiple workflows/branches. Every stable publish must use a version **strictly above the highest version ever published to the listing from any branch** — marketplace versions are monotonic and cannot be unpublished (supersede, never delete). Check what's live first:
```bash
curl -s -X POST "https://marketplace.visualstudio.com/_apis/public/gallery/extensionquery" \
`ext-vscode-ab-package` also enforces this automatically for `publish=true` runs: a preflight job validates the version format (plain `X.Y.Z`) and hard-fails unless it exceeds the live Marketplace version, and the publish job re-checks right before publishing (the approval wait can last days — a legacy hotfix landing in between is caught). Still run the query yourself when *choosing* the version.
2. **Check the flag BEFORE any stable combined publish.** `ext-sdk-bundle-rollout` is **shared between nightly and stable** — the loader sends only a machine id to `/decide`, no channel property, so there is no per-channel targeting. If the flag is high (nightly dogfooding) and you publish stable, stable users get the next bundle at that same percentage. Verify the effective percentage empirically (no PostHog admin needed — sample `/decide` with random ids using the key inlined in any shipped loader):
```bash
node -e '
const KEY = process.argv[1]; // phc_... extracted from a shipped VSIX loader
Flag changes are made in the PostHog UI (Cline project). **0% is the kill switch** — the flag is two-way; there is no separate killswitch flag. Dialing down demotes machines back to legacy on their next window reload.
3. **Ask before pushing** commits or tags. Environment approvals are the maintainer's to give.
4. **Changelog lives at the repo ROOT** (`CHANGELOG.md`), on the branch being released — not `apps/vscode/CHANGELOG.md` (doesn't exist). The legacy and stable workflows hard-fail unless the first heading is exactly `## [<version>]`.
5. **Stuck concurrency groups**: `ext-vscode-ab-package` groups on the version with `cancel-in-progress: false`. Only `publish=true` runs wait on environment approval (build-only rehearsals run ungated to completion), but a publish run left `waiting` still blocks every later dispatch of the same version — cancel it (`gh run cancel <id>`) before re-dispatching.
## Stable release (combined A/B VSIX) — the current stable path
### Pre-flight
```bash
# 1. What's live, and what version comes next (must exceed it — rule 1)
# 2. Flag percentage (rule 2) — decide where it should be for this release
# 3. Legacy tip = what the non-promoted cohort will run; confirm it's the shipped hotfix line
git fetch origin main legacy-extension
git log --oneline -3 origin/legacy-extension
# 4. Cheap local rehearsal of the most likely build failure: the union manifest
# hard-fails if views/viewsContainers/configuration diverged between branches.
git show origin/main:apps/vscode/package.json > /tmp/next.json
git show origin/legacy-extension:apps/vscode/package.json > /tmp/legacy.json
- Add `## [<VERSION>]` entry at the top of root `CHANGELOG.md`.
- Bump `apps/vscode/package.json` to `<VERSION>` so the repo reflects the published line. Side effect: nightly versions become `<major>.<minor>.<unix-ts>` of the new base — harmless (separate listing, still monotonic).
### Dispatch
```bash
gh workflow run ext-vscode-ab-package.yml --ref main \
# publish=false builds an installable .vsix artifact without publishing and
# needs NO environment approval — the ungated build job uploads the artifact
# and the run completes.
gh run list --workflow=ext-vscode-ab-package.yml --limit 1
```
Preflight (version format + monotonicity) and both test suites run first, then the ungated `build` job packages and uploads the VSIX; for `publish=true` the `publish` job then **waits for `Publish` environment approval** (Actions → run → "Review deployments"). Both bundles build the exact revisions their test gates ran against (branch names are resolved once — commits landing on either branch mid-run or during the approval wait are not picked up); `publish=true` is additionally refused for any `next-ref` other than `main` (the bun gate only tests main — non-main next-refs are for build-only artifact rehearsals). Check what a run is waiting on:
```bash
gh api repos/cline/cline/actions/runs/<run-id>/pending_deployments
```
### Post-publish
1. Verify the marketplace serves the new version (query from rule 1) — expect minutes-to-an-hour of validation lag after "Published" appears in the logs. Also verify Open VSX:
2. Tag, GitHub Release (with the .vsix attached), and the Slack release-bot post happen **automatically** after a real publish (all `continue-on-error` — the publish itself already succeeded, so bookkeeping failures leave the run green). Verify they landed; the known failure is the tag push when the built commit touches `.github/workflows/**` (default token cannot create such refs — no grantable permission fixes it). Manual fallback:
```bash
git tag v<VERSION> <main-sha-built> # ask before pushing
A real publish also **hard-fails early** if root `CHANGELOG.md` on the built main revision doesn't start with `## [<VERSION>]` — the release prep PR must be merged before dispatching.
3. Thorough artifact check (`gh run download <run-id>`): union `package.json` is `saoudrizwan.claude-dev@<VERSION>`, `next/package.json` and `legacy/package.json` carry the SAME version, `grep -c 'phc_' extension/extension.js` ≥ 1 (loader key inlined), no leftover `process.env.TELEMETRY_SERVICE_API_KEY` / `process.env.CLINE_ROLLOUT_VARIANT` literals in either bundle's dist (leftovers = a build ran without its env and telemetry is silently dead).
4. Monitor: `extension.rollout.bundle_activated` in `otel.otel_logs` filtered to `extension_version = '<VERSION>'` (stable cohort is cleanly separable — nightly versions are timestamps). Watch the next/legacy ratio and the crash-fallback rate; Metabase dashboards 17 (rollout + task error rate) and 19 (error deep dive). `extension.rollout.loader_decision` (incl. `double_failure`) is PostHog-only, not in ClickHouse.
5. Dial the flag per the rollout plan (e.g. 0% at publish → 1% → up), verifying each change with the probe from rule 2. Announce demotions ahead of time — dialing down also demotes nightly dogfooders unless they set `"cline-nightly.rollout.bundleOverride": "next"`.
### Known caveats of this path
- **`engines.vscode` unions upward** (main's floor wins, e.g. `^1.101.0` vs legacy's `^1.84.0`): users on older VS Code are never offered the combined VSIX. Fail-safe during rollout; must be resolved before 100%.
- A red run can still mean a successful publish on paths that tag (see Gotchas).
gh workflow run ext-vscode-publish-nightly.yml --ref main # real publish
gh workflow run ext-vscode-publish-nightly.yml --ref main -f dry-run=true # artifact only
gh run watch <run-id> --exit-status --interval 60
```
No changelog/version prep — the version is computed. Verify with the marketplace query against `saoudrizwan.cline-nightly`.
**Red run ≠ failed publish**: the final tag-push step fails whenever main's HEAD touches `.github/workflows/**` (default token cannot create such refs). If "Published" appears in the logs, the release went out; push the `nightly-main-<UTC ts>-<sha12>` tag manually with user credentials.
## Legacy hotfix release (and emergency full rollback)
For shipping a fix on the `legacy-extension` branch — or as the **structural rollback** from a bad combined stable VSIX: a standalone legacy publish at a higher version supersedes the combined VSIX entirely (loader and all) for every user. (For "next bundle misbehaving" you don't need this — dial the flag to 0% instead.)
```bash
# On legacy-extension: commit the fix, bump apps/vscode/package.json ABOVE the
# highest version ever published to the listing (rule 1 — including combined
# versions, e.g. combined 4.1.0 live -> hotfix is 4.1.1, not 4.0.13),
# add the matching `## [x.y.z]` entry to root CHANGELOG.md, push.
gh workflow run ext-vscode-publish-legacy.yml --ref main \
npm test suite runs ungated; the publish job waits on the `Publish` environment. This workflow derives + pushes the `v<version>` tag itself and creates the GitHub release — no manual tagging. Publishes to Marketplace **and** Open VSX. The branch is the npm codebase: use `npm`, never `bun`, and expect the old monolith layout (`apps/vscode/src/core/...`).
## Cutover: retiring the A/B machinery (the endgame)
When the next bundle has held at 100% long enough to trust:
1. **Resolve the engines floor**: decide whether stranding VS Code < main's `engines.vscode` on the last combined version is acceptable, or lower main's floor first.
2. Bump `apps/vscode/package.json` on `main` above everything ever published; root `CHANGELOG.md` entry to match (both are enforced by the workflow).
3. Ship standalone from main: `gh workflow run ext-vscode-publish-stable.yml --ref main` — tests main, tags `v<version>` itself, creates the GitHub release, publishes Marketplace + Open VSX.
4. Watch the same rollout telemetry through the transition — `extension_variant` disappears from events as users leave combined builds, which is itself the adoption signal.
5. Only after the standalone version dominates: retire `legacy-extension` (keep for history), delete `ext-vscode-publish-legacy.yml` and `ext-vscode-ab-package.yml`, convert the nightly workflow back to a plain build of main, remove `apps/vscode-rollout/`, and archive the `ext-sdk-bundle-rollout` flag in PostHog (harmless to machines still on a combined VSIX: absent flag fails safe to... nothing changing until they update, but their loader treats a deleted flag as legacy — leave the flag at 100% until combined-VSIX activations flatline, then archive).
6. Update this skill: delete the combined-era sections and keep the standalone flow.
## Gotchas index
- `inputs.*` are empty strings on `schedule` events — preserve `|| 'default'` fallbacks when editing the nightly workflow.
- `bun run package` in `apps/vscode` does not build `@cline/*` workspace deps — fresh checkouts need `bun run build:sdk` first (workflows handle this).
- Job-level `if:` ref checks in workflow YAML are advisory (a dispatched branch runs its own copy of the file); the enforced boundary is each environment's deployment-branch policy in repo settings.
- Marketplace PATs (`VSCE_PAT`/`OVSX_PAT`) are only mounted into publish steps; neither publish workflow has an untrusted trigger surface.
- Environment-approval runs left waiting don't time out quickly — they sit for days and (for ab-package publish runs) block their version's concurrency group.
- Local forcing for manual testing: `CLINE_BUNDLE_OVERRIDE=next|legacy` env (launch VS Code fresh from a terminal) or the `<prefix>.rollout.bundleOverride` setting + reload; both report as `override` in telemetry so they don't pollute cohort data.
Drive and test terminal apps (especially the Cline CLI TUI in apps/cli) through tuistory — named background PTY sessions that agents can read, wait on, snapshot, screenshot, and type into. Like Playwright/tmux for terminals, with reactive waiting instead of blind `sleep`.
Use this skill when you need to:
- Manually test or reproduce bugs in the interactive Cline TUI (`bun run cli -i`) from a headless environment
- Run a dev server or any long-lived/interactive process in the background without hanging your tool call
- Write or extend Playwright-style e2e tests for the TUI (`bun run test:e2e:tuistory` in apps/cli)
- Capture text snapshots or styled PNG screenshots of a TUI screen as evidence
---
# tuistory
[tuistory](https://github.com/remorses/tuistory) wraps any terminal command in a named background PTY session backed by a Ghostty terminal emulator. Agents interact with the session via short CLI calls that return instantly; humans can `tuistory attach` to the same session to watch or intervene. No real terminal or display (`DISPLAY`) is needed — it works fully headless, which makes it the preferred way for cloud agents to exercise the Cline TUI.
It is installed as a devDependency of `@cline/cli`, so the pinned binary resolves when you run from `apps/cli`:
```bash
cd apps/cli
bunx tuistory --help # source of truth for commands, options, and syntax
```
For full upstream docs: `curl -s https://raw.githubusercontent.com/remorses/tuistory/refs/heads/main/README.md`
## Driving the Cline TUI headlessly
Launch the TUI in an isolated environment so you don't touch real user config (`~/.cline`):
-- bun src/index.ts --provider anthropic -m claude-sonnet-4-6 -k test-key
```
The dummy `-k test-key` renders the full chat UI; only an actual agent turn would fail. For recorded LLM turns, use the VCR cassettes described in `apps/cli/src/tests/helpers/env.ts` (`CLINE_VCR=playback` + `CLINE_VCR_CASSETTE`). Real turns need a provider credential (e.g. `ANTHROPIC_API_KEY`, `CLINE_API_KEY`).
Then use an **observe → act → observe** loop:
```bash
# Wait reactively for the chat view — never use sleep
bunx tuistory -s cline wait"What can I do for you?" --timeout 30000
# Act, then always observe the resulting screen state
bunx tuistory -s cline type"/settings"
bunx tuistory -s cline snapshot --trim
bunx tuistory -s cline press enter
bunx tuistory -s cline snapshot --trim
# Styled PNG of the current screen (prints the file path) — good for artifacts
bunx tuistory -s cline screenshot
# Full raw output stream (snapshot shows only the visible screen)
bunx tuistory read -s cline --all
# Tear down a session YOU started (double Ctrl+C exits the TUI cleanly)
bunx tuistory -s cline press ctrl c
bunx tuistory -s cline press ctrl c
bunx tuistory -s cline close
```
## Background processes (instead of tmux)
```bash
bunx tuistory -s my-server -- bun run dev:sidecar # returns immediately
bunx tuistory read -s my-server # new output since last read
bunx tuistory -s my-server restart # after code changes
```
## Key rules
- **Options before `--`, command after.** Everything after the first `--` is passed verbatim to the child: `tuistory -s name --cols 150 -- bun src/index.ts` is correct.
- **Snapshot after every action.** TUIs are stateful; dialogs and errors can render over the view you expect. `snapshot` reflects what the user actually sees (occluded text does not count), unlike grepping the raw stream.
- **Wait, never sleep.** `wait "text"` / `wait "/regex/i"` (case-sensitive by default) reacts as fast as the terminal updates; `wait-idle` when you don't know what to expect. Always pass `--timeout`.
- **Keys land instantly.** Unlike sleep-based scripts, a queued second keypress can leak into the next view (e.g. one Enter both accepts a slash completion and submits it).
- **Never close a session you didn't start.** Sessions are shared with humans (`tuistory attach -s name`) and other agents. Default to leaving sessions running; use `read`/`wait`/`snapshot` to inspect without disrupting.
env: isolatedEnv,// see createCliEnv() in the reference test
cols: 120,
rows: 36,
waitForDataTimeout: 30_000,// CLI cold start compiles a large TS graph
});
awaitsession.waitForText("What can I do for you?",{timeout: 30_000});
constscreen=awaitsession.text({trimEnd: true});// emulated screen state
awaitsession.type("/settings");
awaitsession.press("enter");
session.close();// always close in test teardown
```
Screen-state assertions can check that stale UI is *gone* (`expect(screen).not.toContain(...)`), which stream-grepping harnesses cannot. `session.text({ only: { bold: true } })` filters by style; `session.read()` returns the raw stream since the last read.
- Enter any Vertex model ID by hand, including models the catalog doesn't list yet.
- Support Fable 5 on Vertex.
### Changed
- Show the full model catalog for every Vertex region instead of filtering the picker down to a hardcoded list of global-endpoint models, which lagged behind every model launch. Picking a model the region doesn't serve now fails at request time with recovery guidance in the error row.
- Report Fable 5 cost on Vertex as unknown rather than applying Anthropic's list price, which understated what Vertex actually bills — its rates are region-dependent.
- Make the auto-approve menu the single source of truth for unattended runs and remove the Yolo Mode toggle, which was cosmetic: nothing in the approval path read it. Setups that had Yolo Mode (or auto-approve-all) turned on are migrated to auto-approving every action, so they keep running unattended.
### Fixed
- Respect your configured max output tokens when the compaction summarizer requests a summary.
- Remove the stale "Double-Check Completion" feature tip.
## [4.1.7]
### Added
- Restore the "View Changes" button on completion rows, backed by SDK checkpoints, so you can review everything a task touched from the completion card.
- Bring back a copy button on turn-final response rows.
- Support pre-registered OAuth clients for remote MCP servers, for setups where dynamic client registration isn't available.
### Changed
- Fade the "View Changes" button until changes since the last message are confirmed, and hide it entirely when there is nothing to show.
- Centralize plugin settings and contributions, with host-aware snapshots and atomic plugin toggles.
- Carry execution context in scheduled run reports — readable headers, schedule metadata, durations, and lifecycle error details.
### Fixed
- Preserve prompts queued during a turn when that turn is interrupted: they survive aborts, are drained after a turn aborts itself, and the stop is surfaced instead of the queue being silently dropped.
- Keep session context durable across aborts and hub restarts, so an interrupted session resumes with the state it had.
- Settle the turn phase when a mode switch aborts a running turn.
- Report queued-turn failures as `run.failed` instead of letting them complete silently.
- Keep a hung MCP server from taking down session creation, and give stdio servers that were never configured a 30-second initialize budget instead of blocking indefinitely.
- Surface OAuth authorization for SSE MCP servers on a 401 instead of failing outright.
- Route LiteLLM through Chat Completions instead of the Responses API, fixing requests against LiteLLM proxies.
- Retry network interruptions that happen mid-stream but before any model output, instead of failing the turn.
- Use the configured fetch for Vertex ADC token refreshes, so they work behind proxies and custom transports.
- Include files that were untracked when a snapshot was taken in checkpoint diffs, and pick up checkpoints when git is initialized part-way through a session.
- Fall back to the session cwd or Desktop for @-mention file search in empty windows.
- Never run a foreign compiled plugin-sandbox bootstrap for a source host.
## [4.1.6]
### Added
- Offer `meta/muse-spark-1.2-contributor` on the Cline provider, alongside a refreshed model catalog.
### Fixed
- Attribute error telemetry to the model actually in use for a run, so failures are no longer reported against the wrong model.
## [4.1.5]
### Added
- Explain when a free model promotion ends. Requests to a retired free model now show a dedicated notice with a button to pick another model, instead of a generic error with nothing but a Retry prompt.
### Changed
- Map reasoning settings onto a shared path across AI SDK providers, so effort levels and enable/disable toggles behave consistently (including on Ollama) instead of relying on per-provider overrides.
## [4.1.4]
### Added
- Recognize Chutes as a provider.
- Show skills alongside workflows in the slash command menu, and disambiguate commands that share a name instead of letting one shadow the other.
### Changed
- Remove model-initiated plan-to-act switching. Switching out of plan mode is now driven by you, not by the model deciding mid-turn.
- Hard-block file-editing shell commands in plan mode instead of relying on prompting alone. Read-only investigation still works, but file manipulation, in-place editors, redirection to files, mutating git subcommands, and package installs are refused.
### Fixed
- Stop treating a turn that completes with a plan as a failed turn when a plan-blocked command was its only tool call. The turn no longer ends in the error state with a Retry footer, and toggling to Act correctly re-runs the presented plan instead of appearing to do nothing.
- Show tool paths relative to the workspace in the chat view instead of absolute paths.
- Reset pending attachments when starting a new task, so images from the previous task no longer carry over.
- Surface a clear error when the selected provider has no API key configured, instead of a generic failure.
- Refresh MCP tool and resource lists when a server sends a `list_changed` notification, instead of only showing a toast.
- Show installed plugins under their real package names instead of all appearing as "index".
- Correct the Linux keybinding label in the Plan/Act mode tooltip.
- Recover from running out of context instead of failing with a raw provider error — the run compacts and retries once, and the cases that genuinely cannot be recovered explain why.
- Retry empty model responses on every provider rather than only Ollama, fixing hard "Model returned empty response" failures on OpenRouter, Cline, and OpenAI-compatible endpoints.
- Stop Claude 4.6+ and 5.x models being rejected with "thinking.type.enabled is not supported" when they resolve from the offline catalog or from a hand-typed model id.
- Restore Bedrock prompt caching, which reported zero cache reads and writes because the provider sent a cache format Bedrock discards, and route Bedrock foundation models through geo inference profiles.
- Send `max_completion_tokens` for reasoning models on OpenAI-compatible endpoints, and substitute image content for models without image support instead of failing the request.
- Inherit the MiniMax default model from models.dev, and refresh the bundled catalog, which adds Infomaniak and SCX.ai.
- Report the same provider failure once instead of twice in error telemetry, and rate-limit repeated failures from unattended retry loops.
## [4.1.3]
### Fixed
- Stop the two bundles of the combined rollout package from invalidating each other's Cline account session. A still-open legacy window that refreshed its token after the machine was promoted to the new extension would consume the shared refresh token, producing spurious "Unauthorized" / re-authenticate prompts and unexpected sign-outs. Promoted legacy windows now keep working on their current session and offer a one-time Reload Window prompt instead.
- Fall back to the default Cline model when migrating a setup that references a model id the new extension doesn't recognize, instead of leaving the provider unconfigured.
- Restore reliable checkpoints: checkpoints are created consistently, and restoring one now rewinds the whole workspace rather than a subset of files.
- Keep settings edits that are made before the provider config finishes loading — base URLs, API keys, and the Qwen/Moonshot API line are no longer silently discarded.
- Stop losing keystrokes in custom base URL fields, and keep the custom URL checkbox state after a failed clear.
- Use the AskSage custom API URL at inference time instead of ignoring it.
- Settle a pending tool approval when an edited message replaces the session, so the task no longer hangs waiting on a prompt that is gone.
- Drop attachments from messages that have been edited.
- Complete terminal commands when the shell execution ends, so tasks no longer stall on commands that already finished.
- Include untracked files when generating commit messages.
- Run Windows Store PowerShell profiles correctly.
- Surface the upstream provider error when a gateway-forwarded stream fails, instead of a generic failure.
- Retry empty Ollama responses at the model boundary, and raise the response-start timeout to 5 minutes so cold model loads no longer error out.
- Show proper display names for Cline free models and recommended models in the model picker.
- Preserve video input capability for models that support it.
- Keep the plan/act input border in sync with the actual textarea focus.
## [4.1.2]
### Added
- Show which extension variant is active — "Legacy" or "Next" — next to the version in the settings About page, in both bundles of the combined rollout package.
## [4.1.1]
### Changed
- Remove vestigial MCP server-key machinery from McpHub — native MCP tool calls now route by server name instead of a random in-memory uid, so routing survives restarts and server list changes.
## [4.1.0]
### Changed
- Convert the stable extension to a combined A/B package: one VSIX containing both the current (legacy) extension and the new SDK-based extension, plus a loader that activates exactly one per window via a staged remote rollout. For nearly all users nothing changes — the loader activates the same extension as 4.0.12; a small percentage (starting at 1%) is gradually opted into the SDK-based extension. If the new extension fails to activate, the loader falls back to the current one in the same window. Settings and credentials are shared between the two.
## [4.0.12]
### Added
- Add support for free Cline models, shown as "(free)" in the model picker, with a dedicated error card that includes the reset time when the free limit is reached.
### Fixed
- Keep Claude Code responses that were already streamed when the CLI exits with a max-turns error, instead of discarding a valid response.
## [4.0.11]
### Added
- Add Claude Opus 5 across the Anthropic, Claude Code, Bedrock, Vertex, Cline, and OpenRouter providers, including 1M context window variants.
- Add Moonshot Kimi K3 support.
- Include the host plugin version in telemetry events.
### Fixed
- Correct pricing for the Claude Opus 1M context variants, which overstated costs for requests above 200k tokens.
- Enable native tool calling for Kimi K3 models, fixing empty responses.
## [4.0.10]
### Added
- Add telemetry to track when Cline reaches the consecutive mistake limit.
## [4.0.9]
### Added
- Add GPT-5.6 ChatGPT subscription models.
### Changed
- Soften and shorten the message shown when Cline hits the consecutive mistake limit.
### Fixed
- Handle cumulative usage snapshots from OpenAI-compatible providers so token counts are no longer over-reported.
- Load skills from files saved as UTF-8 with a byte-order mark (BOM).
## [4.0.8]
### Added
- Add more models to the GCP Vertex provider, plus a free-form entry option in the model dropdown for specifying custom Vertex models.
## [4.0.7]
### Added
- Add a ClinePass limit-reached error with a one-click option to switch to Cline usage-based billing.
- Allow selecting Cline free models on the ClinePass provider, organized into Subscribed and Free tabs with model descriptions.
### Changed
- Refine ClinePass onboarding and provider settings copy, and open the "learn more" link via the in-app URL handler.
- Remove the Cline model picker recommendation copy.
### Removed
- Remove all references to GLM 5.1.
## [4.0.6]
### Fixed
- Generalize the model capability warning so it applies more broadly.
## [4.0.5]
### Added
- Add support for Claude Sonnet 5 across the Anthropic, Bedrock, Vertex, Claude Code, SAP AI Core, OpenRouter, and Vercel AI Gateway providers, including model picker and recommended-model updates.
## [4.0.4]
### Changed
- Fully remove the ClinePass feature flag so ClinePass is available everywhere in the UI — onboarding, settings, the welcome promo banner, and the credit-limit "Switch to ClinePass" action.
## [4.0.3]
### Changed
- Enable the ClinePass provider for all users by removing the feature-flag gate that previously fell back to the standard Cline provider.
## [4.0.2]
### Added
- Add reasoning effort support (including `xhigh`) for DeepSeek thinking models.
- Improve the ClinePass provider experience with clearer reasoning controls and model selection.
### Fixed
- Show reasoning effort controls for ClinePass models and align ClinePass model resolution with the rest of the provider.
- Prefer canonical Cline Z.ai model ids and polish ClinePass and Z.ai model metadata.
- Fix environment variable replacement in the webview.
- Default focus chain settings in webview state so the toggle reflects the correct value on load.
## [4.0.1]
### Changed
- Roll the stable VS Code extension back to the pre-SDK-migration codebase to resolve regressions reported in 4.0.0. This release ships the 3.89.2 extension code under a higher version number so existing 4.0.0 users receive the update. SDK-migration work continues separately on `main`.
@@ -7,7 +7,7 @@ We're thrilled you're interested in contributing to Cline. Whether you're fixing
Bug reports help make Cline better for everyone! Before creating a new issue, please [search existing ones](https://github.com/cline/cline/issues) to avoid duplicates. When you're ready to report a bug, head over to our [issues page](https://github.com/cline/cline/issues/new/choose) where you'll find a template to help you with filling out the relevant information.
<blockquote class='warning-note'>
🔐 <b>Important:</b> If you discover a security vulnerability, please use the <a href="https://github.com/cline/cline/security/advisories/new">Github security tool to report it privately</a>.
🔐 <b>Important:</b> If you discover a security vulnerability, please use the <a href="https://github.com/cline/cline/security/advisories/new">GitHub security tool to report it privately</a>.
- Fixed the CLI reconnecting to a stale Hub daemon after an upgrade. Hub daemons now carry a runtime build fingerprint, so an upgraded CLI retires and respawns a daemon still running older code instead of attaching to it (from SDK v0.0.73)
- Fixed compaction being silently skipped on reasoning models. The summarizer no longer hardcodes a 1024-token output cap — it honors your max output tokens setting, defaults to 4096 (lowered when the model reports less), and logs a diagnostic when a summary comes back empty (from SDK v0.0.73)
- Added Fable 5 (`claude-fable-5`) to the Vertex model catalog. Pricing is intentionally omitted because Vertex bills region-dependently, so cost shows as unknown rather than wrong (from SDK v0.0.73)
- Custom Vertex model IDs are now passed through unchanged, routing Claude-style IDs to the Anthropic-on-Vertex path (from SDK v0.0.73)
## 3.0.52
- Added `cline mcp uninstall` for removing an installed MCP server
- Schedules now reuse your saved provider settings instead of needing provider configuration of their own
- Queued messages are legible on light-theme terminals — they were previously rendered in a color that washed out against a light background
- MCP tool results render as readable text in the TUI instead of escaped JSON, and binary payloads survive being expanded instead of being mangled
- Malformed tool input/output payloads no longer break rendering — the formatters degrade gracefully instead of throwing
- Prompts queued during a turn now survive being interrupted: they are preserved across aborts, drained after a turn aborts itself, and the stop is surfaced instead of leaving the queue silently dropped (from SDK v0.0.72)
- Session context stays durable across aborts and hub restarts, so an interrupted session resumes with the state it had (from SDK v0.0.72)
- A hung MCP server no longer takes down session creation, and stdio servers that were never configured get a 30-second initialize budget instead of blocking indefinitely (from SDK v0.0.72)
- Remote SSE MCP servers surface an OAuth authorization prompt on a 401 instead of failing outright, and pre-registered OAuth clients are supported for setups without dynamic client registration (from SDK v0.0.72)
- LiteLLM requests route through Chat Completions instead of the Responses API, fixing calls against LiteLLM proxies (from SDK v0.0.72)
- Network interruptions that happen mid-stream but before any model output are retried instead of failing the turn (from SDK v0.0.72)
- Vertex ADC token refreshes use the configured fetch, so they work behind proxies and custom transports (from SDK v0.0.72)
- Checkpoint diffs include files that were untracked when the snapshot was taken, and checkpoints are picked up when git is initialized part-way through a session (from SDK v0.0.72)
- Scheduled run reports carry execution context — readable headers, schedule metadata, durations, and lifecycle error details (from SDK v0.0.72)
## 3.0.51
- Reasoning effort now applies consistently across providers instead of going through per-provider thinking overrides, including Ollama, and asking for reasoning to be off is respected everywhere (from SDK v0.0.71)
-`meta/muse-spark-1.2-contributor` is now selectable on the Cline provider, alongside a refreshed model catalog (from SDK v0.0.71)
- Error telemetry now reports the model that was actually in use for the run (from SDK v0.0.71)
## 3.0.50
- Added user-selectable color themes to the interactive TUI. Pick one with `/theme`, the command palette, or the Theme row in `/settings` — the picker previews each theme live. Built-in themes are Auto (terminal-adaptive, the default), Cline Dark, Cline Light, Tokyo Night, Gruvbox Dark, Nord, Dracula, Catppuccin Mocha, One Dark, Solarized Dark, and Solarized Light. Named themes paint the background, foreground, accents, syntax highlighting, and diff colors, and `CLINE_THEME` overrides the persisted choice at startup
- The git branch shown below the prompt now updates when you switch branches from another terminal or your editor, instead of showing whatever was checked out when the TUI started
- Telegram slash commands such as `/clear` now reach the connector command host — the Telegram library was intercepting them and they were silently dropped
- Racing connector launches no longer collide: an instance is claimed before it opens socket mode, the hub supervises connector processes, and `doctor`/`connect` skip connectors that are already starting. Connector tools are also enabled by default, and the Slack greeting is no longer replayed on reconnect
- Auto-approval settings are now honored over ACP
- Plan mode now hard-blocks file-editing shell commands instead of relying on prompting alone — `run_commands` stays available for read-only investigation, but file-manipulation commands, in-place editors (`sed -i`, `perl -i`), redirection to files, mutating git subcommands, package installs, and nested command strings (`sh -c`, `eval`, `sudo`) are rejected, on Windows and PowerShell too (from SDK v0.0.70)
- A turn that ends with a completed plan is no longer rendered as a failed turn when a plan-blocked command was its only tool call
- Running out of context is now recovered from instead of failing with a raw provider error: the run force-compacts and retries once, and the cases that genuinely cannot be recovered report why (from SDK v0.0.70)
- Empty model responses are now retried on every provider, not just Ollama — OpenRouter, Cline, and OpenAI-compatible endpoints previously failed the task outright with "Model returned empty response" (from SDK v0.0.70)
- Claude 4.6+ and 5.x models are no longer rejected with "thinking.type.enabled is not supported" when they resolve from the offline catalog or from a hand-typed model id (from SDK v0.0.70)
- Bedrock prompt caching works again — the provider was sending a cache format Bedrock silently discards, so cache reads and writes were always 0 — and Bedrock foundation models are now routed through geo inference profiles (from SDK v0.0.70)
- Reasoning models on OpenAI-compatible endpoints now receive `max_completion_tokens` instead of the rejected `max_tokens`, and requests to models without image support substitute the image content instead of failing (from SDK v0.0.70)
- MiniMax now inherits its default model from models.dev, and the model catalog picked up two new providers, Infomaniak and SCX.ai (from SDK v0.0.70)
- Upgraded the model layer to AI SDK 7 and switched Ollama to the native AI SDK provider (from SDK v0.0.70)
- Error telemetry no longer reports the same provider failure twice, and repeated failures from unattended retry loops are rate-limited (from SDK v0.0.70)
## 3.0.49
-`/undo` works again once the agent has used tools — the checkpoint picker counted tool results as user turns, so restore aborted with "Could not find user message for run N"
- Checkpoints are actually created again; a run-boundary regression meant none were ever recorded in the CLI (from SDK v0.0.69)
- Checkpoint restore is now a full workspace rewind: files Cline created during the task come back at their checkpoint-time content and files created after the checkpoint are removed, while `.gitignore`d paths (build output, `node_modules`, `.env`) are left alone (from SDK v0.0.69)
- After a restore, the rewound message is prefilled as plain text instead of the raw `<user_input mode="act">` envelope
- Ollama's response-start timeout is now 5 minutes instead of 30 seconds, so cold-loading a large local model no longer errors out mid-load (from SDK v0.0.69)
- Empty Ollama responses are now retried instead of failing the task with "Model returned empty response" (from SDK v0.0.69)
- Migrated users whose stored Cline model id isn't in the catalog now fall back to the default model instead of sending an unknown model id on every request (from SDK v0.0.69)
- The ClinePass promo dialog can be dismissed with any key (Enter still opens the subscription page), and it is marked as shown when it appears, so force-quitting no longer replays it on every launch
- Opening a URL no longer crashes the CLI on hosts without an opener binary (headless Linux without `xdg-open`); WSL2 containers now use `xdg-open`, Windows tries the absolute PowerShell path first, and `cline doctor log` converts Linux paths to `\\wsl$` UNC paths
- The hub now restarts through the installed wrapper after a Unix self-update, so npm cannot reuse a deleted cached executable
- ACP: ClinePass is selectable as a provider, organizations can be selected, session resolution and text rendering on session restart are fixed, and agent errors now describe the actual failure
- Provider errors forwarded through the Vercel AI Gateway now surface the real upstream message instead of a raw Zod dump or `[object Object]` (from SDK v0.0.68)
- Cline free models and recommended models now show their real display names in the model picker (from SDK v0.0.68)
- Sessions rooted at the filesystem root (`/`) no longer fail every command (from SDK v0.0.68)
- On Windows, PowerShell commands now travel over UTF-8 stdin, so non-ASCII commands survive the active code page and long commands are not capped by the command-line limit (from SDK v0.0.68)
- The live model catalog no longer drops the video input capability (from SDK v0.0.68)
- Removed the CLI promo code flow
## 3.0.48
-`cline history` now opens inside the existing TUI, with resume and delete actions, instead of rendering a second view in the same process
- Connector threads (Slack, Discord, Telegram, Linear, Google Chat, WhatsApp) now recover when the session they were bound to is gone — the stale binding is dropped and the turn replays against a new session, instead of failing with "session not found" until `threads.json` is edited by hand
-`cline --help` now reports the real default `--config` and `--data-dir` paths
- The per-server `timeout` in `cline_mcp_settings.json` is now honored for `initialize`, `tools/list`, and `tools/call`, so slow MCP servers no longer fail against a hardcoded 5s limit (from SDK v0.0.67)
- Reasoning controls are now routed from the models.dev catalog across providers, with clamped budgets and correct per-provider encoding (from SDK v0.0.67)
- OpenRouter now defaults to `anthropic/claude-sonnet-5` (from SDK v0.0.67)
- Fixed the China and international endpoint toggles being ignored for Qwen, Moonshot, and Z AI (from SDK v0.0.67)
- Legacy API keys are now migrated for every secret-backed provider (from SDK v0.0.67)
- Legacy OpenAI Compatible model-info overrides now survive into the seeded `models.json` (from SDK v0.0.67)
- Fixed auto-compaction state being rejected as stale, which added a redundant summarizer call on every turn past the compaction trigger (from SDK v0.0.67)
- Fixed checkpoint restores across session resumes (from SDK v0.0.67)
- Tool calls that pass line numbers as strings (`insert_line`, `read_files` bounds) are now accepted instead of erroring (from SDK v0.0.67)
- A legacy single-file `.clinerules` no longer aborts the config scan (from SDK v0.0.67)
- Plugins can now emit telemetry through `ctx.telemetry` (from SDK v0.0.67)
## 3.0.47
- Free Cline models are now supported end to end: free models show as "(free)", and hitting the free limit renders a dedicated card with the reset time (from SDK v0.0.66)
-`/settings` general toggles (plan/act mode, tool auto-approve, compaction mode) now persist across restarts
- Upgraded the TUI stack from opentui 0.1.102 to 0.4.3
- Fixed a grey panel left behind on screen after closing a dialog (model picker, help, command palette) — a leftover from the opentui upgrade
- Fixed a React duplicate-key warning when `read_files` listed the same path more than once
- Aborting a task no longer risks killing the shared hub daemon
- Connector status delivery failures are no longer fatal to the turn
- Agentic compaction is now the default context-compaction strategy, with fixes for it silently falling back to basic compaction and for tool-heavy transcripts that could never find a cut point (from SDK v0.0.66)
- Editor edits preserve a file's existing line endings, fixing failed exact-match edits on CRLF files (from SDK v0.0.66)
- Broader built-in provider coverage, now generated from models.dev (from SDK v0.0.66)
- Updated the bundled model catalog (from SDK v0.0.66)
## 3.0.46
- Fixed out-of-credits detection so the CLI reliably recognizes the Cline API's real `insufficient_credits` (402) error and shows the "add credits" card instead of a generic error
@@ -364,6 +367,34 @@ bun run dev -- --interactive --config /tmp/cline-test
Or set `CLINE_FORCE_ONBOARDING=1` to force the onboarding view regardless of existing config.
### Manually testing the TUI (agents / headless environments)
[tuistory](https://github.com/remorses/tuistory) is installed as a devDependency. It wraps the TUI in a named background PTY session that can be scripted from a plain shell — no real terminal or display needed. This is the preferred way for AI agents (or anyone in a headless environment) to poke at the interactive TUI:
# Wait reactively for the chat view (no sleep guessing)
bunx tuistory -s cline wait"What can I do for you?" --timeout 30000
# Interact and inspect
bunx tuistory -s cline type"/settings"
bunx tuistory -s cline press enter
bunx tuistory -s cline snapshot --trim # current screen as text
bunx tuistory -s cline screenshot # current screen as a styled PNG
# A human can watch/drive the same session from another terminal
tuistory attach -s cline
# Tear down
bunx tuistory -s cline close
```
The same engine powers the `test:e2e:tuistory` vitest suite (`src/cli.tuistory.e2e.test.ts`), which uses the programmatic `launchTerminal()` API for assertions against the emulated screen.
it("strips the <user_input> wrapper from replayed user text",()=>{
// Persisted user messages keep their runtime-generated wrapper. Replaying
// it verbatim leaked markup to the client, which rendered the unknown
// element as bare text (a one-word prompt showed up as just its content
// with the wrapper swallowed).
expect(
translateHistoricalMessage({
role:"user",
content:'<user_input mode="act">s</user_input>',
}),
).toEqual([
{
sessionUpdate:"user_message_chunk",
content:{type:"text",text:"s"},
},
]);
expect(
translateHistoricalMessage({
role:"user",
content:[
{
type:"text",
text:'<user_input mode="plan">lets do it</user_input>',
},
],
}),
).toEqual([
{
sessionUpdate:"user_message_chunk",
content:{type:"text",text:"lets do it"},
},
]);
});
it("strips mode notices and formats slash commands for display",()=>{
expect(
translateHistoricalMessage({
role:"user",
content:
'<user_input mode="plan"><mode_notice>The user switched from act mode to plan mode before sending this message.</mode_notice>\nare you okay?</user_input>',
}),
).toEqual([
{
sessionUpdate:"user_message_chunk",
content:{type:"text",text:"are you okay?"},
},
]);
expect(
translateHistoricalMessage({
role:"user",
content:
'<user_command slash="team">spawn a team of agents for the following task: inspect rpc startup</user_command>',
"Stdio MCP install requires a command after the server name, for example: cline mcp install fs -- npx -y @modelcontextprotocol/server-filesystem /tmp",
constSDK_CLINE_PASS_SUBSCRIPTION_MESSAGE=`No access to ClinePass subscription models yet. Subscribe to ClinePass, the low cost open weights model coding plan: ${CLINE_PASS_SUBSCRIPTION_URL}`;
constCLI_CLINE_PASS_SUBSCRIPTION_MESSAGE=`No access to ClinePass subscription models yet. Subscribe to ClinePass, the low cost open weights model coding plan: ${CLI_SUBSCRIPTION_URL}`;
constCLI_CLINE_PASS_SUBSCRIPTION_MESSAGE=`No access to ClinePass subscription models yet. Subscribe to ClinePass, the low cost open weights model coding plan: ${getCliSubscriptionUrl()}`;
constCLINE_PASS_LIMIT_DETAIL_MESSAGE=
"You have reached your 5-hour Clinepass limit. The limit resets in 5h, please try again later.";
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.