* fix(models): keep Sonnet 4.5 as default
* chore(changeset): add release note for Sonnet 4.5 default
* fix(models): remove Sonnet 4.6 from curated model lists
* fix(models): restore Sonnet 4.6 in web recommended list
* feat(cerebras): remove deprecated llama-3.3-70b and qwen-3-32b models
These models have been deprecated from the Cerebras inference platform.
- Remove llama-3.3-70b and qwen-3-32b from cerebrasModels in api.ts
- Update supported models documentation in cerebras.mdx
- Add changeset for the deprecation
* fix: remove stale llama-3.3-70b and qwen-3-32b references from rate limits
Remove dead switch cases in getRateLimits() that referenced deprecated models
no longer present in cerebrasModels.
* feat(cli): add /skills slash command for managing skills
- Add /skills to CLI_ONLY_COMMANDS in slashCommands.ts
- Create SkillsPanelContent component with:
- Display global and workspace skills with toggle indicators
- Enter to use skill (inserts @path into input)
- Space to toggle skill enabled/disabled
- Selectable marketplace link to skills.sh
- Keyboard navigation with arrow keys and vim keys
- Wire up panel in ChatView.tsx
- Add comprehensive tests for keyboard interactions
* refactor(cli): use static skill controller imports
* fix(cli): add React import to skills panel test
* fix(cli): suppress required React import lint in skills test
* fix(cli): harden /skills panel interactions
Revert optimistic skill toggle state when persistence fails, and surface a fallback URL when opening the marketplace fails. Also tighten and extend tests to verify exact marketplace URL handling and rollback behavior.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: disable click-to-set auto-condense threshold and hardcode default
Clicking anywhere on the context window progress bar silently set
autoCondenseThreshold to a value based on click position (e.g. 0.05),
persisting in globalState. This caused compaction to fire at ~10K tokens
instead of the intended ~150K, resulting in ~20 context resets per task.
- Comment out click and keyboard handlers on progress bar (keep components
for future release with proper UX)
- Hardcode threshold to 0.75 default, ignoring corrupted stored values
Co-authored-by: Cursor <cursoragent@cursor.com>
* test: add shouldCompactContextWindow unit tests
Cover threshold math including the accidental low-threshold bug case,
undefined/zero fallbacks, cache token inclusion, and maxAllowedSize cap.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: hardcode autoCondenseThreshold in all remaining callsites
Address Greptile review: SubagentRunner.ts, task/index.ts display
logic, and controller/index.ts webview state all still read the
corrupted value from globalState. Hardcode 0.75 everywhere.
Co-authored-by: Cursor <cursoragent@cursor.com>
* style: remove unnecessary union type on hardcoded threshold
Drop `number | undefined` annotation from the hardcoded 0.75 literal
in SubagentRunner.ts per Greptile review feedback.
Co-authored-by: Cursor <cursoragent@cursor.com>
* refactor: use SETTINGS_DEFAULTS constant, remove commented-out code, clarify test
- Replace hardcoded 0.75 with SETTINGS_DEFAULTS.autoCondenseThreshold
across all 4 callsites for a single source of truth
- Delete commented-out click/keyboard handlers in ContextWindow.tsx,
replace with TODO referencing PR #9348
- Make bug-case test self-documenting by deriving token values from
the threshold calculation
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix: smarter retry for write_to_file missing content parameter (#7998)
Replace generic 'missing parameter' error with progressive guidance when
write_to_file fails due to empty content parameter. This breaks the
infinite retry loop where the model repeatedly attempts the same
write_to_file call that exceeds output token limits.
Changes:
- Add writeToFileMissingContentError() to formatResponse with 3 tiers:
1st failure: suggestions (use skeleton + replace_in_file)
2nd failure: strong directive (stop retrying write_to_file)
3rd+ failure: CRITICAL stop, forces alternative strategies
- Add context window awareness: warns model when >50% context used
- Add getContextUsagePercent() helper to WriteToFileToolHandler
- Add 22 unit tests covering progressive escalation and context awareness
Fixes#7998
* add changeset for write_to_file retry fix
* refactor: simplify write_to_file error handling per review
- Simplify writeToFileMissingContentError to single-tier error following
existing diffError pattern (no progressive escalation)
- Use shared getLastApiReqTotalTokens() for context window awareness
- Remove private getContextUsagePercent() method from handler
- Add proactive skeleton + replace_in_file guidance to write_to_file
tool description for all variants
- Simplify tests to match new API (11 tests)
* test: update system prompt snapshots
* chore: revert write_to_file prompt guidance
* feat: restore progressive 3-tier guidance for write_to_file missing content
Restore the progressive escalation that was removed in dd3c12d4e:
- Tier 1 (1st failure): Gentle suggestions (skeleton + replace_in_file)
- Tier 2 (2nd failure): Strong directive, 'Do NOT attempt full write again'
- Tier 3 (3rd+ failure): CRITICAL stop, forces alternative strategies
- Context window warning when >50% full
- Dynamic UI message: 'Retrying...' vs 'multiple times — different approach'
- 21 tests covering all tiers and context awareness
* nit: extract context window warning threshold to named constant
Also replace emoji with plain text in warning message for consistency.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Robin Newhouse <robin@cline.bot>
Co-authored-by: Cursor <cursoragent@cursor.com>
Updating CHANGELOG.md format
update changelog
update banner and bump version
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* changeset version bump
* v3.62.0 Release Notes
- Banners now display immediately when opening the extension instead of requiring user interaction first
- Resolved 17 security vulnerabilities including high-severity DoS issues in dependencies (body-parser, axios, qs, tar, and others)
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions <github-actions@github.com>
- Fixes for Minimax model family
- Fixes for Response chaining for OpenAI's Responses API
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* chore(evals): reorganize eval structure with purpose-based naming
- Move evals/diff-edits/ → evals/benchmarks/tool-precision/replace-in-file/
- Move evals/cli/ → evals/legacy/cli/ (preserve for reference)
- Create evals/benchmarks/real-world/ directory
- Create evals/benchmarks/coding-exercises/cases/ directory
- Create evals/analysis/ directory structure
Note: No repositories/exercism/ directory found to move.
Skipping pre-commit hook as this is a reorganization of legacy code.
* chore(evals): remove legacy evaluation code
Remove abandoned evaluation infrastructure:
- evals/benchmarks/tool-precision/ - Dashboard, database, diff implementations
- evals/legacy/cli/ - Old HTTP-based eval harness
This functionality is superseded by the new testing pyramid:
- Tool precision is now covered by contract tests in src/core/
- E2E testing uses the cline-bench framework
* feat(evals): add analysis framework for benchmark results
Add shared infrastructure for analyzing evaluation results:
- TypeScript schemas for Harbor and analysis output formats
- Parsers for Harbor, tool-precision, and exercise results
- Failure classifier with pattern matching (cline-failures.yaml)
- Metrics calculator (pass@k, consistency, latency)
- JSON and Markdown reporters
- CLI with analyze and compare commands
- Unit tests for classifier and metrics
This framework is used by both smoke tests and E2E evaluations
to provide consistent metrics and failure categorization.
* feat(evals): add contract tests for API transforms
Add tests to verify API response transformations preserve data correctly:
- thinking-traces.test.ts: Tests thinking block extraction and formatting
- tool-parsing.test.ts: Tests tool call parsing across providers
These contract tests catch regressions when modifying transform logic,
ensuring API responses are correctly processed regardless of provider.
Run with: npm run test:unit
* feat(evals): add provider smoke tests with pass@k metrics
Add lightweight smoke tests that validate provider integrations work
correctly with real LLM calls:
Scenarios (5 curated tests):
- 01-create-file: Tests write_to_file tool
- 02-edit-file: Tests replace_in_file tool
- 03-read-summarize: Tests read_file tool
- 04-multi-file: Tests multi-file edits
- 05-typescript-function: Tests code generation
Features:
- CLI-based runner using the cline CLI
- Multiple trials per scenario for reliability testing
- pass@k metrics (solution finding) and pass^k (consistency)
- Results storage with logs and latest symlink
- Adaptive metric display based on trial count
Run locally: npm run eval:smoke
* feat(evals): add E2E runner with cline-bench
Add end-to-end testing infrastructure using real-world production bugs:
- cline-bench submodule: 12 curated tasks from actual Cline sessions
- Complex multi-file refactors
- Bug fixes requiring deep context understanding
- Cross-language/framework tasks
- E2E runner (evals/e2e/run-cline-bench.ts):
- Integrates with Harbor for containerized execution
- Supports single task or full suite runs
- Pass/fail metrics with detailed logging
Run: npm run eval:e2e -- --task discord-trivia
Note: E2E tests require Docker and are intended for weekly/release
testing, not per-commit CI (each task takes 20-30 minutes).
* feat(evals): add CI workflow and documentation
CI Workflow (.github/workflows/cline-evals-regression.yml):
- Triggers on push/PR to main (src/core, src/shared, proto, evals paths)
- Builds CLI from source with Go 1.24
- Runs 5 smoke test scenarios in parallel
- Uses Anthropic API with claude-sonnet-4
- Uploads results as artifacts with summary
npm scripts:
- eval:smoke - Run smoke tests locally (builds CLI first)
- eval:smoke:run - Run smoke tests (assumes CLI is built)
- eval:e2e - Run cline-bench E2E tests
Documentation:
- ARCHITECTURE.md: Testing pyramid overview with ASCII diagrams
- EVALS_OVERVIEW.md: High-level introduction for mixed audience
- Updated README.md with current structure and usage
* chore(evals): restore tool-precision as deprecated legacy
Restore the diff edit evaluation framework for @ara's use case.
Marked as DEPRECATED - target removal Q2 2026 when cline-bench
is fully operational for model comparison.
Note: Skipping linter as this is legacy code being preserved as-is.
* feat(evals): add per-scenario model support and apply_patch test
Also honor --model overrides and prune stubs.
* chore(evals): update smoke tests for CLI 2.0
- Remove Go setup from workflow (CLI 2.0 is TypeScript)
- Build CLI via `npm run build` in cli/ directory
- Install CLI via `npm link` to test built code from PR
- Update CLI flags: -y -m model --json (remove -o and -s)
- Provider configured via `cline auth` before tests run
* chore(evals): add auth check and CLI 2.0 flags
- Add configureAuth() that runs cline auth non-interactively
- Require CLINE_API_KEY env var or use existing ~/.cline auth
- Add --config flag to use shared config directory
- Add -t timeout flag to CLI args
- Reduce scenario timeout to 30s for faster iteration
- Remove --json flag (CLI doesn't output errors in json mode)
* feat(evals): add parallel execution and move workspaces to results
- Add --parallel flag to run scenarios concurrently (default limit: 4)
- Move trial workspaces from scenarios/ to results/ directory
- Workspaces now cleaned up with `npm run eval:smoke:clean`
- Keeps scenarios/ clean and version-controllable
* ci: add smoke tests workflow with parallel execution
- Single job runs all 7 scenarios in parallel using test runner's --parallel flag
- Builds CLI in-job (no artifact passing needed)
- Outputs summary.md to GitHub step summary
- Syncs package-lock.json for tiktoken/commander deps
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(evals): increase 01-create-file timeout to 120s
The 30s timeout was too short for reliable execution.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore: restore changesets deleted during rebase
These changesets belong to the already-merged CLI fix (#9073)
and should not be deleted by this branch.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore(evals): remove unused dependencies from package.json
Drop execa, node-fetch, ora, sqlite, uuid, yargs and their types.
These were leftovers from the old CLI-based eval runner. The smoke
tests use Node builtins and the tool-precision benchmark only needs
axios, better-sqlite3, chalk, commander, dotenv, tiktoken.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Add TypeScript build info files to .gitignore
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
* feat: add .agents/skills directory support for skill discovery
Add compatibility for the standardized .agents/skills directory pattern,
both globally (~/.agents/skills) and locally (.agents/skills in workspace).
* feat: make .agents/skills the default for new skills
New skills are now created in .agents/skills (local) and ~/.agents/skills
(global) by default. These directories also have highest priority in
skill discovery, overriding skills with the same name from other locations.
* docs: update skills documentation for .agents/skills directories
* refactor skills directory helpers
* changeset version bump
* Updating CHANGELOG.md format
* changeset version bump
* Updating CHANGELOG.md format
* Eve manually updating the banner and the release version
* Manually update the changelog
* Fix GLM 5 model ID in banner
---------
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions <github-actions@github.com>
Co-authored-by: cline-test <132302818+candieduniverse@users.noreply.github.com>
* fix(task): prevent duplicate partial text rows after completion
Avoid adding a new partial text message when the latest text row is already completed with the same content. This stops a presenter race from rendering duplicate streamed text lines for MiniMax-style timing.
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(task): cover duplicate partial text dedupe behavior
Add a Task.say unit test that reproduces the duplicate-partial-after-complete scenario and verifies we skip creating a second text row with identical content.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(claude-code): add opus 4.6 1m model option
* fix(claude-code): support opus[1m] alias and align opus alias
* fix(claude-code): add sonnet[1m] model support
* fix: use vscode.env.asExternalUri for web OAuth callbacks
In VS Code Web (Codespaces, code serve-web), OAuth callbacks using
http://127.0.0.1:PORT break because the extension host runs remotely.
Changes:
- getCallbackUrl now accepts a path parameter
- Desktop: uses vscode://extension-id/path directly
- Web (UIKind.Web): uses vscode.env.asExternalUri() for web-reachable URL
- Updated all callers (/auth, /openrouter, /hicap, /requesty, MCP) to
pass path and use URL+searchParams for proper encoding
- Added regression test asserting web callback URL is not 127.0.0.1
- AuthHandler (localhost HTTP) now only used by CLI/standalone mode
* fix: use URL.searchParams for proper callback URL encoding
Callers were using template literal interpolation to embed callback URLs
into query strings, which breaks when the URL contains special characters
(e.g. from asExternalUri with query params). Use URL+searchParams.set()
which automatically encodes values.
* chore: revert unrelated whitespace change in account.proto
* revert: remove non-essential URL encoding changes in auth callers
Keep only the core fix (getCallbackUrl path parameter + asExternalUri for web).
Revert the URL+searchParams encoding improvement to minimize diff.
* fix: URL-encode callback_url in auth callers, add encoding test
In VS Code Web, callback URLs from asExternalUri can contain their own
query params (?tkn=...&extra=...). String-interpolating them into
callback_url= causes everything after the first & to be parsed as
top-level params, truncating the callback URL.
Use URL + searchParams.set() in openrouter, hicap, and requesty callers.
Replace tautology test with deterministic round-trip encoding assertions.
* fix: use vscode.env.asExternalUri for auth callback URLs in VS Code Web
The OAuth callback redirect was broken in VS Code Web (code serve-web)
environments because the callback URL used a raw vscode:// URI scheme,
which the OS would route to the local desktop VS Code app instead of
the web instance.
This change wraps both getCallbackUrl() and getIdeRedirectUri() with
vscode.env.asExternalUri() which properly transforms URIs based on the
environment:
- Desktop VS Code: unchanged (vscode://...)
- VS Code Remote SSH: adds remote authority for proper routing
- VS Code Web: transforms to HTTPS URL that routes through the web server
Fixes#5109 (remaining callback redirect issue)
Related: #2152
* fix: use HTTP-based auth callback for VS Code Web mode
In VS Code Web (code serve-web), vscode:// URIs redirect to the desktop
app instead of staying in the browser. This change uses AuthHandler
(local HTTP server) for the auth callback in web mode, matching how
CLI/standalone already handles auth.
- getCallbackUrl: use AuthHandler when UIKind.Web
- getIdeRedirectUri: return empty in web mode to avoid vscode:// redirect
* fix: add fallback for openExternal RPC for JetBrains compatibility
The openExternal host bridge RPC is not implemented in the JetBrains
plugin, causing sign-in to fail silently. This adds a fallback to the
'open' npm package when the host RPC fails with UNIMPLEMENTED.
Fixes#9164, #9137, #9138
- Add stdinIsTTY check to shouldUsePlainTextMode() - Ink requires raw mode on stdin
- Only error on empty stdin when no prompt is provided (allows: cline 'prompt' < /dev/null)
- Fixes crash in GitHub Actions and other CI environments
- Cline CLI 2.0 now available. Install with `npm install -g cline`
- Anthopic Opus 4.6
- Minimax-2.1 and Kimi-k2.5 now available for free for a limited time promo
- Codex-5.3 through OpenAI Codex provider
- Fix read file tool to support reading large files
- Fix decimal input crash in OpenAI Compatible price fields (#8129)
- Fix build complete handlers when updating the api config
- Fixed missing provider from list
- Fixed Favorite Icon / Star from getting clipped in the task history view
- Make skills always enabled and remove feature toggle setting
Co-authored-by: Arafatkatze <arafat.da.khan@gmail.com>
* feat: add Claude Opus 4.6 model support with 1M context window
Adds support for Claude Opus 4.6, Anthropic's latest model with:
- 200K base context window with optional 1M context variant
- Tiered pricing for >200K context (2x input/output pricing)
- Extended thinking/reasoning support
- Prompt caching support
Changes:
- Added model definitions for Anthropic, Bedrock, and Vertex providers
- Added OpenRouter 1M variant support
- Updated thinking models lists across all provider UIs
- Added context window switcher for Opus 4.6
- Updated JP cross-region inference models list
* feat: update featured model to Opus 4.6 in model picker
* chore: add changeset for Claude Opus 4.6
* fix: correct Opus 4.6 model IDs (no date suffix)
---------
Co-authored-by: Robin Newhouse <robin@cline.bot>
* fix: use vscode.env.openExternal for auth in remote environments
Fixes#5109
The OAuth authentication flow was broken in VS Code Server and remote
environments because the code used the npm 'open' package directly, which
tries to launch a browser on the server itself (which has no display).
This change routes browser URL opening through VS Code's native
vscode.env.openExternal() API via the HostBridge pattern, which properly
forwards URLs to the user's local machine in remote environments.
Changes:
- Added openExternal RPC to proto/host/env.proto
- Created VS Code handler using vscode.env.openExternal()
- Updated src/utils/env.ts to use HostProvider.env.openExternal()
- Added openExternal to CLI CliEnvServiceClient (uses npm 'open')
- Added openExternal to CLI ACPEnvServiceClient (uses npm 'open')
Related issues: #5394, #2152, #7971
* chore: add changeset for vscode server auth fix
* refactor: extract shared openUrlInBrowser utility for CLI
* add auth option to get API-KEY for hicap from hicap dashboard website
* remove default hicap model selection
* change url hicap get api keys, add useEffect when update hicapApiKey
* add changeset
* fix(cli): prevent hang when spawned without TTY
When the CLI is spawned as a child process without a TTY (e.g., from
spawn() in smoke tests or CI), process.stdin.isTTY is false even when
nothing is piped to stdin. This caused readStdinIfPiped() to wait up
to 5 minutes for input that would never arrive.
Fix by using fs.fstatSync(0) to check if stdin is actually a FIFO
(pipe) or file before waiting. This correctly handles:
- Spawned processes without TTY → returns immediately
- Actual piped input (echo "x" | cline) → waits and reads
- stdin from /dev/null → returns immediately
* chore: add changeset
* test(cli): add tests for stdin type detection
* feat: display markdown table in UI
Simplify the handlePartialBlock method in AttemptCompletionHandler by:
- Removing conditional logic for command vs no-command cases
- Always displaying partial result if present
- Deferring command handling to the final execution step
This fixes an issue where attempt completion response doesn't get streamed to the UI during partial result.
Also replaced react-remark with react-markdown and remark-gfm dependencies to MarkdownBlock in UI for enhanced markdown rendering support with GitHub Flavored Markdown features, including displaying table.
* add changeset
* Update src/core/task/tools/handlers/AttemptCompletionHandler.ts
handlePartialBlock hard-codes the partial flag to true when calling uiHelpers.say(...). For consistency with other tool handlers and to avoid incorrect behavior if this method is ever invoked with a non-partial block, pass block.partial through instead.
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Robin Newhouse <robin@cline.bot>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Fix: decimal input crash in OpenAI Compatible price fields (#8129)
* refactor: use type-safe parsePrice helper for decimal input handling
Replace the `as any` type bypass with a proper parsePrice utility function
that safely handles edge cases (empty string, lone dot, invalid input)
while maintaining type safety. Adds unit tests for the helper.
---------
Co-authored-by: Robin Newhouse <robin@cline.bot>
* feat(skills): Make skills always enabled and remove feature toggle setting
- Remove skillsEnabled from state-keys.ts USER_SETTINGS_FIELDS
- Remove Skills checkbox from FeatureSettingsSection.tsx
- Remove skillsEnabled handling from updateSettings.ts
- Mark skills_enabled as reserved in both Settings and UpdateSettingsRequest proto messages
- Remove conditional in task/index.ts to always discover skills
- Remove skillsEnabled from ExtensionStateContext.tsx default state
- Remove skillsEnabled from ExtensionMessage.ts interface
- Remove skillsEnabled from controller/index.ts state building
- Always show skills tab in ClineRulesToggleModal.tsx
- Remove experimental note from docs/features/skills.mdx
Follows the same pattern as hooks removal (PR #8777).
* fix: Show error message when skill creation fails
Display error to user instead of silently logging when creating a workspace
skill fails (e.g., when no workspace folder is open).