mirror of
https://github.com/cline/cline.git
synced 2026-09-19 02:05:44 +08:00
* chore(evals): reorganize eval structure with purpose-based naming - Move evals/diff-edits/ → evals/benchmarks/tool-precision/replace-in-file/ - Move evals/cli/ → evals/legacy/cli/ (preserve for reference) - Create evals/benchmarks/real-world/ directory - Create evals/benchmarks/coding-exercises/cases/ directory - Create evals/analysis/ directory structure Note: No repositories/exercism/ directory found to move. Skipping pre-commit hook as this is a reorganization of legacy code. * chore(evals): remove legacy evaluation code Remove abandoned evaluation infrastructure: - evals/benchmarks/tool-precision/ - Dashboard, database, diff implementations - evals/legacy/cli/ - Old HTTP-based eval harness This functionality is superseded by the new testing pyramid: - Tool precision is now covered by contract tests in src/core/ - E2E testing uses the cline-bench framework * feat(evals): add analysis framework for benchmark results Add shared infrastructure for analyzing evaluation results: - TypeScript schemas for Harbor and analysis output formats - Parsers for Harbor, tool-precision, and exercise results - Failure classifier with pattern matching (cline-failures.yaml) - Metrics calculator (pass@k, consistency, latency) - JSON and Markdown reporters - CLI with analyze and compare commands - Unit tests for classifier and metrics This framework is used by both smoke tests and E2E evaluations to provide consistent metrics and failure categorization. * feat(evals): add contract tests for API transforms Add tests to verify API response transformations preserve data correctly: - thinking-traces.test.ts: Tests thinking block extraction and formatting - tool-parsing.test.ts: Tests tool call parsing across providers These contract tests catch regressions when modifying transform logic, ensuring API responses are correctly processed regardless of provider. Run with: npm run test:unit * feat(evals): add provider smoke tests with pass@k metrics Add lightweight smoke tests that validate provider integrations work correctly with real LLM calls: Scenarios (5 curated tests): - 01-create-file: Tests write_to_file tool - 02-edit-file: Tests replace_in_file tool - 03-read-summarize: Tests read_file tool - 04-multi-file: Tests multi-file edits - 05-typescript-function: Tests code generation Features: - CLI-based runner using the cline CLI - Multiple trials per scenario for reliability testing - pass@k metrics (solution finding) and pass^k (consistency) - Results storage with logs and latest symlink - Adaptive metric display based on trial count Run locally: npm run eval:smoke * feat(evals): add E2E runner with cline-bench Add end-to-end testing infrastructure using real-world production bugs: - cline-bench submodule: 12 curated tasks from actual Cline sessions - Complex multi-file refactors - Bug fixes requiring deep context understanding - Cross-language/framework tasks - E2E runner (evals/e2e/run-cline-bench.ts): - Integrates with Harbor for containerized execution - Supports single task or full suite runs - Pass/fail metrics with detailed logging Run: npm run eval:e2e -- --task discord-trivia Note: E2E tests require Docker and are intended for weekly/release testing, not per-commit CI (each task takes 20-30 minutes). * feat(evals): add CI workflow and documentation CI Workflow (.github/workflows/cline-evals-regression.yml): - Triggers on push/PR to main (src/core, src/shared, proto, evals paths) - Builds CLI from source with Go 1.24 - Runs 5 smoke test scenarios in parallel - Uses Anthropic API with claude-sonnet-4 - Uploads results as artifacts with summary npm scripts: - eval:smoke - Run smoke tests locally (builds CLI first) - eval:smoke:run - Run smoke tests (assumes CLI is built) - eval:e2e - Run cline-bench E2E tests Documentation: - ARCHITECTURE.md: Testing pyramid overview with ASCII diagrams - EVALS_OVERVIEW.md: High-level introduction for mixed audience - Updated README.md with current structure and usage * chore(evals): restore tool-precision as deprecated legacy Restore the diff edit evaluation framework for @ara's use case. Marked as DEPRECATED - target removal Q2 2026 when cline-bench is fully operational for model comparison. Note: Skipping linter as this is legacy code being preserved as-is. * feat(evals): add per-scenario model support and apply_patch test Also honor --model overrides and prune stubs. * chore(evals): update smoke tests for CLI 2.0 - Remove Go setup from workflow (CLI 2.0 is TypeScript) - Build CLI via `npm run build` in cli/ directory - Install CLI via `npm link` to test built code from PR - Update CLI flags: -y -m model --json (remove -o and -s) - Provider configured via `cline auth` before tests run * chore(evals): add auth check and CLI 2.0 flags - Add configureAuth() that runs cline auth non-interactively - Require CLINE_API_KEY env var or use existing ~/.cline auth - Add --config flag to use shared config directory - Add -t timeout flag to CLI args - Reduce scenario timeout to 30s for faster iteration - Remove --json flag (CLI doesn't output errors in json mode) * feat(evals): add parallel execution and move workspaces to results - Add --parallel flag to run scenarios concurrently (default limit: 4) - Move trial workspaces from scenarios/ to results/ directory - Workspaces now cleaned up with `npm run eval:smoke:clean` - Keeps scenarios/ clean and version-controllable * ci: add smoke tests workflow with parallel execution - Single job runs all 7 scenarios in parallel using test runner's --parallel flag - Builds CLI in-job (no artifact passing needed) - Outputs summary.md to GitHub step summary - Syncs package-lock.json for tiktoken/commander deps Co-authored-by: Cursor <cursoragent@cursor.com> * fix(evals): increase 01-create-file timeout to 120s The 30s timeout was too short for reliable execution. Co-authored-by: Cursor <cursoragent@cursor.com> * chore: restore changesets deleted during rebase These changesets belong to the already-merged CLI fix (#9073) and should not be deleted by this branch. Co-authored-by: Cursor <cursoragent@cursor.com> * chore(evals): remove unused dependencies from package.json Drop execa, node-fetch, ora, sqlite, uuid, yargs and their types. These were leftovers from the old CLI-based eval runner. The smoke tests use Node builtins and the tool-precision benchmark only needs axios, better-sqlite3, chalk, commander, dotenv, tiktoken. Co-authored-by: Cursor <cursoragent@cursor.com> * Add TypeScript build info files to .gitignore --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Saoud Rizwan <7799382+saoudrizwan@users.noreply.github.com>
Cline Evaluation Framework
A layered testing system for measuring Cline's performance at different levels.
Directory Structure
evals/
├── smoke-tests/ # Quick provider validation (minutes)
│ ├── run-smoke-tests.ts
│ └── scenarios/ # 5 curated test scenarios
│
├── e2e/ # Full E2E with cline-bench (hours)
│ └── run-cline-bench.ts
│
├── cline-bench/ # Real-world tasks (git submodule)
│ └── tasks/ # 12 production bug fixes
│
├── analysis/ # Metrics and reporting framework
│ ├── src/
│ │ ├── metrics.ts # pass@k, pass^k calculations
│ │ ├── classifier.ts # Failure pattern matching
│ │ └── reporters/ # Markdown, JSON output
│ └── patterns/
│ └── cline-failures.yaml
│
└── baselines/ # Performance baselines for regression detection
Test Layers
Layer 1: Contract Tests (Unit)
Location: src/core/api/transform/__tests__/
Tests API transform logic without LLM calls:
- Thinking trace preservation
- Tool call parsing (XML, native formats)
- Provider format conversions
npm run test:unit -- --grep "Thinking\|Tool Call"
Layer 2: Smoke Tests (Minutes)
Location: evals/smoke-tests/
Quick validation across providers with real LLM calls:
- 5 curated scenarios
- 3 trials per test for pass@k metrics
- Runs via cline CLI with
-sflags
# Set API key (Cline provider)
export CLINE_API_KEY=sk-...
# Run smoke tests
npm run eval:smoke
# Run specific scenario
npm run eval:smoke -- --scenario 01-create-file
# Run with specific model (overrides per-scenario models)
npm run eval:smoke -- --model anthropic/claude-sonnet-4.5
Layer 3: E2E Tests (Hours)
Location: evals/e2e/ + evals/cline-bench/
Full agent tests on production-grade tasks:
- 12 real-world coding problems
- Docker/Daytona execution via Harbor
- Nightly CI runs
# Prerequisites: Python 3.13, Harbor, Docker
npm run eval:e2e
# Specific task
npm run eval:e2e -- --tasks discord
# Different provider
npm run eval:e2e -- --provider openai --model gpt-4o
Metrics
The framework calculates:
| Metric | Formula | Interpretation |
|---|---|---|
| pass@k | P(≥1 of k passes) | Solution finding capability |
| pass^k | P(all k pass) | Reliability |
| Flakiness | Entropy of pass rate | Consistency |
With 3 trials:
- All pass →
pass(reliable) - All fail →
fail(broken) - Mixed →
flaky(needs investigation)
CI Integration
- PR Gate: Contract tests + smoke tests (fast, ~3min)
- Nightly: E2E tests with cline-bench (not yet implemented, see TODO)
Quick Start
# Run all fast tests
npm run test:unit
npm run eval:smoke
# Run E2E (requires setup)
cd evals/cline-bench
# Follow README.md for Harbor setup
npm run eval:e2e
Adding Tests
Smoke Test Scenario
- Create
evals/smoke-tests/scenarios/<name>/config.json - Add optional
template/directory with starting files - Run to verify:
npm run eval:smoke -- --scenario <name>
Contract Test
- Add to
src/core/api/transform/__tests__/ - Run:
npm run test:unit -- --grep "YourTest"
E2E Task
Contribute to cline/cline-bench
Resources
TODO
- Nightly E2E CI: Add scheduled workflow for cline-bench tests
- Requires: Docker runner, Harbor setup, ~1-2 hour timeout
- Should run on schedule (e.g., nightly) not per-PR
- Separate secrets for E2E environment
- Native tool calling smoke tests: Add CLI support for
native_tool_call_enabledsetting to test Claude 4 with native tools