mirror of
https://github.com/cline/cline.git
synced 2026-09-24 23:20:16 +08:00
* style(docs): update background color scheme to neutral tones Update documentation background colors from purple-tinted theme to neutral gray tones. Changed light mode from lavender (#F0E6FF) to off-white (#fafaf9) and dark mode from pure black (#000000) to dark gray (#0f0f0f) for improved visual consistency. * refactor(docs): remove gradient decoration from theme config Remove the "decoration": "gradient" property from the documentation theme configuration. This simplifies the theme settings by removing the gradient decoration option from the color configuration object. * docs: change documentation font family to Geist Mono Replace Roboto with Geist Mono as the default font family in the documentation configuration. This updates the visual styling of the documentation to use a monospace font, which may improve readability for code-heavy content. * docs: update branding and restructure navigation - Replace robot panel logos with new Cline brand logos - Add icons to navbar links (Docs, GitHub, Discord) - Restructure navigation from groups to tabs format - Add icons to navigation items for improved UX - Include new Docs link in navbar with book icon This update modernizes the documentation appearance and improves navigation hierarchy for better user experience. * docs: restructure navigation with hierarchical groups and pages Restructured documentation navigation from flat menu to organized groups: - Removed redundant "Docs" link from navbar - Migrated from "menu" to "groups/pages" structure - Added comprehensive page organization with nested groups: * Introduction, Getting Started, Features * Prompting Skills, Cline's Tools, Enterprise Solutions * MCP Servers, Provider Configuration - Organized features into logical subgroups (@ Mentions, Commands, Customization, Slash Commands) - Improved documentation discoverability and hierarchy This change provides better content organization and easier navigation for users exploring different aspects of Cline documentation. * docs: remove contextual options from documentation config Remove the contextual configuration section containing the "copy" option from docs.json. This simplifies the documentation configuration by removing unused contextual menu options. * docs(multiroot): improve workspace documentation with limitations and technical details - Add important note about experimental limitations affecting Cline rules and checkpoints - Add "How it works" section explaining automatic workspace detection and tracking - Reorganize technical behavior section with detailed subsections for workspace detection, path resolution, and command execution - Document workspace hint syntax for explicit file references (@workspaceName:path) - Standardize heading capitalization to sentence case for consistency - Improve overall content organization and clarity for better user understanding This update provides users with clearer information about the multiroot feature's current state, its limitations, and how to effectively use workspace hints when working with multiple project folders. * docs: restructure overview page with enhanced visual layout - Convert plain markdown sections to CardGroup and Card components with icons - Add tabbed interface for Plan & Act Mode explanation - Update description from "development assistant" to "coding agent" - Reorganize content for improved readability and visual hierarchy - Enhance feature presentations with icon-based cards Improves user experience by transforming the overview documentation into a more visually appealing and scannable format using modern documentation components. * docs: improve installation guide with enhanced structure and UX Restructure the Cline installation documentation to improve readability and user experience: - Add prominent note highlighting 2-minute installation time - Convert prerequisites into visual card components for better clarity - Transform installation steps into structured Step components for easier following - Add manual installation instructions for JetBrains IDEs - Include feature compatibility accordion for JetBrains users - Enhance visual hierarchy with improved component usage (CardGroup, Steps, Accordion) - Simplify language and improve descriptions throughout This makes the installation process clearer for new users and reduces friction during onboarding. * style(docs): remove text opacity reduction for better readability * docs: refactor model selection guide with visual step-by-step instructions - Replace tab-based layout with linear step-by-step flow - Add screenshots for each configuration step (config, provider, API, model) - Reorganize content structure for improved clarity and user experience - Add quickstart options and streamlined provider recommendations - Improve navigation with visual aids to help users configure Cline faster * docs: add installation screenshots and context management guide * docs: flatten provider config structure in documentation Remove the "Alternative Providers" grouping and move all provider configuration pages (OpenRouter, Cerebras, DeepSeek, Groq, xAI Grok, Mistral AI, Doubao, Fireworks, and ZAI) to the main provider configuration list. This simplifies the documentation navigation by treating all providers equally rather than categorizing some as alternatives. * docs: restructure context management docs and improve content clarity **Changes:** - Reorganized documentation structure by moving context management from `/best-practices` to `/prompting` section for better categorization - Added URL redirect to maintain backward compatibility for old links - Updated navigation references in welcome page to point to new location - Improved readability of context management explanations with more narrative, conversational prose - Enhanced context window documentation by adding cache tokens indicator and using emoji-based formatting for better visual clarity - Streamlined Cline Memory Bank setup instructions from 4 to 3 steps - Updated context bar screenshot to use newer image asset **Why:** Better documentation organization and improved user experience through clearer explanations of how Cline builds and manages context during tasks. * docs(context-management): convert Quick Reference to Info component Replace blockquote formatting with Info component for the Quick Reference section in the context management documentation. This improves visual presentation and maintains consistency with documentation standards. Also removes trailing whitespace at the end of the file for cleaner formatting. * docs: add Cline Enterprise overview and restructure enterprise section - Add comprehensive enterprise overview documentation covering security, governance, observability, and developer experience features - Rename "Enterprise & Security" navigation group to "Enterprise" - Consolidate enterprise documentation by replacing 4 pages with 2: new overview page and security concerns - Document BYOI (Bring Your Own Inference), SSO authentication, and role-based access control capabilities This restructuring provides a clearer entry point for enterprise users and consolidates previously scattered enterprise information into a cohesive overview document. * docs(enterprise): streamline enterprise overview and update font - Change documentation font from Geist Mono to Geist Sans - Add enterprise website link card for detailed information - Remove Developer Experience, Proven at Scale, and Pricing sections - Consolidate Flexible Inference section content - Simplify enterprise overview to focus on core capabilities These changes reduce redundancy by directing users to the enterprise website for pricing and detailed features while keeping the docs focused on technical implementation and core capabilities. * clean-images * Update docs/getting-started/installing-cline.mdx Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Update docs/styles.css Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * docs(cline-cli): add platform availability warning to overview Add a prominent warning callout indicating that Cline CLI is currently in preview and only supports macOS and Linux, with Windows support coming soon. This sets clear expectations for users about platform compatibility. Also remove redundant introductory text in the "What you can build with this" section to improve content clarity. --------- Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
227 lines
7.2 KiB
Plaintext
227 lines
7.2 KiB
Plaintext
---
|
|
title: "Local Models Overview"
|
|
---
|
|
|
|
## Running Models Locally with Cline
|
|
|
|
Run Cline completely offline with genuinely capable models on your own hardware. No API costs, no data leaving your machine, no internet dependency.
|
|
|
|
Local models have reached a turning point where they're now practical for real development work. This guide covers everything you need to know about running Cline with local models.
|
|
|
|
## Quick Start
|
|
|
|
1. **Check your hardware** - 32GB+ RAM minimum
|
|
2. **Choose your runtime** - [LM Studio](/running-models-locally/lm-studio) or [Ollama](/running-models-locally/ollama)
|
|
3. **Download Qwen3 Coder 30B** - The recommended model
|
|
4. **Configure settings** - Enable compact prompts, set max context
|
|
5. **Start coding** - Completely offline
|
|
|
|
## Hardware Requirements
|
|
|
|
Your RAM determines which models you can run effectively:
|
|
|
|
| RAM | Recommended Model | Quantization | Performance Level |
|
|
| --- | --- | --- | --- |
|
|
| 32GB | Qwen3 Coder 30B | 4-bit | Entry-level local coding |
|
|
| 64GB | Qwen3 Coder 30B | 8-bit | Full Cline features |
|
|
| 128GB+ | GLM-4.5-Air | 4-bit | Cloud-competitive performance |
|
|
|
|
## Recommended Models
|
|
|
|
### Primary Recommendation: Qwen3 Coder 30B
|
|
|
|
After extensive testing, **Qwen3 Coder 30B** is the most reliable model under 70B parameters for Cline:
|
|
|
|
- **256K native context window** - Handle entire repositories
|
|
- **Strong tool-use capabilities** - Reliable command execution
|
|
- **Repository-scale understanding** - Maintains context across files
|
|
- **Proven reliability** - Consistent outputs with Cline's tool format
|
|
|
|
Download sizes:
|
|
- 4-bit: ~17GB (recommended for 32GB RAM)
|
|
- 8-bit: ~32GB (recommended for 64GB RAM)
|
|
- 16-bit: ~60GB (requires 128GB+ RAM)
|
|
|
|
### Why Not Smaller Models?
|
|
|
|
Most models under 30B parameters (7B-20B) fail with Cline because they:
|
|
- Produce broken tool-use outputs
|
|
- Refuse to execute commands
|
|
- Can't maintain conversation context
|
|
- Struggle with complex coding tasks
|
|
|
|
## Runtime Options
|
|
|
|
### LM Studio
|
|
- **Pros**: User-friendly GUI, easy model management, built-in server
|
|
- **Cons**: Memory overhead from UI, limited to single model at a time
|
|
- **Best for**: Desktop users who want simplicity
|
|
- [Setup Guide →](/running-models-locally/lm-studio)
|
|
|
|
### Ollama
|
|
- **Pros**: Command-line based, lower memory overhead, scriptable
|
|
- **Cons**: Requires terminal comfort, manual model management
|
|
- **Best for**: Power users and server deployments
|
|
- [Setup Guide →](/running-models-locally/ollama)
|
|
|
|
## Critical Configuration
|
|
|
|
### Required Settings
|
|
|
|
**In Cline:**
|
|
- ✅ Enable "Use Compact Prompt" - Reduces prompt size by 90%
|
|
- ✅ Set appropriate model in settings
|
|
- ✅ Configure Base URL to match your server
|
|
|
|
**In LM Studio:**
|
|
- Context Length: `262144` (maximum)
|
|
- KV Cache Quantization: `OFF` (critical for proper function)
|
|
- Flash Attention: `ON` (if available on your hardware)
|
|
|
|
**In Ollama:**
|
|
- Set context window: `num_ctx 262144`
|
|
- Enable flash attention if supported
|
|
|
|
### Understanding Quantization
|
|
|
|
Quantization reduces model precision to fit on consumer hardware:
|
|
|
|
| Type | Size Reduction | Quality | Use Case |
|
|
| --- | --- | --- | --- |
|
|
| 4-bit | ~75% | Good | Most coding tasks, limited RAM |
|
|
| 8-bit | ~50% | Better | Professional work, more nuance |
|
|
| 16-bit | None | Best | Maximum quality, requires high RAM |
|
|
|
|
### Model Formats
|
|
|
|
**GGUF (Universal)**
|
|
- Works on all platforms (Windows, Linux, Mac)
|
|
- Extensive quantization options
|
|
- Broader tool compatibility
|
|
- Recommended for most users
|
|
|
|
**MLX (Mac only)**
|
|
- Optimized for Apple Silicon (M1/M2/M3)
|
|
- Leverages Metal and AMX acceleration
|
|
- Faster inference on Mac
|
|
- Requires macOS 13+
|
|
|
|
## Performance Expectations
|
|
|
|
### What's Normal
|
|
|
|
- **Initial load time**: 10-30 seconds for model warmup
|
|
- **Token generation**: 5-20 tokens/second on consumer hardware
|
|
- **Context processing**: Slower with large codebases
|
|
- **Memory usage**: Close to your quantization size
|
|
|
|
### Performance Tips
|
|
|
|
1. **Use compact prompts** - Essential for local inference
|
|
2. **Limit context when possible** - Start with smaller windows
|
|
3. **Choose right quantization** - Balance quality vs speed
|
|
4. **Close other applications** - Free up RAM for the model
|
|
5. **Use SSD storage** - Faster model loading
|
|
|
|
## Use Case Comparison
|
|
|
|
### When to Use Local Models
|
|
|
|
✅ **Perfect for:**
|
|
- Offline development environments
|
|
- Privacy-sensitive projects
|
|
- Learning without API costs
|
|
- Unlimited experimentation
|
|
- Air-gapped environments
|
|
- Cost-conscious development
|
|
|
|
### When to Use Cloud Models
|
|
|
|
☁️ **Better for:**
|
|
- Very large codebases (>256K tokens)
|
|
- Multi-hour refactoring sessions
|
|
- Teams needing consistent performance
|
|
- Latest model capabilities
|
|
- Time-critical projects
|
|
|
|
## Troubleshooting
|
|
|
|
### Common Issues & Solutions
|
|
|
|
**"Shell integration unavailable"**
|
|
- Switch to bash in Cline Settings → Terminal → Default Terminal Profile
|
|
- Resolves 90% of terminal integration problems
|
|
|
|
**"No connection could be made"**
|
|
- Verify server is running (LM Studio or Ollama)
|
|
- Check Base URL matches server address
|
|
- Ensure no firewall blocking connection
|
|
- Default ports: LM Studio (1234), Ollama (11434)
|
|
|
|
**Slow or incomplete responses**
|
|
- Normal for local models (5-20 tokens/sec typical)
|
|
- Try smaller quantization (4-bit instead of 8-bit)
|
|
- Enable compact prompts if not already
|
|
- Reduce context window size
|
|
|
|
**Model confusion or errors**
|
|
- Verify KV Cache Quantization is OFF (LM Studio)
|
|
- Ensure compact prompts enabled
|
|
- Check context length set to maximum
|
|
- Confirm sufficient RAM for quantization
|
|
|
|
### Performance Optimization
|
|
|
|
**For faster inference:**
|
|
1. Use 4-bit quantization
|
|
2. Enable Flash Attention
|
|
3. Reduce context window if not needed
|
|
4. Close unnecessary applications
|
|
5. Use NVMe SSD for model storage
|
|
|
|
**For better quality:**
|
|
1. Use 8-bit or higher quantization
|
|
2. Maximize context window
|
|
3. Ensure adequate cooling
|
|
4. Allocate maximum RAM to model
|
|
|
|
## Advanced Configuration
|
|
|
|
### Multi-GPU Setup
|
|
If you have multiple GPUs, you can split model layers:
|
|
- LM Studio: Automatic GPU detection
|
|
- Ollama: Set `num_gpu` parameter
|
|
|
|
### Custom Models
|
|
While Qwen3 Coder 30B is recommended, you can experiment with:
|
|
- DeepSeek Coder V2
|
|
- Codestral 22B
|
|
- StarCoder2 15B
|
|
|
|
Note: These may require additional configuration and testing.
|
|
|
|
## Community & Support
|
|
|
|
- **Discord**: [Join our community](https://discord.gg/cline) for real-time help
|
|
- **Reddit**: [r/cline](https://www.reddit.com/r/CLine/) for discussions
|
|
- **GitHub**: [Report issues](https://github.com/cline/cline/issues)
|
|
|
|
## Next Steps
|
|
|
|
Ready to get started? Choose your path:
|
|
|
|
<CardGroup cols={2}>
|
|
<Card title="LM Studio Setup" icon="desktop" href="/running-models-locally/lm-studio">
|
|
User-friendly GUI approach with detailed configuration guide
|
|
</Card>
|
|
<Card title="Ollama Setup" icon="terminal" href="/running-models-locally/ollama">
|
|
Command-line setup for power users and automation
|
|
</Card>
|
|
</CardGroup>
|
|
|
|
## Summary
|
|
|
|
Local models with Cline are now genuinely practical. While they won't match top-tier cloud APIs in speed, they offer complete privacy, zero costs, and offline capability. With proper configuration and the right hardware, Qwen3 Coder 30B can handle most coding tasks effectively.
|
|
|
|
The key is proper setup: adequate RAM, correct configuration, and realistic expectations. Follow this guide, and you'll have a capable coding assistant running entirely on your hardware.
|