mirror of
https://github.com/aiclientproxy/proxycast.git
synced 2026-09-24 23:10:56 +08:00
feat: add request dedup, response cache, and capability routing metrics to API server
- Add response cache middleware (backed by aster-rust implementation) - Add request deduplication middleware to prevent duplicate in-flight requests - Add capability routing metrics middleware for model/provider fallback tracking - Add idempotency stats (atomic counters) to IdempotencyStore - Expose all four stores in ServerState and surface stats in ServerStatus - Add ResponseCacheSettings to ServerConfig with defaults (600s TTL, 200 max entries) - Add API server observability infrastructure (IdempotencyGuard, RequestDedupGuard, ResponseCacheGuard) - Add inline capability detection for vision/tools/context when routing requests - Register workspace_ensure_ready and workspace_ensure_default_ready commands in runner - Update docs with response_cache configuration example Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 4.6
parent
2d00911c28
commit
0b2d3caa6c
@@ -72,6 +72,25 @@ routing:
|
||||
|
||||
适用:有脚本联动、自动化流程需求的用户。
|
||||
|
||||
## 示例 4:API 缓存策略(高级)
|
||||
|
||||
目标:精确控制哪些响应状态码参与短时缓存(非流式)。
|
||||
|
||||
```yaml
|
||||
server:
|
||||
host: "127.0.0.1"
|
||||
port: 8999
|
||||
api_key: "your-api-key"
|
||||
response_cache:
|
||||
enabled: true
|
||||
ttl_secs: 600
|
||||
max_entries: 200
|
||||
max_body_bytes: 1048576
|
||||
cacheable_status_codes: [200] # 默认仅缓存 200;可按需扩展 [200, 201]
|
||||
```
|
||||
|
||||
适用:对缓存命中与语义一致性有要求的自动化/API 调用场景。
|
||||
|
||||
## 调整顺序建议
|
||||
|
||||
1. 先确认导航与主题
|
||||
|
||||
Reference in New Issue
Block a user