mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
* perf(execution): parallelize preflight gates, cache deployed state, memoize Anthropic client
- Memoize Anthropic + Azure-Anthropic SDK clients (new client-cache.ts) keyed
by apiKey (+beta header; +baseURL/version/pinnedIP for Azure) so HTTP
keep-alive connections are reused instead of a fresh TLS handshake per call.
apiKey is the tenant boundary.
- Parallelize the read-only preflight gates in preprocessing.ts (ban +
subscription, then usage + org-member + rate-limit) while preserving exact
error precedence (ban 403 -> usage 402 -> rate 429) and keeping the sole
write (admission reservation) last.
- Parallelize the independent workflow-state and env-var loads in execution-core.
- Cache deployed workflow state by immutable deploymentVersionId with
deep-clone-on-read, oldest-first eviction, and a 5-min TTL bounding the
credential-mapping edge across ECS tasks.
- Parallelize the independent personal-subscription + membership queries in
getHighestPrioritySubscription.
- BYOK: drop the redundant getWorkspaceById existence check (auth already
validates the workspace); read the key list fresh every call for zero
cross-instance staleness.
Billing/usage/ban/permission reads stay fresh on the primary (no cache, no
replica). Adds tests for every new mechanism and fixes a pre-existing vitest
class-mock incompatibility that had execution-core.test.ts fully red on staging.
* fix(execution): run rate-limit gate only after ban/usage pass
The rate-limit gate is not read-only — checkRateLimitWithSubscription consumes
a token — so running it in parallel with the read-only gates debited rate-limit
quota for requests that the ban (403) or usage (402) gates reject, which the
original sequential flow never did.
Move the rate-limit gate to run sequentially after the ban and usage gates pass,
preserving the read-only gates' parallelism (ban + subscription + usage) and the
exact ban -> usage -> rate precedence. Add regression tests asserting the rate
limiter is not consumed when an earlier gate rejects, and is consumed once when
they pass.
Caught by Cursor Bugbot review.
* chore(execution): trim redundant preflight comments
Tighten the gate overview to match the sequential rate-limit gate and drop
inline notes that duplicated it or the runRateLimitGate doc.
* refactor(cache): address review — idle TTL for client cache, LRUCache for deployed state
- client-cache: add updateAgeOnGet so the TTL is genuinely idle-based (active
clients keep their warm keep-alive connections; the JSDoc now matches behavior).
- deployed-state: replace the hand-rolled Map + manual FIFO eviction/TTL with
LRUCache (real LRU eviction, built-in TTL), matching the effectiveDecryptedEnv
and integration-tool-schema caches. TTL stays absolute (not reset on read) so
the credential-migration remap still propagates across ECS tasks.
Both per review feedback from Greptile.
* test(execution): isolate rate-limit gate test from STEP 7 reservation
The 'consumes the rate-limit gate once' test reached the STEP 7 admission
reservation, which depends on Redis — it passed locally (reserve throws and is
swallowed) but failed in CI (reserve returns not-reserved -> 429). Pass
skipConcurrencyReservation so the test isolates the rate gate deterministically.
* perf(providers): memoize SDK clients where the pool is per-client (bedrock, vllm)
Generalize the Anthropic client cache into one shared memoizer
(providers/client-cache.ts) and apply it only where each new client owns its own
connection pool — so reuse actually keeps connections warm:
- bedrock: AWS SDK clients hold a per-client connection pool (reuse is the AWS
best practice). Keyed by region + credential identity.
- vllm: a pinned endpoint creates its own undici Agent per call; key by the
resolved IP so DNS re-validation still runs each request.
- anthropic + azure-anthropic: migrated onto the shared memoizer.
Deliberately NOT applied to the OpenAI-compatible providers, groq, cerebras, or
google: their SDKs share a process-global keep-alive pool (Node openai-sdk module
singleton agent; anthropic/global undici), so a fresh client per request already
reuses connections and memoization would add complexity with ~no benefit. litellm
uses a plain shared-agent client (no pinning) and is likewise skipped.
Bounded LRU (max 1000, 30m idle TTL) with no close-on-eviction, avoiding the
unbounded-growth and eviction-closes-in-use-client failure modes seen in similar
client caches.
* chore(perf): trim verbose comments to terse why-notes
* chore(perf): drop obvious inline comments, keep nuance as TSDoc
* fix(bedrock): key client cache on full credential, not just access key id
A corrected secret under the same access key id would otherwise keep serving the
stale cached client until TTL/eviction. Caught by Cursor Bugbot.
* test(execution,providers): fix preflight mock reset + isolate provider client cache in tests
- preprocessing.test: re-establish the checkOrgMemberUsageLimit mock in beforeEach
(the only gate mock not re-set). In the full suite its implementation was reset
so the success-path test got undefined -> threw -> 500 -> success:false. Mirrors
how checkServerSideUsageLimits is handled.
- client-cache: add clearProviderClientCacheForTests; call it in the bedrock and
vllm test beforeEach so construction assertions always start from a cache miss
now that those providers memoize their client.
* test(execution): make RateLimiter mock constructable under vitest 4.x
The RateLimiter mock used an arrow factory (vi.fn(() => ({...}))). vitest 4.x
(CI) rejects `new` on an arrow-implemented mock ("not a constructor"); 3.2.4
allowed it. The new rate-gate test is the first to actually `new RateLimiter()`,
so it surfaced the failure only in CI. Switch the mock to a regular function and
drop the speculative beforeEach re-establishments that didn't address it.