mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
* fix(security): meter and throttle the deployed-chat TTS relay POST /api/proxy/tts/stream treated "a live public chat exists" as authorization to spend the platform ElevenLabs key. A public chat id is handed to every visitor, so any anonymous caller could synthesize speech with no length cap, no rate limit and no usage accounting. Bring the relay in line with its STT sibling (/api/speech/token): - Resolve the chat's workspace and bill synthesized characters to that payer via a new `voice-output` usage source, so spend is attributable and counts against the plan's usage limit (402 once exceeded). - Throttle per IP before any database work, and per chat afterwards, to bound both one caller hammering many chats and many callers hammering one chat. - Cap `text` at 2000 characters and allowlist `voiceId`/`modelId`, so the caller can no longer choose an unbounded charge, a premium or cloned voice, or the billing model. - Drop `Access-Control-Allow-Origin: *`, which let any third-party page read the audio; deployed chat and the Office embed are same-origin. * fix(security): correct TTS metering, pricing and body cap Follow-up review of the previous commit found five defects in it: - Usage rows collided. `usage_log.event_key` is unique and inserts are conflict-do-nothing, and the key is derived from the entry's stable fields. With no explicit sourceReference, two synthesis calls of equal character count in the same workspace produced the same key, so every repeat length went unbilled — defeating the metering this change is for. Each call now carries a unique sourceReference. - Priced at $0.10 per 1k characters, twice the published ElevenLabs Flash/Turbo rate of $0.05, which would have overcharged customers 2x. - No body cap, so an anonymous caller could make the route buffer up to the shared 50 MB default before validation. Now 16 KB, as the STT sibling does. - Threshold settlement ran per sentence: several queries and a possible Stripe call on a realtime path. The workflow execution that produced the text already settles the payer. - The per-IP bucket was described as preventing database amplification. getClientIp trusts the leftmost X-Forwarded-For, so an attacker rotates past it; the comment now says the per-chat bucket is load-bearing. * refactor(chat): share the deployed-chat auth gate across voice routes Review of the previous commits surfaced duplication and one more gap: - The TTS and STT routes had grown near-identical copies of the chat auth + payer lookup. Extracted to resolveDeployedChatCaller, so the gate and the payer resolve together and cannot drift per route — that duplication is how the unmetered TTS path shipped in the first place. - Neither copy filtered chat.archivedAt, so an archived chat could still authorize spend against its former owner's workspace. The shared lookup now filters it, fixing both routes at once. Note: not covered by a test — the db chain mock does not evaluate WHERE clauses, so an assertion here could not fail. - Replaced the route's hand-rolled 429 builder with the existing enforceIpRateLimit helper, and added enforceChatRateLimit alongside the per-user/IP/workspace helpers. Gains the standard Retry-After and X-RateLimit-Reset headers plus throttle logging. - Dropped a test that asserted a module the route no longer imports was never called: it could not fail. - Narrowed the contract: unexported the single-use allowlists and dropped .passthrough() now that the body is a closed shape. * fix(security): fail closed when voice-output usage cannot be recorded Review round 1 findings: - A ledger write failure previously logged and streamed the audio anyway, leaving the spend unrecorded and the payer's usage understated. The caller is anonymous, so serving audio we could not charge for is the unmetered spend this route exists to prevent — it now returns 500. - Use generateId() from @sim/utils/id rather than crypto.randomUUID, per the AGENTS.md ID rule. generateId returns a full UUID v4, so the per-call uniqueness the usage_log event_key depends on is unchanged. * fix(chat): split long TTS text so the relay cap cannot drop audio The client sentence-splits on Western `.!?` only, so text that never matches — CJK punctuation, or a list with no terminal punctuation — accumulates and is flushed as one block at the end of the stream. Against the new 2000-character relay cap that block is rejected and the whole message plays no audio, a regression introduced by adding the cap. Split to cap-sized pieces at the single point that enqueues synthesis, so both the per-sentence path and the end-of-stream flush are covered. Prefers a whitespace or CJK punctuation boundary, falling back to a hard cut when a block has none. The server cap stays as the enforcement point. * fix(security): release the vendor stream when metering rejects the request The fail-closed branch returned 500 with the ElevenLabs response body still open, so synthesis and download kept consuming vendor and runtime resources for a caller that was already rejected. Cancel it before returning, and assert the cancellation in the test.