Waleed e1f22bdac9 feat(providers): support large agent-block attachments via Files APIs and remote URLs (#5092)
* feat(providers): support large agent-block attachments via Files APIs and remote URLs

Agent-block file uploads were inlined as base64 with a hard 10MB cap. Files
above the threshold now use each provider's native large-file path:

- OpenAI / Gemini: upload to the provider Files API, reference by file_id/uri
- Anthropic: GA url content-block source (no Files API beta, no upload)
- OpenRouter/Groq/Together/Baseten/xAI/vLLM: remote signed URL in image_url/file
- Limits live per-provider in models.ts; the agent block + /models page reflect them

Files <=10MB keep the identical base64 path (zero regression). Server-only file
handles are stripped from untrusted input to prevent SSRF.

* fix(providers): clear forged file handles for inline providers too

attachLargeFileRemoteUrls early-returned for inline-strategy providers before
clearing server-only handle fields, so a forged remoteUrl on an inline-provider
file could still reach a builder (e.g. buildOpenAICompatibleChatContent for
mistral/ollama). Clear the handles for every provider before the strategy check.

* fix(providers): correct OpenAI expiry serialization and Anthropic large-text-doc handling

- OpenAI upload now uses the SDK (client.files.create) so expires_after is
  serialized as a real nested object; the prior expires_after[anchor] bracket
  FormData keys were ignored by OpenAI's server, leaving files un-expiring.
- Anthropic url document source only supports PDFs/images; large non-PDF text
  docs now throw a clear error instead of emitting an unsupported url source.
- Warn when an oversized file can't be sent because cloud storage is unavailable.

* fix(providers): harden large-file path (SSRF fetch, ceiling gate, per-file UI limit)

- Download files for OpenAI/Gemini uploads via validateUrlWithDNS + IP-pinned
  fetch so a forged URL can't reach internal addresses (covers all callers).
- Reject files above the provider ceiling before downloading/uploading.
- UI now validates each file against the provider's per-file ceiling instead of
  summing all files against it, matching server-side per-file validation.
- Lower Anthropic ceiling to 50MB (documented 32MB request cap / page limits).

* refactor(providers): read files-api upload bytes via storage SDK

Read OpenAI/Gemini upload bytes through downloadFileFromStorage instead of
HTTP-fetching the presigned URL. Removes any server-side URL fetch (no SSRF
vector) and works with internal object storage (e.g. self-hosted MinIO), which
an IP-pinned URL fetch would have blocked.

* docs(providers): clarify files-api bytes are read from storage at upload time

* fix(providers): enforce access checks and strip forged ids in the upload path

uploadLargeFilesToProvider runs on raw request messages for every caller (incl.
the internal providers passthrough), so harden it independently of the agent path:
- verifyFileAccess on each file's storage key before reading its bytes, so a forged
  key can't exfiltrate another user's file.
- clear any inbound providerFileId/providerFileUri up front (legit ids are only set
  by the upload itself), so a forged id can't reference a file in a hosted account.

* fix(providers): resolve UI attachment limit with the same model->provider helper as execution

The file-upload control imported getProviderFromModel from @/providers/models, but
the execution path and every other consumer use the one in @/providers/utils (runtime
registry + reseller patterns). Align the UI so its size cap can't disagree with
server-side validation for reseller or dynamically-listed models.

* test(providers): add new models.ts exports to provider mocks

attachments.ts now reads getProviderFileAttachment / INLINE_ATTACHMENT_MAX_BYTES
from @/providers/models; the provider unit tests that fully mock that module need
both exports or attachments.ts fails to load.

* fix(providers): guard Gemini upload response name before polling

ai.files.upload returns name as string | undefined; guard it (instead of an
as-string cast) so a missing name surfaces a clear error at the upload site
rather than an opaque files.get failure on the first poll.

* fix(uploads): type the file-handle key list so omit preserves UserFile fields

The 'as const' readonly tuple widened omit's K to all keys, collapsing
Omit<UserFile, K> to {} and failing the production build's type check. Declare
the array as Array<keyof handle fields> so K is the precise literal union.

* refactor(providers): run handle-clear + URL-mint in executeProviderRequest for all callers

Move attachLargeFileRemoteUrls out of the agent handler and into
executeProviderRequest (right before uploadLargeFilesToProvider), so every entry
point — including the internal providers passthrough — clears forged handles and
mints/access-checks large-file URLs uniformly. The agent handler now only hydrates
base64; its missing-file guard exempts large files (resolved downstream).

* fix(azure-openai): guard optional attachment dataUrl in inline image part

PreparedProviderAttachment.dataUrl is now optional (large files carry a handle
instead); azure-openai builds chat content inline and assigned it directly to a
required url field, failing the production build's type check.

* fix(providers): upload OpenAI files via multipart and fix Buffer Blob part

The installed openai SDK (4.104) does not type expires_after on files.create, so
upload via POST /v1/files directly with the documented expires_after[...] form
fields (gives the file an auto-expiry). Also wrap the storage Buffer in a
Uint8Array for the Blob, which the production build's stricter lib types require.

These two type errors were masked locally because tsc was OOMing silently without
the type-check script's --max-old-space-size flag.

* fix(providers): forward userId from the providers API to executeProviderRequest

Large-attachment prep now needs request.userId for presigned URLs and access
checks; the authenticated providers proxy has auth.userId but wasn't passing it,
so oversized attachments failed for logged-in callers. Forwarding it makes large
files work there and keeps the access check (verifyFileAccess) intact.

* fix(providers): fail clearly when a large attachment has no cloud storage

The doc claimed a base64 fallback that doesn't exist — above the inline cap there
is no base64, so without cloud storage the file previously reached the builder and
died with a generic read error. Throw a clear 'requires cloud file storage' error
at the point of detection and correct the doc.
2026-06-16 09:16:10 -07:00
2026-06-11 18:13:21 -07:00

Sim Logo

The open-source AI workspace where teams build, deploy, and manage AI agents. Build conversationally, visually, or with code. Connect 1,000+ integrations and every major LLM to automate real work.

Sim.ai Discord Twitter Documentation

Ask DeepWiki Set Up with Cursor

Build everything in Chat

Your AI command center. Describe what you want in plain language. Sim knows your entire workspace and takes action: building agents, running them, querying data, and more.

Sim building and running an agent from chat

Create files and documents

Generate documents, reports, and presentations from a single prompt, grounded in your workspace data.

Sim generating a document from a prompt

Ground agents in your knowledge

Upload documents to a knowledge base and let agents answer questions from your own content.

Creating a knowledge base

Structured data with Tables

A database, built in. Store, query, and wire structured data into agent runs.

Tables view with typed columns

Build visually with Workflows

Prefer a canvas? Design agents block by block in the visual builder, and let Sim generate blocks, wire variables, and fix errors from natural language.

Workflow builder demo

Quickstart

Cloud-hosted: sim.ai

Sim.ai

Self-hosted: NPM Package

npx simstudio

→ http://localhost:3000

Note

Docker must be installed and running on your machine.

Options

Flag Description
-p, --port <port> Port to run Sim on (default 3000)
--no-pull Skip pulling latest Docker images

Self-hosted: Docker Compose

git clone https://github.com/simstudioai/sim.git && cd sim
docker compose -f docker-compose.prod.yml up -d

Open http://localhost:3000

Sim also supports local models via Ollama and vLLM. See the Docker self-hosting docs for setup details.

Self-hosted: Manual Setup

Requirements: Bun, Node.js v20+, PostgreSQL 12+ with pgvector

  1. Clone and install:
git clone https://github.com/simstudioai/sim.git
cd sim
bun install
bun run prepare  # Set up pre-commit hooks
  1. Set up PostgreSQL with pgvector:
docker run --name simstudio-db -e POSTGRES_PASSWORD=your_password -e POSTGRES_DB=simstudio -p 5432:5432 -d pgvector/pgvector:pg17

Or install manually via the pgvector guide.

  1. Configure environment:
cp apps/sim/.env.example apps/sim/.env
# Create your secrets
perl -i -pe "s/your_encryption_key/$(openssl rand -hex 32)/" apps/sim/.env
perl -i -pe "s/your_internal_api_secret/$(openssl rand -hex 32)/" apps/sim/.env
perl -i -pe "s/your_api_encryption_key/$(openssl rand -hex 32)/" apps/sim/.env
# DB configs for migration
cp packages/db/.env.example packages/db/.env
# Edit both .env files to set DATABASE_URL="postgresql://postgres:your_password@localhost:5432/simstudio"
  1. Run migrations:
cd packages/db && bun run db:migrate
  1. Start development servers:
bun run dev:full  # Starts Next.js app and realtime socket server

Or run separately: bun run dev (Next.js) and cd apps/sim && bun run dev:sockets (realtime).

Chat API Keys

Chat is a Sim-managed service. To use Chat on a self-hosted instance:

  • Go to https://sim.ai → Settings → Chat keys and generate a Chat API key
  • Set COPILOT_API_KEY environment variable in your self-hosted apps/sim/.env file to that value

Environment Variables

See the environment variables reference for the full list, or apps/sim/.env.example for defaults.

Tech Stack

Contributing

We welcome contributions! Please see our Contributing Guide for details.

License

This project is licensed under the Apache License 2.0 - see the LICENSE file for details.

Made with ❤️ by the Sim Team

Languages
TypeScript 77%
MDX 20.8%
JavaScript 1.9%
CSS 0.1%