Theodore LiandClaude Fable 5 a028d07e7b improvement(mothership): user_table speed parity — limit bounds, background import/delete/update jobs (#5012)
* improvement(mothership): user_table speed parity — limit bounds, async import/delete/update jobs

- query_rows / filter ops clamp limit to the contract maxes; query_rows
  skips execution metadata.
- import_file / create_from_file (large CSV/TSV) and delete_rows_by_filter
  (>1000 unbounded matches) dispatch background table jobs, claiming the
  per-table job slot; inline paths claim the slot too.
- update_rows_by_filter now escalates the same way: >1000 unbounded matches
  run as a background table job (new 'update' job type + runTableUpdate worker
  + tableUpdateTask), so a broad update on a huge table no longer loads every
  row into the request. Best-effort/non-atomic and skips workflow recompute
  (documented); unique-column patches stay inline. Pagination is limit/offset.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(mothership): trim user_table catalog copy to the essentials

Drop the verbose doomedCount/affectedCount, delete-mask, workflow-recompute,
and unique-column asides from the bulk-op descriptions. The model only needs:
large ops return { jobId }, limit maxes at 1000, one job per table.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* improvement(mothership): make user_table limit cap internal, not model-facing

The model can now pass any limit — no "cannot exceed 1000" rejection. 1000
becomes an internal threshold: query_rows clamps the page to MAX_QUERY_LIMIT
(totalCount signals truncation; the model pages with offset), and bulk filter
ops above the cap run as background jobs.

update_rows_by_filter loads full row data inline, so an explicit limit above
the cap escalates to the background worker with a new maxRows budget (the worker
stops after maxRows; update has no read mask so the cap is exact). delete only
loads ids inline, so an explicit limit (any size) stays inline — only unbounded
deletes use the masked background path, which would over-hide a bounded delete.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* improvement(mothership): bounded delete above the cap runs async, not inline

An explicit delete limit now mirrors update: ≤1000 runs inline, above the cap it
escalates to the background worker honoring the limit via maxRows — instead of
always staying inline. The worker stops after maxRows (per-page fetch capped to
the remaining budget).

Bounded background deletes skip pendingDeleteMask: the filter-based mask hides
every match, which would over-hide the rows beyond the cap the job never deletes.
Unmasked, a bounded delete is eventually consistent like a bounded update (rows
disappear as deleted), and doomedCount is omitted from the payload so the count
isn't double-subtracted.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(mothership): tidy user_table limit/offset param copy

Drop "Any value is allowed" from the limit description and restore the original
offset description.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(tables): skip pendingDeleteMask for bounded background deletes

The bounded-delete commit (f1ee3e9) persisted maxRows and omitted doomedCount
but the pendingDeleteMask guard that makes it work was left uncommitted, so the
shipped mask still hid every filter+cutoff match — over-hiding the rows beyond
maxRows that the job never deletes (they vanished from reads until the job ended,
then reappeared). Return no mask when maxRows is set: a bounded delete is
eventually consistent (rows disappear as deleted), like a bounded update.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(mothership): drop redundant background note from limit arg

The op descriptions already cover background escalation; the limit arg only
needs to say what the param does.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-17 19:34:39 -04:00
2026-06-11 18:13:21 -07:00

Sim Logo

The open-source AI workspace where teams build, deploy, and manage AI agents. Build conversationally, visually, or with code. Connect 1,000+ integrations and every major LLM to automate real work.

Sim.ai Discord Twitter Documentation

Ask DeepWiki Set Up with Cursor

Build everything in Chat

Your AI command center. Describe what you want in plain language. Sim knows your entire workspace and takes action: building agents, running them, querying data, and more.

Sim building and running an agent from chat

Create files and documents

Generate documents, reports, and presentations from a single prompt, grounded in your workspace data.

Sim generating a document from a prompt

Ground agents in your knowledge

Upload documents to a knowledge base and let agents answer questions from your own content.

Creating a knowledge base

Structured data with Tables

A database, built in. Store, query, and wire structured data into agent runs.

Tables view with typed columns

Build visually with Workflows

Prefer a canvas? Design agents block by block in the visual builder, and let Sim generate blocks, wire variables, and fix errors from natural language.

Workflow builder demo

Quickstart

Cloud-hosted: sim.ai

Sim.ai

Self-hosted: NPM Package

npx simstudio

→ http://localhost:3000

Note

Docker must be installed and running on your machine.

Options

Flag Description
-p, --port <port> Port to run Sim on (default 3000)
--no-pull Skip pulling latest Docker images

Self-hosted: Docker Compose

git clone https://github.com/simstudioai/sim.git && cd sim
docker compose -f docker-compose.prod.yml up -d

Open http://localhost:3000

Sim also supports local models via Ollama and vLLM. See the Docker self-hosting docs for setup details.

Self-hosted: Manual Setup

Requirements: Bun, Node.js v20+, PostgreSQL 12+ with pgvector

  1. Clone and install:
git clone https://github.com/simstudioai/sim.git
cd sim
bun install
bun run prepare  # Set up pre-commit hooks
  1. Set up PostgreSQL with pgvector:
docker run --name simstudio-db -e POSTGRES_PASSWORD=your_password -e POSTGRES_DB=simstudio -p 5432:5432 -d pgvector/pgvector:pg17

Or install manually via the pgvector guide.

  1. Configure environment:
cp apps/sim/.env.example apps/sim/.env
# Create your secrets
perl -i -pe "s/your_encryption_key/$(openssl rand -hex 32)/" apps/sim/.env
perl -i -pe "s/your_internal_api_secret/$(openssl rand -hex 32)/" apps/sim/.env
perl -i -pe "s/your_api_encryption_key/$(openssl rand -hex 32)/" apps/sim/.env
# DB configs for migration
cp packages/db/.env.example packages/db/.env
# Edit both .env files to set DATABASE_URL="postgresql://postgres:your_password@localhost:5432/simstudio"
  1. Run migrations:
cd packages/db && bun run db:migrate
  1. Start development servers:
bun run dev:full  # Starts Next.js app and realtime socket server

Or run separately: bun run dev (Next.js) and cd apps/sim && bun run dev:sockets (realtime).

Chat API Keys

Chat is a Sim-managed service. To use Chat on a self-hosted instance:

  • Go to https://sim.ai → Settings → Chat keys and generate a Chat API key
  • Set COPILOT_API_KEY environment variable in your self-hosted apps/sim/.env file to that value

Environment Variables

See the environment variables reference for the full list, or apps/sim/.env.example for defaults.

Tech Stack

Contributing

We welcome contributions! Please see our Contributing Guide for details.

License

This project is licensed under the Apache License 2.0 - see the LICENSE file for details.

Made with ❤️ by the Sim Team

Languages
TypeScript 77%
MDX 20.8%
JavaScript 1.9%
CSS 0.1%