feat(code): cli sandboxes, enterprise timeouts, secrets projections, resolver lift, workflow exec cancellations (#6247)

* feat(code): cli sandboxes, enterprise timeouts, secrets projections, resolver lift

* fix(execution): harden compatibility and secret diagnostics

* fix(execution): harden generated JavaScript literals

* fix(execution): align timeout cleanup semantics

* fix(tables): decouple stale job cleanup

* fix(execution): drain stale workflow backlog

* test(sandbox): make deadline assertions timing-safe

* fix(execution): lock cleanup candidate batches

* fix(execution): preserve cleanup failure metrics

* cancel route fixes

* separate out mship template and func template

* fix

* fix(execution): harden secret projection and block runs

* fix(workflow): validate draft execution state

* run from block ui disabling

* feat(copilot): expose Sim sandboxes to mothership

* feat(copilot): expose sandbox capability catalog in VFS

* Updates

* fix legacy logs showing up

* fix(copilot): keep sandbox config visible

* fix model provenance issues

* fix lint'

* more lint

* more

* test(files): align provenance copy query order

* consolidate migrations, rollout compat

* integration projections

* update skills

* fix

* add provenance linters

* fix: address review and compatibility regressions

* fix: make tool boundary audit Bun 1.3 compatible

---------

Co-authored-by: Siddharth Ganesan <siddharthganesan@gmail.com>
This commit is contained in:
Vikhyath Mondreti
2026-08-05 19:22:04 -07:00
committed by GitHub
co-authored by Siddharth Ganesan
parent 5baa7a41ec
commit 117fe3137b
826 changed files with 95636 additions and 7877 deletions
@@ -84,6 +84,8 @@ return {
};
```
For a secret used as a complete JavaScript expression, prefer the unquoted form, such as `const apiKey = {{OPENWEATHER_API_KEY}};`. Quoted and embedded forms remain supported, including the placeholder embedded in the URL above, `"Bearer {{KEY}}"`, template literals, and JavaScript regex literals. The value is bound separately when the tool executes rather than pasted into its source, so its exact string contents are preserved.
<Callout type="info">
You can also use the AI wand to generate code from a description. Environment variables are referenced with `&#123;&#123;KEY&#125;&#125;` syntax.
</Callout>
@@ -114,7 +116,7 @@ Once created, custom tools appear alongside built-in tools when configuring an A
- **Async/await** — Your code runs in an async context, so you can use `await` directly
- **fetch()** — Make HTTP requests to external APIs
- **Node.js built-ins** — Access to `crypto`, `Buffer`, and other standard modules
- **Environment variables** — Use `{{KEY}}` syntax to inject secrets
- **Environment variables** — Use `{{KEY}}` syntax to bind secrets at execution time without placing plaintext in source code
### Limitations
@@ -155,7 +157,7 @@ From **Settings → Custom Tools** you can:
<FAQ items={[
{ question: "Can I use custom tools in standalone blocks (not agents)?", answer: "No. Custom tools are designed for use within Agent blocks, where the AI model decides when to call them. For deterministic tool execution, use the Function block instead." },
{ question: "How do I pass API keys to my custom tool code?", answer: "Use double curly brace syntax such as {{MY_API_KEY}} for a value saved under Settings → Secrets. The real value is injected at execution time, while exact occurrences are masked in the trace copy. If the tool returns the secret, the raw result still reaches the Agent. See Execution log protection under Secrets for details." },
{ question: "How do I pass API keys to my custom tool code?", answer: "Use double curly brace syntax such as {{MY_API_KEY}} for a value saved under Settings → Secrets. The value is bound outside the source at execution time. Exact occurrences are masked in the trace copy and replaced with the placeholder before a tool result is sent back to the Agent model. Encoded or otherwise transformed values cannot be recognized reliably. See Execution log protection under Secrets for details." },
{ question: "Can I use external npm packages?", answer: "No. Custom tool code runs in a sandboxed environment with access to built-in Node.js modules and fetch(), but not external packages. For complex dependencies, consider calling an external API that wraps the functionality you need." },
{ question: "What's the difference between custom tools and the Function block?", answer: "Custom tools are called by AI agents when they decide the tool is relevant — the agent chooses when to use it. Function blocks run deterministically at a fixed point in the workflow. Use custom tools for agent-driven actions and Function blocks for predictable data transformations." },
{ question: "Are custom tools shared across the workspace?", answer: "Yes. Custom tools are workspace-scoped, so all workspace members can use them in their workflows." },
@@ -77,6 +77,7 @@ result = client.execute_workflow(
- `stream` (bool, optional): Enable streaming responses (default: False)
- `selected_outputs` (list[str], optional): Block outputs to stream in `blockName.attribute` format (e.g., `["agent1.content"]`)
- `async_execution` (bool, optional): Execute asynchronously (default: False)
- `execution_timeout_seconds` (int, optional): Optional server-side async execution cap from 1 to 604800 seconds. Requires `async_execution=True` and cannot extend the account policy.
**Returns:** `WorkflowExecutionResult | AsyncExecutionResult`
@@ -770,4 +771,4 @@ import { FAQ } from '@/components/ui/faq'
{ question: "Can I use the Python SDK as a context manager?", answer: "Yes. The SimStudioClient supports Python's context manager protocol. Use it with the 'with' statement to automatically close the underlying HTTP session when you are done, which is especially useful for scripts that create and discard client instances." },
{ question: "How do I handle different types of errors from the SDK?", answer: "The SDK raises SimStudioError with a code property for API-specific errors. Common error codes are UNAUTHORIZED (invalid API key), TIMEOUT (request timed out), RATE_LIMIT_EXCEEDED (too many requests), USAGE_LIMIT_EXCEEDED (billing limit reached), and EXECUTION_ERROR (workflow failed). Use the error code to implement targeted error handling and recovery logic." },
{ question: "How do I monitor my API usage and remaining quota?", answer: "Use the get_usage_limits() method to check your current usage. It returns sync and async rate limit details (limit, remaining, reset time, whether you are currently limited), plus your current period cost, usage limit, and plan tier. This lets you monitor consumption and alert before hitting limits." },
]} />
]} />
@@ -91,6 +91,7 @@ const result = await client.executeWorkflow('workflow-id', { message: 'Hello, wo
- `stream` (boolean): Enable streaming responses (default: false)
- `selectedOutputs` (string[]): Block outputs to stream in `blockName.attribute` format (e.g., `["agent1.content"]`)
- `async` (boolean): Execute asynchronously (default: false)
- `executionTimeoutSeconds` (number): Optional server-side async execution cap from 1 to 604800 seconds. Requires `async: true` and cannot extend the account policy.
**Returns:** `Promise<WorkflowExecutionResult | AsyncExecutionResult>`
+3 -2
View File
@@ -361,10 +361,11 @@ Storage follows the paid tier: Pro and Pro for Teams share the Pro limit, while
| Plan | Sync | Async |
|------|------|-------|
| **Free** | 5 minutes | 90 minutes |
| **Pro / Max / Team / Enterprise** | 50 minutes | 90 minutes |
| **Pro / Max / Team** | 50 minutes | 90 minutes |
| **Enterprise** | 50 minutes | 90 minutes by default; configurable up to 7 days |
**Sync runs** complete immediately and return results directly. These are triggered via the API with `async: false` (default) or through the UI.
**Async runs** (triggered via API with `async: true`, webhooks, or schedules) run in the background.
**Async runs** (triggered via API with `async: true`, webhooks, schedules, polling triggers, or table workflow columns) run in the background. A direct async API request can shorten the account policy for that request, but cannot extend it.
<Callout type="info">
If a workflow exceeds its time limit, it will be terminated and marked as failed with a timeout error. Design long-running workflows to use async runs or break them into smaller workflows.
@@ -69,10 +69,10 @@ Select the secret you want to use. The reference appears highlighted in blue and
When a saved secret is successfully substituted through a `{{KEY}}` reference, Sim masks exact, case-sensitive occurrences of its resolved value in log-facing content. This includes the editor's live block-log display, Logs Overview input and output, stored execution traces, log-read API responses, and the Logs block's **Get Run Details** output. Function and Agent span inputs, outputs, and errors, Agent thinking, and Agent tool-call arguments, results, and errors are protected. The replacement is normally shown as `{{KEY}}`.
This is an observability projection only. Secret resolution and workflow behavior are unchanged: blocks, tools, models, and downstream steps receive the real runtime value. Stored functional execution data, workflow execution responses, streams, callbacks, block state, and snapshots are not rewritten. Log-facing views and read APIs receive a separate protected copy, so the Logs Overview **Workflow Input** and **Workflow Output** are masked without changing the underlying workflow result.
Secret resolution and functional workflow behavior are unchanged: blocks, tools, and downstream steps receive the real runtime value. Stored functional execution data, workflow execution responses, streams, callbacks, block state, and snapshots are not rewritten. Log-facing views and read APIs receive a separate protected copy, so the Logs Overview **Workflow Input** and **Workflow Output** are masked without changing the underlying workflow result. Model requests receive another protected projection: exact secret values known to the run are replaced with `{{KEY}}` before model-visible messages, prompts, tool arguments, or tool continuations leave Sim.
<Callout type="warn">
Masking is activated only when Sim successfully resolves a value from **Settings → Secrets** through `{{KEY}}`. A hardcoded literal, direct `environmentVariables['KEY']` read, or shell `$KEY` read does not activate it by itself. Once activated, every exact occurrence of that value in the run's log-facing content is masked. Encoded, hashed, or otherwise transformed versions are not matched. Do not deliberately return or print secrets.
Execution-log masking is activated only when Sim successfully resolves a value from **Settings → Secrets** through `{{KEY}}`. A hardcoded literal, direct `environmentVariables['KEY']` read, or shell `$KEY` read does not activate log masking by itself. Model-bound projection also checks the run's authorized secret catalog, including direct reads, but both protections match only exact values. Encoded, hashed, fragmented, or otherwise transformed versions are not matched. Do not deliberately return or print secrets.
</Callout>
### Copilot code execution
@@ -136,7 +136,7 @@ When a workflow runs, secrets resolve in this order:
<FAQ items={[
{ question: "Are my secrets encrypted at rest?", answer: "Yes. Values saved under Secrets are encrypted before being stored in the database." },
{ question: "Can a saved secret still appear in a workflow result?", answer: "Yes. Functional workflow data is not rewritten, so the raw value can still reach downstream blocks, tools, and models and can appear in workflow execution responses, streams, or callbacks if your workflow deliberately returns or prints it. Log-facing views and read APIs receive a protected copy after a successful {{KEY}} substitution. Copilot-visible tool results also mask exact activated values, but transformed values remain outside that protection." },
{ question: "Can a saved secret still appear in a workflow result?", answer: "Yes. Functional workflow data is not rewritten, so the raw value can still reach downstream blocks and tools and can appear in workflow execution responses, streams, or callbacks if your workflow deliberately returns or prints it. Log-facing views and read APIs receive a protected copy after a successful {{KEY}} resolution. Before content is sent to a model, exact values from the run's authorized secret catalog are replaced with placeholders, but encoded or otherwise transformed values remain outside that protection." },
{ question: "What happens if both a workspace secret and a personal secret have the same key name?", answer: "Among secrets available to the execution actor, the workspace secret takes precedence and the personal secret is the fallback. An inaccessible workspace secret does not shadow an authorized personal value." },
{ question: "Who determines which personal secret is used for automated runs?", answer: "For manual runs, the personal secrets of the user who clicked Run are used as fallback. For automated runs triggered by API, webhook, or schedule, the personal secrets of the workflow owner are used instead." },
{ question: "Can I import secrets from a .env file?", answer: "Yes. Paste .env-style content (KEY=VALUE format) into any key or value field and the secrets will be auto-populated. The parser supports export KEY=VALUE, quoted values, and inline comments." },
@@ -53,7 +53,111 @@ The individual flags also work on their own if you would rather opt in one at a
| Sim Mailer inbox | `INBOX_ENABLED` | `NEXT_PUBLIC_INBOX_ENABLED` |
| Sandboxes | `SANDBOXES_ENABLED` | `NEXT_PUBLIC_SANDBOXES_ENABLED` |
Sandboxes also need a remote execution provider, since the deployment builds the images itself: set `E2B_API_KEY`, or `SANDBOX_PROVIDER=daytona` with a `DAYTONA_API_KEY` that has `write:snapshots` and `write:sandboxes`. Without one the settings section appears but builds fail.
Sandboxes also need a remote execution provider and a dedicated Function base
image. Build and configure that base before enabling the UI; custom workspace
sandboxes layer their packages on top of it.
For E2B:
```bash
E2B_API_KEY=... \
bun run apps/sim/scripts/build-function-e2b-template.ts \
--name sim-function
SANDBOX_PROVIDER=e2b
E2B_ENABLED=true
E2B_API_KEY=...
E2B_FUNCTION_TEMPLATE_ID=<sim-function-template>:<sim-function-build-id>
E2B_FUNCTION_TEMPLATE_GENERATION=<release-epoch-ms>
SANDBOXES_ENABLED=true
NEXT_PUBLIC_SANDBOXES_ENABLED=true
```
The builder uses E2B's maintained `code-interpreter-v1` base, assigns a fresh
release generation, and prints both runtime values. `--generation` remains
available for release automation, and `--base-template` accepts an immutable
base override when a deployment deliberately owns one.
For Daytona, use the immutable snapshot ID printed by the builder. The API key needs
`write:snapshots` to build and `write:sandboxes` to execute:
```bash
DAYTONA_API_KEY=... \
bun run apps/sim/scripts/build-function-daytona-snapshot.ts \
--name sim-function-2026-08-03 \
--parity-manifest /tmp/function-sandbox-manifest.json
SANDBOX_PROVIDER=daytona
DAYTONA_API_KEY=...
DAYTONA_FUNCTION_SNAPSHOT_ID=<snapshot-uuid>
SANDBOXES_ENABLED=true
NEXT_PUBLIC_SANDBOXES_ENABLED=true
```
`SANDBOXES_ENABLED` grants the server-side self-hosted entitlement.
`NEXT_PUBLIC_SANDBOXES_ENABLED` exposes remote Python and Shell plus custom
sandbox management in the browser. Set the public flag only after the selected
provider has credentials and a valid immutable Function base configured.
Mothership's `function_execute` and `run_code` tools use Mothership's separate
shell image, including for JavaScript without imports. If the deployment uses
Mothership code tools, also configure the image produced by the Mothership
release process for the selected provider:
```bash
# E2B
MOTHERSHIP_E2B_TEMPLATE_ID=<mothership-shell-template-ref>
# Daytona
DAYTONA_SHELL_SNAPSHOT_ID=<mothership-shell-snapshot-ref>
```
These values are selected only for workflow Copilot and workspace Mothership
code-tool calls. They never replace or act as a fallback for
`E2B_FUNCTION_TEMPLATE_ID` or
`DAYTONA_FUNCTION_SNAPSHOT_ID`; Function blocks and custom workspace sandboxes
continue to use the dedicated Function base.
Use E2B as the release baseline before building or promoting Daytona:
```bash
# 1. Verify the exact E2B Function build and capture its accepted package/runtime surface.
E2B_ENABLED=true \
E2B_API_KEY=... \
E2B_FUNCTION_TEMPLATE_ID=<sim-function-template>:<sim-function-build-id> \
E2B_FUNCTION_TEMPLATE_GENERATION=<release-epoch-ms> \
SANDBOX_PARITY_MANIFEST_OUT=/tmp/function-sandbox-manifest.json \
bun run apps/sim/scripts/verify-sandbox-parity.ts
# 2. Pin Daytona's reconstructed packages to that accepted E2B manifest.
DAYTONA_API_KEY=... \
bun run apps/sim/scripts/build-function-daytona-snapshot.ts \
--name sim-function-2026-08-03 \
--parity-manifest /tmp/function-sandbox-manifest.json
# 3. Verify the immutable Daytona snapshot against the same baseline before promotion.
SANDBOX_PROVIDER=daytona \
DAYTONA_API_KEY=... \
DAYTONA_FUNCTION_SNAPSHOT_ID=<snapshot-uuid> \
SANDBOX_PARITY_MANIFEST_BASELINE=/tmp/function-sandbox-manifest.json \
bun run apps/sim/scripts/verify-sandbox-parity.ts
```
<Callout type="warning">
`E2B_FUNCTION_TEMPLATE_ID` and `DAYTONA_FUNCTION_SNAPSHOT_ID` fail closed when
unset or mutable. The E2B value must be an exact `<template>:<build-id>` ref,
and `E2B_FUNCTION_TEMPLATE_GENERATION` must be the monotonic value printed by
the same build. The Daytona value must be a snapshot ID rather than a name. Sim does not
fall back to `MOTHERSHIP_E2B_TEMPLATE_ID` or
`DAYTONA_SHELL_SNAPSHOT_ID`; Mothership, Function, document, and Pi images have
separate package contracts and release cadences.
</Callout>
<Callout type="info">
Assign every promoted E2B Function base a generation greater than every prior
deployment. A rollback is a new promotion and therefore also needs a new,
higher generation; do not reuse the generation from the older release.
</Callout>
<Callout type="warning">
Data retention is the one feature that deletes data. Its flag controls the cleanup pass, not the settings screen — retention windows are always configurable. Nothing is ever deleted until you enable it, and even then only against windows you configured explicitly. Sim never applies the hosted plan defaults to a self-hosted deployment.
@@ -237,6 +237,21 @@ AZURE_STORAGE_OG_IMAGES_CONTAINER_NAME=og-images
AZURE_STORAGE_WORKSPACE_LOGOS_CONTAINER_NAME=workspace-logos
```
Browser uploads also require an account-level Blob service CORS rule. This rule explicitly allows Sim's create-only `If-None-Match` precondition, Azure upload headers, and multipart `ETag` reads:
```bash
az storage cors add \
--services b \
--methods GET PUT OPTIONS \
--origins "https://your-sim-domain.com" \
--allowed-headers "Content-Type" "If-None-Match" "x-ms-*" \
--exposed-headers "ETag" \
--max-age 3600 \
--account-name mystorageaccount
```
The CORS rule applies to every blob container in the storage account, so it only needs to be added once per account.
A full Helm example lives at `helm/sim/examples/values-azure.yaml`.
## Set up Google Cloud Storage
@@ -276,6 +291,7 @@ cat > /tmp/cors.json <<'EOF'
"responseHeader": [
"Content-Type",
"ETag",
"x-goog-if-generation-match",
"x-goog-meta-originalname",
"x-goog-meta-uploadedat",
"x-goog-meta-purpose",
@@ -283,7 +299,8 @@ cat > /tmp/cors.json <<'EOF'
"x-goog-meta-workspaceid",
"x-goog-meta-folderid",
"x-goog-meta-workflowid",
"x-goog-meta-executionid"
"x-goog-meta-executionid",
"x-goog-meta-simuploadid"
],
"maxAgeSeconds": 3600
}
@@ -297,7 +314,7 @@ done
```
<Callout type="info">
Header names must be listed individually — GCS CORS matches `responseHeader` entries exactly and does not support wildcards like `x-goog-meta-*`. `ETag` is required because large-file multipart uploads read each part's `ETag` from the browser, and CORS hides the header otherwise.
Header names must be listed individually — GCS CORS matches `responseHeader` entries exactly and does not support wildcards like `x-goog-meta-*`. `ETag` is required because large-file multipart uploads read each part's `ETag` from the browser, and CORS hides the header otherwise. `x-goog-if-generation-match` is required by Sim's create-only signed uploads, which prevent a reused upload URL from replacing existing bytes. `x-goog-meta-simuploadid` carries the opaque receipt used to verify an upload after an ambiguous network response.
</Callout>
</Step>
@@ -1,6 +1,6 @@
---
title: Function
description: The Function block runs your JavaScript or Python code as a step and returns what it produces.
description: The Function block runs your JavaScript, Python, or Shell code as a step and returns what it produces.
---
import { Callout } from 'fumadocs-ui/components/callout'
@@ -8,7 +8,7 @@ import { Tab, Tabs } from 'fumadocs-ui/components/tabs'
import { BlockPreview, WorkflowPreview, FUNCTION_RESHAPE_WORKFLOW, FUNCTION_VALIDATE_WORKFLOW } from '@/components/workflow-preview'
import { FAQ } from '@/components/ui/faq'
The **Function block** runs your own JavaScript or Python code as one step of a workflow. Use it to reshape a value, run a calculation, or add logic no other block covers.
The **Function block** runs your own JavaScript, Python, or Shell code as one step of a workflow. Use it to reshape a value, run a calculation, call a CLI, or add logic no other block covers.
<BlockPreview type="function" />
@@ -16,9 +16,9 @@ The **Function block** runs your own JavaScript or Python code as one step of a
### Code
Your code, in JavaScript (the default) or Python — pick the language on the block. Reference an earlier output directly, with no quotes around the tag, and read an environment variable with `{{VAR}}`:
JavaScript is the default. Python and Shell appear when a remote sandbox provider is enabled. Reference an earlier output directly, with no quotes around the tag, and read an environment variable with `{{VAR}}`:
<Tabs items={['JavaScript', 'Python']}>
<Tabs items={['JavaScript', 'Python', 'Shell']}>
<Tab value="JavaScript">
```javascript
const data = <api.data>;
@@ -27,14 +27,51 @@ Your code, in JavaScript (the default) or Python — pick the language on the bl
</Tab>
<Tab value="Python">
```python
import json
data = json.loads('<api.data>')
print(json.dumps([i["id"] for i in data["items"] if i["active"]]))
data = <api.data>
__sim_result__ = [item["id"] for item in data["items"] if item["active"]]
```
</Tab>
<Tab value="Shell">
```bash
set -euo pipefail
active_ids=$(printf '%s' <api.data> | jq -c '[.items[] | select(.active) | .id]')
printf '__SIM_RESULT__=%s\n' "$active_ids"
```
</Tab>
</Tabs>
Return a value in JavaScript with `return`. In Python, print it as JSON to stdout with `print(json.dumps(...))`; the block captures stdout as the result. Your code runs in an async context, so you can `await` directly in JavaScript.
JavaScript returns a value with `return`. Python runs as a normal module: assign the value for downstream blocks to `__sim_result__`. A full script with functions, imports, and an `if __name__ == '__main__':` guard works as written; legacy snippets with a top-level `return` remain supported. `print()` is logging and goes to `stdout` rather than becoming the result.
Shell returns structured data by printing a line beginning with `__SIM_RESULT__=`. A valid JSON payload becomes an object, array, number, boolean, or string; a non-JSON payload becomes a string. The marker line is removed from `stdout`. Other command output remains available in `stdout`.
### Secret placeholders in code
When an environment variable is the complete JavaScript or Python expression, prefer the unquoted form:
```javascript
const apiKey = {{API_KEY}};
```
Existing quoted and embedded forms are also supported, including `"{{API_KEY}}"`, `"Bearer {{API_KEY}}"`, and placeholders in template literals. Function and Custom Tool code use the same compiler at the execution boundary. It binds the secret separately from the source instead of pasting plaintext into your code, so quotes, backslashes, newlines, and string values such as `"123"` and `"true"` retain their exact contents and do not become JavaScript or Python literals of another type.
JavaScript regex literals can contain a placeholder:
```javascript
const matcher = /^{{PATH_PATTERN}}$/i;
```
The value is interpreted as raw regex pattern text and the literal's flags are preserved. If you need the value matched literally rather than as a pattern, escape regex metacharacters before constructing the expression.
In Shell, use the same `{{KEY}}` syntax and write `"{{KEY}}"` when the value should remain one scalar argument. Bare placeholders remain supported for existing workflows and retain Bash's native unquoted behavior, including word splitting and regex-pattern semantics. Placeholders also work inside quoted heredocs without enabling unrelated shell expansion:
```bash
cat > /tmp/request.txt <<'REQUEST'
Authorization: Bearer {{API_KEY}}
Home remains literal: $HOME
REQUEST
```
Sim supplies the rendered heredoc privately while preserving the quoted delimiter's literal `$VAR`, backtick, and command-substitution behavior.
## Outputs
@@ -45,40 +82,47 @@ Return a value in JavaScript with `return`. In Python, print it as JSON to stdou
## Language
JavaScript runs in a fast local sandbox, or in an [E2B](https://e2b.dev) sandbox when your code uses `import` or `require`. Python always runs in the E2B sandbox.
JavaScript without imports runs in a fast local sandbox. JavaScript with `import` or `require`, Python, and Shell run in the configured remote sandbox provider.
| | JavaScript | Python |
| --- | --- | --- |
| **Execution** | Local sandbox (fast), or E2B with imports | Always E2B sandbox |
| **Return a value** | `return { … }` | `print(json.dumps({ … }))` |
| **HTTP requests** | `fetch()` built-in | `requests` or `httpx` |
| **Best for** | quick transforms, JSON | data science, charts, complex math |
| | JavaScript | Python | Shell |
| --- | --- | --- | --- |
| **Execution** | Local when there are no imports; remote with imports | Always remote | Always remote |
| **Return a value** | `return { … }` | Assign `__sim_result__ = { … }` | Print `__SIM_RESULT__={…}` |
| **HTTP requests** | `fetch()` built in | `requests` or `httpx` | `curl` or an installed CLI |
| **Best for** | quick transforms and JSON | scripts, data science, charts, complex math | CLI workflows and system utilities |
<Callout type="info">
Python requires E2B. It is enabled by default on sim.ai; on a self-hosted instance, enable E2B to see Python in the language dropdown. Any figures you generate are captured as images automatically.
Python and Shell require a remote sandbox. They are enabled by default on sim.ai;
on a self-hosted instance, build and configure the provider's dedicated
[Function base](/platform/enterprise/self-hosted) first. Any Python figures you
generate are captured as images automatically.
</Callout>
{/* Package list provenance (verified 2026-06-10): E2B code-interpreter template
requirements (github.com/e2b-dev/code-interpreter, template/requirements.txt)
plus Sim's template additions — pip awscli/yq/csvkit — from
simstudioai/copilot scripts/e2b/template.ts. Update if either source changes. */}
The dedicated Function base has the same runtime and universal package contract
on E2B and Daytona. It includes this data-science stack; use a workspace sandbox
when another dependency must be present:
Beyond the Python standard library, the sandbox ships E2B's data-science stack plus a few Sim additions:
- **Data:** `pandas`, `numpy`, `scipy`, `xarray`, `numba`, `joblib`
- **Data and graphs:** `pandas`, `numpy`, `scipy`, `xarray`, `numba`, `networkx`
- **ML and NLP:** `scikit-learn`, `gensim`, `nltk`, `spacy`, `textblob`
- **Plots and images:** `matplotlib`, `seaborn`, `plotly`, `bokeh`, `pillow`, `opencv-python`, `scikit-image`, `imageio`
- **Plots and images:** `matplotlib`, `seaborn`, `plotly`, `bokeh`, `kaleido`, `pillow`, `opencv-python`, `scikit-image`, `imageio`, `tifffile`
- **Audio:** `librosa`, `soundfile`
- **Web and files:** `requests`, `aiohttp`, `beautifulsoup4`, `openpyxl`, `xlrd`, `python-docx`, `orjson`
- **Math and testing:** `sympy`, `pytest`
- **Sim additions:** `awscli`, `yq`, `csvkit`
- **Web, parsing, and files:** `requests`, `beautifulsoup4`, `lxml`, `openpyxl`, `xlrd`, `python-docx`, `xmltodict`, `PyYAML`, `tomlkit`, `simplejson`, `orjson`, `SQLAlchemy`
- **Math:** `sympy`
- **Application helpers:** `rich`, `typer`, `click`, `tqdm`, `Jinja2`, `pydantic`, `python-dateutil`, `pytz`, `psutil`, `filetype`, `python-slugify`, `parsedatetime`, `pytimeparse`
- **Generic CLIs:** `jq`, `yq`, `csvkit`, `zx`, `xmlstarlet`, `httpie`, `ripgrep`, `fd`, `bat`, `sqlite3`, `tar`, `gzip`, `bzip2`, ZIP/XZ/7z, media/image tools, and standard network utilities
Vendor and service clients such as AWS CLI and GitHub CLI are not part of the
universal base. Add them through the managed catalog. Database clients such as
PostgreSQL, MySQL, and Redis can be added as system packages when available from
the configured Debian repositories.
## Sandboxes
A sandbox is a named dependency set your workspace maintains — a language plus a
list of pip or npm packages. Select one on a Function block and its code can
import everything on that list. Leave it empty and the block runs on the default
image, exactly as before.
A sandbox is a named environment your workspace maintains: a language, a list of
pip or npm dependencies, optional Debian/APT system packages, and optional managed
CLI tools. Select one on a Function block and its code can import the dependencies
and run the commands it declares. Leave the selection empty to use the dedicated
Function base.
Create and edit sandboxes in **Settings → Sandboxes**. Only workspace admins can
create or edit them. On sim.ai they need an active Max or Enterprise plan;
@@ -87,9 +131,17 @@ self-hosted deployments turn them on with `SANDBOXES_ENABLED` (see
hidden when a deployment has no sandbox provider configured.
1. **Name** the sandbox — `bigquery-etl`, `scraping`, whatever the job is.
2. Pick the **language**. A sandbox is language-scoped, so a Python block only
ever lists Python sandboxes.
2. Pick the **language**. This selects pip or npm for the dependency list. Python
and JavaScript blocks list matching sandboxes; Shell can use either kind.
3. Paste your **dependencies**, one per line. Version pins are optional.
4. Add **system packages** by Debian/APT package coordinate, one per line. Use this
for ordinary command-line utilities such as `jq` and `ffmpeg`.
5. Select any **managed CLI tools** that require a specialized, verified installer.
The searchable catalog is grouped by cloud, Kubernetes, infrastructure,
deployment, data and storage, and security tools. Every entry uses a pinned,
integrity-checked vendor artifact.
Dependencies:
```
google-cloud-bigquery==3.25.0
@@ -97,41 +149,99 @@ pyairtable>=3.0
pandas
```
System packages:
```
shellcheck
pandoc
graphviz
```
Then open the block's advanced options and choose the sandbox under **Sandbox**.
In JavaScript the sandbox applies to code that uses `import` or `require` — that is
what sends the block to a remote sandbox in the first place, so a block without them
keeps running locally and ignores the selection. Python always runs remotely, so a
selected sandbox always applies.
The default and custom behavior is intentionally explicit:
- **JavaScript without imports** stays in the local isolated runtime for speed and
ignores the sandbox selection.
- **JavaScript with `import` or `require`** runs remotely. With no selection it
uses the Function base; with a sandbox it gets that sandbox's npm packages and
system packages and managed CLI tools.
- **Python and Shell** always run remotely. With no selection they get only the
Function base; with a sandbox they get its dependencies, system packages, and
managed CLI tools.
When a Function block next loads its sandbox options, a successful lookup that
confirms its selected sandbox was deleted clears that selection. Fetch or auth
failures leave the workflow unchanged.
<Callout type="info">
Two sandboxes with the same language and the same package list share one build,
so duplicating a set costs nothing. Editing a package list starts a new build;
runs already in flight keep using the old one. Deleting a sandbox frees its build
once nothing else uses it.
Two sandboxes with the same language, dependency list, system packages, and managed
CLI tools share one build, so duplicating a set costs nothing. Editing any part of
that install specification starts a new build. Runs already in flight keep using
the old one. Deleting a sandbox frees its build once nothing else uses it.
</Callout>
### Build status
On sim.ai, each dependency set is prebuilt into a reusable image, so runs pay no
install cost. The status row in Settings shows **Queued**, **Building**,
On sim.ai, each sandbox specification is prebuilt into a reusable image, so runs
pay no install cost. The status row in Settings shows **Queued**, **Building**,
**Ready**, or **Failed**. A failed build reports what went wrong — a package that
does not exist, a version that has no match, a resolver conflict — with the
installer log behind a disclosure.
Running a block before its sandbox is **Ready** stops the run and shows you the
status. A failed build is retried periodically on its own; to retry immediately,
save the sandbox again in Settings.
status. A failed build is retried periodically on its own. To retry immediately,
open the three-dot menu at the end of the failed status row and choose **Retry
build**. Save or discard unsaved sandbox edits first so the retry always uses the
visible spec.
On a self-hosted deployment using Daytona, dependencies install inside the
sandbox at the start of every run instead, adding roughly 10–30 seconds per
execution. Prebuilt images require E2B.
On a self-hosted deployment using Daytona, dependencies, system packages, and
managed CLI tools install inside the sandbox at the start of every run instead,
adding startup time per execution. Prebuilt images require E2B.
### System packages, managed CLIs, and one-off installs
Use **System packages** for CLIs available from the sandbox's Debian/APT
repositories. Entries may include an architecture or exact version, such as
`package:architecture=version`. They are validated, deduplicated, and installed
when the sandbox is built or prepared.
The **Managed CLI tools** selector covers CLIs that need more than a normal APT
installation, such as a pinned vendor archive, PATH setup, or command verification.
Service and vendor CLIs are intentionally not part of the universal Function base.
For example, after adding **Google Cloud CLI** to a sandbox, a Shell Function can
run `bq` directly:
```bash
set -euo pipefail
bq query --use_legacy_sql=false --format=json 'SELECT CURRENT_DATE() AS today'
```
For a command needed only once, install it at the start of the Shell script. The
remote sandbox is ephemeral, so the install applies only to that Function call
and adds to its execution time:
```bash
set -euo pipefail
python -m pip install --quiet csvkit
csvcut -n /tmp/input.csv
```
Use a custom sandbox instead when reproducibility or startup time matters. Put
ordinary Debian utilities in **System packages**; use the managed selector only
for one of its specialized installers. Arbitrary install commands are not saved
in a sandbox specification.
### What is allowed
Package names and version specifiers only. URLs, `git+` references, `-e`, local
paths, `--index-url`, and npm aliases are rejected, with the offending line
number reported. A sandbox may declare up to 50 packages.
Dependency fields accept package names and version specifiers only. URLs, `git+`
references, `-e`, local paths, `--index-url`, and npm aliases are rejected, with
the offending line number reported.
System packages use Debian coordinates in the form
`package[:architecture][=version]`. Flags, URLs, paths, whitespace, globs, and
shell syntax are rejected. A sandbox may declare up to 50 dependencies, 50 system
packages, and 10 managed CLI tools.
## Scoping secrets for agent tools
@@ -147,9 +257,10 @@ tool configuration and pick the names the code may read. Two things change:
- Only those secrets are injected. `{{OTHER_SECRET}}` no longer resolves either.
- The selected **names** are added to the tool's description, so the model knows
what it can reference. Values are injected server-side when the Function runs.
If the Function returns a secret, the raw result is still sent to the Agent;
masking affects only the trace and display copy.
what it can reference. Values are bound server-side only when the Function runs;
they are not included in the model request. If a Function result contains an
exact secret value, the Agent model receives `{{NAME}}` in its place. The raw
runtime result and local side effects are not rewritten.
Leaving the default (*All secrets*) resolves the list at run time, so a secret
added next month is included automatically.
@@ -193,24 +304,30 @@ return {
</Tab>
<Tab value="Python">
```python title="loyalty-calculator.py"
import json
def calculate_loyalty(data):
purchase_history = data["purchaseHistory"]
account_age = data["accountAge"]
support_tickets = data["supportTickets"]
data = json.loads('<agent>')
purchase_history = data["purchaseHistory"]
account_age = data["accountAge"]
support_tickets = data["supportTickets"]
total_spent = sum(p["amount"] for p in purchase_history)
purchase_frequency = len(purchase_history) / (account_age / 365)
ticket_ratio = support_tickets["resolved"] / support_tickets["total"]
total_spent = sum(p["amount"] for p in purchase_history)
purchase_frequency = len(purchase_history) / (account_age / 365)
ticket_ratio = support_tickets["resolved"] / support_tickets["total"]
spend_score = min(total_spent / 1000 * 30, 30)
frequency_score = min(purchase_frequency * 20, 40)
support_score = ticket_ratio * 30
loyalty_score = round(spend_score + frequency_score + support_score)
tier = "Platinum" if loyalty_score >= 80 else "Gold" if loyalty_score >= 60 else "Silver"
spend_score = min(total_spent / 1000 * 30, 30)
frequency_score = min(purchase_frequency * 20, 40)
support_score = ticket_ratio * 30
loyalty_score = round(spend_score + frequency_score + support_score)
return {
"customer": data["name"],
"loyaltyScore": loyalty_score,
"loyaltyTier": tier,
}
tier = "Platinum" if loyalty_score >= 80 else "Gold" if loyalty_score >= 60 else "Silver"
print(json.dumps({ "customer": data["name"], "loyaltyScore": loyalty_score, "loyaltyTier": tier }))
if __name__ == "__main__":
__sim_result__ = calculate_loyalty(<agent>)
```
</Tab>
</Tabs>
@@ -233,7 +350,7 @@ const bytes = await sim.files.readBase64Chunk(file, { offset: 0, length: 1024 *
`sim.files.readText`, `readBase64`, and the `…Chunk` variants stream from execution storage under memory caps. `sim.values.read(ref)` and `sim.values.readArray(ref)` read large value and array references. Chunk `offset` and `length` are byte-based, so for exact Unicode parsing prefer smaller structured references. For large generated data, write the result to a file or table with `outputPath`, `outputSandboxPath`, or `outputTable` instead of returning the whole payload inline.
<Callout type="warn">
The lazy `sim.files` and `sim.values` helpers are available only in JavaScript functions without imports. JavaScript with imports, Python, and shell do not support them yet.
The lazy `sim.files` and `sim.values` helpers are available only in JavaScript functions without imports. JavaScript with imports, Python, and Shell do not support them yet.
</Callout>
## Best Practices
@@ -243,13 +360,13 @@ The lazy `sim.files` and `sim.values` helpers are available only in JavaScript f
- **Keep each function focused.** One transform per block is easier to read, test, and debug.
- **Handle errors.** Wrap risky code in `try`/`catch` and return a clear message, or let it throw to the error path.
- **Reference only what you need.** Pull a narrow field rather than a whole large object to keep values out of the request body.
- **Use stdout to debug.** `console.log()` and `print()` land in `<function.stdout>` and the run logs.
- **Use stdout to debug.** `console.log()`, `print()`, and ordinary shell output land in `<function.stdout>` and the run logs.
<FAQ items={[
{ question: "What languages does the Function block support?", answer: "JavaScript and Python. JavaScript is the default. Python requires the E2B feature, since Python always runs in a secure E2B sandbox." },
{ question: "When does code run locally vs. in a sandbox?", answer: "JavaScript without external imports runs in a local isolated sandbox for speed. JavaScript that uses import or require runs in E2B. Python always runs in the E2B sandbox, with or without imports." },
{ question: "What languages does the Function block support?", answer: "JavaScript, Python, and Shell. JavaScript is the default. Python and Shell appear when a remote sandbox provider is enabled." },
{ question: "When does code run locally vs. in a sandbox?", answer: "JavaScript without external imports runs in a local isolated sandbox for speed. JavaScript that uses import or require, Python, and Shell run in the configured remote sandbox." },
{ question: "How do I reference outputs from other blocks inside my code?", answer: "Use angle-bracket syntax directly, like <agent.content> or <api.data>, with no quotes around the tag — Sim replaces it with the real value before execution. For environment variables, use double curly braces: {{API_KEY}}." },
{ question: "What does the Function block return?", answer: "Two outputs: result (the return value of your code, read as <function.result>) and stdout (anything logged with console.log or print, read as <function.stdout>). Include a return statement in JavaScript, or print JSON in Python, to pass data downstream." },
{ question: "Can I make HTTP requests from a Function block?", answer: "Yes. fetch() is available in JavaScript with async/await; libraries like axios are only available when the block has a sandbox selected. In Python, use requests or httpx in the E2B sandbox." },
{ question: "What does the Function block return?", answer: "Two outputs: result and stdout. Use return in JavaScript, assign __sim_result__ in Python, or print an __SIM_RESULT__= marker in Shell to set result. Ordinary console, print, and command output goes to stdout." },
{ question: "Can I make HTTP requests from a Function block?", answer: "Yes. fetch() is available in JavaScript with async/await. In Python, use requests or httpx. In Shell, use curl or a CLI available on the selected sandbox." },
{ question: "Is there a timeout for Function block execution?", answer: "Yes, a configurable execution timeout. If your code exceeds it, the run is terminated and the block reports an error. Keep this in mind for external calls or heavy processing." },
]} />
@@ -216,7 +216,7 @@ Create PR runs in a sandbox image with the Pi CLI, Git, Node.js, and Bun baked i
The template requests **4 vCPU and 8 GB of RAM** (within the per-build maximum on every E2B plan). Sizing is fixed when the template is built — E2B has no per-sandbox override — so changing it means rebuilding the template, not restarting Sim. Sandboxes are billed per second against the resources they are allocated, not the ones they use. Both numbers live in `apps/sim/scripts/pi-sandbox-packages.ts` and are shared with the Daytona snapshot so the failover image cannot drift from the primary.
Sim sizes each Pi sandbox to **the execution's own remaining time**, so a run never holds a sandbox longer than the platform would let it run. The ceiling when there is no deadline to narrow to is the longest execution any plan permits (90 minutes); E2B rejects a create above the session length its plan allows, which is 1 hour on Hobby and 24 hours on Professional. `PI_SANDBOX_LIFETIME_MS` may lower that ceiling but has a **31-minute minimum**; lower values are raised to the minimum. A run that outlives the sandbox loses its work before the push, and an orphaned sandbox — one whose Sim process died mid-run — bills until the lifetime expires, which is why that lifetime tracks the deadline. Babysit Mode uses a second sequential sandbox after the creation sandbox has been destroyed; it is billed while polling checks and reviews. Daytona remains unchanged because its auto-stop setting is inactivity-based rather than an absolute lifetime.
Sim sizes each Pi sandbox to **the execution's own remaining time**, so a run never holds a sandbox longer than the platform would let it run. E2B also imposes a 24-hour continuous-sandbox ceiling on the Professional plan, so one Pi sandbox is capped there even when an Enterprise workflow policy is longer. `PI_SANDBOX_LIFETIME_MS` may lower that ceiling but has a **31-minute minimum**; lower values are raised to the minimum. A run that outlives the sandbox loses its work before the push, and an orphaned sandbox — one whose Sim process died mid-run — bills until the lifetime expires, which is why that lifetime tracks the deadline. Babysit Mode uses a second sequential sandbox after the creation sandbox has been destroyed; it is billed while polling checks and reviews. Daytona remains unchanged because its auto-stop setting is inactivity-based rather than an absolute lifetime.
2. **Bring your own model key.** Set the provider API key in the block's API Key field, or store it in **Settings → BYOK** when the provider supports workspace BYOK.
3. **Create a GitHub token** with permission to clone, push, and open a PR:
- *Fine-grained:* select the repo, then **Contents: Read and write** + **Pull requests: Read and write**.
@@ -282,6 +282,8 @@ The `version` field is part of the external API contract. Treat the reference as
For long-running workflows, async mode returns a job ID immediately so you don't need to hold the connection open. Add the `X-Execution-Mode: async` header to your request. The API returns HTTP 202 with a job ID and status URL. Poll the status URL until the job completes.
To stop an individual async request sooner than the workspace policy, also send `X-Execution-Timeout-Seconds` with an integer from `1` to `604800` (seven days). The effective limit is the smaller of this request value and the account's configured workflow execution timeout. This header is rejected unless `X-Execution-Mode` is `async`. It controls server-side workflow runtime; your HTTP client's own timeout remains separate.
<Tabs items={['Start Job', 'Check Status']}>
<Tab value="Start Job">
```bash
@@ -289,6 +291,7 @@ curl -X POST https://sim.ai/api/workflows/{workflow-id}/execute \
-H "Content-Type: application/json" \
-H "x-api-key: $SIM_API_KEY" \
-H "X-Execution-Mode: async" \
-H "X-Execution-Timeout-Seconds: 3600" \
-d '{ "input": "Process this large dataset" }'
```
@@ -358,9 +361,10 @@ Poll the `statusUrl` from the initial response until the status is `completed` o
| Plan | Sync Limit | Async Limit |
|------|-----------|-------------|
| **Community** | 5 minutes | 90 minutes |
| **Pro / Max / Team / Enterprise** | 50 minutes | 90 minutes |
| **Pro / Max / Team** | 50 minutes | 90 minutes |
| **Enterprise** | 50 minutes | 90 minutes by default; configurable up to 7 days |
If a job exceeds its time limit it is automatically marked as `failed`.
If a job exceeds its time limit it is automatically marked as `failed`. An async API request may use `X-Execution-Timeout-Seconds` to shorten its account policy, but never to extend it.
#### Job Retention
@@ -55,7 +55,7 @@ Here `throwError` fails, so the run leaves through its red error port to `handle
## How long a run can take
A synchronous run is capped at **5 minutes on the free plan** and **50 minutes on paid plans**; an async run gets **90 minutes**. A run that hits the ceiling fails with "Execution timed out." The usual causes are a loop over a large list or a slow agent or API call — split the work, or run it async. Self-hosted deployments can configure these limits.
A synchronous run is capped at **5 minutes on the free plan** and **50 minutes on paid plans**. Async runs get **90 minutes** by default; Enterprise administrators can configure their workflow policy up to **7 days**. Direct async API callers may shorten that policy for an individual request, but cannot extend it. A run that hits the ceiling fails with "Execution timed out." The usual causes are a loop over a large list or a slow agent or API call — split the work, or run it async. Self-hosted deployments can configure these limits.
## Watching a run
@@ -75,7 +75,7 @@ Each run starts fresh from the values defined in the panel. A change made during
Environment variables let you keep sensitive values like API keys and tokens out of your saved workflow configuration. Create them under **Settings → Secrets** by adding a key-value pair.
Reference them with double curly braces in any block field, including Agent system prompts and Function block code. The value is substituted before the block runs. When a saved secret is successfully substituted through `{{KEY}}`, exact occurrences of that value are masked in the live block-log display and stored execution trace without changing the runtime value. See [Execution log protection](/platform/credentials#execution-log-protection) for the exact scope and limitations.
Reference them with double curly braces in any block field, including Agent system prompts and Function block code. Most fields resolve the reference before the block runs. Function and Custom Tool code preserve `{{KEY}}` until their shared compiler runs at the execution boundary, where the value is bound separately from the source. This keeps its exact string contents out of saved and executed source code. When a saved secret is successfully resolved through `{{KEY}}`, exact occurrences of that value are masked in the live block-log display and stored execution trace without changing the runtime value. Exact secret values are also projected back to placeholders before content is sent to a model. See [Execution log protection](/platform/credentials#execution-log-protection) for the exact scope and limitations.
```
{{API_KEY}}