Waleed 7e975e7dae fix(confluence): index mirrored/included page content via rendered view format (#5746)
* fix(confluence): index mirrored/included page content via rendered view format

The KB connector fetched page bodies as body-format=storage, which only
carries unexpanded macro references (Include Page / Excerpt Include). Those
'mirrored' articles were stripped to empty content by htmlToPlainText and never
synced. Switch getDocument to body-format=view (supported on the v2 single-item
page/blogpost GET) so built-in include/excerpt macros render inline and the
included text is indexed.

* fix(confluence): invalidate existing doc hashes on representation change

The version-based contentHash meant already-synced mirrored documents (with
stale empty content) classified as 'unchanged' and never re-hydrated with the
new rendered view content. Embed a body-representation marker in the hash so a
representation change invalidates every previously-synced Confluence document,
forcing a one-time re-hydration that picks up the expanded include/excerpt text.

* feat(connectors): full resync re-hydrates rendered content (transclusions)

Version-based change detection can't see when a Confluence page's rendered view
changes because an *included* page was edited (the container's version doesn't
bump). Add a 'Full resync' path so that drift can be recovered:

- ConnectorMeta.rehydrateOnFullSync flag (set for Confluence)
- on fullSync, classifyExternalDoc promotes unchanged deferred docs to update and
  the hydration guard re-indexes unconditionally, so rendered content is refreshed
- fullSync threaded through the manual sync contract (query param), route, and hook
- 'Sync now' / 'Full resync' dropdown on the connector card

Incremental syncs stay hash-gated and cheap; only the deliberate full resync pays
the re-index cost.

* refactor(connectors): decouple rehydrate from fullSync deletion semantics

The transclusion refresh only needs re-hydration, but reusing the fullSync flag
also activated its deletion-cleanup semantics — which bypass three previously
unreachable safety guards (empty-listing wipe, listingCapped, and the >50%
mass-deletion threshold). Since fullSync had no caller before this PR, the new
'Full resync' button would have exposed all three to any KB editor.

Introduce a dedicated 'rehydrate' request that ONLY forces re-hydration + re-index
of already-synced docs. Listing and deletion reconciliation are identical to a
normal sync (all safety guards stay armed). fullSync's cleanup semantics remain
dormant and untouched.

* refactor(connectors): tidy rehydrate flag + gate Full resync to supported connectors

Cleanup/simplify pass over the connector changes:
- use shared booleanQueryFlagSchema for the rehydrate query param (typed boolean
  at the boundary instead of a hand-rolled 'true'/'false' string enum)
- move rehydrateOnFullSync onto the client-safe ConnectorMeta so the UI can gate on it
- only Confluence (rehydrateOnFullSync) shows the Sync now / Full resync dropdown;
  every other connector keeps its original one-click sync button (Full resync is a
  no-op for them, and this restores the pre-change one-click UX)
- wrap the sync trigger in a span so its tooltip still shows while disabled (cooldown)

* fix(connectors): forward rehydrate flag through the Trigger.dev worker

executeConnectorSyncJob (the production async sync path) destructured only
fullSync from the payload and forwarded only fullSync to executeSync, silently
dropping rehydrate. A manual Full resync would therefore never re-hydrate on the
default Trigger.dev path. Forward rehydrate too.

* fix(connectors): rehydrate forces a full listing so containers aren't omitted

A rehydrate request set forceRehydrate but left listing incremental. For a
connector that is both incremental and rehydrateOnFullSync, an unchanged
container page that transcludes a changed page would be omitted from the
incremental listing and never re-hydrated. Force a full (non-incremental) listing
on rehydrate so every document is seen; deletion-safety guards stay armed (unlike
fullSync). No-op for Confluence, which is already non-incremental.
2026-07-17 15:21:22 -07:00

Sim.ai Documentation Slack X

Ask DeepWiki Set Up with Cursor

Sim — Integrate, Context, Build, and Monitor AI agents

A workspace to build, deploy and manage AI agents and workflows.

Quickstart

Cloud-hosted: sim.ai

Open sim.ai

Self-hosted

npx simstudio

Open http://localhost:3000

Docker must be installed and running. Use -p, --port <port> to run Sim on a different port, or --no-pull to skip pulling the latest Docker images.

The Sim platform — chat on the left, the visual workflow builder on the right

Capabilities

  • Connect 1,000+ integrations and every major LLM
  • Add Slack, Notion, HubSpot, Salesforce, databases, and more
  • Build agents visually, conversationally, or with code
  • Ingest files, knowledge bases, and structured table data
  • Monitor runs, logs, schedules, and workflow activity

One workspace, every surface

Chat and workflows are just the start — tables, files, knowledge, and scheduled tasks all live in the same workspace.

Tables in Sim — structured data your agents can query

Tables — a database, built in

Files in Sim — documents for your team and every agent

Files — one store for your team and every agent

Knowledge bases in Sim — synced docs your agents can search

Knowledge — your agents' memory

Scheduled tasks in Sim — recurring agent runs on a calendar

Scheduled tasks — runs on your schedule

Self-hosting

Docker Compose

git clone https://github.com/simstudioai/sim.git && cd sim
docker compose -f docker-compose.prod.yml up -d

Open http://localhost:3000

Sim also supports local models via Ollama and vLLM. See the Docker self-hosting docs for setup details.

Manual Setup

Requirements: Bun, Node.js v20+, PostgreSQL 12+ with pgvector

  1. Clone and install:
git clone https://github.com/simstudioai/sim.git
cd sim
bun install
bun run prepare  # Set up pre-commit hooks
  1. Set up PostgreSQL with pgvector:
docker run --name simstudio-db -e POSTGRES_PASSWORD=your_password -e POSTGRES_DB=simstudio -p 5432:5432 -d pgvector/pgvector:pg17

Or install manually via the pgvector guide.

  1. Configure environment:
cp apps/sim/.env.example apps/sim/.env
# Create your secrets
perl -i -pe "s/your_encryption_key/$(openssl rand -hex 32)/" apps/sim/.env
perl -i -pe "s/your_internal_api_secret/$(openssl rand -hex 32)/" apps/sim/.env
perl -i -pe "s/your_api_encryption_key/$(openssl rand -hex 32)/" apps/sim/.env
# DB configs for migration
cp packages/db/.env.example packages/db/.env
# Edit both .env files to set DATABASE_URL="postgresql://postgres:your_password@localhost:5432/simstudio"
  1. Run migrations:
cd packages/db && bun run db:migrate
  1. Start development servers:
bun run dev:full  # Starts Next.js app and realtime socket server

Or run separately: bun run dev (Next.js) and cd apps/sim && bun run dev:sockets (realtime).

Chat API Keys

Chat is a Sim-managed service. To use Chat on a self-hosted instance:

  • Go to https://sim.ai → Settings → Chat keys and generate a Chat API key
  • Set COPILOT_API_KEY environment variable in your self-hosted apps/sim/.env file to that value

Environment Variables

See the environment variables reference for the full list, or apps/sim/.env.example for defaults.

Tech Stack

Next.js · Bun · PostgreSQL · Drizzle · Better Auth · Tailwind — and the rest of the stack

Contributing

We welcome contributions! Please see our Contributing Guide for details.

License

This project is licensed under the Apache License 2.0 - see the LICENSE file for details.

Built by the Sim team in San Francisco

Languages
TypeScript 77%
MDX 20.8%
JavaScript 1.9%
CSS 0.1%