mirror of
https://github.com/simstudioai/sim.git
synced 2026-09-24 15:45:35 +08:00
feat(connectors): add 7 knowledge base connectors (Google Forms, Typeform, Azure DevOps, YouTube, JSM, S3, Sentry) (#4880)
* feat(connectors): add 7 knowledge base connectors (Google Forms, Typeform, Azure DevOps, YouTube, JSM, S3, Sentry) * fix(connectors): tighten listingCapped semantics per review (WIQL cap, batch omissions, cap-vs-exhaustion) * fix(connectors): google-forms listingCapped must fire on slice regardless of hitLimit (404-null-filter gap) * fix(connectors): s3 streaming size cap for chunked responses without content-length * fix(connectors): ado byte-exact file content fetch, google-forms hash-poisoning on listing failure * fix(connectors): ado auth-failure deletion guard, jsm last-page slice flag, google-forms response cap in hash * fix(connectors): shared streaming size-cap reader for ado file hydration (promote from s3) * fix(knowledge): flag incomplete listings at engine level when pagination is truncated * fix(connectors): ado flags listing incomplete when a non-empty repo has no resolvable branch * fix(knowledge): engine truncation flag is an absolute deletion block (fullSync cannot override); s3 byte-exact size fallback; ado tsdoc accuracy * improvement(knowledge): extract shouldReconcileDeletions gate as tested pure function, tighten engine comments * test(connectors): mapTags coverage for the 7 new connectors * fix(connectors): ado probes past the wiql 20k cap before flagging; document custom-wiql full-listing behavior * fix(connectors): ado flags partial repo trees when items listing emits a continuation token * fix(connectors): ado discards foreign-phase cursors; google-forms scans all response pages for change detection * fix(connectors): audit fixes across new connectors - registry: register x connector (was dead code, never wired in) - google-docs/google-drive/google-forms: gate deletion reconciliation on Drive incompleteSearch; google-docs also now sets listingCapped on its maxDocs cap path - jsm: add read:jira-user scope so reporter resolves on requests - gong: only set listingCapped on genuine truncation, not exact-cap source exhaustion - gitlab: issues phase switched to keyset pagination (removes ~50k offset ceiling), matching the repo-tree phase - grain: parallelize recording + transcript fetch in getDocument - ashby: document updatedAt-based content-hash limitation for notes/feedback change detection - tests: mapTags coverage for x, granola, greenhouse, fathom, rootly
This commit is contained in:
@@ -463,6 +463,24 @@ const response = await fetchWithRetry(url, { ... }, VALIDATE_RETRY_OPTIONS)
|
||||
|
||||
If `ExternalDocument.sourceUrl` is set, the sync engine stores it on the document record. Always construct the full URL (not a relative path).
|
||||
|
||||
## Capped or Incomplete Listings — `syncContext.listingCapped` (REQUIRED)
|
||||
|
||||
If `listDocuments` can ever return **less than the full source set** on a non-incremental sync — a `maxItems`/`maxDocuments`-style cap, or a transient per-item error that drops a still-existing document from the listing — it MUST set `syncContext.listingCapped = true` when that happens.
|
||||
|
||||
The sync engine reconciles deletions by comparing the full listing against stored documents: anything not seen is **hard-deleted** (sync-engine.ts, gated on `!syncContext?.listingCapped`). A truncated listing without this flag deletes every real document beyond the cap. This was the single most common bug found when auditing connectors — do not omit it.
|
||||
|
||||
```typescript
|
||||
if (hitLimit && syncContext) {
|
||||
syncContext.listingCapped = true
|
||||
}
|
||||
```
|
||||
|
||||
Rules:
|
||||
- Set it when a user-configured cap truncates the listing while more documents exist
|
||||
- Set it when a thrown error caused a still-present document to be skipped during listing
|
||||
- Do NOT set it when the source is genuinely exhausted (deleted documents must still reconcile)
|
||||
- Do NOT set it for intentional scope filters (e.g. a date cutoff) — out-of-scope documents should be reconciled normally
|
||||
|
||||
## Sync Engine Behavior (Do Not Modify)
|
||||
|
||||
The sync engine (`lib/knowledge/connectors/sync-engine.ts`) is connector-agnostic. It:
|
||||
@@ -515,6 +533,7 @@ export const CONNECTOR_REGISTRY: ConnectorRegistry = {
|
||||
- `dependsOn` references selector field IDs (not `canonicalParamId`)
|
||||
- Dependency `canonicalParamId` values exist in `SELECTOR_CONTEXT_FIELDS`
|
||||
- [ ] `listDocuments` handles pagination with metadata-based content hashes
|
||||
- [ ] `syncContext.listingCapped = true` set whenever the listing is truncated (max-items cap or transient per-item error) — required to prevent the engine's deletion reconciliation from removing unseen documents
|
||||
- [ ] `contentDeferred: true` used if content requires per-doc API calls (file download, export, blocks fetch)
|
||||
- [ ] `contentHash` is metadata-based (not content-based) and identical between stub and `getDocument`
|
||||
- [ ] `sourceUrl` set on each ExternalDocument (full URL, not relative)
|
||||
|
||||
@@ -135,6 +135,13 @@ For each API endpoint the connector calls:
|
||||
- [ ] No off-by-one errors in pagination tracking
|
||||
- [ ] The connector does NOT hit known API pagination limits silently (e.g., HubSpot search 10k cap)
|
||||
|
||||
### Deletion-Reconciliation Safety (`listingCapped`) — CRITICAL
|
||||
The sync engine hard-deletes any stored document absent from a full listing. Audit every path where `listDocuments` can return less than the full source set:
|
||||
- [ ] `syncContext.listingCapped = true` is set when a `maxItems`-style cap truncates the listing while more documents exist
|
||||
- [ ] `listingCapped` is set when a transient per-item error drops a still-existing document from the listing
|
||||
- [ ] `listingCapped` is NOT set when the source is genuinely exhausted (deleted documents must reconcile) or for intentional scope filters (date cutoffs)
|
||||
This is the most common connector bug class — verify it explicitly against `sync-engine.ts`'s reconciliation gate.
|
||||
|
||||
### Pagination State Across Pages
|
||||
- [ ] `syncContext` is used to cache state across pages (user names, field maps, instance URLs, portal IDs, etc.)
|
||||
- [ ] Cached state in `syncContext` is correctly initialized on first page and reused on subsequent pages
|
||||
|
||||
@@ -510,6 +510,13 @@ figure[data-rehype-pretty-code-figure],
|
||||
max-width: 480px !important;
|
||||
}
|
||||
|
||||
/* Search dialog overlay + panel must cover the sticky navbar — both default to z-50,
|
||||
and the navbar wins the tie by DOM order, leaving it unblurred above the overlay */
|
||||
.bg-fd-overlay,
|
||||
[role="dialog"][data-state] {
|
||||
z-index: 60 !important;
|
||||
}
|
||||
|
||||
pre {
|
||||
font-size: 0.875rem;
|
||||
line-height: 1.7;
|
||||
|
||||
@@ -14,21 +14,23 @@ Connectors continuously sync documents from external services into your knowledg
|
||||
|
||||
<Image src="/static/connectors/connectors-sources.png" alt="Connect Source picker showing a searchable list of available connectors including Airtable, Asana, Confluence, Discord, Dropbox, Evernote, Fireflies, GitHub, and Gmail" width={800} height={500} />
|
||||
|
||||
Sim ships with 30 built-in connectors:
|
||||
Sim ships with 49 built-in connectors:
|
||||
|
||||
| Category | Connectors |
|
||||
|----------|-----------|
|
||||
| **Productivity** | Notion, Confluence, Asana, Linear, Jira, Google Calendar, Google Sheets |
|
||||
| **Cloud Storage** | Google Drive, Dropbox, OneDrive, SharePoint |
|
||||
| **Documents** | Google Docs, WordPress, Webflow |
|
||||
| **Development** | GitHub |
|
||||
| **Communication** | Slack, Discord, Microsoft Teams, Reddit |
|
||||
| **Productivity** | Notion, Confluence, Asana, Linear, Jira, Jira Service Management, Monday, Google Calendar, Google Sheets, Google Forms, Typeform |
|
||||
| **Cloud Storage** | Google Drive, Dropbox, OneDrive, SharePoint, Amazon S3 |
|
||||
| **Documents** | Google Docs, WordPress, Webflow, DocuSign |
|
||||
| **Development** | GitHub, GitLab, Azure DevOps, Sentry |
|
||||
| **Communication** | Slack, Discord, Microsoft Teams, Reddit, YouTube |
|
||||
| **Email** | Gmail, Outlook |
|
||||
| **CRM** | HubSpot, Salesforce |
|
||||
| **Support** | Intercom, ServiceNow, Zendesk |
|
||||
| **Incident Management** | incident.io, Rootly |
|
||||
| **Data** | Airtable |
|
||||
| **Note-taking** | Evernote, Obsidian |
|
||||
| **Meetings** | Fireflies |
|
||||
| **Meetings** | Zoom, Gong, Grain, Granola, Fathom, Fireflies |
|
||||
| **Recruiting** | Greenhouse, Ashby |
|
||||
|
||||
## Adding a Connector
|
||||
|
||||
@@ -41,13 +43,18 @@ From inside a knowledge base, click **+ New connector** in the top right to open
|
||||
|
||||
Most connectors use **OAuth** — select an existing credential from the dropdown or click **Connect new account** to authorize through the service. Tokens are refreshed automatically.
|
||||
|
||||
A few connectors use **API keys** instead:
|
||||
Other connectors use **API keys** or **personal access tokens** instead. The setup modal tells you which credential each connector expects — for example:
|
||||
|
||||
| Connector | Where to get the key |
|
||||
|-----------|---------------------|
|
||||
| **Evernote** | Developer Token (starts with `S=`) from your Evernote account settings |
|
||||
| **Obsidian** | Install the [Local REST API](https://github.com/coddingtonbear/obsidian-local-rest-api) plugin, then copy the key from its settings |
|
||||
| **Fireflies** | Generate from the Integrations page in your Fireflies account |
|
||||
| **Typeform** | Personal access token from your Typeform account settings |
|
||||
| **Azure DevOps** | Personal access token with Wiki (Read), Work Items (Read), and Code (Read) scopes |
|
||||
| **YouTube** | YouTube Data API key from the Google Cloud Console |
|
||||
| **Amazon S3** | Secret Access Key (the Access Key ID, region, and bucket are entered as config fields) |
|
||||
| **Sentry** | Auth token with `project:read` and `event:read` scopes |
|
||||
|
||||
<Callout type="info">
|
||||
If you rotate an API key in the external service, update it in Sim as well — OAuth tokens refresh automatically, but API keys do not.
|
||||
@@ -63,6 +70,10 @@ Each connector has source-specific fields that control what gets synced. Example
|
||||
- **Notion** — sync an entire workspace, a specific database, or a single page tree
|
||||
- **GitHub** — specify a repository, branch, and optional file extension filter
|
||||
- **Confluence** — enter your Atlassian domain and optionally filter by space key or content type
|
||||
- **Azure DevOps** — choose what to sync (wiki pages, work items, repository files, or all), with optional work item type/state filters, a custom WIQL query, and repository/branch/path filters
|
||||
- **Amazon S3** — point at a bucket with an optional key prefix and a customizable file extension allowlist; S3-compatible stores (Cloudflare R2, MinIO) are supported via a custom endpoint
|
||||
- **YouTube** — sync a channel (by `@handle` or ID) or playlist, with an optional published-after date filter and the option to exclude Shorts
|
||||
- **Sentry** — filter issues by search query (e.g. `is:unresolved`), environment, and time window; self-hosted Sentry is supported via a custom host
|
||||
- **Obsidian** — provide your vault URL (`https://127.0.0.1:27124` by default) and optionally restrict to a folder path
|
||||
- **Fireflies** — optionally filter by host email or cap the number of transcripts synced
|
||||
|
||||
@@ -188,5 +199,5 @@ You can add as many connectors as you need to a single knowledge base. Each mana
|
||||
{ question: "What happens when I delete a connector?", answer: "The connector is removed and future syncs stop. You're given the option to also delete all documents that were synced by that connector. If you don't check that option, they stay in the knowledge base as-is." },
|
||||
{ question: "What does the Disabled status mean?", answer: "After 10 consecutive full-sync failures, the connector is automatically disabled to stop retrying. Reconnect the OAuth account or click Resume to re-enable it." },
|
||||
{ question: "Do metadata tags count against a limit?", answer: "Yes. Tag slots are shared across all documents in a knowledge base — 17 slots total. Multiple connectors draw from the same pool, so plan accordingly if several connectors each auto-populate tags." },
|
||||
{ question: "Do I need to re-authenticate connectors?", answer: "OAuth connectors refresh tokens automatically. API key connectors (Evernote, Obsidian, Fireflies) need manual updates if you rotate the key in the external service." },
|
||||
{ question: "Do I need to re-authenticate connectors?", answer: "OAuth connectors refresh tokens automatically. API key and personal access token connectors need manual updates if you rotate the credential in the external service." },
|
||||
]} />
|
||||
|
||||
@@ -49,7 +49,7 @@ For knowledge bases that should stay current automatically, connectors sync cont
|
||||
|
||||
Connectors are configured through the knowledge base settings, not through Mothership chat. Once connected, all synced content is immediately searchable by Mothership and by any Agent block with the knowledge base attached.
|
||||
|
||||
Sim ships with 30 built-in connectors, including Notion, Google Drive, Slack, GitHub, Confluence, HubSpot, Salesforce, Gmail, and more.
|
||||
Sim ships with 49 built-in connectors, including Notion, Google Drive, Slack, GitHub, Confluence, HubSpot, Salesforce, Gmail, and more.
|
||||
|
||||
Examples of what you can sync:
|
||||
|
||||
|
||||
@@ -298,7 +298,32 @@ function renderFeedbackValue(value: unknown): string {
|
||||
|
||||
/**
|
||||
* Stable, metadata-based content hash for a candidate document. Identical between the
|
||||
* listing stub and the fully-fetched document so unchanged candidates are skipped.
|
||||
* listing stub and the fully-fetched document so unchanged candidates are skipped,
|
||||
* which keeps the `getDocument` re-hydration (notes + feedback fetches) cheap: the
|
||||
* sync engine only re-hydrates a deferred stub when this hash differs from the stored
|
||||
* document's hash (see `lib/knowledge/connectors/sync-engine.ts`).
|
||||
*
|
||||
* Known limitation — notes/feedback freshness depends on `candidate.updatedAt`.
|
||||
* Candidate notes (`candidate.listNotes`) and interview feedback
|
||||
* (`applicationFeedback.list`) are separate Ashby objects, not candidate fields. This
|
||||
* hash is derived solely from the candidate's own `updatedAt`, so a new note or newly
|
||||
* submitted feedback is only re-synced if Ashby advances `candidate.updatedAt` as a
|
||||
* side effect of that write.
|
||||
*
|
||||
* As of this writing Ashby's public API docs do not specify what counts as a
|
||||
* "modification" for `candidate.updatedAt` or for `candidate.list` syncToken
|
||||
* incremental sync, and no third-party ATS-integration vendor (Merge, Nango, Knit)
|
||||
* documents it either — so this behavior is unverified. If Ashby does NOT touch
|
||||
* `candidate.updatedAt` on note/feedback writes, those additions will not be picked up
|
||||
* until some other candidate field changes; a forced full sync re-hydrates everything
|
||||
* regardless. No cheaper listing-time signal exists to fold into this hash: the
|
||||
* `candidate.list` object exposes no note/feedback count, and syncToken carries the
|
||||
* same unspecified change semantics as `updatedAt`.
|
||||
*
|
||||
* Refs:
|
||||
* - https://developers.ashbyhq.com/reference/candidatelist
|
||||
* - https://developers.ashbyhq.com/reference/candidatecreatenote
|
||||
* - https://developers.ashbyhq.com/docs/pagination-and-incremental-sync
|
||||
*/
|
||||
function buildContentHash(id: string, updatedAt: string | null): string {
|
||||
return `ashby:${id}:${updatedAt ?? ''}`
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1 @@
|
||||
export { azureDevopsConnector } from '@/connectors/azure-devops/azure-devops'
|
||||
@@ -470,15 +470,18 @@ async function fetchProject(
|
||||
}
|
||||
|
||||
/**
|
||||
* Encodes the listing cursor. The cursor packs the resource phase (wiki ➜ issues)
|
||||
* and the issues page number so a single sync walks wikis first, then paginates
|
||||
* issues via the X-Next-Page header.
|
||||
* Encodes the listing cursor. The cursor packs the resource phase (repo ➜ wiki ➜
|
||||
* issues) and a per-phase continuation token so a single sync walks the phases in
|
||||
* order. The repository-tree and issues phases both use GitLab keyset pagination
|
||||
* and store the full `rel="next"` URL from the Link header to fetch verbatim.
|
||||
*/
|
||||
interface CursorState {
|
||||
phase: SyncPhase
|
||||
issuePage: number
|
||||
/** Full `rel="next"` URL for the repository-tree keyset page to fetch next. */
|
||||
fileNextUrl?: string
|
||||
/** Full `rel="next"` URL for the issues keyset page to fetch next. */
|
||||
issueNextUrl?: string
|
||||
}
|
||||
|
||||
function encodeCursor(state: CursorState): string {
|
||||
@@ -492,6 +495,7 @@ function decodeCursor(cursor: string | undefined, initialPhase: SyncPhase): Curs
|
||||
phase: SyncPhase
|
||||
issuePage: number
|
||||
fileNextUrl: string
|
||||
issueNextUrl: string
|
||||
}>
|
||||
const phase: SyncPhase =
|
||||
parsed.phase === 'repo' || parsed.phase === 'issues' || parsed.phase === 'wiki'
|
||||
@@ -501,6 +505,7 @@ function decodeCursor(cursor: string | undefined, initialPhase: SyncPhase): Curs
|
||||
phase,
|
||||
issuePage: Number(parsed.issuePage) > 0 ? Number(parsed.issuePage) : 1,
|
||||
fileNextUrl: typeof parsed.fileNextUrl === 'string' ? parsed.fileNextUrl : undefined,
|
||||
issueNextUrl: typeof parsed.issueNextUrl === 'string' ? parsed.issueNextUrl : undefined,
|
||||
}
|
||||
} catch {
|
||||
return { phase: initialPhase, issuePage: 1 }
|
||||
@@ -859,9 +864,9 @@ export const gitlabConnector: ConnectorConfig = {
|
||||
if (state.phase === 'issues') {
|
||||
const params = new URLSearchParams({
|
||||
per_page: String(PAGE_SIZE),
|
||||
page: String(state.issuePage),
|
||||
order_by: 'updated_at',
|
||||
sort: 'desc',
|
||||
pagination: 'keyset',
|
||||
})
|
||||
if (lastSyncAt) params.set('updated_after', lastSyncAt.toISOString())
|
||||
const issueState =
|
||||
@@ -874,11 +879,15 @@ export const gitlabConnector: ConnectorConfig = {
|
||||
typeof sourceConfig.issueMilestone === 'string' ? sourceConfig.issueMilestone.trim() : ''
|
||||
if (issueMilestone) params.set('milestone', issueMilestone)
|
||||
|
||||
const url = `${apiBase}/projects/${encodedProject}/issues?${params.toString()}`
|
||||
if (state.issueNextUrl && !isSameOrigin(state.issueNextUrl, apiBase)) {
|
||||
throw new Error('GitLab pagination cursor points to an unexpected host')
|
||||
}
|
||||
const url =
|
||||
state.issueNextUrl ?? `${apiBase}/projects/${encodedProject}/issues?${params.toString()}`
|
||||
logger.info('Listing GitLab issues', {
|
||||
host,
|
||||
project: encodedProject,
|
||||
page: state.issuePage,
|
||||
continued: Boolean(state.issueNextUrl),
|
||||
incremental: Boolean(lastSyncAt),
|
||||
})
|
||||
|
||||
@@ -909,18 +918,18 @@ export const gitlabConnector: ConnectorConfig = {
|
||||
maxItems,
|
||||
syncContext
|
||||
)
|
||||
if (hitLimit) return { documents: capped, hasMore: false }
|
||||
|
||||
const nextPageHeader = response.headers.get('x-next-page')?.trim()
|
||||
const nextPage = nextPageHeader ? Number(nextPageHeader) : 0
|
||||
const hasMorePages = !hitLimit && Number.isFinite(nextPage) && nextPage > 0
|
||||
|
||||
return {
|
||||
documents: capped,
|
||||
nextCursor: hasMorePages
|
||||
? encodeCursor({ phase: 'issues', issuePage: nextPage })
|
||||
: undefined,
|
||||
hasMore: hasMorePages,
|
||||
const nextLink = parseNextLink(response.headers.get('link'))
|
||||
if (nextLink) {
|
||||
return {
|
||||
documents: capped,
|
||||
nextCursor: encodeCursor({ phase: 'issues', issuePage: 1, issueNextUrl: nextLink }),
|
||||
hasMore: true,
|
||||
}
|
||||
}
|
||||
|
||||
return { documents: capped, hasMore: false }
|
||||
}
|
||||
|
||||
return { documents: [], hasMore: false }
|
||||
|
||||
@@ -417,20 +417,32 @@ export const gongConnector: ConnectorConfig = {
|
||||
|
||||
const prevFetched = (syncContext?.totalDocsFetched as number) ?? 0
|
||||
let documents = allDocuments
|
||||
let capDroppedDocs = false
|
||||
if (maxCalls > 0) {
|
||||
const remaining = Math.max(0, maxCalls - prevFetched)
|
||||
if (allDocuments.length > remaining) {
|
||||
documents = allDocuments.slice(0, remaining)
|
||||
capDroppedDocs = true
|
||||
}
|
||||
}
|
||||
|
||||
const totalFetched = prevFetched + documents.length
|
||||
if (syncContext) syncContext.totalDocsFetched = totalFetched
|
||||
const hitLimit = maxCalls > 0 && totalFetched >= maxCalls
|
||||
if (hitLimit && syncContext) syncContext.listingCapped = true
|
||||
|
||||
const hasMore = !hitLimit && Boolean(nextPageCursor)
|
||||
|
||||
/**
|
||||
* Only flag the listing as capped when the `maxCalls` limit actually
|
||||
* truncated calls that still exist in the source — either by dropping calls
|
||||
* from the current page or by stopping while another page remains. Reaching
|
||||
* the limit exactly at source exhaustion (no dropped calls, no further
|
||||
* cursor) yields a complete listing, so deletion reconciliation must still
|
||||
* run for calls removed in Gong.
|
||||
*/
|
||||
if (syncContext && (capDroppedDocs || (hitLimit && Boolean(nextPageCursor)))) {
|
||||
syncContext.listingCapped = true
|
||||
}
|
||||
|
||||
return {
|
||||
documents,
|
||||
nextCursor: hasMore ? nextPageCursor : undefined,
|
||||
|
||||
@@ -235,13 +235,23 @@ export const googleDocsConnector: ConnectorConfig = {
|
||||
const data = await response.json()
|
||||
const files = (data.files || []) as DriveFile[]
|
||||
|
||||
/**
|
||||
* Drive sets `incompleteSearch` when it could not search every corpus (it
|
||||
* arises with the `allDrives` scope enabled by `includeItemsFromAllDrives`).
|
||||
* A partial listing drops still-existing docs, so reconciliation must be
|
||||
* suppressed to avoid hard-deleting valid documents.
|
||||
*/
|
||||
const incompleteSearch = data.incompleteSearch === true
|
||||
|
||||
const maxDocs = sourceConfig.maxDocs ? Number(sourceConfig.maxDocs) : 0
|
||||
const previouslyFetched = (syncContext?.totalDocsFetched as number) ?? 0
|
||||
|
||||
let documents = files.map(fileToStub)
|
||||
let slicedSome = false
|
||||
if (maxDocs > 0) {
|
||||
const remaining = maxDocs - previouslyFetched
|
||||
if (documents.length > remaining) {
|
||||
slicedSome = true
|
||||
documents = documents.slice(0, remaining)
|
||||
}
|
||||
}
|
||||
@@ -252,6 +262,19 @@ export const googleDocsConnector: ConnectorConfig = {
|
||||
|
||||
const nextPageToken = data.nextPageToken as string | undefined
|
||||
|
||||
/**
|
||||
* Mark the listing as incomplete so the sync engine skips deletion
|
||||
* reconciliation when this page does not represent the full source set:
|
||||
* - `slicedSome`: the page held more docs than the `maxDocs` cap allowed.
|
||||
* - `hitLimit` with a next page: the cap was reached while more pages remain.
|
||||
* - `incompleteSearch`: Drive could not search every corpus, so the page is
|
||||
* partial and may omit still-existing docs.
|
||||
* Reconciliation against any of these would hard-delete valid documents.
|
||||
*/
|
||||
if (syncContext && (slicedSome || (hitLimit && Boolean(nextPageToken)) || incompleteSearch)) {
|
||||
syncContext.listingCapped = true
|
||||
}
|
||||
|
||||
return {
|
||||
documents,
|
||||
nextCursor: hitLimit ? undefined : nextPageToken,
|
||||
|
||||
@@ -268,6 +268,14 @@ export const googleDriveConnector: ConnectorConfig = {
|
||||
const data = await response.json()
|
||||
const files = (data.files || []) as DriveFile[]
|
||||
|
||||
/**
|
||||
* Drive sets `incompleteSearch` when it could not search every corpus (it
|
||||
* arises with the `allDrives` scope enabled by `includeItemsFromAllDrives`).
|
||||
* A partial listing drops still-existing files, so reconciliation must be
|
||||
* suppressed to avoid hard-deleting valid documents.
|
||||
*/
|
||||
const incompleteSearch = data.incompleteSearch === true
|
||||
|
||||
const documents = files
|
||||
.filter((f) => isGoogleWorkspaceFile(f.mimeType) || isSupportedTextFile(f.mimeType))
|
||||
.map(fileToStub)
|
||||
@@ -275,7 +283,7 @@ export const googleDriveConnector: ConnectorConfig = {
|
||||
const totalFetched = previouslyFetched + documents.length
|
||||
if (syncContext) syncContext.totalDocsFetched = totalFetched
|
||||
const hitLimit = maxFiles > 0 && totalFetched >= maxFiles
|
||||
if (hitLimit && syncContext) syncContext.listingCapped = true
|
||||
if (syncContext && (hitLimit || incompleteSearch)) syncContext.listingCapped = true
|
||||
|
||||
const nextPageToken = data.nextPageToken as string | undefined
|
||||
|
||||
|
||||
@@ -0,0 +1,819 @@
|
||||
import { createLogger } from '@sim/logger'
|
||||
import { getErrorMessage, toError } from '@sim/utils/errors'
|
||||
import { GoogleFormsIcon } from '@/components/icons'
|
||||
import { fetchWithRetry, VALIDATE_RETRY_OPTIONS } from '@/lib/knowledge/documents/utils'
|
||||
import type { ConnectorConfig, ExternalDocument, ExternalDocumentList } from '@/connectors/types'
|
||||
import { joinTagArray, parseTagDate } from '@/connectors/utils'
|
||||
|
||||
const logger = createLogger('GoogleFormsConnector')
|
||||
|
||||
const DRIVE_API_BASE = 'https://www.googleapis.com/drive/v3'
|
||||
const FORMS_API_BASE = 'https://forms.googleapis.com/v1'
|
||||
const FORM_MIME_TYPE = 'application/vnd.google-apps.form'
|
||||
const FOLDER_MIME_TYPE = 'application/vnd.google-apps.folder'
|
||||
|
||||
/**
|
||||
* Hard cap on the number of responses appended to a single form document.
|
||||
* Keeps individual documents within a reasonable size for embedding/indexing.
|
||||
*/
|
||||
const MAX_RESPONSES_PER_FORM = 500
|
||||
|
||||
/**
|
||||
* Drive API page size when listing forms. The Drive API caps pageSize at 100.
|
||||
*/
|
||||
const DRIVE_PAGE_SIZE = 100
|
||||
|
||||
/**
|
||||
* Maximum responses returned per Forms API page (API caps and defaults to 5000).
|
||||
*/
|
||||
const RESPONSES_PAGE_SIZE = 5000
|
||||
|
||||
/**
|
||||
* Number of forms whose change indicators are fetched concurrently during
|
||||
* listing. Keeps the Forms API call volume bounded while still parallelizing.
|
||||
*/
|
||||
const LIST_CONCURRENCY = 4
|
||||
|
||||
/**
|
||||
* Content scope for a form document. `both` indexes the form's questions and its
|
||||
* submitted responses; `structure` indexes only the questions (no response reads,
|
||||
* so the responses scope is never exercised for that connector instance).
|
||||
*/
|
||||
type ContentScope = 'both' | 'structure'
|
||||
|
||||
/**
|
||||
* Resolves the content scope from sourceConfig, defaulting to `both`.
|
||||
*/
|
||||
function resolveContentScope(value: unknown): ContentScope {
|
||||
return value === 'structure' ? 'structure' : 'both'
|
||||
}
|
||||
|
||||
/**
|
||||
* Represents a Google Drive file entry for a form, returned by the Drive API.
|
||||
*/
|
||||
interface DriveFormFile {
|
||||
id: string
|
||||
name: string
|
||||
mimeType: string
|
||||
modifiedTime?: string
|
||||
createdTime?: string
|
||||
webViewLink?: string
|
||||
owners?: { displayName?: string; emailAddress?: string }[]
|
||||
trashed?: boolean
|
||||
}
|
||||
|
||||
/**
|
||||
* A single answer entry inside a response answer container.
|
||||
*/
|
||||
interface FormTextAnswer {
|
||||
value?: string
|
||||
}
|
||||
|
||||
/**
|
||||
* A single question's answers within a form response. The Forms API keys the
|
||||
* `answers` map by questionId and stores text values under
|
||||
* `textAnswers.answers[].value`.
|
||||
*/
|
||||
interface FormAnswer {
|
||||
questionId?: string
|
||||
textAnswers?: { answers?: FormTextAnswer[] }
|
||||
}
|
||||
|
||||
/**
|
||||
* A single submitted response to a form.
|
||||
*/
|
||||
interface FormResponse {
|
||||
responseId?: string
|
||||
createTime?: string
|
||||
lastSubmittedTime?: string
|
||||
respondentEmail?: string
|
||||
answers?: Record<string, FormAnswer>
|
||||
}
|
||||
|
||||
/**
|
||||
* Paginated response list from the Forms API.
|
||||
*/
|
||||
interface FormResponseList {
|
||||
responses?: FormResponse[]
|
||||
nextPageToken?: string
|
||||
}
|
||||
|
||||
/**
|
||||
* A question item within a form's structure.
|
||||
*/
|
||||
interface FormQuestionItem {
|
||||
question?: {
|
||||
questionId?: string
|
||||
required?: boolean
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* A single structural item within a form (question, section, image, etc.).
|
||||
*/
|
||||
interface FormItem {
|
||||
itemId?: string
|
||||
title?: string
|
||||
description?: string
|
||||
questionItem?: FormQuestionItem
|
||||
}
|
||||
|
||||
/**
|
||||
* The form structure returned by the Forms API `forms.get` endpoint.
|
||||
*/
|
||||
interface FormStructure {
|
||||
formId?: string
|
||||
info?: {
|
||||
title?: string
|
||||
description?: string
|
||||
documentTitle?: string
|
||||
}
|
||||
items?: FormItem[]
|
||||
revisionId?: string
|
||||
responderUri?: string
|
||||
}
|
||||
|
||||
/**
|
||||
* Lightweight metadata captured during listing, sufficient to build a stub
|
||||
* and detect changes without downloading the full form content.
|
||||
*/
|
||||
interface FormStubInput {
|
||||
file: DriveFormFile
|
||||
formTitle?: string
|
||||
revisionId?: string
|
||||
latestResponseTime?: string
|
||||
contentScope: ContentScope
|
||||
responseCap: number
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves the effective per-form response cap applied when rendering content:
|
||||
* the user-configured `maxResponsesPerForm` clamped to the hard
|
||||
* `MAX_RESPONSES_PER_FORM` ceiling. Part of the content hash so changing the
|
||||
* cap re-syncs every form (the rendered content depends on it).
|
||||
*/
|
||||
function resolveResponseCap(sourceConfig: Record<string, unknown>): number {
|
||||
const configured = parsePositiveInt(sourceConfig.maxResponsesPerForm)
|
||||
return configured > 0 ? Math.min(configured, MAX_RESPONSES_PER_FORM) : MAX_RESPONSES_PER_FORM
|
||||
}
|
||||
|
||||
/**
|
||||
* Parses an optional positive-integer config value, returning 0 when unset/invalid.
|
||||
*/
|
||||
function parsePositiveInt(value: unknown): number {
|
||||
if (value == null || value === '') return 0
|
||||
const num = Number(value)
|
||||
return Number.isNaN(num) || num <= 0 ? 0 : Math.floor(num)
|
||||
}
|
||||
|
||||
/**
|
||||
* Maps a small array over an async worker with a bounded concurrency, preserving
|
||||
* input order in the returned results.
|
||||
*/
|
||||
async function mapWithConcurrency<T, R>(
|
||||
items: T[],
|
||||
limit: number,
|
||||
worker: (item: T, index: number) => Promise<R>
|
||||
): Promise<R[]> {
|
||||
const results = new Array<R>(items.length)
|
||||
let next = 0
|
||||
|
||||
async function run(): Promise<void> {
|
||||
while (next < items.length) {
|
||||
const current = next++
|
||||
results[current] = await worker(items[current], current)
|
||||
}
|
||||
}
|
||||
|
||||
const runners = Array.from({ length: Math.min(limit, items.length) }, run)
|
||||
await Promise.all(runners)
|
||||
return results
|
||||
}
|
||||
|
||||
/**
|
||||
* Fetches the form structure via the Forms API. Returns null on 404 (form
|
||||
* deleted or inaccessible).
|
||||
*/
|
||||
async function fetchFormStructure(
|
||||
accessToken: string,
|
||||
formId: string
|
||||
): Promise<FormStructure | null> {
|
||||
const url = `${FORMS_API_BASE}/forms/${encodeURIComponent(formId)}`
|
||||
const response = await fetchWithRetry(url, {
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
Accept: 'application/json',
|
||||
},
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
if (response.status === 404) return null
|
||||
throw new Error(`Failed to fetch form structure ${formId}: ${response.status}`)
|
||||
}
|
||||
|
||||
return (await response.json()) as FormStructure
|
||||
}
|
||||
|
||||
/**
|
||||
* Result of fetching a form's responses: the collected responses (capped at
|
||||
* `MAX_RESPONSES_PER_FORM` for rendering) plus the greatest submission timestamp
|
||||
* across ALL response pages.
|
||||
*
|
||||
* `latestSubmittedTime` is tracked separately from the capped `responses` so the
|
||||
* content hash computed in getDocument stays identical to the one computed during
|
||||
* listing, which scans the same full set via `fetchLatestResponseTime`. If it
|
||||
* were derived from the capped slice alone, a form with more than
|
||||
* `MAX_RESPONSES_PER_FORM` responses could hash differently between the two paths
|
||||
* and re-sync on every run.
|
||||
*/
|
||||
interface FetchedResponses {
|
||||
responses: FormResponse[]
|
||||
latestSubmittedTime?: string
|
||||
}
|
||||
|
||||
/**
|
||||
* Fetches form responses, retaining up to `MAX_RESPONSES_PER_FORM` for rendering.
|
||||
* Every page is scanned for the latest submission timestamp even after the
|
||||
* render cap is reached — the Forms API does not guarantee response order, so
|
||||
* the newest submission may sit on any page. `fetchLatestResponseTime` scans
|
||||
* the same full set during listing, keeping the content hash identical across
|
||||
* the listing and getDocument paths regardless of form size.
|
||||
*/
|
||||
async function fetchFormResponses(accessToken: string, formId: string): Promise<FetchedResponses> {
|
||||
const collected: FormResponse[] = []
|
||||
let latest = ''
|
||||
let pageToken: string | undefined
|
||||
|
||||
do {
|
||||
const url = new URL(`${FORMS_API_BASE}/forms/${encodeURIComponent(formId)}/responses`)
|
||||
url.searchParams.set('pageSize', String(RESPONSES_PAGE_SIZE))
|
||||
if (pageToken) url.searchParams.set('pageToken', pageToken)
|
||||
|
||||
const response = await fetchWithRetry(url.toString(), {
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
Accept: 'application/json',
|
||||
},
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
throw new Error(`Failed to list responses for form ${formId}: ${response.status}`)
|
||||
}
|
||||
|
||||
const data = (await response.json()) as FormResponseList
|
||||
const responses = data.responses ?? []
|
||||
|
||||
const pageLatest = latestResponseTime(responses)
|
||||
if (pageLatest && pageLatest > latest) latest = pageLatest
|
||||
|
||||
for (const r of responses) {
|
||||
if (collected.length >= MAX_RESPONSES_PER_FORM) break
|
||||
collected.push(r)
|
||||
}
|
||||
|
||||
pageToken = data.nextPageToken
|
||||
} while (pageToken)
|
||||
|
||||
return { responses: collected, latestSubmittedTime: latest || undefined }
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads the latest response submission time for change detection without
|
||||
* retaining responses. Scans every page — the Forms API does not guarantee
|
||||
* response order, so the newest submission may sit on any page. Returns the
|
||||
* greatest `lastSubmittedTime` (falling back to `createTime`), or undefined
|
||||
* when there are none. Throws on a failed read so the caller skips the form
|
||||
* for this run instead of computing a hash from incomplete data — a swallowed
|
||||
* error would poison the stub's content hash and re-process the form on every
|
||||
* sync, while throwing routes into the per-form catch that sets
|
||||
* `skippedOnError` → `listingCapped`.
|
||||
*/
|
||||
async function fetchLatestResponseTime(
|
||||
accessToken: string,
|
||||
formId: string
|
||||
): Promise<string | undefined> {
|
||||
let latest = ''
|
||||
let pageToken: string | undefined
|
||||
|
||||
do {
|
||||
const url = new URL(`${FORMS_API_BASE}/forms/${encodeURIComponent(formId)}/responses`)
|
||||
url.searchParams.set('pageSize', String(RESPONSES_PAGE_SIZE))
|
||||
if (pageToken) url.searchParams.set('pageToken', pageToken)
|
||||
|
||||
const response = await fetchWithRetry(url.toString(), {
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
Accept: 'application/json',
|
||||
},
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
throw new Error(
|
||||
`Failed to read responses for change detection on form ${formId}: ${response.status}`
|
||||
)
|
||||
}
|
||||
|
||||
const data = (await response.json()) as FormResponseList
|
||||
const pageLatest = latestResponseTime(data.responses ?? [])
|
||||
if (pageLatest && pageLatest > latest) latest = pageLatest
|
||||
pageToken = data.nextPageToken
|
||||
} while (pageToken)
|
||||
|
||||
return latest || undefined
|
||||
}
|
||||
|
||||
/**
|
||||
* Returns the greatest submission timestamp across the given responses, or
|
||||
* undefined when the list is empty.
|
||||
*/
|
||||
function latestResponseTime(responses: FormResponse[]): string | undefined {
|
||||
let latest = ''
|
||||
for (const r of responses) {
|
||||
const t = r.lastSubmittedTime || r.createTime || ''
|
||||
if (t > latest) latest = t
|
||||
}
|
||||
return latest || undefined
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the content hash for a form. The hash must change when either the form
|
||||
* structure (revisionId) or, when responses are indexed, the set of responses
|
||||
* (latest submission time) changes. Drive `modifiedTime` alone is insufficient
|
||||
* because new response submissions do not update the form's Drive modifiedTime.
|
||||
* The content scope is part of the hash so that toggling response indexing
|
||||
* forces a re-sync of every document.
|
||||
*/
|
||||
function formContentHash(input: FormStubInput): string {
|
||||
const responsePart =
|
||||
input.contentScope === 'both'
|
||||
? `${input.latestResponseTime ?? ''}:${input.responseCap}`
|
||||
: 'none'
|
||||
return `gforms:${input.file.id}:${input.contentScope}:${input.revisionId ?? ''}:${responsePart}`
|
||||
}
|
||||
|
||||
/**
|
||||
* Creates a lightweight stub from a form's Drive file and change indicators.
|
||||
* Content is deferred and only fetched via getDocument for new/changed forms.
|
||||
*/
|
||||
function formToStub(input: FormStubInput): ExternalDocument {
|
||||
const { file } = input
|
||||
const title = input.formTitle?.trim() || file.name || 'Untitled Form'
|
||||
return {
|
||||
externalId: file.id,
|
||||
title,
|
||||
content: '',
|
||||
contentDeferred: true,
|
||||
mimeType: 'text/plain',
|
||||
sourceUrl: file.webViewLink || `https://docs.google.com/forms/d/${file.id}/edit`,
|
||||
contentHash: formContentHash(input),
|
||||
metadata: {
|
||||
formTitle: title,
|
||||
modifiedTime: file.modifiedTime,
|
||||
createdTime: file.createdTime,
|
||||
latestResponseTime: input.contentScope === 'both' ? input.latestResponseTime : undefined,
|
||||
owners: file.owners?.map((o) => o.displayName || o.emailAddress).filter(Boolean),
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Extracts the answer values for a single question from a response.
|
||||
*/
|
||||
function extractAnswerText(answer: FormAnswer | undefined): string {
|
||||
const values = answer?.textAnswers?.answers
|
||||
?.map((a) => a.value)
|
||||
.filter((v): v is string => typeof v === 'string' && v.trim().length > 0)
|
||||
return values && values.length > 0 ? values.join(', ') : ''
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds a question-id → title map from the form structure, so responses can be
|
||||
* rendered with human-readable question labels instead of opaque IDs.
|
||||
*/
|
||||
function buildQuestionTitleMap(form: FormStructure): Map<string, string> {
|
||||
const map = new Map<string, string>()
|
||||
for (const item of form.items ?? []) {
|
||||
const questionId = item.questionItem?.question?.questionId
|
||||
if (questionId && item.title) {
|
||||
map.set(questionId, item.title)
|
||||
}
|
||||
}
|
||||
return map
|
||||
}
|
||||
|
||||
/**
|
||||
* Renders the full form document: its structure (title, description, questions)
|
||||
* followed by each response's question/answer pairs when responses are included.
|
||||
*/
|
||||
function renderFormDocument(form: FormStructure, responses: FormResponse[]): string {
|
||||
const parts: string[] = []
|
||||
|
||||
const title = form.info?.title || form.info?.documentTitle
|
||||
if (title) parts.push(`# ${title}`)
|
||||
if (form.info?.description?.trim()) parts.push(form.info.description.trim())
|
||||
|
||||
const questionTitles = buildQuestionTitleMap(form)
|
||||
|
||||
const questionLines: string[] = []
|
||||
for (const item of form.items ?? []) {
|
||||
if (!item.title?.trim()) continue
|
||||
const required = item.questionItem?.question?.required ? ' (required)' : ''
|
||||
questionLines.push(`- ${item.title.trim()}${required}`)
|
||||
if (item.description?.trim()) questionLines.push(` ${item.description.trim()}`)
|
||||
}
|
||||
if (questionLines.length > 0) {
|
||||
parts.push('## Questions')
|
||||
parts.push(questionLines.join('\n'))
|
||||
}
|
||||
|
||||
if (responses.length > 0) {
|
||||
parts.push(`## Responses (${responses.length})`)
|
||||
responses.forEach((response, index) => {
|
||||
const responseLines: string[] = []
|
||||
const submitted = response.lastSubmittedTime || response.createTime
|
||||
const header = submitted
|
||||
? `### Response ${index + 1} — ${submitted}`
|
||||
: `### Response ${index + 1}`
|
||||
responseLines.push(header)
|
||||
if (response.respondentEmail) {
|
||||
responseLines.push(`Respondent: ${response.respondentEmail}`)
|
||||
}
|
||||
for (const [questionId, answer] of Object.entries(response.answers ?? {})) {
|
||||
const label = questionTitles.get(questionId) || questionId
|
||||
const value = extractAnswerText(answer)
|
||||
if (value) responseLines.push(`${label}: ${value}`)
|
||||
}
|
||||
parts.push(responseLines.join('\n'))
|
||||
})
|
||||
}
|
||||
|
||||
return parts.join('\n\n').trim()
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the Drive `q` query that selects form files, optionally scoped to a
|
||||
* folder. Single quotes and backslashes in the folder ID are escaped to prevent
|
||||
* query injection.
|
||||
*/
|
||||
function buildDriveQuery(folderId?: string): string {
|
||||
const parts = ['trashed = false', `mimeType = '${FORM_MIME_TYPE}'`]
|
||||
if (folderId?.trim()) {
|
||||
const escaped = folderId.trim().replace(/\\/g, '\\\\').replace(/'/g, "\\'")
|
||||
parts.push(`'${escaped}' in parents`)
|
||||
}
|
||||
return parts.join(' and ')
|
||||
}
|
||||
|
||||
export const googleFormsConnector: ConnectorConfig = {
|
||||
id: 'google_forms',
|
||||
name: 'Google Forms',
|
||||
description: 'Sync Google Forms questions and responses into your knowledge base',
|
||||
version: '1.0.0',
|
||||
icon: GoogleFormsIcon,
|
||||
|
||||
auth: {
|
||||
mode: 'oauth',
|
||||
provider: 'google-forms',
|
||||
requiredScopes: [
|
||||
'https://www.googleapis.com/auth/drive',
|
||||
'https://www.googleapis.com/auth/forms.body',
|
||||
'https://www.googleapis.com/auth/forms.responses.readonly',
|
||||
],
|
||||
},
|
||||
|
||||
configFields: [
|
||||
{
|
||||
id: 'folderId',
|
||||
title: 'Folder ID',
|
||||
type: 'short-input',
|
||||
placeholder: 'e.g. 1aBcDeFgHiJkLmNoPqRsTuVwXyZ (optional)',
|
||||
required: false,
|
||||
description: 'Only sync forms inside this Drive folder. Leave blank to sync all forms.',
|
||||
},
|
||||
{
|
||||
id: 'contentScope',
|
||||
title: 'Content',
|
||||
type: 'dropdown',
|
||||
required: false,
|
||||
options: [
|
||||
{ label: 'Questions & responses', id: 'both' },
|
||||
{ label: 'Questions only', id: 'structure' },
|
||||
],
|
||||
description: 'Whether to index submitted responses alongside each form’s questions.',
|
||||
},
|
||||
{
|
||||
id: 'maxForms',
|
||||
title: 'Max Forms',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
placeholder: 'e.g. 100 (default: unlimited)',
|
||||
},
|
||||
{
|
||||
id: 'maxResponsesPerForm',
|
||||
title: 'Max Responses Per Form',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
mode: 'advanced',
|
||||
placeholder: `e.g. 100 (default: ${MAX_RESPONSES_PER_FORM})`,
|
||||
description: 'Cap on responses indexed per form. Applies only when indexing responses.',
|
||||
},
|
||||
],
|
||||
|
||||
listDocuments: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
cursor?: string,
|
||||
syncContext?: Record<string, unknown>
|
||||
): Promise<ExternalDocumentList> => {
|
||||
const maxForms = parsePositiveInt(sourceConfig.maxForms)
|
||||
const contentScope = resolveContentScope(sourceConfig.contentScope)
|
||||
const responseCap = resolveResponseCap(sourceConfig)
|
||||
const previouslyFetched = (syncContext?.totalDocsFetched as number) ?? 0
|
||||
|
||||
if (maxForms > 0 && previouslyFetched >= maxForms) {
|
||||
return { documents: [], hasMore: false }
|
||||
}
|
||||
|
||||
const folderId = sourceConfig.folderId as string | undefined
|
||||
const queryParams = new URLSearchParams({
|
||||
q: buildDriveQuery(folderId),
|
||||
pageSize: String(DRIVE_PAGE_SIZE),
|
||||
orderBy: 'modifiedTime desc',
|
||||
fields: 'nextPageToken,files(id,name,mimeType,modifiedTime,createdTime,webViewLink,owners)',
|
||||
supportsAllDrives: 'true',
|
||||
includeItemsFromAllDrives: 'true',
|
||||
})
|
||||
if (cursor) queryParams.set('pageToken', cursor)
|
||||
|
||||
const url = `${DRIVE_API_BASE}/files?${queryParams.toString()}`
|
||||
|
||||
logger.info('Listing Google Forms', {
|
||||
folderId: folderId?.trim() || 'all',
|
||||
contentScope,
|
||||
cursor: cursor ?? 'initial',
|
||||
})
|
||||
|
||||
const response = await fetchWithRetry(url, {
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
Accept: 'application/json',
|
||||
},
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
const errorText = await response.text()
|
||||
logger.error('Failed to list Google Forms', { status: response.status, error: errorText })
|
||||
throw new Error(`Failed to list Google Forms: ${response.status}`)
|
||||
}
|
||||
|
||||
const data = await response.json()
|
||||
let files = (data.files || []) as DriveFormFile[]
|
||||
|
||||
/**
|
||||
* Drive sets `incompleteSearch` when it could not search every corpus (it
|
||||
* arises with the `allDrives` scope enabled by `includeItemsFromAllDrives`).
|
||||
* A partial listing drops still-existing forms, so reconciliation must be
|
||||
* suppressed to avoid hard-deleting valid documents.
|
||||
*/
|
||||
const incompleteSearch = data.incompleteSearch === true
|
||||
|
||||
let slicedSome = false
|
||||
if (maxForms > 0) {
|
||||
const remaining = maxForms - previouslyFetched
|
||||
if (files.length > remaining) {
|
||||
slicedSome = true
|
||||
files = files.slice(0, remaining)
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Build stubs with metadata-based change indicators. Each form needs its
|
||||
* revisionId (structure changes) and, when responses are indexed, the latest
|
||||
* response time (new submissions) so the sync engine can detect changes
|
||||
* without downloading full content. Forms are processed with bounded
|
||||
* concurrency; a transient per-form failure is skipped rather than aborting
|
||||
* the whole page, but it is recorded so the listing is marked incomplete.
|
||||
*/
|
||||
let skippedOnError = false
|
||||
const stubs = await mapWithConcurrency(files, LIST_CONCURRENCY, async (file) => {
|
||||
try {
|
||||
const form = await fetchFormStructure(accessToken, file.id)
|
||||
if (!form) return null
|
||||
const latest =
|
||||
contentScope === 'both' ? await fetchLatestResponseTime(accessToken, file.id) : undefined
|
||||
return formToStub({
|
||||
file,
|
||||
formTitle: form.info?.title || form.info?.documentTitle,
|
||||
revisionId: form.revisionId,
|
||||
latestResponseTime: latest,
|
||||
contentScope,
|
||||
responseCap,
|
||||
})
|
||||
} catch (error) {
|
||||
skippedOnError = true
|
||||
logger.warn(`Skipping form during listing: ${file.name} (${file.id})`, {
|
||||
error: toError(error).message,
|
||||
})
|
||||
return null
|
||||
}
|
||||
})
|
||||
|
||||
const documents = stubs.filter((s): s is ExternalDocument => s !== null)
|
||||
|
||||
const totalFetched = previouslyFetched + documents.length
|
||||
if (syncContext) syncContext.totalDocsFetched = totalFetched
|
||||
const hitLimit = maxForms > 0 && totalFetched >= maxForms
|
||||
|
||||
const nextPageToken = data.nextPageToken as string | undefined
|
||||
|
||||
/**
|
||||
* Mark the listing as incomplete so the sync engine skips deletion
|
||||
* reconciliation. Three cases drop still-existing forms from the listing:
|
||||
* - `slicedSome`: this page held more forms than the `maxForms` cap allowed,
|
||||
* so forms beyond the slice were truncated. This is independent of
|
||||
* `hitLimit`, which counts successfully fetched stubs and can fall below
|
||||
* the cap when 404s or errors null out items even though real forms were
|
||||
* sliced off.
|
||||
* - `hitLimit` with a next page: the cap was reached while more pages of
|
||||
* forms remain in the source.
|
||||
* - `skippedOnError`: a transient error dropped a still-present form.
|
||||
* - `incompleteSearch`: Drive could not search every corpus, so the page
|
||||
* itself is partial and may omit still-existing forms.
|
||||
* Deleting any of those would wipe valid documents from the knowledge base.
|
||||
* When the cap merely coincides with source exhaustion (no slice, no next
|
||||
* page), reconciliation stays enabled so deleted forms are cleaned up.
|
||||
*/
|
||||
if (
|
||||
syncContext &&
|
||||
(slicedSome || (hitLimit && Boolean(nextPageToken)) || skippedOnError || incompleteSearch)
|
||||
) {
|
||||
syncContext.listingCapped = true
|
||||
}
|
||||
|
||||
return {
|
||||
documents,
|
||||
nextCursor: hitLimit ? undefined : nextPageToken,
|
||||
hasMore: hitLimit ? false : Boolean(nextPageToken),
|
||||
}
|
||||
},
|
||||
|
||||
getDocument: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
externalId: string
|
||||
): Promise<ExternalDocument | null> => {
|
||||
const contentScope = resolveContentScope(sourceConfig.contentScope)
|
||||
const fields = 'id,name,mimeType,modifiedTime,createdTime,webViewLink,owners,trashed'
|
||||
const metadataUrl = `${DRIVE_API_BASE}/files/${encodeURIComponent(externalId)}?fields=${encodeURIComponent(fields)}&supportsAllDrives=true`
|
||||
|
||||
const metadataResponse = await fetchWithRetry(metadataUrl, {
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
Accept: 'application/json',
|
||||
},
|
||||
})
|
||||
|
||||
if (!metadataResponse.ok) {
|
||||
if (metadataResponse.status === 404) return null
|
||||
throw new Error(`Failed to get form metadata: ${metadataResponse.status}`)
|
||||
}
|
||||
|
||||
const file = (await metadataResponse.json()) as DriveFormFile
|
||||
|
||||
if (file.trashed) return null
|
||||
if (file.mimeType !== FORM_MIME_TYPE) return null
|
||||
|
||||
try {
|
||||
const form = await fetchFormStructure(accessToken, file.id)
|
||||
if (!form) return null
|
||||
|
||||
const responseCap = resolveResponseCap(sourceConfig)
|
||||
const fetched =
|
||||
contentScope === 'both'
|
||||
? await fetchFormResponses(accessToken, file.id)
|
||||
: { responses: [], latestSubmittedTime: undefined }
|
||||
const responses = fetched.responses
|
||||
const cappedResponses =
|
||||
responses.length > responseCap ? responses.slice(0, responseCap) : responses
|
||||
|
||||
const content = renderFormDocument(form, cappedResponses)
|
||||
if (!content.trim()) return null
|
||||
|
||||
const stub = formToStub({
|
||||
file,
|
||||
formTitle: form.info?.title || form.info?.documentTitle,
|
||||
revisionId: form.revisionId,
|
||||
latestResponseTime: fetched.latestSubmittedTime,
|
||||
contentScope,
|
||||
responseCap,
|
||||
})
|
||||
return { ...stub, content, contentDeferred: false }
|
||||
} catch (error) {
|
||||
logger.warn(`Failed to fetch content for form: ${file.name} (${file.id})`, {
|
||||
error: toError(error).message,
|
||||
})
|
||||
return null
|
||||
}
|
||||
},
|
||||
|
||||
validateConfig: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>
|
||||
): Promise<{ valid: boolean; error?: string }> => {
|
||||
const folderId = sourceConfig.folderId as string | undefined
|
||||
const maxForms = sourceConfig.maxForms as string | undefined
|
||||
const maxResponsesPerForm = sourceConfig.maxResponsesPerForm as string | undefined
|
||||
|
||||
if (maxForms && (Number.isNaN(Number(maxForms)) || Number(maxForms) <= 0)) {
|
||||
return { valid: false, error: 'Max forms must be a positive number' }
|
||||
}
|
||||
|
||||
if (
|
||||
maxResponsesPerForm &&
|
||||
(Number.isNaN(Number(maxResponsesPerForm)) || Number(maxResponsesPerForm) <= 0)
|
||||
) {
|
||||
return { valid: false, error: 'Max responses per form must be a positive number' }
|
||||
}
|
||||
|
||||
try {
|
||||
if (folderId?.trim()) {
|
||||
const url = `${DRIVE_API_BASE}/files/${encodeURIComponent(folderId.trim())}?fields=id,name,mimeType&supportsAllDrives=true`
|
||||
const response = await fetchWithRetry(
|
||||
url,
|
||||
{
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
Accept: 'application/json',
|
||||
},
|
||||
},
|
||||
VALIDATE_RETRY_OPTIONS
|
||||
)
|
||||
|
||||
if (!response.ok) {
|
||||
if (response.status === 404) {
|
||||
return { valid: false, error: 'Folder not found. Check the folder ID and permissions.' }
|
||||
}
|
||||
return { valid: false, error: `Failed to access folder: ${response.status}` }
|
||||
}
|
||||
|
||||
const folder = await response.json()
|
||||
if (folder.mimeType !== FOLDER_MIME_TYPE) {
|
||||
return { valid: false, error: 'The provided ID is not a folder' }
|
||||
}
|
||||
} else {
|
||||
const url = `${DRIVE_API_BASE}/files?pageSize=1&q=${encodeURIComponent(`mimeType = '${FORM_MIME_TYPE}'`)}&fields=files(id)&supportsAllDrives=true&includeItemsFromAllDrives=true`
|
||||
const response = await fetchWithRetry(
|
||||
url,
|
||||
{
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
Accept: 'application/json',
|
||||
},
|
||||
},
|
||||
VALIDATE_RETRY_OPTIONS
|
||||
)
|
||||
|
||||
if (!response.ok) {
|
||||
return { valid: false, error: `Failed to access Google Forms: ${response.status}` }
|
||||
}
|
||||
}
|
||||
|
||||
return { valid: true }
|
||||
} catch (error) {
|
||||
return { valid: false, error: getErrorMessage(error, 'Failed to validate configuration') }
|
||||
}
|
||||
},
|
||||
|
||||
tagDefinitions: [
|
||||
{ id: 'formTitle', displayName: 'Form Title', fieldType: 'text' },
|
||||
{ id: 'owners', displayName: 'Owner', fieldType: 'text' },
|
||||
{ id: 'lastModified', displayName: 'Last Modified', fieldType: 'date' },
|
||||
{ id: 'lastResponse', displayName: 'Last Response', fieldType: 'date' },
|
||||
],
|
||||
|
||||
mapTags: (metadata: Record<string, unknown>): Record<string, unknown> => {
|
||||
const result: Record<string, unknown> = {}
|
||||
|
||||
if (typeof metadata.formTitle === 'string' && metadata.formTitle.trim()) {
|
||||
result.formTitle = metadata.formTitle.trim()
|
||||
}
|
||||
|
||||
const owners = joinTagArray(metadata.owners)
|
||||
if (owners) result.owners = owners
|
||||
|
||||
const lastModified = parseTagDate(metadata.modifiedTime)
|
||||
if (lastModified) result.lastModified = lastModified
|
||||
|
||||
const lastResponse = parseTagDate(metadata.latestResponseTime)
|
||||
if (lastResponse) result.lastResponse = lastResponse
|
||||
|
||||
return result
|
||||
},
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
export { googleFormsConnector } from '@/connectors/google-forms/google-forms'
|
||||
@@ -471,10 +471,11 @@ export const grainConnector: ConnectorConfig = {
|
||||
try {
|
||||
if (!externalId) return null
|
||||
|
||||
const recording = await fetchRecording(accessToken, externalId)
|
||||
const [recording, segments] = await Promise.all([
|
||||
fetchRecording(accessToken, externalId),
|
||||
fetchTranscript(accessToken, externalId),
|
||||
])
|
||||
if (!recording) return null
|
||||
|
||||
const segments = await fetchTranscript(accessToken, externalId)
|
||||
if (!segments) return null
|
||||
|
||||
const hasTranscript = segments.some((segment) => segment.text?.trim())
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
export { jsmConnector } from '@/connectors/jsm/jsm'
|
||||
@@ -0,0 +1,687 @@
|
||||
import { createLogger } from '@sim/logger'
|
||||
import { toError } from '@sim/utils/errors'
|
||||
import { JiraServiceManagementIcon } from '@/components/icons'
|
||||
import { fetchWithRetry, VALIDATE_RETRY_OPTIONS } from '@/lib/knowledge/documents/utils'
|
||||
import type { ConnectorConfig, ExternalDocument, ExternalDocumentList } from '@/connectors/types'
|
||||
import { parseTagDate } from '@/connectors/utils'
|
||||
import { extractAdfText, getJiraCloudId } from '@/tools/jira/utils'
|
||||
import { getJsmApiBaseUrl, getJsmHeaders } from '@/tools/jsm/utils'
|
||||
|
||||
const logger = createLogger('JsmConnector')
|
||||
|
||||
const PAGE_SIZE = 50
|
||||
|
||||
/**
|
||||
* Allowed `requestStatus` filter values for `GET /rest/servicedeskapi/request`.
|
||||
* When omitted, the JSM API defaults to `ALL_REQUESTS`.
|
||||
*/
|
||||
const VALID_REQUEST_STATUS = ['OPEN_REQUESTS', 'CLOSED_REQUESTS', 'ALL_REQUESTS'] as const
|
||||
type JsmRequestStatus = (typeof VALID_REQUEST_STATUS)[number]
|
||||
|
||||
/**
|
||||
* Allowed `requestOwnership` filter values for `GET /rest/servicedeskapi/request`.
|
||||
*
|
||||
* This param scopes results to the OAuth user's relationship to each request. When
|
||||
* omitted, the JSM API defaults to `OWNED_REQUESTS` — i.e. only requests the
|
||||
* authenticated user reported. For a knowledge-base sync the user almost always
|
||||
* wants every request in the service desk, so the connector defaults this to
|
||||
* `ALL_REQUESTS` (which the JSM API treats as "owned + participated") rather than
|
||||
* relying on the API's narrower default.
|
||||
*/
|
||||
const VALID_REQUEST_OWNERSHIP = ['OWNED_REQUESTS', 'PARTICIPATED_REQUESTS', 'ALL_REQUESTS'] as const
|
||||
type JsmRequestOwnership = (typeof VALID_REQUEST_OWNERSHIP)[number]
|
||||
|
||||
/**
|
||||
* Which comments to include in synced documents.
|
||||
*/
|
||||
const VALID_COMMENT_SCOPE = ['none', 'public', 'all'] as const
|
||||
type JsmCommentScope = (typeof VALID_COMMENT_SCOPE)[number]
|
||||
|
||||
/**
|
||||
* A JSM date object as returned by the Service Desk REST API. The same shape is
|
||||
* used for `createdDate`, `currentStatus.statusDate`, and comment `created`.
|
||||
*/
|
||||
interface JsmDate {
|
||||
iso8601?: string
|
||||
friendly?: string
|
||||
epochMillis?: number
|
||||
}
|
||||
|
||||
/**
|
||||
* Subset of a JSM customer request returned by `GET /request` and
|
||||
* `GET /request/{issueIdOrKey}`. Only the fields the connector reads are typed.
|
||||
*/
|
||||
interface JsmRequest {
|
||||
issueId?: string
|
||||
issueKey?: string
|
||||
requestTypeId?: string
|
||||
serviceDeskId?: string
|
||||
createdDate?: JsmDate
|
||||
currentStatus?: {
|
||||
status?: string
|
||||
statusCategory?: string
|
||||
statusDate?: JsmDate
|
||||
}
|
||||
reporter?: {
|
||||
displayName?: string
|
||||
emailAddress?: string
|
||||
}
|
||||
requestFieldValues?: Array<{
|
||||
fieldId?: string
|
||||
label?: string
|
||||
value?: unknown
|
||||
renderedValue?: unknown
|
||||
}>
|
||||
_links?: {
|
||||
web?: string
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* A single comment on a JSM request. The JSM API returns the comment `body` as a
|
||||
* plain string containing Jira wiki markup (not an ADF document), so no rich-text
|
||||
* extraction is required.
|
||||
*/
|
||||
interface JsmComment {
|
||||
id?: string
|
||||
body?: string
|
||||
public?: boolean
|
||||
author?: {
|
||||
displayName?: string
|
||||
}
|
||||
created?: JsmDate
|
||||
}
|
||||
|
||||
/**
|
||||
* Paginated envelope shared by every JSM Service Desk list endpoint.
|
||||
*/
|
||||
interface JsmPage<T> {
|
||||
values?: T[]
|
||||
size?: number
|
||||
isLastPage?: boolean
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads the resolved sync options off the raw `sourceConfig`, normalizing
|
||||
* enum-like fields to their valid set and clamping the numeric cap. Centralized
|
||||
* so `listDocuments`, `getDocument`, and `validateConfig` agree on defaults.
|
||||
*/
|
||||
function resolveOptions(sourceConfig: Record<string, unknown>): {
|
||||
requestStatus: JsmRequestStatus
|
||||
requestOwnership: JsmRequestOwnership
|
||||
requestTypeId: string
|
||||
searchTerm: string
|
||||
commentScope: JsmCommentScope
|
||||
maxRequests: number
|
||||
} {
|
||||
const requestStatus = VALID_REQUEST_STATUS.includes(
|
||||
sourceConfig.requestStatus as JsmRequestStatus
|
||||
)
|
||||
? (sourceConfig.requestStatus as JsmRequestStatus)
|
||||
: 'ALL_REQUESTS'
|
||||
|
||||
const requestOwnership = VALID_REQUEST_OWNERSHIP.includes(
|
||||
sourceConfig.requestOwnership as JsmRequestOwnership
|
||||
)
|
||||
? (sourceConfig.requestOwnership as JsmRequestOwnership)
|
||||
: 'ALL_REQUESTS'
|
||||
|
||||
const commentScope = VALID_COMMENT_SCOPE.includes(sourceConfig.comments as JsmCommentScope)
|
||||
? (sourceConfig.comments as JsmCommentScope)
|
||||
: 'public'
|
||||
|
||||
const requestTypeId =
|
||||
typeof sourceConfig.requestTypeId === 'string' ? sourceConfig.requestTypeId.trim() : ''
|
||||
const searchTerm =
|
||||
typeof sourceConfig.searchTerm === 'string' ? sourceConfig.searchTerm.trim() : ''
|
||||
|
||||
const parsedMax = sourceConfig.maxRequests ? Number(sourceConfig.maxRequests) : 0
|
||||
const maxRequests = Number.isFinite(parsedMax) && parsedMax > 0 ? Math.floor(parsedMax) : 0
|
||||
|
||||
return { requestStatus, requestOwnership, requestTypeId, searchTerm, commentScope, maxRequests }
|
||||
}
|
||||
|
||||
/**
|
||||
* Extracts a plain-text value for a given request field id (e.g. `summary`,
|
||||
* `description`) from a request's `requestFieldValues`. The JSM API returns
|
||||
* `value` either as a plain string (wiki markup) or, for some rich-text fields,
|
||||
* as an ADF document — both are handled.
|
||||
*/
|
||||
function getFieldText(request: JsmRequest, fieldId: string): string {
|
||||
const field = request.requestFieldValues?.find((f) => f.fieldId === fieldId)
|
||||
if (!field) return ''
|
||||
const { value } = field
|
||||
if (typeof value === 'string') return value
|
||||
if (value && typeof value === 'object') {
|
||||
const adf = extractAdfText(value)
|
||||
if (adf) return adf
|
||||
}
|
||||
return ''
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves the best available "change indicator" timestamp for a request.
|
||||
*
|
||||
* The JSM list endpoint does NOT return an updated/last-modified field — only
|
||||
* `createdDate` and `currentStatus.statusDate` are present. We use
|
||||
* `statusDate` (the time the request last changed status) when available, and
|
||||
* fall back to `createdDate`. This is the change signal encoded into the
|
||||
* contentHash. Note: edits that do not change status (e.g. a new comment) are
|
||||
* not reflected here, so such changes may not trigger a re-sync.
|
||||
*/
|
||||
function getChangeIndicator(request: JsmRequest): string {
|
||||
const statusDate = request.currentStatus?.statusDate
|
||||
if (statusDate?.epochMillis != null) return String(statusDate.epochMillis)
|
||||
if (statusDate?.iso8601) return statusDate.iso8601
|
||||
const created = request.createdDate
|
||||
if (created?.epochMillis != null) return String(created.epochMillis)
|
||||
if (created?.iso8601) return created.iso8601
|
||||
return ''
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds a stub ExternalDocument from a request returned by the list endpoint.
|
||||
* Content is deferred — description and comments require a per-request API call
|
||||
* fetched lazily in `getDocument`. The contentHash is metadata-only so it is
|
||||
* identical whether produced here or in `getDocument`.
|
||||
*/
|
||||
function requestToStub(request: JsmRequest, domain: string): ExternalDocument {
|
||||
const issueId = String(request.issueId ?? '')
|
||||
const issueKey = request.issueKey ?? issueId
|
||||
const summary = getFieldText(request, 'summary') || 'Untitled'
|
||||
const status = request.currentStatus?.status
|
||||
|
||||
const bareDomain = domain
|
||||
.trim()
|
||||
.replace(/^https?:\/\//i, '')
|
||||
.replace(/\/+$/, '')
|
||||
|
||||
return {
|
||||
externalId: issueId,
|
||||
title: `${issueKey}: ${summary}`,
|
||||
content: '',
|
||||
contentDeferred: true,
|
||||
mimeType: 'text/plain',
|
||||
sourceUrl: request._links?.web || `https://${bareDomain}/browse/${issueKey}`,
|
||||
contentHash: `jsm:${issueId}:${getChangeIndicator(request)}`,
|
||||
metadata: {
|
||||
issueKey,
|
||||
requestTypeId: request.requestTypeId,
|
||||
serviceDeskId: request.serviceDeskId,
|
||||
status,
|
||||
reporter: request.reporter?.displayName,
|
||||
created: request.createdDate?.iso8601,
|
||||
/**
|
||||
* The list endpoint has no true "last updated" field; `statusDate` is the
|
||||
* closest available signal (time of last status change). Mapped to the
|
||||
* `updated` tag and documented as such.
|
||||
*/
|
||||
statusDate: request.currentStatus?.statusDate?.iso8601,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Renders a readable plain-text document from a fully-fetched request and its
|
||||
* comments. Includes summary, description, reporter, status, and comment thread.
|
||||
*/
|
||||
function buildContent(request: JsmRequest, comments: JsmComment[]): string {
|
||||
const parts: string[] = []
|
||||
|
||||
const summary = getFieldText(request, 'summary')
|
||||
if (summary) parts.push(summary)
|
||||
|
||||
const description = getFieldText(request, 'description')
|
||||
if (description) parts.push(description)
|
||||
|
||||
const status = request.currentStatus?.status
|
||||
if (status) parts.push(`Status: ${status}`)
|
||||
|
||||
const reporter = request.reporter?.displayName
|
||||
if (reporter) parts.push(`Reporter: ${reporter}`)
|
||||
|
||||
if (comments.length > 0) {
|
||||
parts.push('Comments:')
|
||||
for (const comment of comments) {
|
||||
const body = (comment.body ?? '').trim()
|
||||
if (!body) continue
|
||||
const author = comment.author?.displayName
|
||||
parts.push(author ? `${author}: ${body}` : body)
|
||||
}
|
||||
}
|
||||
|
||||
return parts.join('\n\n').trim()
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves and caches the Jira cloud ID for a domain across a sync run.
|
||||
*/
|
||||
async function resolveCloudId(
|
||||
domain: string,
|
||||
accessToken: string,
|
||||
syncContext?: Record<string, unknown>
|
||||
): Promise<string> {
|
||||
const cached = syncContext?.cloudId as string | undefined
|
||||
if (cached) return cached
|
||||
const cloudId = await getJiraCloudId(domain, accessToken)
|
||||
if (syncContext) syncContext.cloudId = cloudId
|
||||
return cloudId
|
||||
}
|
||||
|
||||
/**
|
||||
* Fetches comments for a request, following offset pagination until the API
|
||||
* signals `isLastPage`. When `publicOnly` is true the `public=true` filter is
|
||||
* applied so internal/agent-only comments are excluded.
|
||||
*/
|
||||
async function fetchComments(
|
||||
baseUrl: string,
|
||||
accessToken: string,
|
||||
issueIdOrKey: string,
|
||||
publicOnly: boolean
|
||||
): Promise<JsmComment[]> {
|
||||
const comments: JsmComment[] = []
|
||||
let start = 0
|
||||
|
||||
while (true) {
|
||||
const params = new URLSearchParams({
|
||||
start: String(start),
|
||||
limit: String(PAGE_SIZE),
|
||||
})
|
||||
/**
|
||||
* The JSM comment endpoint exposes `public` and `internal` as independent
|
||||
* inclusion filters that both default to `true`. Requesting public-only
|
||||
* therefore requires explicitly disabling `internal` — passing `public=true`
|
||||
* alone would still return agent-only/internal comments.
|
||||
*/
|
||||
if (publicOnly) {
|
||||
params.append('public', 'true')
|
||||
params.append('internal', 'false')
|
||||
}
|
||||
const url = `${baseUrl}/request/${encodeURIComponent(issueIdOrKey)}/comment?${params.toString()}`
|
||||
|
||||
const response = await fetchWithRetry(url, {
|
||||
method: 'GET',
|
||||
headers: getJsmHeaders(accessToken),
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
logger.warn('Failed to fetch JSM comments', {
|
||||
issueIdOrKey,
|
||||
status: response.status,
|
||||
})
|
||||
break
|
||||
}
|
||||
|
||||
const data = (await response.json()) as JsmPage<JsmComment>
|
||||
const values = data.values ?? []
|
||||
comments.push(...values)
|
||||
|
||||
if (data.isLastPage || values.length === 0) break
|
||||
start += values.length
|
||||
}
|
||||
|
||||
return comments
|
||||
}
|
||||
|
||||
export const jsmConnector: ConnectorConfig = {
|
||||
id: 'jsm',
|
||||
name: 'Jira Service Management',
|
||||
description: 'Sync service desk requests from Jira Service Management into your knowledge base',
|
||||
version: '1.0.0',
|
||||
icon: JiraServiceManagementIcon,
|
||||
|
||||
auth: {
|
||||
mode: 'oauth',
|
||||
provider: 'jira',
|
||||
requiredScopes: [
|
||||
'read:servicedesk:jira-service-management',
|
||||
'read:request:jira-service-management',
|
||||
'read:request.comment:jira-service-management',
|
||||
'read:request.status:jira-service-management',
|
||||
/**
|
||||
* Requests embed a `reporter` user object whose `displayName` is surfaced
|
||||
* in document content and the Reporter tag. Atlassian only populates
|
||||
* embedded user data when the user-read scope is granted, so request it
|
||||
* here. Present in the `jira` OAuth provider config as `read:jira-user`.
|
||||
*/
|
||||
'read:jira-user',
|
||||
'offline_access',
|
||||
],
|
||||
},
|
||||
|
||||
configFields: [
|
||||
{
|
||||
id: 'domain',
|
||||
title: 'Jira Domain',
|
||||
type: 'short-input',
|
||||
placeholder: 'yoursite.atlassian.net',
|
||||
required: true,
|
||||
},
|
||||
{
|
||||
id: 'serviceDeskSelector',
|
||||
title: 'Service Desk',
|
||||
type: 'selector',
|
||||
selectorKey: 'jsm.serviceDesks',
|
||||
canonicalParamId: 'serviceDeskId',
|
||||
mode: 'basic',
|
||||
dependsOn: ['domain'],
|
||||
placeholder: 'Select a service desk',
|
||||
required: true,
|
||||
},
|
||||
{
|
||||
id: 'serviceDeskId',
|
||||
title: 'Service Desk ID',
|
||||
type: 'short-input',
|
||||
canonicalParamId: 'serviceDeskId',
|
||||
mode: 'advanced',
|
||||
placeholder: 'e.g. 1, 2',
|
||||
required: true,
|
||||
},
|
||||
{
|
||||
id: 'requestTypeSelector',
|
||||
title: 'Request Type',
|
||||
type: 'selector',
|
||||
selectorKey: 'jsm.requestTypes',
|
||||
canonicalParamId: 'requestTypeId',
|
||||
mode: 'basic',
|
||||
dependsOn: ['domain', 'serviceDeskSelector'],
|
||||
placeholder: 'All request types',
|
||||
required: false,
|
||||
},
|
||||
{
|
||||
id: 'requestTypeId',
|
||||
title: 'Request Type ID',
|
||||
type: 'short-input',
|
||||
canonicalParamId: 'requestTypeId',
|
||||
mode: 'advanced',
|
||||
placeholder: 'e.g. 10 (leave blank for all)',
|
||||
required: false,
|
||||
},
|
||||
{
|
||||
id: 'requestStatus',
|
||||
title: 'Request Status',
|
||||
type: 'dropdown',
|
||||
required: false,
|
||||
options: [
|
||||
{ label: 'All requests', id: 'ALL_REQUESTS' },
|
||||
{ label: 'Open requests', id: 'OPEN_REQUESTS' },
|
||||
{ label: 'Closed requests', id: 'CLOSED_REQUESTS' },
|
||||
],
|
||||
},
|
||||
{
|
||||
id: 'requestOwnership',
|
||||
title: 'Request Ownership',
|
||||
type: 'dropdown',
|
||||
required: false,
|
||||
description:
|
||||
'Which requests the connected account can see. "Owned + participated" is the broadest scope a customer token can sync.',
|
||||
options: [
|
||||
{ label: 'Owned + participated', id: 'ALL_REQUESTS' },
|
||||
{ label: 'Owned only', id: 'OWNED_REQUESTS' },
|
||||
{ label: 'Participated only', id: 'PARTICIPATED_REQUESTS' },
|
||||
],
|
||||
},
|
||||
{
|
||||
id: 'comments',
|
||||
title: 'Include Comments',
|
||||
type: 'dropdown',
|
||||
required: false,
|
||||
description: 'Comments require an extra API call per request during sync.',
|
||||
options: [
|
||||
{ label: 'Public comments only', id: 'public' },
|
||||
{ label: 'All comments (incl. internal)', id: 'all' },
|
||||
{ label: 'No comments', id: 'none' },
|
||||
],
|
||||
},
|
||||
{
|
||||
id: 'searchTerm',
|
||||
title: 'Search Filter',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
placeholder: 'e.g. password reset (optional)',
|
||||
},
|
||||
{
|
||||
id: 'maxRequests',
|
||||
title: 'Max Requests',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
placeholder: 'e.g. 500 (default: unlimited)',
|
||||
},
|
||||
],
|
||||
|
||||
listDocuments: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
cursor?: string,
|
||||
syncContext?: Record<string, unknown>
|
||||
): Promise<ExternalDocumentList> => {
|
||||
const domain = sourceConfig.domain as string
|
||||
const serviceDeskId = sourceConfig.serviceDeskId as string
|
||||
|
||||
if (!domain || !serviceDeskId) {
|
||||
throw new Error('Domain and service desk ID are required')
|
||||
}
|
||||
|
||||
const { requestStatus, requestOwnership, requestTypeId, searchTerm, maxRequests } =
|
||||
resolveOptions(sourceConfig)
|
||||
|
||||
const cloudId = await resolveCloudId(domain, accessToken, syncContext)
|
||||
const baseUrl = getJsmApiBaseUrl(cloudId)
|
||||
|
||||
/**
|
||||
* `start|collected` is encoded in the cursor so the maxRequests cap holds
|
||||
* across pages even if syncContext is not threaded through by the caller.
|
||||
*/
|
||||
let start = 0
|
||||
let collectedSoFar = (syncContext?.collectedCount as number | undefined) ?? 0
|
||||
if (cursor) {
|
||||
const sep = cursor.indexOf('|')
|
||||
if (sep > 0) {
|
||||
const parsedStart = Number(cursor.slice(0, sep))
|
||||
const parsedCount = Number(cursor.slice(sep + 1))
|
||||
if (Number.isFinite(parsedStart) && parsedStart >= 0) start = parsedStart
|
||||
if (Number.isFinite(parsedCount) && parsedCount >= 0) collectedSoFar = parsedCount
|
||||
} else {
|
||||
const parsedStart = Number(cursor)
|
||||
if (Number.isFinite(parsedStart) && parsedStart >= 0) start = parsedStart
|
||||
}
|
||||
}
|
||||
|
||||
const remaining = maxRequests > 0 ? Math.max(0, maxRequests - collectedSoFar) : PAGE_SIZE
|
||||
if (maxRequests > 0 && remaining === 0) {
|
||||
return { documents: [], hasMore: false }
|
||||
}
|
||||
|
||||
const params = new URLSearchParams({
|
||||
serviceDeskId,
|
||||
requestStatus,
|
||||
start: String(start),
|
||||
limit: String(Math.min(PAGE_SIZE, remaining)),
|
||||
})
|
||||
params.append('requestOwnership', requestOwnership)
|
||||
if (requestTypeId) params.append('requestTypeId', requestTypeId)
|
||||
if (searchTerm) params.append('searchTerm', searchTerm)
|
||||
|
||||
const url = `${baseUrl}/request?${params.toString()}`
|
||||
|
||||
logger.info('Listing JSM requests', {
|
||||
serviceDeskId,
|
||||
requestStatus,
|
||||
requestOwnership,
|
||||
hasCursor: Boolean(cursor),
|
||||
})
|
||||
|
||||
const response = await fetchWithRetry(url, {
|
||||
method: 'GET',
|
||||
headers: getJsmHeaders(accessToken),
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
const errorText = await response.text()
|
||||
logger.error('Failed to list JSM requests', { status: response.status, error: errorText })
|
||||
throw new Error(`Failed to list JSM requests: ${response.status}`)
|
||||
}
|
||||
|
||||
const data = (await response.json()) as JsmPage<JsmRequest>
|
||||
let requests = data.values ?? []
|
||||
|
||||
let slicedSome = false
|
||||
if (maxRequests > 0 && requests.length > remaining) {
|
||||
slicedSome = true
|
||||
requests = requests.slice(0, remaining)
|
||||
}
|
||||
|
||||
const documents = requests.map((request) => requestToStub(request, domain))
|
||||
|
||||
const newCollected = collectedSoFar + requests.length
|
||||
if (syncContext) syncContext.collectedCount = newCollected
|
||||
|
||||
const reachedCap = maxRequests > 0 && newCollected >= maxRequests
|
||||
|
||||
/**
|
||||
* When `maxRequests` truncates the listing before the source is exhausted,
|
||||
* flag the run as capped so the sync engine skips deletion reconciliation —
|
||||
* otherwise unseen requests beyond the cap would be deleted on every sync.
|
||||
* `slicedSome` covers truncation on the final page: requests dropped from
|
||||
* this page still exist even when `isLastPage` is true. (The requested
|
||||
* `limit` never exceeds the remaining budget, so a slice should be
|
||||
* impossible — this is defense in depth against the API over-returning.)
|
||||
*/
|
||||
if (((reachedCap && !data.isLastPage) || slicedSome) && syncContext) {
|
||||
syncContext.listingCapped = true
|
||||
}
|
||||
|
||||
const hasMore = !data.isLastPage && requests.length > 0 && !reachedCap
|
||||
const nextStart = start + requests.length
|
||||
|
||||
return {
|
||||
documents,
|
||||
nextCursor: hasMore ? `${nextStart}|${newCollected}` : undefined,
|
||||
hasMore,
|
||||
}
|
||||
},
|
||||
|
||||
getDocument: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
externalId: string,
|
||||
syncContext?: Record<string, unknown>
|
||||
): Promise<ExternalDocument | null> => {
|
||||
const domain = sourceConfig.domain as string
|
||||
const { commentScope } = resolveOptions(sourceConfig)
|
||||
const cloudId = await resolveCloudId(domain, accessToken, syncContext)
|
||||
const baseUrl = getJsmApiBaseUrl(cloudId)
|
||||
|
||||
const requestUrl = `${baseUrl}/request/${encodeURIComponent(externalId)}?expand=status`
|
||||
const response = await fetchWithRetry(requestUrl, {
|
||||
method: 'GET',
|
||||
headers: getJsmHeaders(accessToken),
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
if (response.status === 404) return null
|
||||
if (response.status === 401 || response.status === 403) {
|
||||
logger.warn('Access denied fetching JSM request', { externalId, status: response.status })
|
||||
return null
|
||||
}
|
||||
throw new Error(`Failed to get JSM request: ${response.status}`)
|
||||
}
|
||||
|
||||
const request = (await response.json()) as JsmRequest
|
||||
|
||||
const comments =
|
||||
commentScope === 'none'
|
||||
? []
|
||||
: await fetchComments(baseUrl, accessToken, externalId, commentScope === 'public')
|
||||
|
||||
const stub = requestToStub(request, domain)
|
||||
const content = buildContent(request, comments)
|
||||
|
||||
return {
|
||||
...stub,
|
||||
content,
|
||||
contentDeferred: false,
|
||||
}
|
||||
},
|
||||
|
||||
validateConfig: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>
|
||||
): Promise<{ valid: boolean; error?: string }> => {
|
||||
const domain = sourceConfig.domain as string
|
||||
const serviceDeskId = sourceConfig.serviceDeskId as string
|
||||
|
||||
if (!domain || !serviceDeskId) {
|
||||
return { valid: false, error: 'Domain and service desk ID are required' }
|
||||
}
|
||||
|
||||
if (sourceConfig.maxRequests) {
|
||||
const max = Number(sourceConfig.maxRequests)
|
||||
if (Number.isNaN(max) || max <= 0) {
|
||||
return { valid: false, error: 'Max requests must be a positive number' }
|
||||
}
|
||||
}
|
||||
|
||||
try {
|
||||
const cloudId = await getJiraCloudId(domain, accessToken, VALIDATE_RETRY_OPTIONS)
|
||||
const baseUrl = getJsmApiBaseUrl(cloudId)
|
||||
const url = `${baseUrl}/servicedesk/${encodeURIComponent(serviceDeskId)}`
|
||||
|
||||
const response = await fetchWithRetry(
|
||||
url,
|
||||
{
|
||||
method: 'GET',
|
||||
headers: getJsmHeaders(accessToken),
|
||||
},
|
||||
VALIDATE_RETRY_OPTIONS
|
||||
)
|
||||
|
||||
if (!response.ok) {
|
||||
if (response.status === 404) {
|
||||
return { valid: false, error: `Service desk "${serviceDeskId}" not found` }
|
||||
}
|
||||
if (response.status === 401 || response.status === 403) {
|
||||
return {
|
||||
valid: false,
|
||||
error: 'Access denied. Check the connected account has access to this service desk.',
|
||||
}
|
||||
}
|
||||
const errorText = await response.text()
|
||||
return { valid: false, error: `Failed to validate: ${response.status} - ${errorText}` }
|
||||
}
|
||||
|
||||
return { valid: true }
|
||||
} catch (error) {
|
||||
return { valid: false, error: toError(error).message || 'Failed to validate configuration' }
|
||||
}
|
||||
},
|
||||
|
||||
tagDefinitions: [
|
||||
{ id: 'status', displayName: 'Status', fieldType: 'text' },
|
||||
{ id: 'requestTypeId', displayName: 'Request Type', fieldType: 'text' },
|
||||
{ id: 'reporter', displayName: 'Reporter', fieldType: 'text' },
|
||||
{ id: 'created', displayName: 'Created', fieldType: 'date' },
|
||||
{ id: 'updated', displayName: 'Last Status Change', fieldType: 'date' },
|
||||
],
|
||||
|
||||
mapTags: (metadata: Record<string, unknown>): Record<string, unknown> => {
|
||||
const result: Record<string, unknown> = {}
|
||||
|
||||
if (typeof metadata.status === 'string') result.status = metadata.status
|
||||
if (typeof metadata.requestTypeId === 'string') result.requestTypeId = metadata.requestTypeId
|
||||
if (typeof metadata.reporter === 'string') result.reporter = metadata.reporter
|
||||
|
||||
const created = parseTagDate(metadata.created)
|
||||
if (created) result.created = created
|
||||
|
||||
/**
|
||||
* The list endpoint exposes no true last-updated field; `statusDate` (time
|
||||
* of last status change) is the closest available signal and surfaces under
|
||||
* the "Last Status Change" tag.
|
||||
*/
|
||||
const statusDate = parseTagDate(metadata.statusDate)
|
||||
if (statusDate) result.updated = statusDate
|
||||
|
||||
return result
|
||||
},
|
||||
}
|
||||
@@ -11,6 +11,18 @@ vi.mock('@/components/icons', () => ({
|
||||
NotionIcon: () => null,
|
||||
GoogleDriveIcon: () => null,
|
||||
AirtableIcon: () => null,
|
||||
SentryIcon: () => null,
|
||||
TypeformIcon: () => null,
|
||||
YouTubeIcon: () => null,
|
||||
JiraServiceManagementIcon: () => null,
|
||||
S3Icon: () => null,
|
||||
GoogleFormsIcon: () => null,
|
||||
AzureDevOpsIcon: () => null,
|
||||
xIcon: () => null,
|
||||
GranolaIcon: () => null,
|
||||
GreenhouseIcon: () => null,
|
||||
FathomIcon: () => null,
|
||||
RootlyIcon: () => null,
|
||||
}))
|
||||
vi.mock('@/lib/knowledge/documents/utils', () => ({
|
||||
fetchWithRetry: vi.fn(),
|
||||
@@ -18,14 +30,37 @@ vi.mock('@/lib/knowledge/documents/utils', () => ({
|
||||
}))
|
||||
vi.mock('@/tools/jira/utils', () => ({ extractAdfText: vi.fn(), getJiraCloudId: vi.fn() }))
|
||||
vi.mock('@/tools/confluence/utils', () => ({ getConfluenceCloudId: vi.fn() }))
|
||||
vi.mock('@/tools/jsm/utils', () => ({
|
||||
getJsmApiBaseUrl: vi.fn(),
|
||||
getJsmFormsApiBaseUrl: vi.fn(),
|
||||
getJsmHeaders: vi.fn(),
|
||||
}))
|
||||
vi.mock('@/tools/s3/utils', () => ({
|
||||
encodeS3PathComponent: vi.fn(),
|
||||
getSignatureKey: vi.fn(),
|
||||
parseS3Uri: vi.fn(),
|
||||
generatePresignedUrl: vi.fn(),
|
||||
}))
|
||||
|
||||
import { airtableConnector } from '@/connectors/airtable/airtable'
|
||||
import { azureDevopsConnector } from '@/connectors/azure-devops/azure-devops'
|
||||
import { confluenceConnector } from '@/connectors/confluence/confluence'
|
||||
import { fathomConnector } from '@/connectors/fathom/fathom'
|
||||
import { githubConnector } from '@/connectors/github/github'
|
||||
import { googleDriveConnector } from '@/connectors/google-drive/google-drive'
|
||||
import { googleFormsConnector } from '@/connectors/google-forms/google-forms'
|
||||
import { granolaConnector } from '@/connectors/granola/granola'
|
||||
import { greenhouseConnector } from '@/connectors/greenhouse/greenhouse'
|
||||
import { jiraConnector } from '@/connectors/jira/jira'
|
||||
import { jsmConnector } from '@/connectors/jsm/jsm'
|
||||
import { linearConnector } from '@/connectors/linear/linear'
|
||||
import { notionConnector } from '@/connectors/notion/notion'
|
||||
import { rootlyConnector } from '@/connectors/rootly/rootly'
|
||||
import { s3Connector } from '@/connectors/s3/s3'
|
||||
import { sentryConnector } from '@/connectors/sentry/sentry'
|
||||
import { typeformConnector } from '@/connectors/typeform/typeform'
|
||||
import { xConnector } from '@/connectors/x/x'
|
||||
import { youtubeConnector } from '@/connectors/youtube/youtube'
|
||||
|
||||
const ISO_DATE = '2025-06-15T10:30:00.000Z'
|
||||
|
||||
@@ -388,3 +423,714 @@ describe('Airtable mapTags', () => {
|
||||
expect(result).toEqual({ createdTime: new Date(ISO_DATE) })
|
||||
})
|
||||
})
|
||||
|
||||
describe('Sentry mapTags', () => {
|
||||
const mapTags = sentryConnector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
level: 'error',
|
||||
status: 'unresolved',
|
||||
count: 1234,
|
||||
firstSeen: '2025-01-01T00:00:00.000Z',
|
||||
lastSeen: ISO_DATE,
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
level: 'error',
|
||||
status: 'unresolved',
|
||||
count: 1234,
|
||||
firstSeen: new Date('2025-01-01T00:00:00.000Z'),
|
||||
lastSeen: new Date(ISO_DATE),
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
level: 123,
|
||||
status: null,
|
||||
count: 'not-a-number',
|
||||
firstSeen: 'bad-date',
|
||||
lastSeen: 99999,
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips blank string fields', () => {
|
||||
const result = mapTags({ level: ' ', status: '' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('converts string count to number', () => {
|
||||
const result = mapTags({ count: '42' })
|
||||
expect(result).toEqual({ count: 42 })
|
||||
})
|
||||
|
||||
it.concurrent('maps count of zero', () => {
|
||||
const result = mapTags({ count: 0 })
|
||||
expect(result).toEqual({ count: 0 })
|
||||
})
|
||||
})
|
||||
|
||||
describe('Typeform mapTags', () => {
|
||||
const mapTags = typeformConnector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
formTitle: 'Customer Survey',
|
||||
platform: 'web',
|
||||
submittedAt: ISO_DATE,
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
formTitle: 'Customer Survey',
|
||||
platform: 'web',
|
||||
submittedAt: new Date(ISO_DATE),
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
formTitle: 123,
|
||||
platform: null,
|
||||
submittedAt: 99999,
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips submittedAt when date is invalid', () => {
|
||||
const result = mapTags({ submittedAt: 'not-a-date' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips empty string fields', () => {
|
||||
const result = mapTags({ formTitle: '', platform: '' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
})
|
||||
|
||||
describe('YouTube mapTags', () => {
|
||||
const mapTags = youtubeConnector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
channelTitle: 'Tech Channel',
|
||||
publishedAt: ISO_DATE,
|
||||
duration: 'PT10M30S',
|
||||
tags: ['tutorial', 'coding'],
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
channelTitle: 'Tech Channel',
|
||||
publishedAt: new Date(ISO_DATE),
|
||||
duration: 'PT10M30S',
|
||||
tags: 'tutorial, coding',
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
channelTitle: 123,
|
||||
publishedAt: 99999,
|
||||
duration: null,
|
||||
tags: 'not-an-array',
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips publishedAt when date is invalid', () => {
|
||||
const result = mapTags({ publishedAt: 'bad-date' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips blank string fields', () => {
|
||||
const result = mapTags({ channelTitle: ' ', duration: '' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips tags when array is empty', () => {
|
||||
const result = mapTags({ tags: [] })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
})
|
||||
|
||||
describe('JSM mapTags', () => {
|
||||
const mapTags = jsmConnector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
status: 'Waiting for support',
|
||||
requestTypeId: 'rt-123',
|
||||
reporter: 'Carol',
|
||||
created: '2025-01-01T00:00:00.000Z',
|
||||
statusDate: ISO_DATE,
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
status: 'Waiting for support',
|
||||
requestTypeId: 'rt-123',
|
||||
reporter: 'Carol',
|
||||
created: new Date('2025-01-01T00:00:00.000Z'),
|
||||
updated: new Date(ISO_DATE),
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
status: 123,
|
||||
requestTypeId: null,
|
||||
reporter: true,
|
||||
created: 'bad-date',
|
||||
statusDate: 99999,
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips created when date is invalid', () => {
|
||||
const result = mapTags({ created: 'not-a-date' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('maps statusDate to updated key', () => {
|
||||
const result = mapTags({ statusDate: ISO_DATE })
|
||||
expect(result).toEqual({ updated: new Date(ISO_DATE) })
|
||||
expect(result).not.toHaveProperty('statusDate')
|
||||
})
|
||||
})
|
||||
|
||||
describe('S3 mapTags', () => {
|
||||
const mapTags = s3Connector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
prefix: 'documents/reports/',
|
||||
fileSize: 2048,
|
||||
lastModified: ISO_DATE,
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
prefix: 'documents/reports/',
|
||||
fileSize: 2048,
|
||||
lastModified: new Date(ISO_DATE),
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
prefix: 123,
|
||||
fileSize: 'not-a-number',
|
||||
lastModified: 99999,
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips prefix when empty string', () => {
|
||||
const result = mapTags({ prefix: '' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips lastModified when date is invalid', () => {
|
||||
const result = mapTags({ lastModified: 'bad-date' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('converts string fileSize to number', () => {
|
||||
const result = mapTags({ fileSize: '512' })
|
||||
expect(result).toEqual({ fileSize: 512 })
|
||||
})
|
||||
|
||||
it.concurrent('maps fileSize of zero', () => {
|
||||
const result = mapTags({ fileSize: 0 })
|
||||
expect(result).toEqual({ fileSize: 0 })
|
||||
})
|
||||
})
|
||||
|
||||
describe('Google Forms mapTags', () => {
|
||||
const mapTags = googleFormsConnector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
formTitle: 'Feedback Form',
|
||||
owners: ['Alice', 'Bob'],
|
||||
modifiedTime: ISO_DATE,
|
||||
latestResponseTime: '2025-01-01T00:00:00.000Z',
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
formTitle: 'Feedback Form',
|
||||
owners: 'Alice, Bob',
|
||||
lastModified: new Date(ISO_DATE),
|
||||
lastResponse: new Date('2025-01-01T00:00:00.000Z'),
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
formTitle: 123,
|
||||
owners: 'not-an-array',
|
||||
modifiedTime: 99999,
|
||||
latestResponseTime: false,
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('trims formTitle in output', () => {
|
||||
const result = mapTags({ formTitle: ' Feedback Form ' })
|
||||
expect(result).toEqual({ formTitle: 'Feedback Form' })
|
||||
})
|
||||
|
||||
it.concurrent('skips formTitle when blank', () => {
|
||||
const result = mapTags({ formTitle: ' ' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips owners when array is empty', () => {
|
||||
const result = mapTags({ owners: [] })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('maps modifiedTime to lastModified key', () => {
|
||||
const result = mapTags({ modifiedTime: ISO_DATE })
|
||||
expect(result).toEqual({ lastModified: new Date(ISO_DATE) })
|
||||
expect(result).not.toHaveProperty('modifiedTime')
|
||||
})
|
||||
|
||||
it.concurrent('maps latestResponseTime to lastResponse key', () => {
|
||||
const result = mapTags({ latestResponseTime: ISO_DATE })
|
||||
expect(result).toEqual({ lastResponse: new Date(ISO_DATE) })
|
||||
expect(result).not.toHaveProperty('latestResponseTime')
|
||||
})
|
||||
})
|
||||
|
||||
describe('Azure DevOps mapTags', () => {
|
||||
const mapTags = azureDevopsConnector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
kind: 'workItem',
|
||||
wikiName: 'Engineering Wiki',
|
||||
workItemType: 'Bug',
|
||||
state: 'Active',
|
||||
areaPath: 'Project\\Team',
|
||||
tags: ['frontend', 'urgent'],
|
||||
repository: 'owner/repo',
|
||||
path: 'src/index.ts',
|
||||
changedDate: ISO_DATE,
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
kind: 'workItem',
|
||||
wikiName: 'Engineering Wiki',
|
||||
workItemType: 'Bug',
|
||||
state: 'Active',
|
||||
areaPath: 'Project\\Team',
|
||||
tags: 'frontend, urgent',
|
||||
repository: 'owner/repo',
|
||||
path: 'src/index.ts',
|
||||
changedDate: new Date(ISO_DATE),
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
kind: 123,
|
||||
wikiName: null,
|
||||
workItemType: true,
|
||||
state: [],
|
||||
areaPath: 99999,
|
||||
tags: 'not-an-array',
|
||||
repository: false,
|
||||
path: 42,
|
||||
changedDate: 'bad-date',
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips empty string fields except kind', () => {
|
||||
const result = mapTags({
|
||||
kind: '',
|
||||
wikiName: '',
|
||||
workItemType: '',
|
||||
state: '',
|
||||
areaPath: '',
|
||||
repository: '',
|
||||
path: '',
|
||||
})
|
||||
expect(result).toEqual({ kind: '' })
|
||||
})
|
||||
|
||||
it.concurrent('skips tags when array is empty', () => {
|
||||
const result = mapTags({ tags: [] })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips changedDate when date is invalid', () => {
|
||||
const result = mapTags({ changedDate: 'garbage' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
})
|
||||
|
||||
describe('X mapTags', () => {
|
||||
const mapTags = xConnector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
author: 'jack',
|
||||
createdAt: ISO_DATE,
|
||||
likeCount: 1500,
|
||||
retweetCount: 300,
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
author: 'jack',
|
||||
createdAt: new Date(ISO_DATE),
|
||||
likeCount: 1500,
|
||||
retweetCount: 300,
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
author: 123,
|
||||
createdAt: 99999,
|
||||
likeCount: 'not-a-number',
|
||||
retweetCount: 'nope',
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips createdAt when date is invalid', () => {
|
||||
const result = mapTags({ createdAt: 'not-a-date' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('converts string counts to numbers', () => {
|
||||
const result = mapTags({ likeCount: '42', retweetCount: '7' })
|
||||
expect(result).toEqual({ likeCount: 42, retweetCount: 7 })
|
||||
})
|
||||
|
||||
it.concurrent('maps counts of zero', () => {
|
||||
const result = mapTags({ likeCount: 0, retweetCount: 0 })
|
||||
expect(result).toEqual({ likeCount: 0, retweetCount: 0 })
|
||||
})
|
||||
})
|
||||
|
||||
describe('Granola mapTags', () => {
|
||||
const mapTags = granolaConnector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
title: 'Weekly Sync',
|
||||
owner: 'Alice',
|
||||
attendees: ['Alice', 'Bob'],
|
||||
folders: ['Team', 'Projects'],
|
||||
meeting: 'Q3 Planning',
|
||||
noteDate: ISO_DATE,
|
||||
meetingDate: '2025-01-01T00:00:00.000Z',
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
title: 'Weekly Sync',
|
||||
owner: 'Alice',
|
||||
attendees: 'Alice, Bob',
|
||||
folders: 'Team, Projects',
|
||||
meeting: 'Q3 Planning',
|
||||
noteDate: new Date(ISO_DATE),
|
||||
meetingDate: new Date('2025-01-01T00:00:00.000Z'),
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
title: 123,
|
||||
owner: null,
|
||||
attendees: 'not-an-array',
|
||||
folders: 'not-an-array',
|
||||
meeting: true,
|
||||
noteDate: 99999,
|
||||
meetingDate: false,
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('trims text fields in output', () => {
|
||||
const result = mapTags({ title: ' Weekly Sync ', owner: ' Alice ', meeting: ' Q3 ' })
|
||||
expect(result).toEqual({ title: 'Weekly Sync', owner: 'Alice', meeting: 'Q3' })
|
||||
})
|
||||
|
||||
it.concurrent('skips blank text fields', () => {
|
||||
const result = mapTags({ title: ' ', owner: '', meeting: ' ' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips array fields when empty', () => {
|
||||
const result = mapTags({ attendees: [], folders: [] })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips noteDate when date is invalid', () => {
|
||||
const result = mapTags({ noteDate: 'not-a-date' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips meetingDate when date is invalid', () => {
|
||||
const result = mapTags({ meetingDate: 'garbage' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
})
|
||||
|
||||
describe('Greenhouse mapTags', () => {
|
||||
const mapTags = greenhouseConnector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
candidateName: 'Jane Doe',
|
||||
company: 'Acme',
|
||||
title: 'Engineer',
|
||||
recruiter: 'Alice',
|
||||
coordinator: 'Bob',
|
||||
source: 'LinkedIn',
|
||||
applicationCount: 3,
|
||||
updatedAt: ISO_DATE,
|
||||
lastActivity: '2025-01-01T00:00:00.000Z',
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
candidateName: 'Jane Doe',
|
||||
company: 'Acme',
|
||||
title: 'Engineer',
|
||||
recruiter: 'Alice',
|
||||
coordinator: 'Bob',
|
||||
source: 'LinkedIn',
|
||||
applicationCount: 3,
|
||||
updatedAt: new Date(ISO_DATE),
|
||||
lastActivity: new Date('2025-01-01T00:00:00.000Z'),
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
candidateName: 123,
|
||||
company: null,
|
||||
title: true,
|
||||
recruiter: [],
|
||||
coordinator: false,
|
||||
source: 99999,
|
||||
applicationCount: 'not-a-number',
|
||||
updatedAt: 12345,
|
||||
lastActivity: 'bad-date',
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('trims text fields in output', () => {
|
||||
const result = mapTags({ candidateName: ' Jane Doe ', company: ' Acme ' })
|
||||
expect(result).toEqual({ candidateName: 'Jane Doe', company: 'Acme' })
|
||||
})
|
||||
|
||||
it.concurrent('skips blank text fields', () => {
|
||||
const result = mapTags({ candidateName: ' ', company: '', source: ' ' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('maps applicationCount of zero', () => {
|
||||
const result = mapTags({ applicationCount: 0 })
|
||||
expect(result).toEqual({ applicationCount: 0 })
|
||||
})
|
||||
|
||||
it.concurrent('skips applicationCount when string', () => {
|
||||
const result = mapTags({ applicationCount: '3' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips updatedAt when date is invalid', () => {
|
||||
const result = mapTags({ updatedAt: 'not-a-date' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips lastActivity when date is invalid', () => {
|
||||
const result = mapTags({ lastActivity: 'garbage' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
})
|
||||
|
||||
describe('Fathom mapTags', () => {
|
||||
const mapTags = fathomConnector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
title: 'Sales Call',
|
||||
recordedByEmail: 'john@example.com',
|
||||
recordedByName: 'John Smith',
|
||||
team: 'Sales',
|
||||
meetingType: 'external',
|
||||
transcriptLanguage: 'en',
|
||||
durationSeconds: 1800,
|
||||
meetingDate: ISO_DATE,
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
title: 'Sales Call',
|
||||
recordedByEmail: 'john@example.com',
|
||||
recordedByName: 'John Smith',
|
||||
team: 'Sales',
|
||||
meetingType: 'external',
|
||||
transcriptLanguage: 'en',
|
||||
durationSeconds: 1800,
|
||||
meetingDate: new Date(ISO_DATE),
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
title: 123,
|
||||
recordedByEmail: null,
|
||||
recordedByName: true,
|
||||
team: [],
|
||||
meetingType: false,
|
||||
transcriptLanguage: 99999,
|
||||
durationSeconds: 'not-a-number',
|
||||
meetingDate: 12345,
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips blank string fields', () => {
|
||||
const result = mapTags({ title: ' ', team: '', transcriptLanguage: ' ' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('converts string durationSeconds to number', () => {
|
||||
const result = mapTags({ durationSeconds: '900' })
|
||||
expect(result).toEqual({ durationSeconds: 900 })
|
||||
})
|
||||
|
||||
it.concurrent('maps durationSeconds of zero', () => {
|
||||
const result = mapTags({ durationSeconds: 0 })
|
||||
expect(result).toEqual({ durationSeconds: 0 })
|
||||
})
|
||||
|
||||
it.concurrent('skips meetingDate when date is invalid', () => {
|
||||
const result = mapTags({ meetingDate: 'not-a-date' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
})
|
||||
|
||||
describe('Rootly mapTags', () => {
|
||||
const mapTags = rootlyConnector.mapTags!
|
||||
|
||||
it.concurrent('maps all fields when present', () => {
|
||||
const result = mapTags({
|
||||
status: 'resolved',
|
||||
severityName: 'SEV1',
|
||||
kind: 'incident',
|
||||
services: ['api', 'web'],
|
||||
teams: ['platform'],
|
||||
environments: ['production'],
|
||||
labels: ['platform:osx'],
|
||||
incidentDate: ISO_DATE,
|
||||
resolvedDate: '2025-01-01T00:00:00.000Z',
|
||||
})
|
||||
|
||||
expect(result).toEqual({
|
||||
status: 'resolved',
|
||||
severity: 'SEV1',
|
||||
kind: 'incident',
|
||||
services: 'api, web',
|
||||
teams: 'platform',
|
||||
environments: 'production',
|
||||
labels: 'platform:osx',
|
||||
incidentDate: new Date(ISO_DATE),
|
||||
resolvedDate: new Date('2025-01-01T00:00:00.000Z'),
|
||||
})
|
||||
})
|
||||
|
||||
it.concurrent('returns empty object for empty metadata', () => {
|
||||
expect(mapTags({})).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips fields with wrong types', () => {
|
||||
const result = mapTags({
|
||||
status: 123,
|
||||
severityName: null,
|
||||
severityLevel: true,
|
||||
kind: [],
|
||||
services: 'not-an-array',
|
||||
teams: 'not-an-array',
|
||||
environments: 'not-an-array',
|
||||
labels: 'not-an-array',
|
||||
incidentDate: 99999,
|
||||
resolvedDate: false,
|
||||
})
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('falls back to severityLevel when severityName is absent', () => {
|
||||
const result = mapTags({ severityLevel: 'sev0' })
|
||||
expect(result).toEqual({ severity: 'sev0' })
|
||||
})
|
||||
|
||||
it.concurrent('prefers severityName over severityLevel', () => {
|
||||
const result = mapTags({ severityName: 'Critical', severityLevel: 'sev0' })
|
||||
expect(result).toEqual({ severity: 'Critical' })
|
||||
})
|
||||
|
||||
it.concurrent('skips array fields when empty', () => {
|
||||
const result = mapTags({ services: [], teams: [], environments: [], labels: [] })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips incidentDate when date is invalid', () => {
|
||||
const result = mapTags({ incidentDate: 'not-a-date' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
|
||||
it.concurrent('skips resolvedDate when date is invalid', () => {
|
||||
const result = mapTags({ resolvedDate: 'garbage' })
|
||||
expect(result).toEqual({})
|
||||
})
|
||||
})
|
||||
|
||||
@@ -1,6 +1,7 @@
|
||||
import { airtableConnector } from '@/connectors/airtable'
|
||||
import { asanaConnector } from '@/connectors/asana'
|
||||
import { ashbyConnector } from '@/connectors/ashby'
|
||||
import { azureDevopsConnector } from '@/connectors/azure-devops'
|
||||
import { confluenceConnector } from '@/connectors/confluence'
|
||||
import { discordConnector } from '@/connectors/discord'
|
||||
import { docusignConnector } from '@/connectors/docusign'
|
||||
@@ -15,6 +16,7 @@ import { gongConnector } from '@/connectors/gong'
|
||||
import { googleCalendarConnector } from '@/connectors/google-calendar'
|
||||
import { googleDocsConnector } from '@/connectors/google-docs'
|
||||
import { googleDriveConnector } from '@/connectors/google-drive'
|
||||
import { googleFormsConnector } from '@/connectors/google-forms'
|
||||
import { googleSheetsConnector } from '@/connectors/google-sheets'
|
||||
import { grainConnector } from '@/connectors/grain'
|
||||
import { granolaConnector } from '@/connectors/granola'
|
||||
@@ -23,6 +25,7 @@ import { hubspotConnector } from '@/connectors/hubspot'
|
||||
import { incidentioConnector } from '@/connectors/incidentio'
|
||||
import { intercomConnector } from '@/connectors/intercom'
|
||||
import { jiraConnector } from '@/connectors/jira'
|
||||
import { jsmConnector } from '@/connectors/jsm'
|
||||
import { linearConnector } from '@/connectors/linear'
|
||||
import { microsoftTeamsConnector } from '@/connectors/microsoft-teams'
|
||||
import { mondayConnector } from '@/connectors/monday'
|
||||
@@ -32,13 +35,18 @@ import { onedriveConnector } from '@/connectors/onedrive'
|
||||
import { outlookConnector } from '@/connectors/outlook'
|
||||
import { redditConnector } from '@/connectors/reddit'
|
||||
import { rootlyConnector } from '@/connectors/rootly'
|
||||
import { s3Connector } from '@/connectors/s3'
|
||||
import { salesforceConnector } from '@/connectors/salesforce'
|
||||
import { sentryConnector } from '@/connectors/sentry'
|
||||
import { servicenowConnector } from '@/connectors/servicenow'
|
||||
import { sharepointConnector } from '@/connectors/sharepoint'
|
||||
import { slackConnector } from '@/connectors/slack'
|
||||
import { typeformConnector } from '@/connectors/typeform'
|
||||
import type { ConnectorRegistry } from '@/connectors/types'
|
||||
import { webflowConnector } from '@/connectors/webflow'
|
||||
import { wordpressConnector } from '@/connectors/wordpress'
|
||||
import { xConnector } from '@/connectors/x'
|
||||
import { youtubeConnector } from '@/connectors/youtube'
|
||||
import { zendeskConnector } from '@/connectors/zendesk'
|
||||
import { zoomConnector } from '@/connectors/zoom'
|
||||
|
||||
@@ -46,6 +54,7 @@ export const CONNECTOR_REGISTRY: ConnectorRegistry = {
|
||||
airtable: airtableConnector,
|
||||
asana: asanaConnector,
|
||||
ashby: ashbyConnector,
|
||||
azure_devops: azureDevopsConnector,
|
||||
confluence: confluenceConnector,
|
||||
discord: discordConnector,
|
||||
docusign: docusignConnector,
|
||||
@@ -60,6 +69,7 @@ export const CONNECTOR_REGISTRY: ConnectorRegistry = {
|
||||
google_calendar: googleCalendarConnector,
|
||||
google_docs: googleDocsConnector,
|
||||
google_drive: googleDriveConnector,
|
||||
google_forms: googleFormsConnector,
|
||||
google_sheets: googleSheetsConnector,
|
||||
grain: grainConnector,
|
||||
granola: granolaConnector,
|
||||
@@ -68,6 +78,7 @@ export const CONNECTOR_REGISTRY: ConnectorRegistry = {
|
||||
incidentio: incidentioConnector,
|
||||
intercom: intercomConnector,
|
||||
jira: jiraConnector,
|
||||
jsm: jsmConnector,
|
||||
linear: linearConnector,
|
||||
microsoft_teams: microsoftTeamsConnector,
|
||||
monday: mondayConnector,
|
||||
@@ -77,12 +88,17 @@ export const CONNECTOR_REGISTRY: ConnectorRegistry = {
|
||||
outlook: outlookConnector,
|
||||
reddit: redditConnector,
|
||||
rootly: rootlyConnector,
|
||||
s3: s3Connector,
|
||||
salesforce: salesforceConnector,
|
||||
sentry: sentryConnector,
|
||||
servicenow: servicenowConnector,
|
||||
sharepoint: sharepointConnector,
|
||||
slack: slackConnector,
|
||||
typeform: typeformConnector,
|
||||
webflow: webflowConnector,
|
||||
wordpress: wordpressConnector,
|
||||
x: xConnector,
|
||||
youtube: youtubeConnector,
|
||||
zendesk: zendeskConnector,
|
||||
zoom: zoomConnector,
|
||||
}
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
export { s3Connector } from '@/connectors/s3/s3'
|
||||
@@ -0,0 +1,730 @@
|
||||
import crypto from 'crypto'
|
||||
import { createLogger } from '@sim/logger'
|
||||
import { getErrorMessage, toError } from '@sim/utils/errors'
|
||||
import { S3Icon } from '@/components/icons'
|
||||
import { fetchWithRetry, VALIDATE_RETRY_OPTIONS } from '@/lib/knowledge/documents/utils'
|
||||
import type { ConnectorConfig, ExternalDocument, ExternalDocumentList } from '@/connectors/types'
|
||||
import { parseTagDate, readBodyWithLimit } from '@/connectors/utils'
|
||||
import { encodeS3PathComponent, getSignatureKey } from '@/tools/s3/utils'
|
||||
|
||||
const logger = createLogger('S3Connector')
|
||||
|
||||
/** Maximum object size to sync. Larger objects are skipped during listing. */
|
||||
const MAX_FILE_SIZE = 10 * 1024 * 1024 // 10 MB
|
||||
|
||||
/** Number of objects requested per ListObjectsV2 page (S3 caps at 1000). */
|
||||
const LIST_MAX_KEYS = 1000
|
||||
|
||||
/**
|
||||
* Default set of file extensions considered safely text-extractable. Objects
|
||||
* with any other extension (or no extension) are skipped, since their content
|
||||
* cannot be reliably decoded to plain text. Users can override this list via
|
||||
* the `extensions` config field.
|
||||
*/
|
||||
const DEFAULT_EXTENSIONS = new Set([
|
||||
'txt',
|
||||
'md',
|
||||
'markdown',
|
||||
'csv',
|
||||
'tsv',
|
||||
'json',
|
||||
'jsonl',
|
||||
'ndjson',
|
||||
'html',
|
||||
'htm',
|
||||
'xml',
|
||||
'yaml',
|
||||
'yml',
|
||||
'log',
|
||||
'rtf',
|
||||
])
|
||||
|
||||
/**
|
||||
* A single object entry parsed out of a ListObjectsV2 XML response.
|
||||
*/
|
||||
interface S3ObjectEntry {
|
||||
key: string
|
||||
etag: string
|
||||
lastModified: string
|
||||
size: number
|
||||
}
|
||||
|
||||
/**
|
||||
* A parsed custom S3-compatible endpoint (Cloudflare R2, MinIO, etc.).
|
||||
*
|
||||
* `host` is the bare hostname, `hostHeader` is the value used both as the wire
|
||||
* `Host` header and in the SigV4 canonical headers — it includes the port when
|
||||
* a non-default port is configured (e.g. `localhost:9000`). When the endpoint
|
||||
* uses the scheme's default port (443 for https, 80 for http) the port is
|
||||
* omitted from `hostHeader`, matching what the HTTP client sends on the wire.
|
||||
*/
|
||||
interface S3Endpoint {
|
||||
scheme: 'http' | 'https'
|
||||
host: string
|
||||
hostHeader: string
|
||||
}
|
||||
|
||||
/**
|
||||
* AWS credentials and target resource resolved from sourceConfig + access token.
|
||||
*
|
||||
* When `endpoint` is present the connector targets an S3-compatible store using
|
||||
* path-style addressing (`{endpoint}/{bucket}/{key}`). When absent it targets
|
||||
* AWS S3 using virtual-hosted-style addressing
|
||||
* (`{bucket}.s3.{region}.amazonaws.com`), preserving the original behavior.
|
||||
*/
|
||||
interface S3Context {
|
||||
accessKeyId: string
|
||||
secretAccessKey: string
|
||||
region: string
|
||||
bucket: string
|
||||
endpoint?: S3Endpoint
|
||||
}
|
||||
|
||||
/**
|
||||
* Parses the comma-separated `extensions` config override into a normalized set
|
||||
* (lowercased, no leading dot). Returns the built-in default set when the
|
||||
* override is empty or contains no usable entries.
|
||||
*/
|
||||
function resolveExtensions(raw: unknown): Set<string> {
|
||||
if (typeof raw !== 'string') return DEFAULT_EXTENSIONS
|
||||
const exts = raw
|
||||
.split(',')
|
||||
.map((e) => e.trim().toLowerCase().replace(/^\./, ''))
|
||||
.filter(Boolean)
|
||||
return exts.length > 0 ? new Set(exts) : DEFAULT_EXTENSIONS
|
||||
}
|
||||
|
||||
/**
|
||||
* Extracts the lowercased file extension from an object key, or '' if none.
|
||||
*/
|
||||
function getExtension(key: string): string {
|
||||
const lastSegment = key.split('/').pop() ?? ''
|
||||
const dotIndex = lastSegment.lastIndexOf('.')
|
||||
if (dotIndex <= 0 || dotIndex === lastSegment.length - 1) return ''
|
||||
return lastSegment.slice(dotIndex + 1).toLowerCase()
|
||||
}
|
||||
|
||||
/**
|
||||
* Returns true when the object key ends in one of the allowed text extensions.
|
||||
*/
|
||||
function isSupportedKey(key: string, allowedExtensions: Set<string>): boolean {
|
||||
return allowedExtensions.has(getExtension(key))
|
||||
}
|
||||
|
||||
/**
|
||||
* Returns true when the host is a loopback address for which plain `http://`
|
||||
* is tolerated (local MinIO development). Any other host must use `https://` so
|
||||
* that credentials are never transmitted over cleartext.
|
||||
*/
|
||||
function isLoopbackHost(host: string): boolean {
|
||||
const bare = host.replace(/^\[|\]$/g, '')
|
||||
return bare === 'localhost' || bare === '127.0.0.1' || bare === '::1'
|
||||
}
|
||||
|
||||
/**
|
||||
* Parses and validates a custom S3-compatible endpoint string.
|
||||
*
|
||||
* Accepts a full origin such as `https://accountid.r2.cloudflarestorage.com` or
|
||||
* `http://localhost:9000`. Trailing slashes are stripped. Throws when the value
|
||||
* is not a valid URL, carries a path/query/fragment beyond `/` (which would
|
||||
* corrupt the path-style canonical URI), or uses plain `http://` against a
|
||||
* non-loopback host.
|
||||
*
|
||||
* The returned `hostHeader` includes the port only when it differs from the
|
||||
* scheme default, matching the `Host` header the HTTP client emits — this keeps
|
||||
* the SigV4 canonical Host byte-identical to the wire Host.
|
||||
*/
|
||||
function parseEndpoint(raw: string): S3Endpoint {
|
||||
let url: URL
|
||||
try {
|
||||
url = new URL(raw)
|
||||
} catch {
|
||||
throw new Error('Endpoint must be a valid URL, e.g. https://accountid.r2.cloudflarestorage.com')
|
||||
}
|
||||
|
||||
if (url.protocol !== 'https:' && url.protocol !== 'http:') {
|
||||
throw new Error('Endpoint must use http:// or https://')
|
||||
}
|
||||
const scheme = url.protocol === 'https:' ? 'https' : 'http'
|
||||
|
||||
if (url.username || url.password) {
|
||||
throw new Error('Endpoint must not contain credentials')
|
||||
}
|
||||
if (url.search || url.hash) {
|
||||
throw new Error('Endpoint must not contain a query string or fragment')
|
||||
}
|
||||
const path = url.pathname.replace(/\/+$/, '')
|
||||
if (path !== '') {
|
||||
throw new Error('Endpoint must not contain a path — provide only the host, e.g. https://host')
|
||||
}
|
||||
|
||||
const host = url.hostname
|
||||
if (!host) throw new Error('Endpoint is missing a host')
|
||||
if (scheme === 'http' && !isLoopbackHost(host)) {
|
||||
throw new Error(
|
||||
'Plain http:// endpoints are only allowed for localhost — use https:// otherwise'
|
||||
)
|
||||
}
|
||||
|
||||
const defaultPort = scheme === 'https' ? '443' : '80'
|
||||
const port = url.port && url.port !== defaultPort ? url.port : ''
|
||||
const hostHeader = port ? `${host}:${port}` : host
|
||||
|
||||
return { scheme, host, hostHeader }
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves AWS credentials and the target bucket from the connector's
|
||||
* sourceConfig and the encrypted secret (delivered as accessToken). When an
|
||||
* `endpoint` is configured it is parsed/validated into an {@link S3Endpoint} so
|
||||
* the connector targets an S3-compatible store via path-style addressing.
|
||||
*/
|
||||
function resolveContext(accessToken: string, sourceConfig: Record<string, unknown>): S3Context {
|
||||
const accessKeyId = ((sourceConfig.accessKeyId as string) ?? '').trim()
|
||||
const region = ((sourceConfig.region as string) ?? '').trim()
|
||||
const bucket = ((sourceConfig.bucket as string) ?? '').trim()
|
||||
const secretAccessKey = (accessToken ?? '').trim()
|
||||
const rawEndpoint = ((sourceConfig.endpoint as string) ?? '').trim()
|
||||
|
||||
if (!accessKeyId) throw new Error('Missing AWS Access Key ID')
|
||||
if (!secretAccessKey) throw new Error('Missing AWS Secret Access Key')
|
||||
if (!region) throw new Error('Missing AWS region')
|
||||
if (!bucket) throw new Error('Missing S3 bucket name')
|
||||
|
||||
const endpoint = rawEndpoint ? parseEndpoint(rawEndpoint) : undefined
|
||||
|
||||
return { accessKeyId, secretAccessKey, region, bucket, endpoint }
|
||||
}
|
||||
|
||||
/**
|
||||
* Returns the SigV4 canonical Host header for the request. For AWS this is the
|
||||
* virtual-hosted-style host; for a custom endpoint it is the endpoint host
|
||||
* (with port when non-default).
|
||||
*/
|
||||
function resolveHost(ctx: S3Context): string {
|
||||
return ctx.endpoint ? ctx.endpoint.hostHeader : `${ctx.bucket}.s3.${ctx.region}.amazonaws.com`
|
||||
}
|
||||
|
||||
/**
|
||||
* Returns the request scheme: always `https` for AWS, or the endpoint scheme
|
||||
* (which may be `http` for local MinIO) for a custom endpoint.
|
||||
*/
|
||||
function resolveScheme(ctx: S3Context): string {
|
||||
return ctx.endpoint ? ctx.endpoint.scheme : 'https'
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the canonical URI for an object key.
|
||||
*
|
||||
* AWS (virtual-hosted-style): `/{key}` — the bucket lives in the host.
|
||||
* Custom endpoint (path-style): `/{bucket}/{key}` — the bucket is the first
|
||||
* path segment. Both the bucket and key are percent-encoded per AWS UriEncode
|
||||
* rules while preserving `/` separators via {@link encodeS3PathComponent}.
|
||||
*/
|
||||
function buildObjectPath(ctx: S3Context, key: string): string {
|
||||
const encodedKey = encodeS3PathComponent(key)
|
||||
return ctx.endpoint ? `/${encodeS3PathComponent(ctx.bucket)}/${encodedKey}` : `/${encodedKey}`
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the canonical URI for a bucket-level (ListObjectsV2) request.
|
||||
*
|
||||
* AWS (virtual-hosted-style): `/`.
|
||||
* Custom endpoint (path-style): `/{bucket}/`.
|
||||
*/
|
||||
function buildBucketPath(ctx: S3Context): string {
|
||||
return ctx.endpoint ? `/${encodeS3PathComponent(ctx.bucket)}/` : '/'
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the full request URL from the canonical path and an optional canonical
|
||||
* query string. The path passed here is the same canonical, percent-encoded
|
||||
* string used to compute the SigV4 signature, so the signed URI and the wire
|
||||
* URI are byte-identical.
|
||||
*/
|
||||
function buildUrl(ctx: S3Context, encodedPath: string, canonicalQueryString: string): string {
|
||||
const base = `${resolveScheme(ctx)}://${resolveHost(ctx)}${encodedPath}`
|
||||
return canonicalQueryString ? `${base}?${canonicalQueryString}` : base
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds SigV4 request headers for an S3 REST call.
|
||||
*
|
||||
* `canonicalQueryString` must be the already-sorted, percent-encoded query
|
||||
* string (empty for GetObject) — the caller builds the request URL from this
|
||||
* exact same string so the signed query and the wire query are byte-identical
|
||||
* (the classic continuation-token signing mismatch cannot occur here).
|
||||
* `encodedPath` is the canonical URI path starting with '/' (virtual-hosted
|
||||
* `/{key}` for AWS, path-style `/{bucket}/{key}` for custom endpoints). The
|
||||
* canonical Host header is resolved via {@link resolveHost} and includes the
|
||||
* port for non-default custom-endpoint ports, exactly matching the wire Host.
|
||||
* Reuses {@link getSignatureKey} from the s3 tool utilities.
|
||||
*
|
||||
* The signed headers embed `x-amz-date` and are reused verbatim across
|
||||
* `fetchWithRetry` attempts. S3 allows a 15-minute clock-skew window; the
|
||||
* retry helper's worst-case total backoff (~31s default, ~10s in validate) is
|
||||
* far inside that window, so a stale timestamp never triggers
|
||||
* RequestTimeTooSkewed.
|
||||
*/
|
||||
function buildSignedHeaders(
|
||||
ctx: S3Context,
|
||||
method: 'GET',
|
||||
encodedPath: string,
|
||||
canonicalQueryString: string
|
||||
): Record<string, string> {
|
||||
const date = new Date()
|
||||
const amzDate = date.toISOString().replace(/[:-]|\.\d{3}/g, '')
|
||||
const dateStamp = amzDate.slice(0, 8)
|
||||
|
||||
const host = resolveHost(ctx)
|
||||
const payloadHash = crypto.createHash('sha256').update('').digest('hex')
|
||||
|
||||
const canonicalHeaders =
|
||||
`host:${host}\n` + `x-amz-content-sha256:${payloadHash}\n` + `x-amz-date:${amzDate}\n`
|
||||
const signedHeaders = 'host;x-amz-content-sha256;x-amz-date'
|
||||
|
||||
const canonicalRequest = `${method}\n${encodedPath}\n${canonicalQueryString}\n${canonicalHeaders}\n${signedHeaders}\n${payloadHash}`
|
||||
|
||||
const algorithm = 'AWS4-HMAC-SHA256'
|
||||
const credentialScope = `${dateStamp}/${ctx.region}/s3/aws4_request`
|
||||
const stringToSign = `${algorithm}\n${amzDate}\n${credentialScope}\n${crypto
|
||||
.createHash('sha256')
|
||||
.update(canonicalRequest)
|
||||
.digest('hex')}`
|
||||
|
||||
const signingKey = getSignatureKey(ctx.secretAccessKey, dateStamp, ctx.region, 's3')
|
||||
const signature = crypto.createHmac('sha256', signingKey).update(stringToSign).digest('hex')
|
||||
|
||||
const authorizationHeader = `${algorithm} Credential=${ctx.accessKeyId}/${credentialScope}, SignedHeaders=${signedHeaders}, Signature=${signature}`
|
||||
|
||||
return {
|
||||
Host: host,
|
||||
'X-Amz-Content-Sha256': payloadHash,
|
||||
'X-Amz-Date': amzDate,
|
||||
Authorization: authorizationHeader,
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Percent-encodes a query parameter name or value per AWS SigV4 canonical rules
|
||||
* (every byte except the unreserved set `A-Za-z0-9-_.~` is encoded).
|
||||
* `encodeURIComponent` leaves `!'()*` unencoded, so those are encoded here.
|
||||
*/
|
||||
function encodeQueryValue(value: string): string {
|
||||
return encodeURIComponent(value).replace(
|
||||
/[!'()*]/g,
|
||||
(c) => `%${c.charCodeAt(0).toString(16).toUpperCase()}`
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the canonical (sorted, percent-encoded) query string for a
|
||||
* ListObjectsV2 request. Keys are sorted lexicographically after encoding and
|
||||
* each name/value pair is encoded individually.
|
||||
*/
|
||||
function buildListQueryString(params: Record<string, string>): string {
|
||||
return Object.keys(params)
|
||||
.sort()
|
||||
.map((key) => `${encodeQueryValue(key)}=${encodeQueryValue(params[key])}`)
|
||||
.join('&')
|
||||
}
|
||||
|
||||
/**
|
||||
* Decodes XML entities found in S3 response text values. `&` is decoded
|
||||
* last so sequences like `&lt;` resolve to `<` rather than `<`.
|
||||
*/
|
||||
function decodeXmlEntities(value: string): string {
|
||||
return value
|
||||
.replace(/</g, '<')
|
||||
.replace(/>/g, '>')
|
||||
.replace(/"/g, '"')
|
||||
.replace(/'/g, "'")
|
||||
.replace(/&/g, '&')
|
||||
}
|
||||
|
||||
/**
|
||||
* Normalizes an ETag from either a ListObjectsV2 XML `<ETag>` element or a
|
||||
* GetObject response header into a stable bare token used in the content hash.
|
||||
*
|
||||
* Strips surrounding double quotes and a leading weak-validator prefix (`W/`).
|
||||
* AWS S3 always returns strong, quoted ETags (including the multipart `-N`
|
||||
* suffix) identically from List and Get, but S3-compatible stores (MinIO, R2)
|
||||
* are not contractually bound to that and could emit a weak ETag on one path
|
||||
* and a strong one on the other. Normalizing both ends keeps the
|
||||
* `s3:{key}:{etag}` hash invariant between the listing stub and the hydrated
|
||||
* document so unchanged objects are not re-uploaded every sync.
|
||||
*/
|
||||
function normalizeEtag(raw: string): string {
|
||||
return raw.replace(/^W\//, '').replace(/"/g, '')
|
||||
}
|
||||
|
||||
/**
|
||||
* Decodes a URL-encoded object key returned when `encoding-type=url` is set.
|
||||
* Falls back to the raw value if decoding fails (malformed percent sequence).
|
||||
*/
|
||||
function decodeObjectKey(value: string): string {
|
||||
try {
|
||||
return decodeURIComponent(value)
|
||||
} catch {
|
||||
return value
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Extracts the text content of the first matching XML tag within a fragment.
|
||||
*/
|
||||
function extractTag(fragment: string, tag: string): string | undefined {
|
||||
const match = fragment.match(new RegExp(`<${tag}>([\\s\\S]*?)</${tag}>`))
|
||||
return match ? decodeXmlEntities(match[1]) : undefined
|
||||
}
|
||||
|
||||
/**
|
||||
* Parses a ListObjectsV2 XML response into object entries plus pagination state.
|
||||
*
|
||||
* The request is always made with `encoding-type=url`, so the per-`Key` values
|
||||
* are percent-encoded in the XML (safe for the regex parser even when keys
|
||||
* contain XML-hostile bytes such as `&`, `<`, or ASCII control characters).
|
||||
* Each `Key` is XML-entity-decoded then URL-decoded back to its true value.
|
||||
* `NextContinuationToken` is opaque and is not affected by `encoding-type`, so
|
||||
* it is used verbatim.
|
||||
*/
|
||||
function parseListResponse(xml: string): {
|
||||
objects: S3ObjectEntry[]
|
||||
isTruncated: boolean
|
||||
nextContinuationToken?: string
|
||||
} {
|
||||
const objects: S3ObjectEntry[] = []
|
||||
|
||||
for (const match of xml.matchAll(/<Contents>([\s\S]*?)<\/Contents>/g)) {
|
||||
const block = match[1]
|
||||
const rawKey = extractTag(block, 'Key')
|
||||
if (!rawKey) continue
|
||||
const key = decodeObjectKey(rawKey)
|
||||
|
||||
const etag = normalizeEtag(extractTag(block, 'ETag') ?? '')
|
||||
const lastModified = extractTag(block, 'LastModified') ?? ''
|
||||
const size = Number(extractTag(block, 'Size') ?? '0')
|
||||
|
||||
objects.push({ key, etag, lastModified, size: Number.isNaN(size) ? 0 : size })
|
||||
}
|
||||
|
||||
const isTruncated = extractTag(xml, 'IsTruncated') === 'true'
|
||||
const nextContinuationToken = extractTag(xml, 'NextContinuationToken')
|
||||
|
||||
return { objects, isTruncated, nextContinuationToken }
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds a metadata stub for an S3 object. The content hash combines the key
|
||||
* and ETag — S3's ETag changes whenever object content changes, making it an
|
||||
* ideal change indicator. Used by both listDocuments and getDocument to
|
||||
* guarantee identical hashes.
|
||||
*/
|
||||
function objectToStub(ctx: S3Context, entry: S3ObjectEntry): ExternalDocument {
|
||||
const title = entry.key.split('/').pop() || entry.key
|
||||
const prefix = entry.key.includes('/') ? entry.key.slice(0, entry.key.lastIndexOf('/')) : ''
|
||||
|
||||
return {
|
||||
externalId: entry.key,
|
||||
title,
|
||||
content: '',
|
||||
contentDeferred: true,
|
||||
mimeType: 'text/plain',
|
||||
sourceUrl: buildUrl(ctx, buildObjectPath(ctx, entry.key), ''),
|
||||
contentHash: `s3:${entry.key}:${entry.etag}`,
|
||||
metadata: {
|
||||
key: entry.key,
|
||||
prefix,
|
||||
etag: entry.etag,
|
||||
lastModified: entry.lastModified,
|
||||
fileSize: entry.size,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Performs a single ListObjectsV2 page request and returns the parsed result.
|
||||
*/
|
||||
async function listObjectsPage(
|
||||
ctx: S3Context,
|
||||
prefix: string,
|
||||
continuationToken: string | undefined,
|
||||
retryOptions?: Parameters<typeof fetchWithRetry>[2],
|
||||
maxKeys: number = LIST_MAX_KEYS
|
||||
): Promise<{ objects: S3ObjectEntry[]; isTruncated: boolean; nextContinuationToken?: string }> {
|
||||
const queryParams: Record<string, string> = {
|
||||
'list-type': '2',
|
||||
'encoding-type': 'url',
|
||||
'max-keys': String(maxKeys),
|
||||
}
|
||||
if (prefix) queryParams.prefix = prefix
|
||||
if (continuationToken) queryParams['continuation-token'] = continuationToken
|
||||
|
||||
const canonicalQueryString = buildListQueryString(queryParams)
|
||||
const bucketPath = buildBucketPath(ctx)
|
||||
const headers = buildSignedHeaders(ctx, 'GET', bucketPath, canonicalQueryString)
|
||||
|
||||
const url = buildUrl(ctx, bucketPath, canonicalQueryString)
|
||||
|
||||
const response = await fetchWithRetry(url, { method: 'GET', headers }, retryOptions)
|
||||
|
||||
if (!response.ok) {
|
||||
const errorText = await response.text()
|
||||
throw new Error(`S3 ListObjectsV2 failed: ${response.status} ${errorText}`)
|
||||
}
|
||||
|
||||
const xml = await response.text()
|
||||
return parseListResponse(xml)
|
||||
}
|
||||
|
||||
export const s3Connector: ConnectorConfig = {
|
||||
id: 's3',
|
||||
name: 'Amazon S3',
|
||||
description:
|
||||
'Sync text-based objects from Amazon S3 or any S3-compatible store (Cloudflare R2, MinIO) into your knowledge base',
|
||||
version: '1.1.0',
|
||||
icon: S3Icon,
|
||||
|
||||
auth: {
|
||||
mode: 'apiKey',
|
||||
label: 'Secret Access Key',
|
||||
placeholder: 'Enter your AWS Secret Access Key',
|
||||
},
|
||||
|
||||
configFields: [
|
||||
{
|
||||
id: 'accessKeyId',
|
||||
title: 'Access Key ID',
|
||||
type: 'short-input',
|
||||
placeholder: 'e.g. AKIAIOSFODNN7EXAMPLE',
|
||||
required: true,
|
||||
},
|
||||
{
|
||||
id: 'region',
|
||||
title: 'Region',
|
||||
type: 'short-input',
|
||||
placeholder: 'e.g. us-east-1 (use auto for Cloudflare R2)',
|
||||
required: true,
|
||||
description:
|
||||
'AWS region for the bucket. For Cloudflare R2 use "auto"; for MinIO use the region the server is configured with (commonly us-east-1).',
|
||||
},
|
||||
{
|
||||
id: 'bucket',
|
||||
title: 'Bucket',
|
||||
type: 'short-input',
|
||||
placeholder: 'e.g. my-bucket',
|
||||
required: true,
|
||||
},
|
||||
{
|
||||
id: 'endpoint',
|
||||
title: 'Custom Endpoint',
|
||||
type: 'short-input',
|
||||
placeholder: 'https://accountid.r2.cloudflarestorage.com (optional — leave empty for AWS S3)',
|
||||
required: false,
|
||||
description:
|
||||
'S3-compatible endpoint for Cloudflare R2, MinIO, etc. Leave empty for AWS S3. Uses path-style addressing. Plain http:// is only allowed for localhost.',
|
||||
},
|
||||
{
|
||||
id: 'prefix',
|
||||
title: 'Prefix',
|
||||
type: 'short-input',
|
||||
placeholder: 'e.g. docs/ (optional)',
|
||||
required: false,
|
||||
description: 'Only sync objects whose key starts with this prefix',
|
||||
},
|
||||
{
|
||||
id: 'extensions',
|
||||
title: 'File Extensions',
|
||||
type: 'short-input',
|
||||
placeholder: 'e.g. txt, md, csv (optional)',
|
||||
required: false,
|
||||
description:
|
||||
'Comma-separated list of file extensions to sync. Leave blank to use the built-in text formats.',
|
||||
},
|
||||
{
|
||||
id: 'maxObjects',
|
||||
title: 'Max Objects',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
placeholder: 'e.g. 500 (default: unlimited)',
|
||||
description: 'Stop syncing after this many objects',
|
||||
},
|
||||
],
|
||||
|
||||
listDocuments: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
cursor?: string,
|
||||
syncContext?: Record<string, unknown>
|
||||
): Promise<ExternalDocumentList> => {
|
||||
const ctx = resolveContext(accessToken, sourceConfig)
|
||||
const prefix = ((sourceConfig.prefix as string) ?? '').trim()
|
||||
const allowedExtensions = resolveExtensions(sourceConfig.extensions)
|
||||
|
||||
const maxObjects = sourceConfig.maxObjects ? Number(sourceConfig.maxObjects) : 0
|
||||
const previouslyFetched = (syncContext?.totalDocsFetched as number) ?? 0
|
||||
|
||||
if (maxObjects > 0 && previouslyFetched >= maxObjects) {
|
||||
return { documents: [], hasMore: false }
|
||||
}
|
||||
|
||||
logger.info('Listing S3 objects', { bucket: ctx.bucket, prefix, cursor: cursor ?? 'initial' })
|
||||
|
||||
const { objects, isTruncated, nextContinuationToken } = await listObjectsPage(
|
||||
ctx,
|
||||
prefix,
|
||||
cursor
|
||||
)
|
||||
|
||||
let documents = objects
|
||||
.filter((entry) => isSupportedKey(entry.key, allowedExtensions))
|
||||
.filter((entry) => entry.size > 0 && entry.size <= MAX_FILE_SIZE)
|
||||
.map((entry) => objectToStub(ctx, entry))
|
||||
|
||||
let slicedSome = false
|
||||
if (maxObjects > 0) {
|
||||
const remaining = maxObjects - previouslyFetched
|
||||
if (documents.length > remaining) {
|
||||
slicedSome = true
|
||||
documents = documents.slice(0, remaining)
|
||||
}
|
||||
}
|
||||
|
||||
const totalFetched = previouslyFetched + documents.length
|
||||
if (syncContext) syncContext.totalDocsFetched = totalFetched
|
||||
const hitLimit = maxObjects > 0 && totalFetched >= maxObjects
|
||||
const moreAvailable = slicedSome || (isTruncated && Boolean(nextContinuationToken))
|
||||
if (hitLimit && moreAvailable && syncContext) syncContext.listingCapped = true
|
||||
|
||||
return {
|
||||
documents,
|
||||
nextCursor: hitLimit ? undefined : isTruncated ? nextContinuationToken : undefined,
|
||||
hasMore: hitLimit ? false : isTruncated && Boolean(nextContinuationToken),
|
||||
}
|
||||
},
|
||||
|
||||
getDocument: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
externalId: string
|
||||
): Promise<ExternalDocument | null> => {
|
||||
const ctx = resolveContext(accessToken, sourceConfig)
|
||||
const key = externalId
|
||||
|
||||
try {
|
||||
const encodedPath = buildObjectPath(ctx, key)
|
||||
const headers = buildSignedHeaders(ctx, 'GET', encodedPath, '')
|
||||
const url = buildUrl(ctx, encodedPath, '')
|
||||
|
||||
const response = await fetchWithRetry(url, { method: 'GET', headers })
|
||||
|
||||
if (response.status === 404) return null
|
||||
if (!response.ok) {
|
||||
const errorText = await response.text()
|
||||
throw new Error(`S3 GetObject failed: ${response.status} ${errorText}`)
|
||||
}
|
||||
|
||||
const etag = normalizeEtag(response.headers.get('etag') ?? '')
|
||||
const lastModified = response.headers.get('last-modified') ?? ''
|
||||
const declaredLength = Number(response.headers.get('content-length') ?? '')
|
||||
|
||||
if (declaredLength > MAX_FILE_SIZE) {
|
||||
logger.warn('Skipping oversized S3 object', { key, size: declaredLength })
|
||||
return null
|
||||
}
|
||||
|
||||
const body = await readBodyWithLimit(response, MAX_FILE_SIZE)
|
||||
if (body === null) {
|
||||
logger.warn('Skipping oversized S3 object (size cap exceeded while streaming)', { key })
|
||||
return null
|
||||
}
|
||||
const content = body.toString('utf-8')
|
||||
if (!content.trim()) return null
|
||||
|
||||
const entry: S3ObjectEntry = {
|
||||
key,
|
||||
etag,
|
||||
lastModified,
|
||||
size:
|
||||
Number.isNaN(declaredLength) || declaredLength <= 0 ? body.byteLength : declaredLength,
|
||||
}
|
||||
const stub = objectToStub(ctx, entry)
|
||||
return { ...stub, content, contentDeferred: false }
|
||||
} catch (error) {
|
||||
logger.warn('Failed to get S3 object', { key, error: toError(error).message })
|
||||
return null
|
||||
}
|
||||
},
|
||||
|
||||
validateConfig: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>
|
||||
): Promise<{ valid: boolean; error?: string }> => {
|
||||
let ctx: S3Context
|
||||
try {
|
||||
ctx = resolveContext(accessToken, sourceConfig)
|
||||
} catch (error) {
|
||||
return { valid: false, error: getErrorMessage(error, 'Invalid configuration') }
|
||||
}
|
||||
|
||||
const maxObjects = sourceConfig.maxObjects as string | undefined
|
||||
if (maxObjects && (Number.isNaN(Number(maxObjects)) || Number(maxObjects) <= 0)) {
|
||||
return { valid: false, error: 'Max objects must be a positive number' }
|
||||
}
|
||||
|
||||
const prefix = ((sourceConfig.prefix as string) ?? '').trim()
|
||||
|
||||
try {
|
||||
await listObjectsPage(ctx, prefix, undefined, VALIDATE_RETRY_OPTIONS, 1)
|
||||
return { valid: true }
|
||||
} catch (error) {
|
||||
const message = getErrorMessage(error, 'Failed to validate configuration')
|
||||
const lower = message.toLowerCase()
|
||||
if (
|
||||
lower.includes('permanentredirect') ||
|
||||
lower.includes('authorizationheadermalformed') ||
|
||||
lower.includes(' 301 ')
|
||||
) {
|
||||
return {
|
||||
valid: false,
|
||||
error:
|
||||
'Wrong region for this bucket. Update the region to match where the bucket lives (or use "auto" for Cloudflare R2).',
|
||||
}
|
||||
}
|
||||
if (lower.includes('403') || lower.includes('accessdenied') || lower.includes('signature')) {
|
||||
return {
|
||||
valid: false,
|
||||
error: 'Access denied. Check the access key, secret key, and bucket permissions.',
|
||||
}
|
||||
}
|
||||
if (lower.includes('404') || lower.includes('nosuchbucket')) {
|
||||
return { valid: false, error: 'Bucket not found. Check the bucket name and region.' }
|
||||
}
|
||||
return { valid: false, error: message }
|
||||
}
|
||||
},
|
||||
|
||||
tagDefinitions: [
|
||||
{ id: 'prefix', displayName: 'Folder', fieldType: 'text' },
|
||||
{ id: 'fileSize', displayName: 'Size (bytes)', fieldType: 'number' },
|
||||
{ id: 'lastModified', displayName: 'Last Modified', fieldType: 'date' },
|
||||
],
|
||||
|
||||
mapTags: (metadata: Record<string, unknown>): Record<string, unknown> => {
|
||||
const result: Record<string, unknown> = {}
|
||||
|
||||
if (typeof metadata.prefix === 'string' && metadata.prefix.length > 0) {
|
||||
result.prefix = metadata.prefix
|
||||
}
|
||||
|
||||
if (metadata.fileSize != null) {
|
||||
const num = Number(metadata.fileSize)
|
||||
if (!Number.isNaN(num)) result.fileSize = num
|
||||
}
|
||||
|
||||
const lastModified = parseTagDate(metadata.lastModified)
|
||||
if (lastModified) result.lastModified = lastModified
|
||||
|
||||
return result
|
||||
},
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
export { sentryConnector } from '@/connectors/sentry/sentry'
|
||||
@@ -0,0 +1,736 @@
|
||||
import { createLogger } from '@sim/logger'
|
||||
import { getErrorMessage, toError } from '@sim/utils/errors'
|
||||
import { SentryIcon } from '@/components/icons'
|
||||
import { fetchWithRetry, VALIDATE_RETRY_OPTIONS } from '@/lib/knowledge/documents/utils'
|
||||
import type { ConnectorConfig, ExternalDocument, ExternalDocumentList } from '@/connectors/types'
|
||||
import { parseTagDate } from '@/connectors/utils'
|
||||
|
||||
const logger = createLogger('SentryConnector')
|
||||
|
||||
const DEFAULT_HOST = 'sentry.io'
|
||||
const ISSUES_PER_PAGE = 100
|
||||
|
||||
/**
|
||||
* Default issue search query.
|
||||
*
|
||||
* Reconciliation semantics: the sync engine hard-deletes any previously-synced
|
||||
* document whose `externalId` is absent from a full (non-capped) listing pass.
|
||||
* With the default `is:unresolved` query this means an issue that is resolved,
|
||||
* ignored/muted, or aged out of the query window will fall out of the listing
|
||||
* and be removed from the knowledge base on the next full sync. That is the
|
||||
* intended semantic — the KB tracks the *currently matching* issue set, not a
|
||||
* permanent archive. Users who want resolved issues retained should widen the
|
||||
* query (e.g. drop `is:unresolved`). When `maxIssues` caps the listing, the
|
||||
* engine sets `listingCapped` and skips deletion, so capped runs never remove
|
||||
* unseen issues.
|
||||
*/
|
||||
const DEFAULT_QUERY = 'is:unresolved'
|
||||
|
||||
/**
|
||||
* Allowed `statsPeriod` values for the project issues list endpoint. Sentry's
|
||||
* project issues endpoint only honors `24h` (default) or `14d` for its timeline
|
||||
* stats; an empty value disables the stats window. Other periods (e.g. `90d`)
|
||||
* are accepted by the organization issues endpoint but not this one, so they are
|
||||
* rejected during validation to avoid a silently-ignored filter.
|
||||
*/
|
||||
const ALLOWED_STATS_PERIODS = new Set(['24h', '14d'])
|
||||
|
||||
/**
|
||||
* Metadata block on a Sentry issue, carrying the human-readable error type/value.
|
||||
*/
|
||||
interface SentryIssueMetadata {
|
||||
type?: string
|
||||
value?: string
|
||||
function?: string
|
||||
title?: string
|
||||
}
|
||||
|
||||
/**
|
||||
* A single issue (error group) returned by the issues list/detail endpoints.
|
||||
*/
|
||||
interface SentryIssue {
|
||||
id: string
|
||||
shortId?: string
|
||||
title?: string
|
||||
culprit?: string | null
|
||||
permalink?: string
|
||||
logger?: string | null
|
||||
level?: string
|
||||
status?: string
|
||||
platform?: string | null
|
||||
type?: string | null
|
||||
metadata?: SentryIssueMetadata
|
||||
/** Sentry returns the event count as a string (e.g. "12"), not a number. */
|
||||
count?: string
|
||||
userCount?: number
|
||||
firstSeen?: string
|
||||
lastSeen?: string
|
||||
}
|
||||
|
||||
/**
|
||||
* One entry inside a Sentry event. Entries carry the structured payload (exception,
|
||||
* breadcrumbs, request, message) keyed by `type`, with the shape under `data` varying
|
||||
* per entry type.
|
||||
*/
|
||||
interface SentryEventEntry {
|
||||
type?: string
|
||||
data?: unknown
|
||||
}
|
||||
|
||||
/**
|
||||
* A key/value tag pair attached to a Sentry event.
|
||||
*/
|
||||
interface SentryEventTag {
|
||||
key?: string
|
||||
value?: string
|
||||
}
|
||||
|
||||
/**
|
||||
* The latest event for an issue, used to enrich the synced document with the concrete
|
||||
* message, exception detail, and tags from the most recent occurrence.
|
||||
*/
|
||||
interface SentryEvent {
|
||||
id?: string
|
||||
eventID?: string
|
||||
message?: string
|
||||
title?: string
|
||||
culprit?: string | null
|
||||
platform?: string | null
|
||||
dateCreated?: string
|
||||
metadata?: SentryIssueMetadata
|
||||
entries?: SentryEventEntry[]
|
||||
tags?: SentryEventTag[]
|
||||
}
|
||||
|
||||
/**
|
||||
* The shape of an exception entry's `data` payload: a list of exception values, each
|
||||
* with a type, message, and an optional rendered stack frame list.
|
||||
*/
|
||||
interface SentryExceptionData {
|
||||
values?: {
|
||||
type?: string
|
||||
value?: string
|
||||
module?: string
|
||||
stacktrace?: {
|
||||
frames?: {
|
||||
filename?: string
|
||||
function?: string
|
||||
lineNo?: number
|
||||
module?: string
|
||||
}[]
|
||||
}
|
||||
}[]
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolved connector source configuration after normalization.
|
||||
*/
|
||||
interface SentrySourceConfig {
|
||||
/** Bare host (no protocol, no trailing slash), e.g. `sentry.io` or a self-hosted host. */
|
||||
host: string
|
||||
/** REST API base, e.g. `https://sentry.io/api/0`. */
|
||||
apiBase: string
|
||||
organization: string
|
||||
project: string
|
||||
query: string
|
||||
statsPeriod: string
|
||||
environment: string
|
||||
maxIssues: number
|
||||
}
|
||||
|
||||
/**
|
||||
* Normalizes the host config value: trims whitespace, strips any protocol prefix,
|
||||
* trailing slashes, and a pasted `/api` or `/api/0` suffix (the connector appends
|
||||
* `/api/0` itself), and falls back to sentry.io when empty. Genuine path prefixes
|
||||
* (e.g. `company.com/sentry` for subpath self-hosted installs) are preserved.
|
||||
*/
|
||||
function normalizeHost(rawHost: unknown): string {
|
||||
const host = typeof rawHost === 'string' ? rawHost.trim() : ''
|
||||
if (!host) return DEFAULT_HOST
|
||||
return host
|
||||
.replace(/^https?:\/\//i, '')
|
||||
.replace(/\/+$/, '')
|
||||
.replace(/\/api(\/0)?$/i, '')
|
||||
.replace(/\/+$/, '')
|
||||
.trim()
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads and normalizes the connector source configuration once per call.
|
||||
*/
|
||||
function readSourceConfig(sourceConfig: Record<string, unknown>): SentrySourceConfig {
|
||||
const host = normalizeHost(sourceConfig.baseUrl)
|
||||
const organization =
|
||||
typeof sourceConfig.organization === 'string' ? sourceConfig.organization.trim() : ''
|
||||
const project = typeof sourceConfig.project === 'string' ? sourceConfig.project.trim() : ''
|
||||
const query =
|
||||
typeof sourceConfig.query === 'string' && sourceConfig.query.trim()
|
||||
? sourceConfig.query.trim()
|
||||
: DEFAULT_QUERY
|
||||
const statsPeriod =
|
||||
typeof sourceConfig.statsPeriod === 'string' ? sourceConfig.statsPeriod.trim() : ''
|
||||
const environment =
|
||||
typeof sourceConfig.environment === 'string' ? sourceConfig.environment.trim() : ''
|
||||
const maxIssues = sourceConfig.maxIssues ? Number(sourceConfig.maxIssues) : 0
|
||||
|
||||
return {
|
||||
host,
|
||||
apiBase: `https://${host}/api/0`,
|
||||
organization,
|
||||
project,
|
||||
query,
|
||||
statsPeriod,
|
||||
environment,
|
||||
maxIssues,
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the standard JSON request headers carrying the Sentry auth token.
|
||||
*/
|
||||
function authHeaders(accessToken: string): Record<string, string> {
|
||||
return {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
'Content-Type': 'application/json',
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads the `cursor` of the `rel="next"` link from a Sentry `Link` header.
|
||||
*
|
||||
* Sentry paginates via the `Link` header: each link is annotated with `rel`,
|
||||
* `results`, and `cursor` attributes, e.g.
|
||||
* `<https://…/issues/?cursor=0:100:0>; rel="next"; results="true"; cursor="0:100:0"`.
|
||||
* A further page exists only when the `next` link reports `results="true"`; when it
|
||||
* reports `results="false"` (or the header is absent) the cursor points at an empty
|
||||
* page and pagination must stop. The cursor is read from the `cursor="…"` attribute,
|
||||
* which is the canonical token Sentry expects echoed back on the next request.
|
||||
*/
|
||||
function parseNextCursor(linkHeader: string | null): string | undefined {
|
||||
if (!linkHeader) return undefined
|
||||
|
||||
for (const part of linkHeader.split(',')) {
|
||||
if (!/rel="next"/.test(part)) continue
|
||||
if (!/results="true"/.test(part)) return undefined
|
||||
const cursorMatch = part.match(/cursor="([^"]*)"/)
|
||||
if (cursorMatch) return cursorMatch[1]
|
||||
return undefined
|
||||
}
|
||||
|
||||
return undefined
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the metadata-based content hash for an issue.
|
||||
*
|
||||
* The hash combines the issue id, its status, and `lastSeen`. `lastSeen` advances every
|
||||
* time a new event lands on the group, which is exactly when the latest-event content can
|
||||
* change — so it captures content freshness without hashing the downloaded body. `status`
|
||||
* is included so resolve/ignore transitions also re-sync. `count` is deliberately omitted:
|
||||
* it changes on every single occurrence and would churn the document on each event even
|
||||
* when `lastSeen` already moved, providing no extra signal over `lastSeen`.
|
||||
*
|
||||
* The hash is derived purely from issue metadata present on both the list stub and the
|
||||
* getDocument detail fetch, so both paths produce an identical hash for the same issue
|
||||
* snapshot. If a fresh event lands between listing and hydration, `lastSeen` advances and
|
||||
* getDocument computes a newer hash; the sync engine stores that newer hash, which the next
|
||||
* list pass reproduces — so the document converges without churn.
|
||||
*/
|
||||
function buildContentHash(issue: SentryIssue): string {
|
||||
return `sentry:${issue.id}:${issue.status ?? ''}:${issue.lastSeen ?? ''}`
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the document title, preferring the issue title and falling back to the
|
||||
* metadata type/value or short id.
|
||||
*/
|
||||
function buildTitle(issue: SentryIssue): string {
|
||||
const title = issue.title?.trim()
|
||||
if (title) return title
|
||||
|
||||
const metaType = issue.metadata?.type?.trim()
|
||||
const metaValue = issue.metadata?.value?.trim()
|
||||
if (metaType && metaValue) return `${metaType}: ${metaValue}`
|
||||
return metaType || issue.shortId || `Issue ${issue.id}`
|
||||
}
|
||||
|
||||
/**
|
||||
* Collects the source-specific metadata fed to mapTags. Shared between the list stub and
|
||||
* getDocument so tag values stay consistent regardless of which path produced the doc.
|
||||
*/
|
||||
function buildMetadata(issue: SentryIssue): Record<string, unknown> {
|
||||
return {
|
||||
level: issue.level,
|
||||
status: issue.status,
|
||||
firstSeen: issue.firstSeen,
|
||||
lastSeen: issue.lastSeen,
|
||||
count: issue.count != null ? Number(issue.count) : undefined,
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Creates a lightweight document stub from a list entry. No per-issue API calls — the
|
||||
* latest-event content is deferred to getDocument and only fetched for new/changed issues.
|
||||
*/
|
||||
function issueToStub(issue: SentryIssue): ExternalDocument {
|
||||
return {
|
||||
externalId: issue.id,
|
||||
title: buildTitle(issue),
|
||||
content: '',
|
||||
contentDeferred: true,
|
||||
mimeType: 'text/plain',
|
||||
sourceUrl: issue.permalink || undefined,
|
||||
contentHash: buildContentHash(issue),
|
||||
metadata: buildMetadata(issue),
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Renders the exception entry of a latest event into readable lines: each exception's
|
||||
* type/value plus a compact, top-down stack frame list.
|
||||
*/
|
||||
function formatException(data: SentryExceptionData): string[] {
|
||||
const lines: string[] = []
|
||||
|
||||
for (const value of data.values ?? []) {
|
||||
const header = [value.type, value.value].filter(Boolean).join(': ')
|
||||
if (header) lines.push(header)
|
||||
|
||||
const frames = value.stacktrace?.frames ?? []
|
||||
for (const frame of frames.slice().reverse()) {
|
||||
const location = [frame.module || frame.filename, frame.function].filter(Boolean).join(' in ')
|
||||
const lineNo = frame.lineNo != null ? `:${frame.lineNo}` : ''
|
||||
if (location) lines.push(` at ${location}${lineNo}`)
|
||||
}
|
||||
}
|
||||
|
||||
return lines
|
||||
}
|
||||
|
||||
/**
|
||||
* Formats an issue and its latest event into a single plain-text document covering the
|
||||
* title, culprit, counts, the latest event's message/exception, and event tags.
|
||||
*/
|
||||
function formatIssueContent(issue: SentryIssue, event: SentryEvent | null): string {
|
||||
const parts: string[] = []
|
||||
|
||||
parts.push(`Issue: ${buildTitle(issue)}`)
|
||||
if (issue.shortId) parts.push(`Short ID: ${issue.shortId}`)
|
||||
if (issue.culprit) parts.push(`Culprit: ${issue.culprit}`)
|
||||
if (issue.level) parts.push(`Level: ${issue.level}`)
|
||||
if (issue.status) parts.push(`Status: ${issue.status}`)
|
||||
if (issue.platform) parts.push(`Platform: ${issue.platform}`)
|
||||
if (issue.count) parts.push(`Events: ${issue.count}`)
|
||||
if (issue.userCount != null) parts.push(`Users affected: ${issue.userCount}`)
|
||||
if (issue.firstSeen) parts.push(`First seen: ${issue.firstSeen}`)
|
||||
if (issue.lastSeen) parts.push(`Last seen: ${issue.lastSeen}`)
|
||||
|
||||
if (event) {
|
||||
const message = event.message?.trim() || event.title?.trim()
|
||||
if (message) {
|
||||
parts.push('')
|
||||
parts.push('--- Latest Event ---')
|
||||
if (event.dateCreated) parts.push(`Occurred: ${event.dateCreated}`)
|
||||
parts.push(message)
|
||||
}
|
||||
|
||||
const exceptionEntry = event.entries?.find((entry) => entry.type === 'exception')
|
||||
if (exceptionEntry?.data) {
|
||||
const exceptionLines = formatException(exceptionEntry.data as SentryExceptionData)
|
||||
if (exceptionLines.length > 0) {
|
||||
parts.push('')
|
||||
parts.push('--- Exception ---')
|
||||
parts.push(...exceptionLines)
|
||||
}
|
||||
}
|
||||
|
||||
const tagLines = (event.tags ?? [])
|
||||
.map((tag) => (tag.key && tag.value ? `${tag.key}: ${tag.value}` : undefined))
|
||||
.filter((line): line is string => Boolean(line))
|
||||
if (tagLines.length > 0) {
|
||||
parts.push('')
|
||||
parts.push('--- Tags ---')
|
||||
parts.push(...tagLines)
|
||||
}
|
||||
}
|
||||
|
||||
return parts.join('\n').trim()
|
||||
}
|
||||
|
||||
/**
|
||||
* Fetches the latest event for an issue. Returns null when the issue has no events or the
|
||||
* request fails, so the document still syncs with its list-level summary.
|
||||
*
|
||||
* Uses the organization-scoped event endpoint
|
||||
* `/api/0/organizations/{org}/issues/{id}/events/latest/`, which is the documented path
|
||||
* and works for both sentry.io and self-hosted installs.
|
||||
*/
|
||||
async function fetchLatestEvent(
|
||||
apiBase: string,
|
||||
organization: string,
|
||||
accessToken: string,
|
||||
issueId: string
|
||||
): Promise<SentryEvent | null> {
|
||||
const url = `${apiBase}/organizations/${encodeURIComponent(organization)}/issues/${encodeURIComponent(issueId)}/events/latest/`
|
||||
|
||||
const response = await fetchWithRetry(url, {
|
||||
method: 'GET',
|
||||
headers: authHeaders(accessToken),
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
if (response.status !== 404) {
|
||||
logger.warn('Failed to fetch latest Sentry event', { issueId, status: response.status })
|
||||
}
|
||||
return null
|
||||
}
|
||||
|
||||
return (await response.json()) as SentryEvent
|
||||
}
|
||||
|
||||
export const sentryConnector: ConnectorConfig = {
|
||||
id: 'sentry',
|
||||
name: 'Sentry',
|
||||
description: 'Sync issues and errors from Sentry into your knowledge base',
|
||||
version: '1.0.0',
|
||||
icon: SentryIcon,
|
||||
|
||||
auth: {
|
||||
mode: 'apiKey',
|
||||
label: 'Auth Token',
|
||||
placeholder: 'Enter your Sentry auth token',
|
||||
},
|
||||
|
||||
configFields: [
|
||||
{
|
||||
id: 'baseUrl',
|
||||
title: 'Sentry URL',
|
||||
type: 'short-input',
|
||||
placeholder: 'sentry.io',
|
||||
required: false,
|
||||
mode: 'advanced',
|
||||
description:
|
||||
'Host of your Sentry install. Leave blank for sentry.io. Set this for self-hosted Sentry (e.g. sentry.mycompany.com).',
|
||||
},
|
||||
{
|
||||
id: 'organization',
|
||||
title: 'Organization Slug',
|
||||
type: 'short-input',
|
||||
placeholder: 'e.g. my-org',
|
||||
required: true,
|
||||
description: 'The slug of your Sentry organization.',
|
||||
},
|
||||
{
|
||||
id: 'project',
|
||||
title: 'Project Slug',
|
||||
type: 'short-input',
|
||||
placeholder: 'e.g. my-project',
|
||||
required: true,
|
||||
description: 'The slug of the project whose issues should be synced.',
|
||||
},
|
||||
{
|
||||
id: 'query',
|
||||
title: 'Search Query',
|
||||
type: 'short-input',
|
||||
placeholder: `e.g. ${DEFAULT_QUERY}`,
|
||||
required: false,
|
||||
description:
|
||||
'Sentry search query to filter issues (e.g. "is:unresolved level:error environment:production"). Defaults to "is:unresolved".',
|
||||
},
|
||||
{
|
||||
id: 'environment',
|
||||
title: 'Environment',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
mode: 'advanced',
|
||||
placeholder: 'e.g. production',
|
||||
description: 'Only sync issues seen in this environment. Leave blank for all environments.',
|
||||
},
|
||||
{
|
||||
id: 'statsPeriod',
|
||||
title: 'Stats Period',
|
||||
type: 'dropdown',
|
||||
required: false,
|
||||
mode: 'advanced',
|
||||
options: [
|
||||
{ label: 'Sentry default (24h)', id: '' },
|
||||
{ label: 'Last 24 hours', id: '24h' },
|
||||
{ label: 'Last 14 days', id: '14d' },
|
||||
],
|
||||
description: 'Time window for the issue stats Sentry computes on the project issues list.',
|
||||
},
|
||||
{
|
||||
id: 'maxIssues',
|
||||
title: 'Max Issues',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
placeholder: 'e.g. 500 (default: unlimited)',
|
||||
description: 'Cap the number of issues synced. Leave empty to sync all matching issues.',
|
||||
},
|
||||
],
|
||||
|
||||
listDocuments: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
cursor?: string,
|
||||
syncContext?: Record<string, unknown>
|
||||
): Promise<ExternalDocumentList> => {
|
||||
const { apiBase, organization, project, query, statsPeriod, environment, maxIssues } =
|
||||
readSourceConfig(sourceConfig)
|
||||
|
||||
if (!organization || !project) {
|
||||
throw new Error('Organization and project slugs are required')
|
||||
}
|
||||
|
||||
/*
|
||||
* Uses the project issues list endpoint
|
||||
* `/api/0/projects/{org}/{project}/issues/`. This endpoint is deprecated in favor of
|
||||
* `/api/0/organizations/{org}/issues/?project=<id>`, but the organization endpoint
|
||||
* filters by numeric project ID rather than slug — a UX regression for a connector
|
||||
* keyed on the human-readable project slug. The project endpoint remains functional
|
||||
* and slug-addressable, so it is retained deliberately for the listing path. Issue
|
||||
* detail and latest-event fetches use the organization-scoped paths.
|
||||
*/
|
||||
const url = new URL(
|
||||
`${apiBase}/projects/${encodeURIComponent(organization)}/${encodeURIComponent(project)}/issues/`
|
||||
)
|
||||
url.searchParams.set('query', query)
|
||||
url.searchParams.set('limit', String(ISSUES_PER_PAGE))
|
||||
if (statsPeriod) url.searchParams.set('statsPeriod', statsPeriod)
|
||||
if (environment) url.searchParams.set('environment', environment)
|
||||
if (cursor) url.searchParams.set('cursor', cursor)
|
||||
|
||||
logger.info('Listing Sentry issues', {
|
||||
organization,
|
||||
project,
|
||||
cursor: cursor ?? 'initial',
|
||||
maxIssues,
|
||||
})
|
||||
|
||||
const response = await fetchWithRetry(url.toString(), {
|
||||
method: 'GET',
|
||||
headers: authHeaders(accessToken),
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
const errorText = await response.text().catch(() => '')
|
||||
logger.error('Failed to list Sentry issues', {
|
||||
status: response.status,
|
||||
error: errorText.slice(0, 500),
|
||||
})
|
||||
throw new Error(`Failed to list Sentry issues: ${response.status}`)
|
||||
}
|
||||
|
||||
const issues = ((await response.json()) as SentryIssue[]).filter((issue) => Boolean(issue.id))
|
||||
|
||||
const prevFetched = (syncContext?.totalDocsFetched as number) ?? 0
|
||||
let documents = issues.map(issueToStub)
|
||||
let slicedSome = false
|
||||
if (maxIssues > 0) {
|
||||
const remaining = Math.max(0, maxIssues - prevFetched)
|
||||
if (documents.length > remaining) {
|
||||
slicedSome = true
|
||||
documents = documents.slice(0, remaining)
|
||||
}
|
||||
}
|
||||
|
||||
const totalFetched = prevFetched + documents.length
|
||||
if (syncContext) syncContext.totalDocsFetched = totalFetched
|
||||
const hitLimit = maxIssues > 0 && totalFetched >= maxIssues
|
||||
|
||||
const nextCursor = parseNextCursor(response.headers.get('Link'))
|
||||
if (hitLimit && (slicedSome || Boolean(nextCursor)) && syncContext) {
|
||||
syncContext.listingCapped = true
|
||||
}
|
||||
const hasMore = !hitLimit && Boolean(nextCursor)
|
||||
|
||||
return {
|
||||
documents,
|
||||
nextCursor: hasMore ? nextCursor : undefined,
|
||||
hasMore,
|
||||
}
|
||||
},
|
||||
|
||||
getDocument: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
externalId: string
|
||||
): Promise<ExternalDocument | null> => {
|
||||
try {
|
||||
if (!externalId) return null
|
||||
|
||||
const { apiBase, organization } = readSourceConfig(sourceConfig)
|
||||
if (!organization) return null
|
||||
|
||||
const url = `${apiBase}/organizations/${encodeURIComponent(organization)}/issues/${encodeURIComponent(externalId)}/`
|
||||
|
||||
const response = await fetchWithRetry(url, {
|
||||
method: 'GET',
|
||||
headers: authHeaders(accessToken),
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
if (response.status === 404 || response.status === 410) return null
|
||||
throw new Error(`Failed to fetch Sentry issue: ${response.status}`)
|
||||
}
|
||||
|
||||
const issue = (await response.json()) as SentryIssue
|
||||
if (!issue?.id) return null
|
||||
|
||||
const event = await fetchLatestEvent(apiBase, organization, accessToken, issue.id)
|
||||
const content = formatIssueContent(issue, event)
|
||||
if (!content.trim()) return null
|
||||
|
||||
return {
|
||||
externalId: issue.id,
|
||||
title: buildTitle(issue),
|
||||
content,
|
||||
contentDeferred: false,
|
||||
mimeType: 'text/plain',
|
||||
sourceUrl: issue.permalink || undefined,
|
||||
contentHash: buildContentHash(issue),
|
||||
metadata: buildMetadata(issue),
|
||||
}
|
||||
} catch (error) {
|
||||
logger.warn('Failed to get Sentry issue', {
|
||||
externalId,
|
||||
error: toError(error).message,
|
||||
})
|
||||
return null
|
||||
}
|
||||
},
|
||||
|
||||
validateConfig: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>
|
||||
): Promise<{ valid: boolean; error?: string }> => {
|
||||
const { apiBase, organization, project, statsPeriod, maxIssues, host } =
|
||||
readSourceConfig(sourceConfig)
|
||||
|
||||
if (!organization) {
|
||||
return { valid: false, error: 'Organization slug is required' }
|
||||
}
|
||||
if (!project) {
|
||||
return { valid: false, error: 'Project slug is required' }
|
||||
}
|
||||
|
||||
if (statsPeriod && !ALLOWED_STATS_PERIODS.has(statsPeriod)) {
|
||||
return { valid: false, error: 'Stats period must be 24h or 14d' }
|
||||
}
|
||||
|
||||
const rawMax = sourceConfig.maxIssues as string | undefined
|
||||
if (rawMax && (Number.isNaN(maxIssues) || maxIssues < 0)) {
|
||||
return { valid: false, error: 'Max issues must be a non-negative number' }
|
||||
}
|
||||
|
||||
try {
|
||||
/*
|
||||
* Probe the project detail endpoint first. This exercises the `project:read`
|
||||
* scope and the project-scoped path style, and gives a precise "not found"
|
||||
* message when the org or project slug is wrong.
|
||||
*/
|
||||
const projectResponse = await fetchWithRetry(
|
||||
`${apiBase}/projects/${encodeURIComponent(organization)}/${encodeURIComponent(project)}/`,
|
||||
{
|
||||
method: 'GET',
|
||||
headers: authHeaders(accessToken),
|
||||
},
|
||||
VALIDATE_RETRY_OPTIONS
|
||||
)
|
||||
|
||||
if (!projectResponse.ok) {
|
||||
if (projectResponse.status === 401 || projectResponse.status === 403) {
|
||||
return { valid: false, error: 'Invalid auth token or insufficient permissions' }
|
||||
}
|
||||
if (projectResponse.status === 404) {
|
||||
return {
|
||||
valid: false,
|
||||
error: `Organization or project not found on ${host}`,
|
||||
}
|
||||
}
|
||||
const errorText = await projectResponse.text().catch(() => '')
|
||||
return {
|
||||
valid: false,
|
||||
error: `Sentry access failed: ${projectResponse.status}${errorText ? ` — ${errorText.slice(0, 200)}` : ''}`,
|
||||
}
|
||||
}
|
||||
|
||||
/*
|
||||
* Probe the issues-list endpoint with a single-result page. The project
|
||||
* detail probe above only proves `project:read`, but every sync operation —
|
||||
* `listDocuments` and the org-scoped `getDocument`/latest-event hydration —
|
||||
* needs `event:read`. A token scoped to `project:read` only would pass the
|
||||
* first probe yet fail at hydration time, so this second probe forces a
|
||||
* misconfigured token to fail fast at save time. It is slug-addressable and
|
||||
* cheap (one issue, no stats window).
|
||||
*/
|
||||
const issuesProbeUrl = new URL(
|
||||
`${apiBase}/projects/${encodeURIComponent(organization)}/${encodeURIComponent(project)}/issues/`
|
||||
)
|
||||
issuesProbeUrl.searchParams.set('query', DEFAULT_QUERY)
|
||||
issuesProbeUrl.searchParams.set('limit', '1')
|
||||
|
||||
const issuesResponse = await fetchWithRetry(
|
||||
issuesProbeUrl.toString(),
|
||||
{
|
||||
method: 'GET',
|
||||
headers: authHeaders(accessToken),
|
||||
},
|
||||
VALIDATE_RETRY_OPTIONS
|
||||
)
|
||||
|
||||
if (!issuesResponse.ok) {
|
||||
if (issuesResponse.status === 401 || issuesResponse.status === 403) {
|
||||
return {
|
||||
valid: false,
|
||||
error:
|
||||
'Auth token cannot read issues. The token needs the "event:read" scope (in addition to "project:read").',
|
||||
}
|
||||
}
|
||||
const errorText = await issuesResponse.text().catch(() => '')
|
||||
return {
|
||||
valid: false,
|
||||
error: `Sentry issue access failed: ${issuesResponse.status}${errorText ? ` — ${errorText.slice(0, 200)}` : ''}`,
|
||||
}
|
||||
}
|
||||
|
||||
return { valid: true }
|
||||
} catch (error) {
|
||||
const message = getErrorMessage(error, 'Failed to validate configuration')
|
||||
return { valid: false, error: message }
|
||||
}
|
||||
},
|
||||
|
||||
tagDefinitions: [
|
||||
{ id: 'level', displayName: 'Level', fieldType: 'text' },
|
||||
{ id: 'status', displayName: 'Status', fieldType: 'text' },
|
||||
{ id: 'count', displayName: 'Event Count', fieldType: 'number' },
|
||||
{ id: 'firstSeen', displayName: 'First Seen', fieldType: 'date' },
|
||||
{ id: 'lastSeen', displayName: 'Last Seen', fieldType: 'date' },
|
||||
],
|
||||
|
||||
mapTags: (metadata: Record<string, unknown>): Record<string, unknown> => {
|
||||
const result: Record<string, unknown> = {}
|
||||
|
||||
if (typeof metadata.level === 'string' && metadata.level.trim()) {
|
||||
result.level = metadata.level
|
||||
}
|
||||
|
||||
if (typeof metadata.status === 'string' && metadata.status.trim()) {
|
||||
result.status = metadata.status
|
||||
}
|
||||
|
||||
if (metadata.count != null) {
|
||||
const num = Number(metadata.count)
|
||||
if (!Number.isNaN(num)) result.count = num
|
||||
}
|
||||
|
||||
const firstSeen = parseTagDate(metadata.firstSeen)
|
||||
if (firstSeen) result.firstSeen = firstSeen
|
||||
|
||||
const lastSeen = parseTagDate(metadata.lastSeen)
|
||||
if (lastSeen) result.lastSeen = lastSeen
|
||||
|
||||
return result
|
||||
},
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
export { typeformConnector } from '@/connectors/typeform/typeform'
|
||||
@@ -0,0 +1,605 @@
|
||||
import { createLogger } from '@sim/logger'
|
||||
import { getErrorMessage, toError } from '@sim/utils/errors'
|
||||
import { TypeformIcon } from '@/components/icons'
|
||||
import { fetchWithRetry, VALIDATE_RETRY_OPTIONS } from '@/lib/knowledge/documents/utils'
|
||||
import type { ConnectorConfig, ExternalDocument, ExternalDocumentList } from '@/connectors/types'
|
||||
import { parseTagDate } from '@/connectors/utils'
|
||||
|
||||
const logger = createLogger('TypeformConnector')
|
||||
|
||||
const TYPEFORM_API_BASE = 'https://api.typeform.com'
|
||||
/** Typeform allows page_size up to 1000; 100 keeps per-batch memory bounded. */
|
||||
const RESPONSES_PER_PAGE = 100
|
||||
|
||||
/**
|
||||
* Allowed `response_type` filter values per the Responses API. `completed` is the
|
||||
* API default; `all` is a connector-local sentinel that omits the filter so every
|
||||
* response type (`started`, `partial`, `completed`) is returned.
|
||||
*/
|
||||
type ResponseTypeChoice = 'completed' | 'partial' | 'all'
|
||||
|
||||
/**
|
||||
* A single field definition from the Typeform form structure.
|
||||
*/
|
||||
interface TypeformField {
|
||||
id: string
|
||||
ref?: string
|
||||
title?: string
|
||||
type?: string
|
||||
}
|
||||
|
||||
/**
|
||||
* The relevant subset of a Typeform form definition.
|
||||
*/
|
||||
interface TypeformFormDefinition {
|
||||
id: string
|
||||
title?: string
|
||||
fields?: TypeformField[]
|
||||
_links?: { display?: string }
|
||||
}
|
||||
|
||||
/**
|
||||
* A single answer within a Typeform response. Only the value-bearing keys for
|
||||
* each answer `type` are declared explicitly; the remainder are optional.
|
||||
*/
|
||||
interface TypeformAnswer {
|
||||
field?: { id?: string; type?: string; ref?: string }
|
||||
type?: string
|
||||
text?: string
|
||||
email?: string
|
||||
url?: string
|
||||
phone_number?: string
|
||||
file_url?: string
|
||||
number?: number
|
||||
boolean?: boolean
|
||||
date?: string
|
||||
choice?: { label?: string; other?: string }
|
||||
choices?: { labels?: string[]; other?: string }
|
||||
payment?: { amount?: string; last4?: string; name?: string; success?: boolean }
|
||||
}
|
||||
|
||||
/**
|
||||
* A single Typeform response item.
|
||||
*
|
||||
* `token` is the cursor field consumed by the `before`/`after` query params, while
|
||||
* `response_id` is the identifier consumed by `included_response_ids`. They are
|
||||
* distinct values, so both are tracked: the externalId is keyed off `response_id`
|
||||
* (used by getDocument), the pagination cursor off `token`.
|
||||
*/
|
||||
interface TypeformResponseItem {
|
||||
response_id?: string
|
||||
token: string
|
||||
landing_id?: string
|
||||
landed_at?: string
|
||||
submitted_at?: string
|
||||
metadata?: {
|
||||
platform?: string
|
||||
browser?: string
|
||||
referer?: string
|
||||
}
|
||||
answers?: TypeformAnswer[] | null
|
||||
hidden?: Record<string, unknown> | null
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads the `response_type` choice from sourceConfig, defaulting to `completed`.
|
||||
*/
|
||||
function getResponseTypeChoice(sourceConfig: Record<string, unknown>): ResponseTypeChoice {
|
||||
const value =
|
||||
typeof sourceConfig.responseType === 'string' ? sourceConfig.responseType.trim() : ''
|
||||
if (value === 'partial' || value === 'all') return value
|
||||
return 'completed'
|
||||
}
|
||||
|
||||
/**
|
||||
* Appends the `response_type` filter to a query string for a given choice. `all`
|
||||
* omits the parameter so every type is returned; `partial` requests both partial
|
||||
* and completed so partially-answered submissions are included alongside finished
|
||||
* ones.
|
||||
*/
|
||||
function appendResponseType(params: URLSearchParams, choice: ResponseTypeChoice): void {
|
||||
if (choice === 'completed') params.append('response_type', 'completed')
|
||||
else if (choice === 'partial') params.append('response_type', 'partial,completed')
|
||||
}
|
||||
|
||||
/**
|
||||
* Renders a single answer's value into a human-readable string.
|
||||
*/
|
||||
function renderAnswerValue(answer: TypeformAnswer): string {
|
||||
switch (answer.type) {
|
||||
case 'text':
|
||||
return answer.text ?? ''
|
||||
case 'email':
|
||||
return answer.email ?? ''
|
||||
case 'url':
|
||||
return answer.url ?? ''
|
||||
case 'phone_number':
|
||||
return answer.phone_number ?? ''
|
||||
case 'file_url':
|
||||
return answer.file_url ?? ''
|
||||
case 'number':
|
||||
return answer.number != null ? String(answer.number) : ''
|
||||
case 'boolean':
|
||||
return answer.boolean != null ? (answer.boolean ? 'Yes' : 'No') : ''
|
||||
case 'date':
|
||||
return answer.date ?? ''
|
||||
case 'choice': {
|
||||
const parts = [answer.choice?.label, answer.choice?.other].filter(Boolean)
|
||||
return parts.join(', ')
|
||||
}
|
||||
case 'choices': {
|
||||
const labels = Array.isArray(answer.choices?.labels) ? (answer.choices?.labels ?? []) : []
|
||||
const parts = [...labels]
|
||||
if (answer.choices?.other) parts.push(answer.choices.other)
|
||||
return parts.join(', ')
|
||||
}
|
||||
case 'payment':
|
||||
return answer.payment?.amount != null ? String(answer.payment.amount) : ''
|
||||
default:
|
||||
return ''
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds a map of field id to its human-readable question title from a form definition.
|
||||
*/
|
||||
function buildFieldTitleMap(form: TypeformFormDefinition): Map<string, string> {
|
||||
const map = new Map<string, string>()
|
||||
for (const field of form.fields ?? []) {
|
||||
if (field.id) map.set(field.id, field.title || field.id)
|
||||
}
|
||||
return map
|
||||
}
|
||||
|
||||
/**
|
||||
* Renders a Typeform response as readable "Question: Answer" plain text.
|
||||
*/
|
||||
function renderResponseContent(
|
||||
form: TypeformFormDefinition,
|
||||
response: TypeformResponseItem,
|
||||
fieldTitles: Map<string, string>
|
||||
): string {
|
||||
const parts: string[] = []
|
||||
|
||||
if (form.title) parts.push(`Form: ${form.title}`)
|
||||
if (response.submitted_at) parts.push(`Submitted: ${response.submitted_at}`)
|
||||
parts.push('')
|
||||
|
||||
const answers = Array.isArray(response.answers) ? response.answers : []
|
||||
for (const answer of answers) {
|
||||
const fieldId = answer.field?.id
|
||||
const question = (fieldId && fieldTitles.get(fieldId)) || fieldId || 'Answer'
|
||||
const value = renderAnswerValue(answer)
|
||||
parts.push(`${question}: ${value}`)
|
||||
}
|
||||
|
||||
if (response.hidden && Object.keys(response.hidden).length > 0) {
|
||||
parts.push('')
|
||||
parts.push('--- Hidden Fields ---')
|
||||
for (const [key, val] of Object.entries(response.hidden)) {
|
||||
parts.push(`${key}: ${String(val)}`)
|
||||
}
|
||||
}
|
||||
|
||||
return parts.join('\n')
|
||||
}
|
||||
|
||||
/**
|
||||
* Derives the stable external identifier for a response. Prefers `response_id`
|
||||
* (the identifier `included_response_ids` filters on, so getDocument can fetch the
|
||||
* exact response) and falls back to `token` when `response_id` is absent.
|
||||
*/
|
||||
function getResponseExternalId(response: TypeformResponseItem): string {
|
||||
return response.response_id || response.token
|
||||
}
|
||||
|
||||
/**
|
||||
* Produces the metadata-based content hash for a response. Responses are immutable
|
||||
* once submitted, so `submitted_at` is a stable change key. For not-yet-submitted
|
||||
* (started/partial) responses, `landed_at` is used as the fallback indicator.
|
||||
*/
|
||||
function getResponseContentHash(response: TypeformResponseItem): string {
|
||||
const indicator = response.submitted_at || response.landed_at || ''
|
||||
return `typeform:${getResponseExternalId(response)}:${indicator}`
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds a full ExternalDocument from a rendered response.
|
||||
*/
|
||||
function responseToDocument(
|
||||
form: TypeformFormDefinition,
|
||||
response: TypeformResponseItem,
|
||||
fieldTitles: Map<string, string>
|
||||
): ExternalDocument {
|
||||
const externalId = getResponseExternalId(response)
|
||||
const submittedAt = response.submitted_at
|
||||
const displayUrl = form._links?.display
|
||||
|
||||
return {
|
||||
externalId,
|
||||
title: `${form.title || 'Typeform'} — ${submittedAt || response.landed_at || externalId}`,
|
||||
content: renderResponseContent(form, response, fieldTitles),
|
||||
contentDeferred: false,
|
||||
mimeType: 'text/plain',
|
||||
sourceUrl: displayUrl || undefined,
|
||||
contentHash: getResponseContentHash(response),
|
||||
metadata: {
|
||||
formId: form.id,
|
||||
formTitle: form.title,
|
||||
submittedAt,
|
||||
landedAt: response.landed_at,
|
||||
platform: response.metadata?.platform,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Fetches a form definition, caching it in syncContext keyed by form id so a
|
||||
* single sync run fetches each form's structure only once.
|
||||
*/
|
||||
async function getFormDefinition(
|
||||
accessToken: string,
|
||||
formId: string,
|
||||
syncContext?: Record<string, unknown>,
|
||||
retryOptions?: Parameters<typeof fetchWithRetry>[2]
|
||||
): Promise<TypeformFormDefinition> {
|
||||
const cacheKey = `form:${formId}`
|
||||
const cached = syncContext?.[cacheKey] as TypeformFormDefinition | undefined
|
||||
if (cached) return cached
|
||||
|
||||
const response = await fetchWithRetry(
|
||||
`${TYPEFORM_API_BASE}/forms/${encodeURIComponent(formId)}`,
|
||||
{
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
Accept: 'application/json',
|
||||
},
|
||||
},
|
||||
retryOptions
|
||||
)
|
||||
|
||||
if (!response.ok) {
|
||||
throw new Error(`Failed to fetch Typeform form ${formId}: ${response.status}`)
|
||||
}
|
||||
|
||||
const form = (await response.json()) as TypeformFormDefinition
|
||||
if (syncContext) syncContext[cacheKey] = form
|
||||
return form
|
||||
}
|
||||
|
||||
export const typeformConnector: ConnectorConfig = {
|
||||
id: 'typeform',
|
||||
name: 'Typeform',
|
||||
description: 'Sync form responses from Typeform into your knowledge base',
|
||||
version: '1.0.0',
|
||||
icon: TypeformIcon,
|
||||
|
||||
auth: {
|
||||
mode: 'apiKey',
|
||||
label: 'Personal Access Token',
|
||||
placeholder: 'Enter your Typeform personal access token',
|
||||
},
|
||||
|
||||
/**
|
||||
* Incremental sync narrows the listing to responses submitted after the last
|
||||
* sync via the `since` filter (inclusive, matched against `submitted_at` for
|
||||
* completed responses). Responses are immutable, so reconciliation by content
|
||||
* hash skips anything already indexed.
|
||||
*/
|
||||
supportsIncrementalSync: true,
|
||||
|
||||
configFields: [
|
||||
{
|
||||
id: 'formId',
|
||||
title: 'Form ID',
|
||||
type: 'short-input',
|
||||
placeholder: 'e.g. abc123XYZ',
|
||||
required: true,
|
||||
description: 'The Typeform form whose responses you want to sync',
|
||||
},
|
||||
{
|
||||
id: 'responseType',
|
||||
title: 'Responses',
|
||||
type: 'dropdown',
|
||||
required: false,
|
||||
options: [
|
||||
{ label: 'Completed only', id: 'completed' },
|
||||
{ label: 'Partial & completed', id: 'partial' },
|
||||
{ label: 'All (including started)', id: 'all' },
|
||||
],
|
||||
description: 'Which responses to sync by completion status. Defaults to completed only.',
|
||||
},
|
||||
{
|
||||
id: 'since',
|
||||
title: 'Submitted After',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
mode: 'advanced',
|
||||
placeholder: 'e.g. 2024-01-01T00:00:00Z',
|
||||
description: 'Only sync responses submitted on or after this date (ISO 8601, UTC).',
|
||||
},
|
||||
{
|
||||
id: 'until',
|
||||
title: 'Submitted Before',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
mode: 'advanced',
|
||||
placeholder: 'e.g. 2024-12-31T23:59:59Z',
|
||||
description: 'Only sync responses submitted on or before this date (ISO 8601, UTC).',
|
||||
},
|
||||
{
|
||||
id: 'query',
|
||||
title: 'Search Filter',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
mode: 'advanced',
|
||||
placeholder: 'e.g. acme',
|
||||
description:
|
||||
'Only sync responses containing this text in any answer, hidden field, or variable.',
|
||||
},
|
||||
{
|
||||
id: 'maxResponses',
|
||||
title: 'Max Responses',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
placeholder: 'e.g. 500 (default: unlimited)',
|
||||
},
|
||||
],
|
||||
|
||||
listDocuments: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
cursor?: string,
|
||||
syncContext?: Record<string, unknown>,
|
||||
lastSyncAt?: Date
|
||||
): Promise<ExternalDocumentList> => {
|
||||
const formId = (sourceConfig.formId as string)?.trim()
|
||||
if (!formId) {
|
||||
throw new Error('Form ID is required')
|
||||
}
|
||||
const maxResponses = sourceConfig.maxResponses ? Number(sourceConfig.maxResponses) : 0
|
||||
|
||||
const form = await getFormDefinition(accessToken, formId, syncContext)
|
||||
const fieldTitles = buildFieldTitleMap(form)
|
||||
|
||||
const queryParams = new URLSearchParams()
|
||||
queryParams.append('page_size', String(RESPONSES_PER_PAGE))
|
||||
appendResponseType(queryParams, getResponseTypeChoice(sourceConfig))
|
||||
|
||||
const since = typeof sourceConfig.since === 'string' ? sourceConfig.since.trim() : ''
|
||||
const until = typeof sourceConfig.until === 'string' ? sourceConfig.until.trim() : ''
|
||||
const search = typeof sourceConfig.query === 'string' ? sourceConfig.query.trim() : ''
|
||||
if (until) queryParams.append('until', until)
|
||||
if (search) queryParams.append('query', search)
|
||||
|
||||
/**
|
||||
* `since` from the user config wins; otherwise incremental sync derives it
|
||||
* from lastSyncAt. `since` narrows the set by submission date while `before`
|
||||
* (token paging) walks it newest-to-oldest; the two compose — only `sort` is
|
||||
* mutually exclusive with `before`/`after`, which this connector never sets.
|
||||
*/
|
||||
if (since) queryParams.append('since', since)
|
||||
else if (lastSyncAt) queryParams.append('since', lastSyncAt.toISOString())
|
||||
|
||||
if (cursor) {
|
||||
queryParams.append('before', cursor)
|
||||
}
|
||||
|
||||
const url = `${TYPEFORM_API_BASE}/forms/${encodeURIComponent(formId)}/responses?${queryParams.toString()}`
|
||||
|
||||
logger.info('Listing Typeform responses', {
|
||||
formId,
|
||||
before: cursor,
|
||||
incremental: Boolean(lastSyncAt),
|
||||
})
|
||||
|
||||
const response = await fetchWithRetry(url, {
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
Accept: 'application/json',
|
||||
},
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
const errorText = await response.text().catch(() => '')
|
||||
logger.error('Failed to list Typeform responses', {
|
||||
formId,
|
||||
status: response.status,
|
||||
error: errorText.slice(0, 500),
|
||||
})
|
||||
throw new Error(`Failed to list Typeform responses: ${response.status}`)
|
||||
}
|
||||
|
||||
const data = (await response.json()) as { items?: TypeformResponseItem[] }
|
||||
const items = Array.isArray(data.items) ? data.items.filter((item) => item?.token) : []
|
||||
|
||||
const prevTotal = (syncContext?.totalDocsFetched as number) ?? 0
|
||||
|
||||
/**
|
||||
* Trim the page to the remaining `maxResponses` budget so the cap is honored
|
||||
* exactly rather than overshooting by up to a full page. The `before` cursor
|
||||
* is still derived from the untrimmed page below, but it is unused once the
|
||||
* cap is hit because `hasMore` becomes false.
|
||||
*/
|
||||
let cappedItems = items
|
||||
let slicedSome = false
|
||||
if (maxResponses > 0) {
|
||||
const remaining = Math.max(0, maxResponses - prevTotal)
|
||||
if (items.length > remaining) {
|
||||
slicedSome = true
|
||||
cappedItems = items.slice(0, remaining)
|
||||
}
|
||||
}
|
||||
|
||||
const documents: ExternalDocument[] = cappedItems.map((item) =>
|
||||
responseToDocument(form, item, fieldTitles)
|
||||
)
|
||||
|
||||
const totalFetched = prevTotal + documents.length
|
||||
if (syncContext) syncContext.totalDocsFetched = totalFetched
|
||||
const hitLimit = maxResponses > 0 && totalFetched >= maxResponses
|
||||
|
||||
/**
|
||||
* The `before` cursor is the response `token` (not `response_id`). Each full
|
||||
* page advances to the oldest token seen so the next request pages strictly
|
||||
* older responses. A short page or a missing token ends pagination, which also
|
||||
* guards against an infinite loop if the API ever repeats a cursor.
|
||||
*/
|
||||
const lastItem = items[items.length - 1]
|
||||
const nextCursor = lastItem?.token
|
||||
const sourceHasMore = items.length === RESPONSES_PER_PAGE && Boolean(nextCursor)
|
||||
|
||||
/**
|
||||
* Signal a truncated listing so the engine skips deletion reconciliation —
|
||||
* but only when the cap actually dropped responses (this page was sliced, or
|
||||
* the source had more pages). If the cap merely coincides with source
|
||||
* exhaustion, reconciliation stays enabled so deleted responses are cleaned up.
|
||||
*/
|
||||
if (hitLimit && (slicedSome || sourceHasMore) && syncContext) {
|
||||
syncContext.listingCapped = true
|
||||
}
|
||||
|
||||
const hasMore = !hitLimit && sourceHasMore
|
||||
|
||||
return {
|
||||
documents,
|
||||
nextCursor: hasMore ? nextCursor : undefined,
|
||||
hasMore,
|
||||
}
|
||||
},
|
||||
|
||||
getDocument: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
externalId: string,
|
||||
syncContext?: Record<string, unknown>
|
||||
): Promise<ExternalDocument | null> => {
|
||||
const formId = (sourceConfig.formId as string)?.trim()
|
||||
if (!formId || !externalId) return null
|
||||
|
||||
try {
|
||||
const form = await getFormDefinition(accessToken, formId, syncContext)
|
||||
const fieldTitles = buildFieldTitleMap(form)
|
||||
|
||||
/**
|
||||
* `included_response_ids` filters by `response_id`, matching the externalId
|
||||
* minted in listDocuments. The configured response_type is forwarded so a
|
||||
* partial response stays fetchable (the endpoint defaults to completed-only,
|
||||
* which would otherwise exclude it).
|
||||
*/
|
||||
const params = new URLSearchParams()
|
||||
params.append('included_response_ids', externalId)
|
||||
appendResponseType(params, getResponseTypeChoice(sourceConfig))
|
||||
|
||||
const url = `${TYPEFORM_API_BASE}/forms/${encodeURIComponent(formId)}/responses?${params.toString()}`
|
||||
const response = await fetchWithRetry(url, {
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
Accept: 'application/json',
|
||||
},
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
if (response.status === 404) return null
|
||||
throw new Error(`Failed to fetch Typeform response ${externalId}: ${response.status}`)
|
||||
}
|
||||
|
||||
const data = (await response.json()) as { items?: TypeformResponseItem[] }
|
||||
const item = Array.isArray(data.items)
|
||||
? data.items.find((candidate) => getResponseExternalId(candidate) === externalId)
|
||||
: undefined
|
||||
if (!item) return null
|
||||
|
||||
return responseToDocument(form, item, fieldTitles)
|
||||
} catch (error) {
|
||||
logger.warn('Failed to get Typeform response', {
|
||||
externalId,
|
||||
error: toError(error).message,
|
||||
})
|
||||
return null
|
||||
}
|
||||
},
|
||||
|
||||
validateConfig: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>
|
||||
): Promise<{ valid: boolean; error?: string }> => {
|
||||
const formId = (sourceConfig.formId as string)?.trim()
|
||||
if (!formId) {
|
||||
return { valid: false, error: 'Form ID is required' }
|
||||
}
|
||||
|
||||
const maxResponses = sourceConfig.maxResponses as string | undefined
|
||||
if (maxResponses && (Number.isNaN(Number(maxResponses)) || Number(maxResponses) <= 0)) {
|
||||
return { valid: false, error: 'Max responses must be a positive number' }
|
||||
}
|
||||
|
||||
const since = typeof sourceConfig.since === 'string' ? sourceConfig.since.trim() : ''
|
||||
if (since && Number.isNaN(new Date(since).getTime())) {
|
||||
return { valid: false, error: '"Submitted After" must be a valid ISO 8601 date' }
|
||||
}
|
||||
|
||||
const until = typeof sourceConfig.until === 'string' ? sourceConfig.until.trim() : ''
|
||||
if (until && Number.isNaN(new Date(until).getTime())) {
|
||||
return { valid: false, error: '"Submitted Before" must be a valid ISO 8601 date' }
|
||||
}
|
||||
|
||||
if (since && until && new Date(since).getTime() > new Date(until).getTime()) {
|
||||
return { valid: false, error: '"Submitted After" must not be later than "Submitted Before"' }
|
||||
}
|
||||
|
||||
try {
|
||||
const response = await fetchWithRetry(
|
||||
`${TYPEFORM_API_BASE}/forms/${encodeURIComponent(formId)}`,
|
||||
{
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
Accept: 'application/json',
|
||||
},
|
||||
},
|
||||
VALIDATE_RETRY_OPTIONS
|
||||
)
|
||||
|
||||
if (response.status === 401 || response.status === 403) {
|
||||
return { valid: false, error: 'Invalid or unauthorized Typeform personal access token' }
|
||||
}
|
||||
if (response.status === 404) {
|
||||
return { valid: false, error: `Form not found: ${formId}` }
|
||||
}
|
||||
if (!response.ok) {
|
||||
return { valid: false, error: `Failed to validate Typeform form: ${response.status}` }
|
||||
}
|
||||
|
||||
return { valid: true }
|
||||
} catch (error) {
|
||||
return { valid: false, error: getErrorMessage(error, 'Failed to validate configuration') }
|
||||
}
|
||||
},
|
||||
|
||||
tagDefinitions: [
|
||||
{ id: 'formTitle', displayName: 'Form Title', fieldType: 'text' },
|
||||
{ id: 'platform', displayName: 'Platform', fieldType: 'text' },
|
||||
{ id: 'submittedAt', displayName: 'Submitted At', fieldType: 'date' },
|
||||
],
|
||||
|
||||
mapTags: (metadata: Record<string, unknown>): Record<string, unknown> => {
|
||||
const result: Record<string, unknown> = {}
|
||||
|
||||
if (typeof metadata.formTitle === 'string' && metadata.formTitle) {
|
||||
result.formTitle = metadata.formTitle
|
||||
}
|
||||
|
||||
if (typeof metadata.platform === 'string' && metadata.platform) {
|
||||
result.platform = metadata.platform
|
||||
}
|
||||
|
||||
const submittedAt = parseTagDate(metadata.submittedAt)
|
||||
if (submittedAt) result.submittedAt = submittedAt
|
||||
|
||||
return result
|
||||
},
|
||||
}
|
||||
@@ -78,3 +78,36 @@ export function parseMultiValue(value: unknown): string[] {
|
||||
}
|
||||
return []
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads a response body into a Buffer while enforcing a hard byte cap. The
|
||||
* declared `content-length` header cannot be trusted as the sole guard —
|
||||
* chunked transfer encoding may omit it entirely — so bytes are accumulated
|
||||
* from the stream and reading aborts as soon as the cap is exceeded, ensuring
|
||||
* an oversized (or hostile) body is never fully buffered into memory.
|
||||
* Returns null when the cap is exceeded.
|
||||
*/
|
||||
export async function readBodyWithLimit(
|
||||
response: Response,
|
||||
maxBytes: number
|
||||
): Promise<Buffer | null> {
|
||||
if (!response.body) {
|
||||
const buffer = Buffer.from(await response.arrayBuffer())
|
||||
return buffer.byteLength > maxBytes ? null : buffer
|
||||
}
|
||||
|
||||
const reader = response.body.getReader()
|
||||
const chunks: Uint8Array[] = []
|
||||
let total = 0
|
||||
while (true) {
|
||||
const { done, value } = await reader.read()
|
||||
if (done) break
|
||||
total += value.byteLength
|
||||
if (total > maxBytes) {
|
||||
await reader.cancel().catch(() => {})
|
||||
return null
|
||||
}
|
||||
chunks.push(value)
|
||||
}
|
||||
return Buffer.concat(chunks)
|
||||
}
|
||||
|
||||
@@ -0,0 +1 @@
|
||||
export { xConnector } from '@/connectors/x/x'
|
||||
@@ -0,0 +1,628 @@
|
||||
import { createLogger } from '@sim/logger'
|
||||
import { getErrorMessage, toError } from '@sim/utils/errors'
|
||||
import { xIcon } from '@/components/icons'
|
||||
import { fetchWithRetry, VALIDATE_RETRY_OPTIONS } from '@/lib/knowledge/documents/utils'
|
||||
import type { ConnectorConfig, ExternalDocument, ExternalDocumentList } from '@/connectors/types'
|
||||
import { parseMultiValue, parseTagDate } from '@/connectors/utils'
|
||||
|
||||
const logger = createLogger('XConnector')
|
||||
|
||||
const X_API_BASE = 'https://api.x.com/2'
|
||||
const DEFAULT_MAX_POSTS = 200
|
||||
/** Max page size accepted by the timeline, mentions, bookmarks, and likes endpoints. */
|
||||
const POSTS_PER_PAGE = 100
|
||||
/**
|
||||
* Minimum `max_results` accepted by the user-tweets, mentions, and liked-tweets
|
||||
* endpoints. The bookmarks endpoint is the sole exception and accepts a minimum of 1.
|
||||
*/
|
||||
const MIN_PAGE_SIZE = 5
|
||||
/**
|
||||
* `edit_history_tweet_ids` is requested explicitly (it is not a default field) so the
|
||||
* content hash can key on edit-history length and detect edits.
|
||||
*/
|
||||
const TWEET_FIELDS = 'created_at,public_metrics,text,edit_history_tweet_ids'
|
||||
|
||||
/**
|
||||
* Sync mode determines which timeline the connector reads.
|
||||
* - `me`: the authenticated user's own posts (GET /2/users/:id/tweets)
|
||||
* - `user`: another account's posts by username (GET /2/users/:id/tweets)
|
||||
* - `mentions`: posts mentioning the authenticated user (GET /2/users/:id/mentions)
|
||||
* - `bookmarks`: the authenticated user's bookmarks (GET /2/users/:id/bookmarks)
|
||||
* - `likes`: posts the authenticated user has liked (GET /2/users/:id/liked_tweets)
|
||||
*/
|
||||
type SyncMode = 'me' | 'user' | 'mentions' | 'bookmarks' | 'likes'
|
||||
|
||||
/** Modes whose endpoint supports the `exclude=retweets,replies` parameter. */
|
||||
const EXCLUDE_CAPABLE_MODES: ReadonlySet<SyncMode> = new Set<SyncMode>(['me', 'user'])
|
||||
/** Modes whose endpoint supports the `start_time` / `end_time` parameters. */
|
||||
const DATE_RANGE_CAPABLE_MODES: ReadonlySet<SyncMode> = new Set<SyncMode>([
|
||||
'me',
|
||||
'user',
|
||||
'mentions',
|
||||
])
|
||||
|
||||
interface XPublicMetrics {
|
||||
retweet_count?: number
|
||||
reply_count?: number
|
||||
like_count?: number
|
||||
quote_count?: number
|
||||
}
|
||||
|
||||
interface XTweet {
|
||||
id: string
|
||||
text: string
|
||||
created_at?: string
|
||||
author_id?: string
|
||||
public_metrics?: XPublicMetrics
|
||||
edit_history_tweet_ids?: string[]
|
||||
}
|
||||
|
||||
interface XUser {
|
||||
id: string
|
||||
name?: string
|
||||
username?: string
|
||||
}
|
||||
|
||||
interface XListResponse {
|
||||
data?: XTweet[]
|
||||
includes?: { users?: XUser[] }
|
||||
meta?: { next_token?: string; result_count?: number }
|
||||
errors?: Array<{ detail?: string; title?: string }>
|
||||
}
|
||||
|
||||
interface XSingleResponse {
|
||||
data?: XTweet
|
||||
includes?: { users?: XUser[] }
|
||||
errors?: Array<{ detail?: string; title?: string }>
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves the configured sync mode, defaulting to the authenticated user's
|
||||
* own posts.
|
||||
*/
|
||||
function resolveSyncMode(sourceConfig: Record<string, unknown>): SyncMode {
|
||||
const mode = sourceConfig.syncMode
|
||||
if (mode === 'user' || mode === 'mentions' || mode === 'bookmarks' || mode === 'likes') {
|
||||
return mode
|
||||
}
|
||||
return 'me'
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads a boolean toggle from a dropdown config field that stores 'true' / 'false'
|
||||
* strings. Falls back to `defaultValue` when unset or unrecognized.
|
||||
*/
|
||||
function readBooleanOption(value: unknown, defaultValue: boolean): boolean {
|
||||
if (value === 'true' || value === true) return true
|
||||
if (value === 'false' || value === false) return false
|
||||
return defaultValue
|
||||
}
|
||||
|
||||
/**
|
||||
* Parses the configured usernames into a normalized, deduplicated handle list.
|
||||
*
|
||||
* Handles are lowercased and stripped of a leading `@` before deduplication so
|
||||
* that `jack`, `@jack`, and `Jack` collapse to a single entry — avoiding a
|
||||
* duplicate user-id lookup and a redundant `userIndex` slot in the packed
|
||||
* pagination cursor. Both `validateConfig` and `listDocuments` call this so the
|
||||
* cursor's `userIndex` stays aligned to the same array across pages.
|
||||
*/
|
||||
function parseUsernames(value: unknown): string[] {
|
||||
const seen = new Set<string>()
|
||||
const out: string[] = []
|
||||
for (const raw of parseMultiValue(value)) {
|
||||
const handle = raw.replace(/^@/, '').toLowerCase()
|
||||
if (!handle || seen.has(handle)) continue
|
||||
seen.add(handle)
|
||||
out.push(handle)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads and trims a string config field, returning undefined when blank.
|
||||
*/
|
||||
function readTrimmed(value: unknown): string | undefined {
|
||||
if (typeof value !== 'string') return undefined
|
||||
const trimmed = value.trim()
|
||||
return trimmed.length > 0 ? trimmed : undefined
|
||||
}
|
||||
|
||||
/**
|
||||
* Performs an authenticated GET against the X API v2 and returns the parsed JSON.
|
||||
*/
|
||||
async function xApiGet(
|
||||
path: string,
|
||||
accessToken: string,
|
||||
params?: Record<string, string>,
|
||||
retryOptions?: Parameters<typeof fetchWithRetry>[2]
|
||||
): Promise<unknown> {
|
||||
const queryParams = params ? `?${new URLSearchParams(params).toString()}` : ''
|
||||
const url = `${X_API_BASE}${path}${queryParams}`
|
||||
|
||||
const response = await fetchWithRetry(
|
||||
url,
|
||||
{
|
||||
method: 'GET',
|
||||
headers: {
|
||||
Authorization: `Bearer ${accessToken}`,
|
||||
'Content-Type': 'application/json',
|
||||
},
|
||||
},
|
||||
retryOptions
|
||||
)
|
||||
|
||||
if (!response.ok) {
|
||||
const body = await response.text().catch(() => '')
|
||||
throw new Error(`X API HTTP error: ${response.status} ${response.statusText} ${body}`.trim())
|
||||
}
|
||||
|
||||
return response.json()
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves the authenticated user's numeric ID via GET /2/users/me.
|
||||
*/
|
||||
async function resolveMyUserId(
|
||||
accessToken: string,
|
||||
retryOptions?: Parameters<typeof fetchWithRetry>[2]
|
||||
): Promise<string> {
|
||||
const data = (await xApiGet('/users/me', accessToken, undefined, retryOptions)) as {
|
||||
data?: { id?: string }
|
||||
}
|
||||
const id = data.data?.id
|
||||
if (!id) throw new Error('Failed to resolve authenticated user ID')
|
||||
return id
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves a public username to its numeric user ID via
|
||||
* GET /2/users/by/username/:username.
|
||||
*/
|
||||
async function resolveUsernameId(
|
||||
accessToken: string,
|
||||
username: string,
|
||||
retryOptions?: Parameters<typeof fetchWithRetry>[2]
|
||||
): Promise<string> {
|
||||
const handle = username.trim().replace(/^@/, '')
|
||||
const data = (await xApiGet(
|
||||
`/users/by/username/${encodeURIComponent(handle)}`,
|
||||
accessToken,
|
||||
undefined,
|
||||
retryOptions
|
||||
)) as { data?: { id?: string }; errors?: Array<{ detail?: string }> }
|
||||
const id = data.data?.id
|
||||
if (!id) {
|
||||
throw new Error(data.errors?.[0]?.detail || `User @${handle} not found`)
|
||||
}
|
||||
return id
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds a deterministic, metadata-based content hash for a tweet.
|
||||
*
|
||||
* Tweets are immutable outside the brief post-publish edit window; an edit
|
||||
* appends a new ID to `edit_history_tweet_ids`. We therefore key the hash on
|
||||
* the edit-history length when present (so edits are detected as changes), and
|
||||
* fall back to `created_at` when the field is absent.
|
||||
*/
|
||||
function tweetContentHash(tweet: XTweet): string {
|
||||
const historyLength = Array.isArray(tweet.edit_history_tweet_ids)
|
||||
? tweet.edit_history_tweet_ids.length
|
||||
: undefined
|
||||
const changeIndicator = historyLength ?? tweet.created_at ?? ''
|
||||
return `x:${tweet.id}:${changeIndicator}`
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the canonical source URL for a tweet. When the author's username is
|
||||
* unknown, falls back to the username-agnostic permalink which X redirects.
|
||||
*/
|
||||
function tweetSourceUrl(tweetId: string, username?: string): string {
|
||||
if (username) return `https://x.com/${username}/status/${tweetId}`
|
||||
return `https://x.com/i/web/status/${tweetId}`
|
||||
}
|
||||
|
||||
/**
|
||||
* Derives a short title from the tweet text (first line, truncated).
|
||||
*/
|
||||
function tweetTitle(text: string): string {
|
||||
const firstLine = text.split('\n')[0].trim()
|
||||
if (!firstLine) return 'Tweet'
|
||||
return firstLine.length > 80 ? `${firstLine.slice(0, 77)}...` : firstLine
|
||||
}
|
||||
|
||||
/**
|
||||
* Converts a tweet (and its resolved author) into an ExternalDocument with
|
||||
* inline content — the list API returns full text, so no deferral is needed.
|
||||
*
|
||||
* The author is the actual tweet author resolved from the `author_id` expansion,
|
||||
* not the credential owner — important for bookmarks and likes, where most posts
|
||||
* belong to other accounts.
|
||||
*/
|
||||
function tweetToDocument(tweet: XTweet, author?: XUser): ExternalDocument {
|
||||
const metrics = tweet.public_metrics ?? {}
|
||||
return {
|
||||
externalId: tweet.id,
|
||||
title: tweetTitle(tweet.text),
|
||||
content: tweet.text,
|
||||
mimeType: 'text/plain',
|
||||
sourceUrl: tweetSourceUrl(tweet.id, author?.username),
|
||||
contentHash: tweetContentHash(tweet),
|
||||
metadata: {
|
||||
author: author?.username ?? author?.name ?? undefined,
|
||||
authorName: author?.name ?? undefined,
|
||||
createdAt: tweet.created_at ?? undefined,
|
||||
likeCount: metrics.like_count ?? 0,
|
||||
retweetCount: metrics.retweet_count ?? 0,
|
||||
replyCount: metrics.reply_count ?? 0,
|
||||
quoteCount: metrics.quote_count ?? 0,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Maps tweets from a list response to documents, joining each tweet to its
|
||||
* author via the `includes.users` expansion (matched on `author_id`).
|
||||
*/
|
||||
function mapTweets(response: XListResponse): ExternalDocument[] {
|
||||
const usersById = new Map<string, XUser>()
|
||||
for (const user of response.includes?.users ?? []) {
|
||||
usersById.set(user.id, user)
|
||||
}
|
||||
const tweets = response.data ?? []
|
||||
return tweets.map((tweet) => tweetToDocument(tweet, usersById.get(tweet.author_id ?? '')))
|
||||
}
|
||||
|
||||
/**
|
||||
* Returns the API path for a given mode and resolved user ID.
|
||||
*/
|
||||
function listPathForMode(mode: SyncMode, userId: string): string {
|
||||
switch (mode) {
|
||||
case 'bookmarks':
|
||||
return `/users/${userId}/bookmarks`
|
||||
case 'likes':
|
||||
return `/users/${userId}/liked_tweets`
|
||||
case 'mentions':
|
||||
return `/users/${userId}/mentions`
|
||||
default:
|
||||
return `/users/${userId}/tweets`
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the query string for the active listing endpoint. `pageSize` is the
|
||||
* per-request `max_results`, already clamped to the endpoint's valid range and
|
||||
* to any remaining cap. `exclude` and date-range params are only attached for
|
||||
* the modes whose endpoint supports them.
|
||||
*/
|
||||
function buildListParams(
|
||||
sourceConfig: Record<string, unknown>,
|
||||
mode: SyncMode,
|
||||
pageSize: number,
|
||||
cursor?: string
|
||||
): Record<string, string> {
|
||||
const params: Record<string, string> = {
|
||||
max_results: String(pageSize),
|
||||
'tweet.fields': TWEET_FIELDS,
|
||||
expansions: 'author_id',
|
||||
'user.fields': 'name,username',
|
||||
}
|
||||
|
||||
if (EXCLUDE_CAPABLE_MODES.has(mode)) {
|
||||
const includeReplies = readBooleanOption(sourceConfig.includeReplies, false)
|
||||
const includeRetweets = readBooleanOption(sourceConfig.includeRetweets, false)
|
||||
const exclude: string[] = []
|
||||
if (!includeRetweets) exclude.push('retweets')
|
||||
if (!includeReplies) exclude.push('replies')
|
||||
if (exclude.length > 0) params.exclude = exclude.join(',')
|
||||
}
|
||||
|
||||
if (DATE_RANGE_CAPABLE_MODES.has(mode)) {
|
||||
const startTime = readTrimmed(sourceConfig.startTime)
|
||||
const endTime = readTrimmed(sourceConfig.endTime)
|
||||
if (startTime) params.start_time = startTime
|
||||
if (endTime) params.end_time = endTime
|
||||
}
|
||||
|
||||
if (cursor) params.pagination_token = cursor
|
||||
return params
|
||||
}
|
||||
|
||||
/**
|
||||
* Clamps the requested page size to the endpoint's valid range and to the number
|
||||
* of posts still needed under the cap. The user-tweets, mentions, and liked-tweets
|
||||
* endpoints require `max_results` ≥ 5; only bookmarks accepts ≥ 1. We always request
|
||||
* at least the endpoint minimum (over-fetch on the final page is trimmed afterward).
|
||||
*/
|
||||
function resolvePageSize(mode: SyncMode, remaining: number): number {
|
||||
const floor = mode === 'bookmarks' ? 1 : MIN_PAGE_SIZE
|
||||
if (remaining <= 0) return POSTS_PER_PAGE
|
||||
return Math.max(floor, Math.min(POSTS_PER_PAGE, remaining))
|
||||
}
|
||||
|
||||
export const xConnector: ConnectorConfig = {
|
||||
id: 'x',
|
||||
name: 'X',
|
||||
description: 'Sync posts from X (formerly Twitter) into your knowledge base',
|
||||
version: '1.0.0',
|
||||
icon: xIcon,
|
||||
|
||||
auth: {
|
||||
mode: 'oauth',
|
||||
provider: 'x',
|
||||
requiredScopes: ['tweet.read', 'users.read', 'bookmark.read', 'like.read', 'offline.access'],
|
||||
},
|
||||
|
||||
configFields: [
|
||||
{
|
||||
id: 'syncMode',
|
||||
title: 'Sync Mode',
|
||||
type: 'dropdown',
|
||||
required: false,
|
||||
description: 'Which posts to sync into the knowledge base',
|
||||
options: [
|
||||
{ label: 'My posts', id: 'me' },
|
||||
{ label: 'Another user', id: 'user' },
|
||||
{ label: 'My mentions', id: 'mentions' },
|
||||
{ label: 'My bookmarks', id: 'bookmarks' },
|
||||
{ label: 'My likes', id: 'likes' },
|
||||
],
|
||||
},
|
||||
{
|
||||
id: 'username',
|
||||
title: 'Username(s)',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
multi: true,
|
||||
placeholder: 'e.g. jack, xdevelopers (required for "Another user")',
|
||||
description:
|
||||
'One or more X usernames to sync posts from (comma-separated). Only used when Sync Mode is "Another user".',
|
||||
},
|
||||
{
|
||||
id: 'includeReplies',
|
||||
title: 'Include Replies',
|
||||
type: 'dropdown',
|
||||
required: false,
|
||||
options: [
|
||||
{ label: 'Exclude replies', id: 'false' },
|
||||
{ label: 'Include replies', id: 'true' },
|
||||
],
|
||||
description: 'Whether to include reply posts. Applies to "My posts" and "Another user".',
|
||||
},
|
||||
{
|
||||
id: 'includeRetweets',
|
||||
title: 'Include Retweets',
|
||||
type: 'dropdown',
|
||||
required: false,
|
||||
options: [
|
||||
{ label: 'Exclude retweets', id: 'false' },
|
||||
{ label: 'Include retweets', id: 'true' },
|
||||
],
|
||||
description: 'Whether to include retweets. Applies to "My posts" and "Another user".',
|
||||
},
|
||||
{
|
||||
id: 'startTime',
|
||||
title: 'Start Time',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
mode: 'advanced',
|
||||
placeholder: 'e.g. 2024-01-01T00:00:00Z',
|
||||
description:
|
||||
'Oldest post time (ISO 8601 UTC). Applies to posts and mentions; ignored for bookmarks and likes.',
|
||||
},
|
||||
{
|
||||
id: 'endTime',
|
||||
title: 'End Time',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
mode: 'advanced',
|
||||
placeholder: 'e.g. 2024-12-31T23:59:59Z',
|
||||
description:
|
||||
'Newest post time (ISO 8601 UTC). Applies to posts and mentions; ignored for bookmarks and likes.',
|
||||
},
|
||||
{
|
||||
id: 'maxPosts',
|
||||
title: 'Max Posts',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
placeholder: `e.g. 100 (default: ${DEFAULT_MAX_POSTS})`,
|
||||
description:
|
||||
'Maximum number of posts to sync (across all configured users). Posts beyond this limit are not deleted from the knowledge base; X also only exposes a limited recent window (≈3,200 timeline posts, ≈800 bookmarks), so posts that age out of that window are removed on the next sync.',
|
||||
},
|
||||
],
|
||||
|
||||
listDocuments: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
cursor?: string,
|
||||
syncContext?: Record<string, unknown>
|
||||
): Promise<ExternalDocumentList> => {
|
||||
const mode = resolveSyncMode(sourceConfig)
|
||||
const maxPosts = sourceConfig.maxPosts ? Number(sourceConfig.maxPosts) : DEFAULT_MAX_POSTS
|
||||
|
||||
const collectedSoFar = (syncContext?.collected as number) ?? 0
|
||||
if (maxPosts > 0 && collectedSoFar >= maxPosts) {
|
||||
return { documents: [], hasMore: false }
|
||||
}
|
||||
|
||||
// For the multi-username "user" mode, walk one username per cursor cycle. The
|
||||
// cursor packs the username index and that user's pagination token; the shared
|
||||
// cap is enforced across all users via syncContext.collected.
|
||||
const usernames = mode === 'user' ? parseUsernames(sourceConfig.username) : []
|
||||
if (mode === 'user' && usernames.length === 0) {
|
||||
throw new Error('Username is required when Sync Mode is "Another user"')
|
||||
}
|
||||
|
||||
let userIndex = 0
|
||||
let pageToken = cursor
|
||||
if (mode === 'user' && cursor) {
|
||||
const sep = cursor.indexOf(':')
|
||||
if (sep >= 0) {
|
||||
userIndex = Number(cursor.slice(0, sep)) || 0
|
||||
const token = cursor.slice(sep + 1)
|
||||
pageToken = token.length > 0 ? token : undefined
|
||||
}
|
||||
}
|
||||
|
||||
// Resolve the target user ID. For `user` mode it depends on the current index
|
||||
// (resolved per page, cheap); for self-modes it is cached on syncContext.
|
||||
let userId: string
|
||||
if (mode === 'user') {
|
||||
userId = await resolveUsernameId(accessToken, usernames[userIndex])
|
||||
} else {
|
||||
userId = (syncContext?.userId as string | undefined) ?? (await resolveMyUserId(accessToken))
|
||||
if (syncContext) syncContext.userId = userId
|
||||
}
|
||||
|
||||
const remaining = maxPosts > 0 ? maxPosts - collectedSoFar : 0
|
||||
const pageSize = resolvePageSize(mode, remaining)
|
||||
const path = listPathForMode(mode, userId)
|
||||
const params = buildListParams(sourceConfig, mode, pageSize, pageToken)
|
||||
|
||||
logger.info('Syncing X posts', { mode, userId, userIndex, maxPosts })
|
||||
|
||||
const response = (await xApiGet(path, accessToken, params)) as XListResponse
|
||||
if (response.errors?.length && !response.data) {
|
||||
throw new Error(response.errors[0]?.detail || response.errors[0]?.title || 'X API error')
|
||||
}
|
||||
|
||||
let documents = mapTweets(response)
|
||||
|
||||
if (maxPosts > 0 && collectedSoFar + documents.length > maxPosts) {
|
||||
documents = documents.slice(0, maxPosts - collectedSoFar)
|
||||
}
|
||||
const newCollected = collectedSoFar + documents.length
|
||||
if (syncContext) syncContext.collected = newCollected
|
||||
|
||||
const capReached = maxPosts > 0 && newCollected >= maxPosts
|
||||
const nextToken = response.meta?.next_token
|
||||
|
||||
// Advance pagination: continue the current user's pages, else move to the next
|
||||
// username (user mode), else stop.
|
||||
if (capReached) {
|
||||
// We stopped before exhausting the source, so the listing is incomplete:
|
||||
// older previously-synced posts may still exist beyond the `maxPosts` cap.
|
||||
// Flag the sync as capped so the engine skips deletion reconciliation and
|
||||
// does not soft-delete posts that simply fell outside this run's window.
|
||||
// A forced full sync bypasses this guard and reconciles normally.
|
||||
if (syncContext) syncContext.listingCapped = true
|
||||
return { documents, hasMore: false }
|
||||
}
|
||||
|
||||
if (mode === 'user') {
|
||||
if (nextToken) {
|
||||
return { documents, nextCursor: `${userIndex}:${nextToken}`, hasMore: true }
|
||||
}
|
||||
const nextUserIndex = userIndex + 1
|
||||
if (nextUserIndex < usernames.length) {
|
||||
return { documents, nextCursor: `${nextUserIndex}:`, hasMore: true }
|
||||
}
|
||||
return { documents, hasMore: false }
|
||||
}
|
||||
|
||||
return {
|
||||
documents,
|
||||
nextCursor: nextToken ?? undefined,
|
||||
hasMore: Boolean(nextToken),
|
||||
}
|
||||
},
|
||||
|
||||
getDocument: async (
|
||||
accessToken: string,
|
||||
_sourceConfig: Record<string, unknown>,
|
||||
externalId: string
|
||||
): Promise<ExternalDocument | null> => {
|
||||
try {
|
||||
const response = (await xApiGet(`/tweets/${encodeURIComponent(externalId)}`, accessToken, {
|
||||
'tweet.fields': TWEET_FIELDS,
|
||||
expansions: 'author_id',
|
||||
'user.fields': 'name,username',
|
||||
})) as XSingleResponse
|
||||
|
||||
const tweet = response.data
|
||||
if (!tweet) return null
|
||||
|
||||
const author = response.includes?.users?.find((u) => u.id === tweet.author_id)
|
||||
return tweetToDocument(tweet, author)
|
||||
} catch (error) {
|
||||
logger.warn('Failed to get X tweet document', {
|
||||
externalId,
|
||||
error: toError(error).message,
|
||||
})
|
||||
return null
|
||||
}
|
||||
},
|
||||
|
||||
validateConfig: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>
|
||||
): Promise<{ valid: boolean; error?: string }> => {
|
||||
const mode = resolveSyncMode(sourceConfig)
|
||||
const usernames = mode === 'user' ? parseUsernames(sourceConfig.username) : []
|
||||
const maxPosts = sourceConfig.maxPosts as string | undefined
|
||||
|
||||
if (mode === 'user' && usernames.length === 0) {
|
||||
return { valid: false, error: 'Username is required when Sync Mode is "Another user"' }
|
||||
}
|
||||
|
||||
if (maxPosts && (Number.isNaN(Number(maxPosts)) || Number(maxPosts) <= 0)) {
|
||||
return { valid: false, error: 'Max posts must be a positive number' }
|
||||
}
|
||||
|
||||
const startTime = readTrimmed(sourceConfig.startTime)
|
||||
if (startTime && Number.isNaN(new Date(startTime).getTime())) {
|
||||
return { valid: false, error: 'Start Time must be a valid ISO 8601 timestamp' }
|
||||
}
|
||||
const endTime = readTrimmed(sourceConfig.endTime)
|
||||
if (endTime && Number.isNaN(new Date(endTime).getTime())) {
|
||||
return { valid: false, error: 'End Time must be a valid ISO 8601 timestamp' }
|
||||
}
|
||||
|
||||
try {
|
||||
await resolveMyUserId(accessToken, VALIDATE_RETRY_OPTIONS)
|
||||
|
||||
if (mode === 'user') {
|
||||
for (const username of usernames) {
|
||||
await resolveUsernameId(accessToken, username, VALIDATE_RETRY_OPTIONS)
|
||||
}
|
||||
}
|
||||
|
||||
return { valid: true }
|
||||
} catch (error) {
|
||||
return { valid: false, error: getErrorMessage(error, 'Failed to validate configuration') }
|
||||
}
|
||||
},
|
||||
|
||||
tagDefinitions: [
|
||||
{ id: 'author', displayName: 'Author', fieldType: 'text' },
|
||||
{ id: 'createdAt', displayName: 'Created Date', fieldType: 'date' },
|
||||
{ id: 'likeCount', displayName: 'Like Count', fieldType: 'number' },
|
||||
{ id: 'retweetCount', displayName: 'Retweet Count', fieldType: 'number' },
|
||||
],
|
||||
|
||||
mapTags: (metadata: Record<string, unknown>): Record<string, unknown> => {
|
||||
const result: Record<string, unknown> = {}
|
||||
|
||||
if (typeof metadata.author === 'string') {
|
||||
result.author = metadata.author
|
||||
}
|
||||
|
||||
const createdAt = parseTagDate(metadata.createdAt)
|
||||
if (createdAt) {
|
||||
result.createdAt = createdAt
|
||||
}
|
||||
|
||||
if (metadata.likeCount != null) {
|
||||
const num = Number(metadata.likeCount)
|
||||
if (!Number.isNaN(num)) result.likeCount = num
|
||||
}
|
||||
|
||||
if (metadata.retweetCount != null) {
|
||||
const num = Number(metadata.retweetCount)
|
||||
if (!Number.isNaN(num)) result.retweetCount = num
|
||||
}
|
||||
|
||||
return result
|
||||
},
|
||||
}
|
||||
@@ -0,0 +1 @@
|
||||
export { youtubeConnector } from '@/connectors/youtube/youtube'
|
||||
@@ -0,0 +1,650 @@
|
||||
import { createLogger } from '@sim/logger'
|
||||
import { getErrorMessage, toError } from '@sim/utils/errors'
|
||||
import { YouTubeIcon } from '@/components/icons'
|
||||
import { fetchWithRetry, VALIDATE_RETRY_OPTIONS } from '@/lib/knowledge/documents/utils'
|
||||
import type { ConnectorConfig, ExternalDocument, ExternalDocumentList } from '@/connectors/types'
|
||||
import { joinTagArray, parseTagDate } from '@/connectors/utils'
|
||||
|
||||
const logger = createLogger('YouTubeConnector')
|
||||
|
||||
const YOUTUBE_API_BASE = 'https://www.googleapis.com/youtube/v3'
|
||||
|
||||
/** Max videos fetched per `playlistItems.list` page (YouTube hard limit is 50). */
|
||||
const PAGE_SIZE = 50
|
||||
|
||||
/** Videos shorter than this (seconds) are treated as Shorts when the exclude filter is on. */
|
||||
const SHORTS_MAX_DURATION_SECONDS = 60
|
||||
|
||||
/**
|
||||
* Minimal `playlistItems.list` item shape we consume.
|
||||
* `contentDetails.videoId` is the stable video identifier; `snippet.resourceId.videoId`
|
||||
* is used as a fallback for older API responses.
|
||||
*
|
||||
* `snippet.publishedAt` is the time the item was ADDED to the playlist, whereas
|
||||
* `contentDetails.videoPublishedAt` is the time the VIDEO was published to YouTube.
|
||||
* These differ for hand-curated playlists, so only `videoPublishedAt` is used for the
|
||||
* change-detection hash (it matches `videos.list` `snippet.publishedAt`).
|
||||
*/
|
||||
interface PlaylistItem {
|
||||
contentDetails?: { videoId?: string; videoPublishedAt?: string }
|
||||
snippet?: {
|
||||
title?: string
|
||||
publishedAt?: string
|
||||
channelTitle?: string
|
||||
videoOwnerChannelTitle?: string
|
||||
resourceId?: { videoId?: string }
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Minimal `videos.list` item shape we consume in `getDocument`.
|
||||
*/
|
||||
interface VideoItem {
|
||||
id?: string
|
||||
snippet?: {
|
||||
title?: string
|
||||
description?: string
|
||||
publishedAt?: string
|
||||
channelTitle?: string
|
||||
tags?: string[]
|
||||
categoryId?: string
|
||||
}
|
||||
contentDetails?: { duration?: string }
|
||||
status?: { privacyStatus?: string }
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves the API key from the access token the sync engine provides.
|
||||
* In `apiKey` mode the engine decrypts the stored key and passes it as `accessToken`.
|
||||
*/
|
||||
function getApiKey(accessToken: string): string {
|
||||
return accessToken.trim()
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the change-detection hash for a video.
|
||||
*
|
||||
* The hash is keyed on the video's own publish time (`videos.list` `snippet.publishedAt`
|
||||
* / playlistItem `contentDetails.videoPublishedAt`), which is identical on both the
|
||||
* listing stub and the hydrated document — guaranteeing the stub/getDocument hash
|
||||
* invariant. The playlist-item "added at" time (`snippet.publishedAt`) is deliberately
|
||||
* NOT used, since `getDocument` (via `videos.list`) cannot reproduce it.
|
||||
*
|
||||
* YouTube exposes no field that reliably changes when a video's title/description is
|
||||
* edited, so edits to already-synced videos are not detected — only new videos are
|
||||
* picked up. This is a known limitation of the API-key data surface.
|
||||
*/
|
||||
function buildContentHash(videoId: string, videoPublishedAt: string): string {
|
||||
return `youtube:${videoId}:${videoPublishedAt}`
|
||||
}
|
||||
|
||||
/**
|
||||
* Parses an ISO 8601 duration (e.g. `PT1M30S`, `PT2H`, `P1DT2H`) into total seconds.
|
||||
* Returns null when the value is missing or unparseable.
|
||||
*/
|
||||
function parseIso8601Duration(value: string | undefined): number | null {
|
||||
if (!value) return null
|
||||
const match = value.match(/^P(?:(\d+)D)?(?:T(?:(\d+)H)?(?:(\d+)M)?(?:(\d+)S)?)?$/)
|
||||
if (!match) return null
|
||||
const [, days, hours, minutes, seconds] = match
|
||||
if (!days && !hours && !minutes && !seconds) return null
|
||||
return (
|
||||
Number(days ?? 0) * 86400 +
|
||||
Number(hours ?? 0) * 3600 +
|
||||
Number(minutes ?? 0) * 60 +
|
||||
Number(seconds ?? 0)
|
||||
)
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves a channel reference to its "uploads" playlist ID via `channels.list`.
|
||||
*
|
||||
* Accepts a `UC…` channel ID, an `@handle` (resolved with `forHandle`), or a legacy
|
||||
* username (resolved with `forUsername`). Returns null when the channel is missing or
|
||||
* has no uploads playlist.
|
||||
*/
|
||||
async function resolveUploadsPlaylistId(
|
||||
apiKey: string,
|
||||
channelRef: string,
|
||||
retryOptions?: Parameters<typeof fetchWithRetry>[2]
|
||||
): Promise<string | null> {
|
||||
const ref = channelRef.trim()
|
||||
if (!ref) return null
|
||||
|
||||
const params = new URLSearchParams({ part: 'contentDetails', key: apiKey })
|
||||
if (ref.startsWith('@')) {
|
||||
params.set('forHandle', ref)
|
||||
} else if (/^UC[\w-]{20,}$/.test(ref)) {
|
||||
params.set('id', ref)
|
||||
} else {
|
||||
params.set('forUsername', ref)
|
||||
}
|
||||
|
||||
const url = `${YOUTUBE_API_BASE}/channels?${params.toString()}`
|
||||
|
||||
const response = await fetchWithRetry(
|
||||
url,
|
||||
{ method: 'GET', headers: { Accept: 'application/json' } },
|
||||
retryOptions
|
||||
)
|
||||
|
||||
if (!response.ok) {
|
||||
const errorText = await response.text().catch(() => '')
|
||||
logger.error('Failed to resolve channel uploads playlist', {
|
||||
channelRef: ref,
|
||||
status: response.status,
|
||||
error: errorText.slice(0, 500),
|
||||
})
|
||||
throw new Error(`Failed to resolve channel: ${response.status}`)
|
||||
}
|
||||
|
||||
const data = await response.json()
|
||||
const items = (data.items ?? []) as Array<{
|
||||
contentDetails?: { relatedPlaylists?: { uploads?: string } }
|
||||
}>
|
||||
return items[0]?.contentDetails?.relatedPlaylists?.uploads ?? null
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves the effective playlist ID to sync from sourceConfig, and whether the source
|
||||
* is a channel's reverse-chronological uploads playlist (which enables early-stop for
|
||||
* the `publishedAfter` filter). A `playlistId` takes precedence over a `channelId`.
|
||||
*/
|
||||
async function resolvePlaylistId(
|
||||
apiKey: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
retryOptions?: Parameters<typeof fetchWithRetry>[2]
|
||||
): Promise<{ playlistId: string | null; isUploadsPlaylist: boolean }> {
|
||||
const playlistId = (sourceConfig.playlistId as string | undefined)?.trim()
|
||||
if (playlistId) return { playlistId, isUploadsPlaylist: false }
|
||||
|
||||
const channelId = (sourceConfig.channelId as string | undefined)?.trim()
|
||||
if (channelId) {
|
||||
const resolved = await resolveUploadsPlaylistId(apiKey, channelId, retryOptions)
|
||||
return { playlistId: resolved, isUploadsPlaylist: resolved != null }
|
||||
}
|
||||
|
||||
return { playlistId: null, isUploadsPlaylist: false }
|
||||
}
|
||||
|
||||
/**
|
||||
* Extracts the video ID from a playlist item, preferring the stable
|
||||
* `contentDetails.videoId` over the legacy `snippet.resourceId.videoId`.
|
||||
*/
|
||||
function getVideoId(item: PlaylistItem): string {
|
||||
return item.contentDetails?.videoId ?? item.snippet?.resourceId?.videoId ?? ''
|
||||
}
|
||||
|
||||
/**
|
||||
* Reads the optional `publishedAfter` cutoff from sourceConfig as a timestamp (ms),
|
||||
* or null when unset/invalid.
|
||||
*/
|
||||
function getPublishedAfter(sourceConfig: Record<string, unknown>): number | null {
|
||||
const raw = (sourceConfig.publishedAfter as string | undefined)?.trim()
|
||||
if (!raw) return null
|
||||
const ms = new Date(raw).getTime()
|
||||
return Number.isNaN(ms) ? null : ms
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds a metadata-only stub from a playlist item.
|
||||
*
|
||||
* Duration/tags/category are not available on `playlistItems.list` — they are populated
|
||||
* during hydration in `getDocument` via `videos.list`. The content hash uses the video's
|
||||
* publish time only, so it is identical between this stub and the hydrated document.
|
||||
*/
|
||||
function itemToStub(item: PlaylistItem): ExternalDocument | null {
|
||||
const videoId = getVideoId(item)
|
||||
if (!videoId) return null
|
||||
|
||||
const snippet = item.snippet ?? {}
|
||||
const videoPublishedAt = item.contentDetails?.videoPublishedAt ?? ''
|
||||
const channelTitle = snippet.videoOwnerChannelTitle ?? snippet.channelTitle ?? ''
|
||||
|
||||
return {
|
||||
externalId: videoId,
|
||||
title: snippet.title || 'Untitled',
|
||||
content: '',
|
||||
contentDeferred: true,
|
||||
mimeType: 'text/plain',
|
||||
sourceUrl: `https://www.youtube.com/watch?v=${videoId}`,
|
||||
contentHash: buildContentHash(videoId, videoPublishedAt),
|
||||
metadata: {
|
||||
channelTitle,
|
||||
publishedAt: videoPublishedAt,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Batch-fetches full video resources for the given IDs via `videos.list`.
|
||||
*
|
||||
* `videos.list` accepts up to 50 comma-separated IDs and costs a flat 1 quota unit per
|
||||
* call regardless of ID count, so a single call covers a full `playlistItems.list` page.
|
||||
* Videos that are private, deleted, or region-blocked are simply absent from the response.
|
||||
*/
|
||||
async function fetchVideosByIds(
|
||||
apiKey: string,
|
||||
videoIds: string[],
|
||||
retryOptions?: Parameters<typeof fetchWithRetry>[2]
|
||||
): Promise<Map<string, VideoItem>> {
|
||||
const result = new Map<string, VideoItem>()
|
||||
if (videoIds.length === 0) return result
|
||||
|
||||
const url = `${YOUTUBE_API_BASE}/videos?part=snippet,contentDetails,status&id=${encodeURIComponent(
|
||||
videoIds.join(',')
|
||||
)}&key=${encodeURIComponent(apiKey)}`
|
||||
|
||||
const response = await fetchWithRetry(
|
||||
url,
|
||||
{ method: 'GET', headers: { Accept: 'application/json' } },
|
||||
retryOptions
|
||||
)
|
||||
|
||||
if (!response.ok) {
|
||||
const errorText = await response.text().catch(() => '')
|
||||
logger.error('Failed to batch-fetch YouTube videos', {
|
||||
count: videoIds.length,
|
||||
status: response.status,
|
||||
error: errorText.slice(0, 500),
|
||||
})
|
||||
throw new Error(`Failed to batch-fetch YouTube videos: ${response.status}`)
|
||||
}
|
||||
|
||||
const data = await response.json()
|
||||
const items = (data.items ?? []) as VideoItem[]
|
||||
for (const item of items) {
|
||||
if (item.id) result.set(item.id, item)
|
||||
}
|
||||
return result
|
||||
}
|
||||
|
||||
/**
|
||||
* Builds the full document for a video, combining title and description as plain-text
|
||||
* content. Returns null for unlisted/private/deleted videos, and (when configured) for
|
||||
* Shorts shorter than 60 seconds.
|
||||
*
|
||||
* Captions/transcripts are intentionally not fetched: `captions.download` requires OAuth
|
||||
* as the video owner, which the API-key auth surface cannot provide. Content is therefore
|
||||
* the video title plus description only.
|
||||
*/
|
||||
function videoToDocument(video: VideoItem, excludeShorts: boolean): ExternalDocument | null {
|
||||
const videoId = video.id
|
||||
if (!videoId) return null
|
||||
|
||||
const privacyStatus = video.status?.privacyStatus
|
||||
if (privacyStatus && privacyStatus !== 'public' && privacyStatus !== 'unlisted') {
|
||||
return null
|
||||
}
|
||||
|
||||
if (excludeShorts) {
|
||||
const seconds = parseIso8601Duration(video.contentDetails?.duration)
|
||||
if (seconds != null && seconds > 0 && seconds < SHORTS_MAX_DURATION_SECONDS) {
|
||||
return null
|
||||
}
|
||||
}
|
||||
|
||||
const snippet = video.snippet ?? {}
|
||||
const title = snippet.title || 'Untitled'
|
||||
const description = snippet.description ?? ''
|
||||
const publishedAt = snippet.publishedAt ?? ''
|
||||
const content = description.trim() ? `${title}\n\n${description}` : title
|
||||
const tags = Array.isArray(snippet.tags) ? snippet.tags : []
|
||||
|
||||
return {
|
||||
externalId: videoId,
|
||||
title,
|
||||
content,
|
||||
contentDeferred: false,
|
||||
mimeType: 'text/plain',
|
||||
sourceUrl: `https://www.youtube.com/watch?v=${videoId}`,
|
||||
contentHash: buildContentHash(videoId, publishedAt),
|
||||
metadata: {
|
||||
channelTitle: snippet.channelTitle ?? '',
|
||||
publishedAt,
|
||||
duration: video.contentDetails?.duration ?? '',
|
||||
categoryId: snippet.categoryId ?? '',
|
||||
tags,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
export const youtubeConnector: ConnectorConfig = {
|
||||
id: 'youtube',
|
||||
name: 'YouTube',
|
||||
description: 'Sync videos from a YouTube channel or playlist into your knowledge base',
|
||||
version: '1.0.0',
|
||||
icon: YouTubeIcon,
|
||||
|
||||
auth: {
|
||||
mode: 'apiKey',
|
||||
label: 'YouTube Data API Key',
|
||||
placeholder: 'Enter your YouTube Data API v3 key',
|
||||
},
|
||||
|
||||
configFields: [
|
||||
{
|
||||
id: 'channelId',
|
||||
title: 'Channel',
|
||||
type: 'short-input',
|
||||
placeholder: 'e.g. @mkbhd or UCXXXXXXXXXXXXXXXXXXXXXX',
|
||||
required: false,
|
||||
description:
|
||||
'Channel handle (@name), channel ID (starts with "UC"), or legacy username. Syncs the channel\'s uploaded videos.',
|
||||
},
|
||||
{
|
||||
id: 'playlistId',
|
||||
title: 'Playlist ID',
|
||||
type: 'short-input',
|
||||
placeholder: 'e.g. PLXXXXXXXXXXXXXXXX',
|
||||
required: false,
|
||||
description: 'Playlist ID. Takes precedence over Channel when both are set.',
|
||||
},
|
||||
{
|
||||
id: 'publishedAfter',
|
||||
title: 'Published After',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
mode: 'advanced',
|
||||
placeholder: 'e.g. 2024-01-01',
|
||||
description:
|
||||
'Only sync videos published on or after this date (ISO 8601, e.g. 2024-01-01). Applies to the video publish date.',
|
||||
},
|
||||
{
|
||||
id: 'excludeShorts',
|
||||
title: 'Exclude Shorts',
|
||||
type: 'dropdown',
|
||||
required: false,
|
||||
mode: 'advanced',
|
||||
options: [
|
||||
{ label: 'Include Shorts', id: 'false' },
|
||||
{ label: 'Exclude Shorts (< 60s)', id: 'true' },
|
||||
],
|
||||
description: 'Skip videos shorter than 60 seconds (Shorts).',
|
||||
},
|
||||
{
|
||||
id: 'maxVideos',
|
||||
title: 'Max Videos',
|
||||
type: 'short-input',
|
||||
required: false,
|
||||
placeholder: 'e.g. 500 (default: unlimited)',
|
||||
},
|
||||
],
|
||||
|
||||
listDocuments: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
cursor?: string,
|
||||
syncContext?: Record<string, unknown>
|
||||
): Promise<ExternalDocumentList> => {
|
||||
const apiKey = getApiKey(accessToken)
|
||||
|
||||
const maxVideos = sourceConfig.maxVideos ? Number(sourceConfig.maxVideos) : 0
|
||||
const previouslyFetched = (syncContext?.totalDocsFetched as number) ?? 0
|
||||
|
||||
if (maxVideos > 0 && previouslyFetched >= maxVideos) {
|
||||
return { documents: [], hasMore: false }
|
||||
}
|
||||
|
||||
const cachedPlaylistId = syncContext?.resolvedPlaylistId as string | undefined
|
||||
let playlistId: string | null = cachedPlaylistId ?? null
|
||||
let isUploadsPlaylist = (syncContext?.isUploadsPlaylist as boolean | undefined) ?? false
|
||||
|
||||
if (!playlistId) {
|
||||
const resolved = await resolvePlaylistId(apiKey, sourceConfig)
|
||||
playlistId = resolved.playlistId
|
||||
isUploadsPlaylist = resolved.isUploadsPlaylist
|
||||
if (syncContext) {
|
||||
if (playlistId) syncContext.resolvedPlaylistId = playlistId
|
||||
syncContext.isUploadsPlaylist = isUploadsPlaylist
|
||||
}
|
||||
}
|
||||
|
||||
if (!playlistId) {
|
||||
throw new Error('No playlistId or channelId configured, or channel has no uploads playlist')
|
||||
}
|
||||
|
||||
const publishedAfter = getPublishedAfter(sourceConfig)
|
||||
|
||||
const remaining = maxVideos > 0 ? maxVideos - previouslyFetched : 0
|
||||
const effectivePageSize = maxVideos > 0 ? Math.min(PAGE_SIZE, remaining) : PAGE_SIZE
|
||||
|
||||
const queryParams = new URLSearchParams({
|
||||
part: 'snippet,contentDetails',
|
||||
playlistId,
|
||||
maxResults: String(effectivePageSize),
|
||||
key: apiKey,
|
||||
})
|
||||
if (cursor) queryParams.set('pageToken', cursor)
|
||||
|
||||
const url = `${YOUTUBE_API_BASE}/playlistItems?${queryParams.toString()}`
|
||||
|
||||
logger.info('Listing YouTube playlist items', { playlistId, cursor: cursor ?? 'initial' })
|
||||
|
||||
const response = await fetchWithRetry(url, {
|
||||
method: 'GET',
|
||||
headers: { Accept: 'application/json' },
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
const errorText = await response.text().catch(() => '')
|
||||
logger.error('Failed to list YouTube playlist items', {
|
||||
playlistId,
|
||||
status: response.status,
|
||||
error: errorText.slice(0, 500),
|
||||
})
|
||||
throw new Error(`Failed to list YouTube playlist items: ${response.status}`)
|
||||
}
|
||||
|
||||
const data = await response.json()
|
||||
const items = (data.items ?? []) as PlaylistItem[]
|
||||
const excludeShorts = String(sourceConfig.excludeShorts ?? '') === 'true'
|
||||
|
||||
const keptItems: PlaylistItem[] = []
|
||||
let stopEarly = false
|
||||
|
||||
for (const item of items) {
|
||||
if (!getVideoId(item)) continue
|
||||
|
||||
if (publishedAfter != null) {
|
||||
const videoPublishedAt = item.contentDetails?.videoPublishedAt
|
||||
const ms = videoPublishedAt ? new Date(videoPublishedAt).getTime() : Number.NaN
|
||||
if (!Number.isNaN(ms) && ms < publishedAfter) {
|
||||
// Uploads playlists are reverse-chronological by publish date, so once we
|
||||
// cross the cutoff no later item can qualify — stop paginating. For arbitrary
|
||||
// playlists we only filter per-item (order is not guaranteed).
|
||||
if (isUploadsPlaylist) {
|
||||
stopEarly = true
|
||||
break
|
||||
}
|
||||
continue
|
||||
}
|
||||
}
|
||||
|
||||
keptItems.push(item)
|
||||
}
|
||||
|
||||
let documents: ExternalDocument[] = []
|
||||
|
||||
if (excludeShorts && keptItems.length > 0) {
|
||||
// When excluding Shorts we must know each video's duration, which is not exposed on
|
||||
// `playlistItems.list`. Resolve it here with a single batched `videos.list` call
|
||||
// (1 quota unit per page) and emit FULLY-HYDRATED documents. This is deliberate:
|
||||
// emitting deferred stubs for Shorts would make every excluded Short re-list as a
|
||||
// brand-new doc on every sync (it is never persisted), re-hydrating to null forever.
|
||||
// Filtering at listing time bounds the cost to one batched call per page per sync.
|
||||
const videoMap = await fetchVideosByIds(apiKey, keptItems.map(getVideoId))
|
||||
for (const item of keptItems) {
|
||||
const video = videoMap.get(getVideoId(item))
|
||||
// Absent from `videos.list` => private/deleted/region-blocked. Drop it instead of
|
||||
// emitting a stub that would re-hydrate to null on every sync.
|
||||
if (!video) continue
|
||||
const doc = videoToDocument(video, true)
|
||||
if (doc) documents.push(doc)
|
||||
}
|
||||
} else {
|
||||
for (const item of keptItems) {
|
||||
const stub = itemToStub(item)
|
||||
if (stub) documents.push(stub)
|
||||
}
|
||||
}
|
||||
|
||||
const totalFetched = previouslyFetched + documents.length
|
||||
if (syncContext) syncContext.totalDocsFetched = totalFetched
|
||||
|
||||
const hitMax = maxVideos > 0 && totalFetched >= maxVideos
|
||||
if (hitMax && maxVideos > 0) {
|
||||
const overflow = totalFetched - maxVideos
|
||||
if (overflow > 0) documents = documents.slice(0, documents.length - overflow)
|
||||
if (syncContext) syncContext.totalDocsFetched = maxVideos
|
||||
}
|
||||
|
||||
const nextPageToken = data.nextPageToken as string | undefined
|
||||
|
||||
// When the `maxVideos` cap stops the listing before the source is exhausted, mark the
|
||||
// listing as capped so the sync engine does not delete still-present-but-unlisted
|
||||
// videos from the knowledge base. `stopEarly` (publishedAfter cutoff) is NOT a cap —
|
||||
// every remaining video is older than the cutoff and intentionally out of scope, so
|
||||
// those should reconcile (delete) normally.
|
||||
if (hitMax && Boolean(nextPageToken) && syncContext) {
|
||||
syncContext.listingCapped = true
|
||||
}
|
||||
|
||||
const hasMore = !hitMax && !stopEarly && Boolean(nextPageToken)
|
||||
|
||||
return {
|
||||
documents,
|
||||
nextCursor: hasMore ? nextPageToken : undefined,
|
||||
hasMore,
|
||||
}
|
||||
},
|
||||
|
||||
getDocument: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>,
|
||||
externalId: string
|
||||
): Promise<ExternalDocument | null> => {
|
||||
const apiKey = getApiKey(accessToken)
|
||||
const excludeShorts = String(sourceConfig.excludeShorts ?? '') === 'true'
|
||||
|
||||
const url = `${YOUTUBE_API_BASE}/videos?part=snippet,contentDetails,status&id=${encodeURIComponent(externalId)}&key=${encodeURIComponent(apiKey)}`
|
||||
|
||||
try {
|
||||
const response = await fetchWithRetry(url, {
|
||||
method: 'GET',
|
||||
headers: { Accept: 'application/json' },
|
||||
})
|
||||
|
||||
if (!response.ok) {
|
||||
if (response.status === 403 || response.status === 404) return null
|
||||
throw new Error(`Failed to get YouTube video: ${response.status}`)
|
||||
}
|
||||
|
||||
const data = await response.json()
|
||||
const items = (data.items ?? []) as VideoItem[]
|
||||
const video = items[0]
|
||||
|
||||
// An empty items array means the video is deleted, private, or region-blocked.
|
||||
if (!video) return null
|
||||
|
||||
return videoToDocument(video, excludeShorts)
|
||||
} catch (error) {
|
||||
logger.warn(`Failed to fetch YouTube video ${externalId}`, { error: toError(error).message })
|
||||
return null
|
||||
}
|
||||
},
|
||||
|
||||
validateConfig: async (
|
||||
accessToken: string,
|
||||
sourceConfig: Record<string, unknown>
|
||||
): Promise<{ valid: boolean; error?: string }> => {
|
||||
const apiKey = getApiKey(accessToken)
|
||||
if (!apiKey) {
|
||||
return { valid: false, error: 'A YouTube Data API key is required' }
|
||||
}
|
||||
|
||||
const channelId = (sourceConfig.channelId as string | undefined)?.trim()
|
||||
const playlistId = (sourceConfig.playlistId as string | undefined)?.trim()
|
||||
if (!channelId && !playlistId) {
|
||||
return { valid: false, error: 'Provide a channel or a playlistId' }
|
||||
}
|
||||
|
||||
const maxVideos = sourceConfig.maxVideos as string | undefined
|
||||
if (maxVideos && (Number.isNaN(Number(maxVideos)) || Number(maxVideos) <= 0)) {
|
||||
return { valid: false, error: 'Max videos must be a positive number' }
|
||||
}
|
||||
|
||||
const publishedAfterRaw = (sourceConfig.publishedAfter as string | undefined)?.trim()
|
||||
if (publishedAfterRaw && Number.isNaN(new Date(publishedAfterRaw).getTime())) {
|
||||
return { valid: false, error: 'Published After must be a valid date (e.g. 2024-01-01)' }
|
||||
}
|
||||
|
||||
try {
|
||||
const resolvedPlaylistId = playlistId
|
||||
? playlistId
|
||||
: await resolveUploadsPlaylistId(apiKey, channelId as string, VALIDATE_RETRY_OPTIONS)
|
||||
|
||||
if (!resolvedPlaylistId) {
|
||||
return { valid: false, error: 'Channel not found or has no uploaded videos' }
|
||||
}
|
||||
|
||||
const url = `${YOUTUBE_API_BASE}/playlistItems?part=id&maxResults=1&playlistId=${encodeURIComponent(resolvedPlaylistId)}&key=${encodeURIComponent(apiKey)}`
|
||||
const response = await fetchWithRetry(
|
||||
url,
|
||||
{ method: 'GET', headers: { Accept: 'application/json' } },
|
||||
VALIDATE_RETRY_OPTIONS
|
||||
)
|
||||
|
||||
if (!response.ok) {
|
||||
if (response.status === 403) {
|
||||
return {
|
||||
valid: false,
|
||||
error:
|
||||
'API key rejected. Check that the key is valid, has no HTTP referrer/IP restrictions (server-side use requires an unrestricted or IP-allowed key), and that your daily quota is not exhausted.',
|
||||
}
|
||||
}
|
||||
if (response.status === 404) {
|
||||
return { valid: false, error: 'Playlist not found. Check the playlist or channel ID.' }
|
||||
}
|
||||
return { valid: false, error: `Failed to access YouTube: ${response.status}` }
|
||||
}
|
||||
|
||||
return { valid: true }
|
||||
} catch (error) {
|
||||
return { valid: false, error: getErrorMessage(error, 'Failed to validate configuration') }
|
||||
}
|
||||
},
|
||||
|
||||
tagDefinitions: [
|
||||
{ id: 'channelTitle', displayName: 'Channel', fieldType: 'text' },
|
||||
{ id: 'publishedAt', displayName: 'Published Date', fieldType: 'date' },
|
||||
{ id: 'duration', displayName: 'Duration', fieldType: 'text' },
|
||||
{ id: 'tags', displayName: 'Tags', fieldType: 'text' },
|
||||
],
|
||||
|
||||
/**
|
||||
* Maps document metadata to tag slots. `duration` and `tags` are only present after
|
||||
* hydration in `getDocument`; on the listing stub they are absent and simply skipped
|
||||
* by the guards below. The sync engine only runs `mapTags` on add/update (after
|
||||
* hydration), so durations/tags are populated when tags are actually written.
|
||||
*/
|
||||
mapTags: (metadata: Record<string, unknown>): Record<string, unknown> => {
|
||||
const result: Record<string, unknown> = {}
|
||||
|
||||
if (typeof metadata.channelTitle === 'string' && metadata.channelTitle.trim()) {
|
||||
result.channelTitle = metadata.channelTitle
|
||||
}
|
||||
|
||||
const publishedAt = parseTagDate(metadata.publishedAt)
|
||||
if (publishedAt) result.publishedAt = publishedAt
|
||||
|
||||
if (typeof metadata.duration === 'string' && metadata.duration.trim()) {
|
||||
result.duration = metadata.duration
|
||||
}
|
||||
|
||||
const tags = joinTagArray(metadata.tags)
|
||||
if (tags) result.tags = tags
|
||||
|
||||
return result
|
||||
},
|
||||
}
|
||||
@@ -37,6 +37,46 @@ vi.mock('@/connectors/registry', () => ({
|
||||
},
|
||||
}))
|
||||
|
||||
describe('shouldReconcileDeletions', () => {
|
||||
it('runs on a clean full listing', async () => {
|
||||
const { shouldReconcileDeletions } = await import('@/lib/knowledge/connectors/sync-engine')
|
||||
|
||||
expect(shouldReconcileDeletions(false, {}, undefined)).toBe(true)
|
||||
expect(shouldReconcileDeletions(false, undefined, undefined)).toBe(true)
|
||||
})
|
||||
|
||||
it('never runs on incremental syncs', async () => {
|
||||
const { shouldReconcileDeletions } = await import('@/lib/knowledge/connectors/sync-engine')
|
||||
|
||||
expect(shouldReconcileDeletions(true, {}, undefined)).toBe(false)
|
||||
expect(shouldReconcileDeletions(true, {}, true)).toBe(false)
|
||||
expect(shouldReconcileDeletions(true, { listingCapped: true }, true)).toBe(false)
|
||||
})
|
||||
|
||||
it('skips when a connector capped the listing', async () => {
|
||||
const { shouldReconcileDeletions } = await import('@/lib/knowledge/connectors/sync-engine')
|
||||
|
||||
expect(shouldReconcileDeletions(false, { listingCapped: true }, undefined)).toBe(false)
|
||||
expect(shouldReconcileDeletions(false, { listingCapped: true }, false)).toBe(false)
|
||||
})
|
||||
|
||||
it('lets a forced fullSync override a connector cap', async () => {
|
||||
const { shouldReconcileDeletions } = await import('@/lib/knowledge/connectors/sync-engine')
|
||||
|
||||
expect(shouldReconcileDeletions(false, { listingCapped: true }, true)).toBe(true)
|
||||
})
|
||||
|
||||
it('never runs when the engine truncated pagination, even on a forced fullSync', async () => {
|
||||
const { shouldReconcileDeletions } = await import('@/lib/knowledge/connectors/sync-engine')
|
||||
|
||||
expect(shouldReconcileDeletions(false, { listingTruncated: true }, undefined)).toBe(false)
|
||||
expect(shouldReconcileDeletions(false, { listingTruncated: true }, true)).toBe(false)
|
||||
expect(
|
||||
shouldReconcileDeletions(false, { listingCapped: true, listingTruncated: true }, true)
|
||||
).toBe(false)
|
||||
})
|
||||
})
|
||||
|
||||
describe('resolveTagMapping', () => {
|
||||
beforeEach(() => {
|
||||
vi.clearAllMocks()
|
||||
|
||||
@@ -128,6 +128,27 @@ async function completeSyncLog(
|
||||
.where(eq(knowledgeConnectorSyncLog.id, syncLogId))
|
||||
}
|
||||
|
||||
/**
|
||||
* Decides whether deletion reconciliation may run for a sync.
|
||||
*
|
||||
* Reconciliation hard-deletes every stored document absent from the listing,
|
||||
* so it must only run against a complete source set:
|
||||
* - never on incremental syncs (they list only changed documents)
|
||||
* - never when the engine truncated pagination (`listingTruncated`) — a forced
|
||||
* fullSync cannot fix truncation, so it cannot override it
|
||||
* - not when a connector capped its listing (`listingCapped`), unless a forced
|
||||
* fullSync deliberately overrides the cap to reconcile the capped scope
|
||||
*/
|
||||
export function shouldReconcileDeletions(
|
||||
isIncremental: boolean | undefined,
|
||||
syncContext: Record<string, unknown> | undefined,
|
||||
fullSync: boolean | undefined
|
||||
): boolean {
|
||||
if (isIncremental) return false
|
||||
if (syncContext?.listingTruncated) return false
|
||||
return !syncContext?.listingCapped || Boolean(fullSync)
|
||||
}
|
||||
|
||||
/**
|
||||
* Resolves tag values from connector metadata using the connector's mapTags function.
|
||||
* Translates semantic keys returned by mapTags to actual DB slots using the
|
||||
@@ -415,6 +436,22 @@ export async function executeSync(
|
||||
hasMore = page.hasMore
|
||||
}
|
||||
|
||||
if (hasMore) {
|
||||
/**
|
||||
* Pagination stopped before source exhaustion (MAX_PAGES or a missing
|
||||
* cursor), so the listing is incomplete. `listingTruncated` blocks
|
||||
* deletion reconciliation absolutely — unlike connector-set
|
||||
* `listingCapped`, it cannot be overridden by a forced fullSync, since
|
||||
* re-running one truncates identically.
|
||||
*/
|
||||
syncContext.listingCapped = true
|
||||
syncContext.listingTruncated = true
|
||||
logger.warn('Pagination ended before source exhaustion; skipping deletion reconciliation', {
|
||||
connectorId,
|
||||
docsSoFar: externalDocs.length,
|
||||
})
|
||||
}
|
||||
|
||||
logger.info(`Fetched ${externalDocs.length} documents from ${connectorConfig.name}`, {
|
||||
connectorId,
|
||||
})
|
||||
@@ -635,9 +672,7 @@ export async function executeSync(
|
||||
}
|
||||
}
|
||||
|
||||
// Reconcile deletions for non-incremental syncs that returned ALL docs.
|
||||
// Skip when listing was capped (maxFiles/maxThreads) — unseen docs may still exist in the source.
|
||||
if (!isIncremental && (!syncContext?.listingCapped || options?.fullSync)) {
|
||||
if (shouldReconcileDeletions(isIncremental, syncContext, options?.fullSync)) {
|
||||
const removedIds = existingDocs
|
||||
.filter((d) => d.externalId && !seenExternalIds.has(d.externalId))
|
||||
.map((d) => d.id)
|
||||
|
||||
Reference in New Issue
Block a user