feat(textract): migrate to AWS SDK, add AnalyzeExpense and AnalyzeID (#5456)

* feat(textract): migrate to AWS SDK, add AnalyzeExpense and AnalyzeID

- Replace hand-rolled AWS SigV4 signing with @aws-sdk/client-textract, matching sibling AWS integrations (secrets_manager, s3, sts)
- Add Analyze Expense operation (invoice/receipt structured extraction) via AnalyzeExpense/StartExpenseAnalysis+GetExpenseAnalysis
- Add Analyze Identity Document operation (AnalyzeID) with optional back-of-ID page
- Add an operation selector to the textract_v2 block; defaults to the existing Analyze Document behavior for backward compatibility
- Add tests for tool body/response mapping and route-level AWS response normalization

* fix(textract): forward URL documents and fix ambiguous error status

- textract_analyze_expense/textract_analyze_id now fall back to filePath/filePathBack when the document input is a URL string rather than an uploaded file object, so advanced "File reference" URL inputs actually reach the API (Cursor Bugbot)
- mapTextractSdkError defaults to 500 (not 400) when the AWS SDK error has no HTTP status, since that implies a server-side/network failure rather than a bad request (Greptile)

* fix(textract): stop stale processingMode from hiding ID document fields

- Front-document fields (fileUpload/fileReference) are shared across all 3 operations; gate them with a values-aware condition so switching to Analyze Identity Document keeps them visible even if a stale processingMode='async' is left over from a previous operation
- S3 URI field now also requires operation !== 'analyze_id', since that operation never supports S3 input

* fix(textract): preserve first-page metadata across async pagination

pollTextractJob's merge callbacks spread only the latest page, dropping
any field (DocumentMetadata, model version) the first page had but a
follow-up NextToken page omits. Merge accumulated first so later pages
only override fields they actually return.

* fix(textract): pass through the real upstream status for filePath fetch failures

fetchDocumentBytes hardcoded 400 for any non-OK response from a document
URL, masking transient 5xx failures from the document host as client
errors and blocking tool-execution retries. Use the actual response
status instead.

* chore(textract): drop redundant inline comments
This commit is contained in:
Waleed
2026-07-06 19:42:36 -07:00
committed by GitHub
parent 017e8ca68f
commit b686111082
18 changed files with 2170 additions and 609 deletions
+2 -2
View File
@@ -9,8 +9,8 @@ const QUERY_HOOKS_DIR = path.join(ROOT, 'apps/sim/hooks/queries')
const SELECTOR_HOOKS_DIR = path.join(ROOT, 'apps/sim/hooks/selectors')
const BASELINE = {
totalRoutes: 904,
zodRoutes: 904,
totalRoutes: 906,
zodRoutes: 906,
nonZodRoutes: 0,
} as const