feat(firecrawl): add parse operation and revert short-input selection style (#4340)

* feat(firecrawl): add parse operation and revert short-input selection style

* chore(firecrawl): regenerate docs and integrations data for parse

* fix(firecrawl): forward firecrawl error body in parse route response

* fix(firecrawl): add pricing config to parse tool hosting
This commit is contained in:
Waleed
2026-04-29 12:04:12 -07:00
committed by GitHub
parent 8d042f7d20
commit 7d8ec24768
10 changed files with 611 additions and 8 deletions
@@ -234,4 +234,48 @@ Autonomous web data extraction agent. Searches and gathers information based on
| `expiresAt` | string | Timestamp when the results expire \(24 hours\) |
| `sources` | object | Array of source URLs used by the agent |
### `firecrawl_parse`
Parse uploaded documents (PDF, DOCX, HTML, etc.) into clean markdown using Firecrawl. Supports .html, .htm, .pdf, .docx, .doc, .odt, .rtf, .xlsx, .xls.
#### Input
| Parameter | Type | Required | Description |
| --------- | ---- | -------- | ----------- |
| `file` | file | Yes | Document file to be parsed |
| `formats` | array | No | Output formats to return \(e.g., \["markdown"\]\). Defaults to markdown. |
| `onlyMainContent` | boolean | No | Exclude headers, navs, footers. Defaults to true. |
| `includeTags` | array | No | HTML tags to include |
| `excludeTags` | array | No | HTML tags to exclude |
| `timeout` | number | No | Timeout in milliseconds \(max 300000\). Defaults to 30000. |
| `parsers` | array | No | Parser configuration \(e.g., \[\{ "type": "pdf" \}\]\) |
| `removeBase64Images` | boolean | No | Remove base64 images, keep alt text. Defaults to true. |
| `blockAds` | boolean | No | Block ads and popups. Defaults to true. |
| `proxy` | string | No | Proxy mode: "basic" or "auto" |
| `zeroDataRetention` | boolean | No | Enable zero data retention. Defaults to false. |
| `apiKey` | string | Yes | Firecrawl API key |
| `rateLimit` | string | No | No description |
#### Output
| Parameter | Type | Description |
| --------- | ---- | ----------- |
| `markdown` | string | Parsed document content in markdown format |
| `summary` | string | Generated summary of the document |
| `html` | string | Processed HTML content |
| `rawHtml` | string | Unprocessed raw HTML content |
| `screenshot` | string | Screenshot URL or base64 \(when requested\) |
| `links` | array | URLs discovered in the document |
| `metadata` | object | Document metadata |
| ↳ `title` | string | Document title |
| ↳ `description` | string | Document description |
| ↳ `language` | string | Document language code |
| ↳ `sourceURL` | string | Source URL |
| ↳ `url` | string | Final URL |
| ↳ `keywords` | string | Document keywords |
| ↳ `statusCode` | number | HTTP status code |
| ↳ `contentType` | string | Document content type |
| ↳ `error` | string | Error message if parse failed |
| `warning` | string | Warning message from the parse operation |
@@ -256,8 +256,6 @@ Create a new database in Notion with custom properties
### `notion_add_database_row`
Add a new row to a Notion database with specified properties
#### Input
| Parameter | Type | Required | Description |