Files
sim/apps
Waleed 441004ad24 improvement(file-parsers): bound PDF text extraction (#6425)
* improvement(file-parsers): bound PDF text extraction

Extract page text through pdf.js's streaming API with page, character, and
wall-clock budgets instead of buffering the whole document, so extraction
memory stays bounded regardless of input. Release the document proxy when
done, and route output through sanitizeTextForUTF8 like the other parsers.

* fix(file-parsers): only flag truncation when PDF text is actually dropped
2026-08-08 12:21:07 -07:00
..