diff --git a/MANUAL.txt b/MANUAL.txt index 63b99ceb0..328adfd89 100644 --- a/MANUAL.txt +++ b/MANUAL.txt @@ -3585,8 +3585,7 @@ Enabling this extension with `context` output will produce markup suitable for the production of tagged PDFs. This includes additional markers for paragraphs and alternative markup for emphasized text. The `emphasis-command` template variable is set -if the extension is enabled. Combine this with the `pdfa` variable -to generate accessible PDFs. +if the extension is enabled. # Pandoc's Markdown @@ -7182,6 +7181,80 @@ Some document formats also include a unique identifier. For EPUB, this can be set explicitly by setting the `identifier` metadata field (see [EPUB Metadata], above). +# Accessible PDFs and PDF archiving standards + +PDF is a flexible format, and using PDF in certain contexts +requires additional conventions. For example, PDFs are not +accessible by default, they define how characters are placed on a +page but do not contain semantic information on the content by +default. However, it is possible to generate accessible PDFs, +which use tagging to add semantic information to the document. + +Pandoc's default method to generate PDF output is via LaTeX. +Tagging support in LaTeX is in development and not readily +available, so PDFs generated in this way will always be untagged +and not accessible. Alternative engines must be used to generate +accessible PDFs. + +The PDF standards PDF/A and PDF/UA define further restrictions +intended to optimize PDFs for archiving and accessibility. Tagging +is commonly used in combination with these standards to ensure +best results. + +Note, however, that standard compliance depends on many things, +including the colorspace of embedded images. Pandoc cannot check +this, and external programs must be used to ensure that generated +PDFs are in compliance. + +## ConTeXt + +ConTeXt always produces tagged PDFs, but the quality depends on +the input. The default ConTeXt markup generated by pandoc is +optimized for readability and reuse, not tagging. Enable the +[`tagging`](#extension--tagging) format extension to force markup +that is optimized for tagging. This can be combined with the +`pdfa` variable to generate standard-compliant PDFs. E.g.: + + pandoc --to=context+tagging -V pdfa=3a + +A recent `context` version should be used, as older versions +contained a bug that lead to invalid PDF metadata. + +## WeasyPrint + +The HTML-based engine WeasyPrint includes experimental support for +PDF/A and PDF/UA since version 57. Tagged PDFs can created with + + pandoc --pdf-engine=weasyprint \ + --pdf-engine-opt=--pdf-variant=pdf/ua-1 ... + +The feature is experimental and standard compliance should not be +assumed. + +## Prince XML + +The non-free HTML-to-PDf converter `prince` has extensive support +for various PDF standards as well as tagging. E.g.: + + pandoc --pdf-engine=prince \ + --pdf-engine-opt=--tagged-pdf ... + +See the prince documentation for more info. + +## Word Processors + +Word processors like LibreOffice and MS Word can also be used to +generate standardized and tagged PDF output. Pandoc does not +support direct conversions via these tools. Pandoc can convert a +document to a `docx` or `odt` file, which can then be opened and +converted to PDF with the respective word processor. See the +documentation for [Word][word-accessible-pdfs] and +[LibreOffice][lo-pdf-export]. + +[word-accessible-pdfs]: https://support.microsoft.com/en-us/office/create-accessible-pdfs-064625e0-56ea-4e16-ad71-3aa33bb4b7ed +[lo-pdf-export]: https://help.libreoffice.org/7.1/en-US/text/shared/01/ref_pdf_export_general.html + + # Running pandoc as a web server If you rename (or symlink) the pandoc executable to