...strings that are numbers beginning with 0 or ending with 0 and
having a decimal point. Otherwise they will read as YAML numbers
and potentially be modified (e.g. 3.10 -> 3.1).
It would be better to fix this on the reader side, but since
we use a standard YAML parser it's hard to see how.
Closes#11715.
Headings without an explicit label can now be assigned automatic
identifiers based on the heading text, making them linkable in a
generated table of contents. The extension is available for the
typst reader but is off by default; enable it with
`-f typst+auto_identifiers`. The related `gfm_auto_identifiers` and
`ascii_identifiers` extensions are also made available.
Closes#11041.
Text.Pandoc.Readers.Typst.Parsing: PState gains sOptions,
sIdentifiers, and sLogMessages fields, and now has
HasReaderOptions, HasIdentifierList, and HasLogMessages
instances, allowing reuse of the shared registerHeader. (Not an
API change.)
Co-Authored-By: Claude <noreply@anthropic.com>
When a reference document uses a non-standard namespace prefix for the
WordprocessingML namespace (e.g. `ns0` instead of `w`), `sectPr` elements
copied from the reference would retain the non-`w` prefix, producing
malformed XML in the output document. Similarly,
`extractPageLayout` only matched elements with prefix `w`, missing
`sectPr` elements with other prefixes. This is fixed by matching on the
namespace URI rather than the prefix, and normalizing the prefix to `w`
on all elements and attributes copied from reference-doc `sectPr`.
Some new tests have been added, and the test suite has been streamlined
using helper functions.
Word represents "restart numbering" on a style-based list by pointing
only the first item of the restarted list at a new `numId` that shares the
original list's abstract numbering definition but carries a
`w:startOverride`; the remaining items keep using the original `numId`.
Pandoc keyed list continuation and grouping on the `numId`, so the
restarted items continued the stale count from the earlier list (and
were split into a separate ordered list with the wrong start).
Key continuation and grouping off the `abstractNumId` instead (the real
running counter in Word), and treat `startOverride` as a restart that
resets the count.
Closes#8367.
Co-Authored-By: Claude <noreply@anthropic.com>
Word-95-style numbered and bulleted lists are encoded with a
`{\pntext ...}` auto-number destination at the start of each list
paragraph plus a `{\*\pn ...}` destination describing the numbering,
rather than the modern `\listtext`/`\listtable` mechanism. Two problems:
1. The `\pntext` marker text ("1.", "·", etc.) was captured as the
paragraph's first text run, which sits before the paragraph's
`\ls`/`\ilvl`, so emitBlocks (which reads list properties from the first
run) misclassified the paragraph as an ordinary paragraph. The first
item of each list therefore came out as a stray paragraph.
2. The numbering style was ignored, so numbered lists defaulted to
bullets.
Treat `\pntext` like `\listtext`: drop its visible marker text and flag the
start of a new list item. Parse the `{\*\pn ...}` destination
(`\pnlvlbody`/`\pnlvlblt`, `\pndec`, `\pnucltr`, `\pnlcltr`,
`\pnucrm`, `\pnlcrm`, `\pnstart`, and the `\ls`/`\ilvl` keys it
carries) into the list override table so numbered lists are
emitted as ordered lists with the right number style.
`\pn` is a paragraph property that remains in effect until reset
by `\pard`, auto-numbering every paragraph in scope. Track
this (`sPnActive`) so each paragraph becomes its own list item,
rather than merging markerless continuation paragraphs as is done
for modern `\listtext` lists.
Co-Authored-By: Claude <noreply@anthropic.com>
Lists were written as plain paragraphs with a literal marker and a
hanging indent, carrying none of the `\ls`/`\ilvl` references or the
`\listtable`/`\listoverridetable` that the RTF list model (and pandoc's
own reader) expect, so lists could not round-trip.
Walk the document tagging each list with a unique id and nesting level
(`prepareLists`), build a `\listtable`/`\listoverridetable` describing
each list (`listTableRTF`, exposed via the `listtable` template
variable), and render every list paragraph with its `\ls`/`\ilvl`
reference. The first paragraph of each item gets a `{\listtext}`
marker; continuation paragraphs keep the reference but omit it, so
multi-paragraph items round-trip.
Co-Authored-By: Claude <noreply@anthropic.com>
The reader treated every list paragraph as a new list item, so a list
item containing several paragraphs was split into one item per
paragraph. In the RTF list model a new item is marked by a
`{\listtext ...}` destination group; a list paragraph lacking one is a
continuation of the current item.
Track whether a `{\listtext}` group was seen (`sListText`) and, in
`emitBlocks`, append a continuation paragraph (no `\listtext`) to the
current item rather than starting a new one.
Adds a `list_multiparagraph` reader test.
Co-Authored-By: Claude <noreply@anthropic.com>
Previously we failed to detect horizontal rules when the
underlying XML contained text elements for semantically
insignificant whitespace. We also failed to handle the case
where a paragraph contains text and has a bottom border
property: this seems to be what Word does when you insert
a horizontal rule, e.g. by typing a series of hyphens or #
characters.
Closes#11689.
Closes#11682.
Fixes two bugs that caused a simple table to be read as a cascade of
nested tables:
* `\plain` reset the entire property record (including the in-table
flag) via `const def`. When a cell paragraph used `\plain` after
`\intbl`, the in-table flag was cleared, so the next `\cell` closed
the partially-built table and embedded it inside a new cell. Per the
RTF spec, `\plain` should reset only character formatting, so it now
preserves paragraph/context properties (in-table, list level, outline
level, hyperlink, anchor).
* `\row` was ignored and only `\trowd` started a new row. Real-world
RTF often emits a single `\trowd` and separates rows with `\row`, so
every cell ended up in one row. Both `\trowd` and `\row` now begin a
fresh row (only when the current one has cells), and empty trailing
rows are dropped when closing the table.
Co-Authored-By: Claude <noreply@anthropic.com>
Add support for the `auto_identifiers`, `gfm_auto_identifiers`, and
`ascii_identifiers` extensions in the man reader. Section headings
parsed from .SH and .SS macros now receive auto-generated id
attributes when the extension is enabled, enabling `--toc` to
produce working anchor links.
- Add `autoIdExtensions` to default man extensions [behavior change]
- Add `HasReaderOptions`, `HasLogMessages` and `HasIdentifierList` to
`ManState` to run `registerHeader`
Closes#8852.
Like the other table syntaxes (pipe, simple, and multiline tables) and
block-level constructs generally, a grid table may now be indented by up
to three spaces and still be recognized as a table. Previously the
grid-table parser required the table to begin at the left margin, so an
indented grid table was parsed as a paragraph.
The leading indentation is stripped uniformly from each line before the
table is parsed, so an indented grid table produces the same AST as its
non-indented equivalent.
Adds a command test.
Previously the OpenDocument writer emitted a fresh automatic style
(L1..Ln, P1..Pn, T1..Tn) for nearly every list, list-item paragraph,
block quote, preformatted block, and inline text style. This produced
large ODT files, made `--reference-doc` customization ineffective (the
user's predefined styles were never referenced), and gave each list its
own indentation independent of any containing block quote.
This commit teaches the writer to reference the predefined styles that
LibreOffice ships and that pandoc's reference.odt now exports:
- Bullet lists use `List_20_1`; ordered lists with default start and
decimal format use `Numbering_20_1`. Non-default ordered lists
generate a single named override style (`Pandoc_Numbering_N`)
memoised by (ListNumberStyle, ListNumberDelim); a non-default start
value with the default format is expressed via `text:start-value`
on the `text:list` element instead of a new style.
- List-item paragraphs use `List_20_Bullet[_Tight]` and
`List_20_Number[_Tight]`. The Tight variants are pandoc-specific
(zero top/bottom margin) and are injected into the user's
reference.odt if missing, just like the Skylighting token styles.
- Block quotes use the predefined `Quotations` paragraph style
directly. Nested block quotes use a single automatic style that
inherits from Quotations and only adds extra margin-left, so a list
inside a block quote now inherits its container's indent (#2747).
- Preformatted blocks use `Preformatted_20_Text` directly.
- Emphasis, Strong, Strikeout, Subscript, Superscript and Code spans
use the predefined `Emphasis`, `Strong_20_Emphasis`, `Strikeout`,
`Subscript`, `Superscript` and `Source_20_Text` text styles.
- `paraStyle`/`paraStyleFromParent` no longer emit a wrapper automatic
style when its only attribute would be `parent-style-name`; the
parent name is returned directly.
Closes#9136.
Closes#5086.
Closes#2747.
Closes#3426.
Closes#7336.
Co-authored by: Claude Opus 4.7.
We parse these as DefinitionList items, but we previously
sometimes stopped prematurely in including material in the
definition. We should include everything until we hit a new
indentation-changing macro.
Closes#11668.
This change ensures that raw content marked `epub2` will appear in (only) EPUBv2 output
and content marked `epub3` will appear in (only) EPUBv3 output.
When parsing an inline note (`^[...]`) inside a quoted span,
`stateQuoteContext` was still set to `InSingleQuote`/`InDoubleQuote`,
so quotes within notes failed to parse as `Quoted` nodes.
Fix this by wrapping the note body parser in
`withQuoteContext NoQuote`.
Closes#11613.
Styles unconditionally emits css that uses screen-only properties.
Paged-media engines (weasyprint, prince, pagedjs) have no viewport
and issue warnings. Fix it to hide these properties from the engines.
Closes#11524.
This allows one to pass parameters to typst, which are available
at `sys.inputs`, just as `typst` itself does with its `--input`
option.
[API changes]
* ReaderOptions has a new field `readerTypstInputs`.
* Opt has a new field `optTypstInputs`.
Closes#11588.
When a heading contained a footnote, processing the footnote's block
content would consume the stFirstPara flag, causing the following
paragraph to incorrectly receive BodyText style instead of
FirstParagraph. Fix by saving and restoring stFirstPara around
footnote block processing.
Closes#11573.
Co-Authored-By: Claude <noreply@anthropic.com>
MediaBag test used `inDirectory` (which calls `setCurrentDirectory`),
changing the process-wide CWD and causing other parallel tests
to fail with "does not exist" errors on relative paths.
Replace with absolute paths so the test no longer changes CWD.
Closes#11566.
Co-Authored-By: Claude <noreply@anthropic.com>
For paragraphs beginning with literal `#`, `*`, or other things
that would otherwise produce lists, we insert `<nowiki></nowiki>`
rather than (as previously) `\`. `\` does not work to escape
these.
Closes#11563.
- Properly handle the case where the first item is an indented
code block. (Closes #11542.)
- Use correct indentation when `four_space_rule` extension is
disabled.
Previously, when a w:p paragraph contained runs with textboxes,
the entire paragraph was replaced by just the textbox content,
discarding all other runs (including image-bearing runs).
Now we walk the paragraph's children in order, grouping
non-textbox content into copies of the original w:p and splicing
unwrapped textbox content in place, preserving the original order.
We also treat text inside a textbox containing an image as a
figure caption. (One often finds captioned images of this
kind in docx files.)
Closes#11510.
Closes#6893.
Closes#11412.
Closes#5394.
Closes#9633.
Co-Authored-By: Claude <noreply@anthropic.com>