fix(confluence): preserve panel/callout macro semantics through sync (#5896)

* fix(confluence): preserve panel/callout macro semantics through sync

Confluence's rendered view HTML wraps Info/Note/Warning/Tip and custom Panel
macros in divs whose class/color convey meaning that the shared
htmlToPlainText tag-stripper discards along with the tags — a red "do not
use" warning panel becomes indistinguishable from a plain paragraph once
flattened, so RAG has no signal that a bullet under it is an exclusion rule
rather than a normal one.

Adds preserveConfluenceCallouts, a Confluence-specific pre-pass that rewrites
each detected panel into a single bracketed label (e.g. "[WARNING] Do NOT use
this form for: GitLab") before the generic plain-text conversion runs, so the
callout semantic survives both the tag strip and htmlToPlainText's trailing
whitespace collapse. Bumps the connector's content-representation marker so
already-synced pages get one automatic re-hydration under the new extraction,
rather than silently keeping their stale flattened content until their next
edit.

* fix(confluence): preserve word boundaries when extracting callout body text

Greptile P1: cheerio's .text() concatenates every descendant text node with
no separator, so pulling a macro body's text in one call fused adjacent
blocks together (e.g. a paragraph ending in "for:" immediately followed by a
list item "GitLab" became "for:GitLab"), corrupting the exact word boundaries
RAG chunking and keyword matching depend on.

extractBlockJoinedText now extracts each paragraph/list-item/heading/cell/quote
individually and joins them with a single space, keeping every block's text
intact and properly separated, matching how htmlToPlainText already treats
the rest of the page.

* fix(confluence): fix nested-block duplication in callout text extraction

Greptile P1: filtering the found blocks to only top-level ones still wasn't
enough — a nested block (an outer <li> containing its own nested <ul><li>, a
<td> containing a <blockquote>) matched the selector once, but its .text()
call recurses into and flattens its own matched descendants with no
separator, reproducing the exact word-fusion bug one level deeper (and any
duplicate-selection would have double-counted the same text).

Replaces the block-selector approach with a recursive text-node walk:
extractBlockJoinedText now visits every text node individually and joins them
all with a single space, so word boundaries are preserved at every nesting
depth with no double-counting, matching the pattern html-parser.ts already
uses elsewhere in this codebase for the same class of problem.

* fix(confluence): apply the same word-boundary-safe extraction to panel headers

Greptile: panelHeader text extraction was left on the plain .text() call
while panel/macro body extraction was already fixed to use
extractBlockJoinedText, so a rich multi-node header (e.g. <b>Warning:</b>
followed by a sibling <span>) could still fuse into "Warning:Do not use"
with no space. Panel headers now go through the same recursive text-node
walk as bodies, for consistency across every text extraction in this file.

* fix(confluence): distinguish inline formatting from block boundaries in extraction

Greptile: the recursive text-node walk unconditionally inserted a space
between every text node, which fixed block-boundary fusion but broke
genuinely inline-formatted text — "un<b>believe</b>able" became
"un believe able" and "Hello<b>!</b>" became "Hello !", corrupting valid
callout content on its way into the index.

Adds an INLINE_FORMATTING_TAGS allowlist (b, strong, i, em, span, a, etc.):
text flowing through those tags accumulates with no artificial separator,
preserving exact source adjacency, while every other tag boundary (p, li,
td, headings, br, ...) still flushes to a new segment — a block always
implies a break even with no literal whitespace in the source, but an inline
tag never does. Fixed one test that had encoded the old, incorrect
expectation for two genuinely adjacent inline tags with no source whitespace
between them, and added regression tests for mid-word inline formatting,
punctuation attached to an inline tag, and a header with real source spacing.

* fix(confluence): process nested panels/macros innermost-first

Cursor: processing matches in document order (outermost first) read a
nested, not-yet-converted panel/macro as plain body text before it ever got
its own bracketed label, silently dropping the inner callout's type. Worse,
an untitled outer panel's `.find('.panelHeader')` could reach past its own
missing header into a nested panel's header and adopt it as its own title.

Replaces the two independent .each() passes with a loop that converts only
"leaf" macros (no remaining nested macro/panel inside them) and repeats
until none are left. This processes innermost-first, so a nested macro is
already its own bracketed <p> by the time its parent's body/header text is
read, and an untitled outer panel's .find() can no longer reach a header
that isn't its own, since a leaf by definition has no nested panel left to
reach into.

* test(confluence): add explicit regression for the exact reported <br> repro

Formalizes an explicit test for the exact <br>-separated string Greptile's
review cited as broken (verified manually not to reproduce, but wasn't
directly asserted in the suite before this).
This commit is contained in:
Waleed
2026-07-23 14:08:09 -07:00
committed by GitHub
parent 84e7aad633
commit 5c427795a2
2 changed files with 407 additions and 5 deletions
@@ -2,7 +2,12 @@
* @vitest-environment node
*/
import { describe, expect, it } from 'vitest'
import { escapeCql, isCurrentContent } from '@/connectors/confluence/confluence'
import {
escapeCql,
isCurrentContent,
preserveConfluenceCallouts,
} from '@/connectors/confluence/confluence'
import { htmlToPlainText } from '@/connectors/utils'
describe('escapeCql', () => {
it.concurrent('returns plain strings unchanged', () => {
@@ -48,3 +53,248 @@ describe('isCurrentContent', () => {
expect(isCurrentContent({ id: '1', status: 'deleted' })).toBe(false)
})
})
describe('preserveConfluenceCallouts', () => {
it.concurrent('handles empty content', () => {
expect(preserveConfluenceCallouts('')).toBe('')
})
it.concurrent('leaves content with no macros unchanged', () => {
const html = '<p>Just a normal paragraph.</p>'
expect(preserveConfluenceCallouts(html)).toContain('Just a normal paragraph.')
})
it.concurrent('labels a built-in warning macro and keeps its body', () => {
const html =
'<div class="confluence-information-macro confluence-information-macro-warning">' +
'<span class="aui-icon aui-icon-small aui-iconfont-warning confluence-information-macro-icon"></span>' +
'<div class="confluence-information-macro-body"><p>Do NOT use this form for GitLab access.</p></div>' +
'</div>'
const result = preserveConfluenceCallouts(html)
expect(result).toContain('[WARNING]')
expect(result).toContain('Do NOT use this form for GitLab access.')
})
it.concurrent('labels a built-in info macro', () => {
const html =
'<div class="confluence-information-macro confluence-information-macro-information">' +
'<div class="confluence-information-macro-body"><p>Heads up.</p></div>' +
'</div>'
expect(preserveConfluenceCallouts(html)).toContain('[INFO] Heads up.')
})
it.concurrent('labels a built-in note macro', () => {
const html =
'<div class="confluence-information-macro confluence-information-macro-note">' +
'<div class="confluence-information-macro-body"><p>See also.</p></div>' +
'</div>'
expect(preserveConfluenceCallouts(html)).toContain('[NOTE] See also.')
})
it.concurrent('labels a built-in tip macro', () => {
const html =
'<div class="confluence-information-macro confluence-information-macro-tip">' +
'<div class="confluence-information-macro-body"><p>Pro tip.</p></div>' +
'</div>'
expect(preserveConfluenceCallouts(html)).toContain('[TIP] Pro tip.')
})
it.concurrent('labels a generic custom-colored Panel macro using its header title', () => {
const html =
'<div class="panel" style="border-width: 1px;">' +
'<div class="panelHeader" style="background-color: #ffebe6;"><b>Do NOT use this form for:</b></div>' +
'<div class="panelContent"><p>GitLab access requests go to the private channel instead.</p></div>' +
'</div>'
const result = preserveConfluenceCallouts(html)
expect(result).toContain('[CALLOUT: Do NOT use this form for:]')
expect(result).toContain('GitLab access requests go to the private channel instead.')
})
it.concurrent('preserves word boundaries between a block header and its own content', () => {
const html =
'<div class="panel"><div class="panelHeader"><b>Warning:</b></div>' +
'<div class="panelContent"><p>See replacement form.</p></div></div>'
const result = preserveConfluenceCallouts(html)
expect(result).toContain('[CALLOUT: Warning:] See replacement form.')
})
it.concurrent(
'keeps a rich header with a real source space intact, without adding a second one',
() => {
const html =
'<div class="panel"><div class="panelHeader"><b>Warning:</b> <span>Do not use</span></div>' +
'<div class="panelContent"><p>See replacement form.</p></div></div>'
const result = preserveConfluenceCallouts(html)
expect(result).toContain('[CALLOUT: Warning: Do not use]')
}
)
it.concurrent('falls back to a bare CALLOUT label when a Panel macro has no header text', () => {
const html =
'<div class="panel"><div class="panelContent"><p>Untitled panel body.</p></div></div>'
const result = preserveConfluenceCallouts(html)
expect(result).toContain('[CALLOUT]')
expect(result).toContain('Untitled panel body.')
})
it.concurrent(
'keeps the exclusion marker attached to its content through htmlToPlainText, even across surrounding whitespace collapse',
() => {
const html =
'<p>Intro paragraph.</p>\n\n' +
'<div class="confluence-information-macro confluence-information-macro-warning">' +
'<div class="confluence-information-macro-body"><p>Do NOT use this form for:</p>' +
'<ul><li>GitLab</li></ul></div>' +
'</div>\n\n' +
'<p>Trailing paragraph.</p>'
const plainText = htmlToPlainText(preserveConfluenceCallouts(html))
expect(plainText).toContain('[WARNING] Do NOT use this form for: GitLab')
expect(plainText).toContain('Intro paragraph.')
expect(plainText).toContain('Trailing paragraph.')
}
)
it.concurrent(
'does not fuse adjacent paragraph and list-item text together (word-boundary regression)',
() => {
const html =
'<div class="confluence-information-macro confluence-information-macro-warning">' +
'<div class="confluence-information-macro-body">' +
'<p>Do NOT use this form for:</p>' +
'<ul><li>GitLab</li><li>ServiceNow</li></ul>' +
'</div></div>'
const result = preserveConfluenceCallouts(html)
expect(result).not.toContain('for:GitLab')
expect(result).not.toContain('GitLabServiceNow')
expect(result).toContain('Do NOT use this form for: GitLab ServiceNow')
}
)
it.concurrent(
'preserves word boundaries across multiple paragraphs in a generic Panel macro',
() => {
const html =
'<div class="panel"><div class="panelContent">' +
'<p>First sentence.</p><p>Second sentence.</p>' +
'</div></div>'
const result = preserveConfluenceCallouts(html)
expect(result).toContain('First sentence. Second sentence.')
expect(result).not.toContain('sentence.Second')
}
)
it.concurrent(
'does not duplicate or fuse text from a nested list inside a callout body (nesting regression)',
() => {
const html =
'<div class="confluence-information-macro confluence-information-macro-note">' +
'<div class="confluence-information-macro-body">' +
'<ul><li>Outer item' +
'<ul><li>Nested item A</li><li>Nested item B</li></ul>' +
'</li><li>Outer item two</li></ul>' +
'</div></div>'
const result = preserveConfluenceCallouts(html)
// Each nested <li>'s text must appear exactly once, not duplicated by the
// outer <li> also being matched and its .text() recursing into it.
const occurrences = (result.match(/Nested item A/g) ?? []).length
expect(occurrences).toBe(1)
expect(result).not.toContain('Nested item ANested item B')
expect(result).toContain('Outer item Nested item A Nested item B Outer item two')
}
)
it.concurrent('does not fuse text from a blockquote nested inside a table cell', () => {
const html =
'<div class="panel"><div class="panelContent">' +
'<table><tr><td>Cell text<blockquote><p>quoted text</p></blockquote>after quote</td></tr></table>' +
'</div></div>'
const result = preserveConfluenceCallouts(html)
expect(result).not.toContain('quotedtext')
expect(result).not.toContain('textafter')
expect(result).toContain('Cell text quoted text after quote')
})
it.concurrent(
'does not inject an artificial space into inline-formatted text mid-word (inline vs. block regression)',
() => {
const html =
'<div class="panel"><div class="panelContent">' +
'<p>This is un<b>believe</b>able.</p>' +
'</div></div>'
const result = preserveConfluenceCallouts(html)
expect(result).not.toContain('un believe able')
expect(result).toContain('This is unbelieveable.')
}
)
it.concurrent('does not inject a space before punctuation carried by an inline tag', () => {
const html =
'<div class="confluence-information-macro confluence-information-macro-warning">' +
'<div class="confluence-information-macro-body"><p>Do not proceed<b>!</b></p></div>' +
'</div>'
const result = preserveConfluenceCallouts(html)
expect(result).not.toContain('proceed !')
expect(result).toContain('[WARNING] Do not proceed!')
})
it.concurrent('keeps natural word spacing when inline tags wrap a whole word', () => {
const html =
'<div class="panel"><div class="panelContent">' +
'<p>Do <b>NOT</b> use this form.</p>' +
'</div></div>'
const result = preserveConfluenceCallouts(html)
expect(result).toContain('Do NOT use this form.')
})
it.concurrent(
'labels a nested panel-in-panel with both its own and its parent label (nesting regression)',
() => {
const html =
'<div class="panel"><div class="panelHeader"><b>Outer</b></div><div class="panelContent">' +
'<div class="panel"><div class="panelHeader"><b>Inner</b></div>' +
'<div class="panelContent"><p>inner body</p></div></div>' +
'</div></div>'
const result = preserveConfluenceCallouts(html)
expect(result).toContain('[CALLOUT: Outer]')
expect(result).toContain('[CALLOUT: Inner] inner body')
}
)
it.concurrent(
'labels a nested info-macro inside a panel with its own type instead of dropping it',
() => {
const html =
'<div class="panel"><div class="panelContent">' +
'<div class="confluence-information-macro confluence-information-macro-warning">' +
'<div class="confluence-information-macro-body"><p>Do not use this.</p></div></div>' +
'</div></div>'
const result = preserveConfluenceCallouts(html)
expect(result).toContain('[WARNING] Do not use this.')
}
)
it.concurrent(
"does not let an untitled outer panel adopt a nested panel's header as its own",
() => {
const html =
'<div class="panel"><div class="panelContent">' +
'<div class="panel"><div class="panelHeader"><b>Inner title</b></div>' +
'<div class="panelContent"><p>inner body</p></div></div>' +
'</div></div>'
const result = preserveConfluenceCallouts(html)
// The outer panel has no header of its own — it must fall back to a
// bare [CALLOUT], not steal "Inner title" from the nested panel.
expect(result).toContain('[CALLOUT] [CALLOUT: Inner title] inner body')
}
)
it.concurrent('does not fuse text on either side of a <br> line break', () => {
const html =
'<div class="confluence-information-macro confluence-information-macro-warning">' +
'<div class="confluence-information-macro-body"><p>Do NOT use this form for:<br>GitLab</p></div>' +
'</div>'
const result = preserveConfluenceCallouts(html)
expect(result).not.toContain('for:GitLab')
expect(result).toContain('[WARNING] Do NOT use this form for: GitLab')
})
})
+156 -4
View File
@@ -1,5 +1,6 @@
import { createLogger } from '@sim/logger'
import { toError } from '@sim/utils/errors'
import * as cheerio from 'cheerio'
import { fetchWithRetry, VALIDATE_RETRY_OPTIONS } from '@/lib/knowledge/documents/utils'
import { confluenceConnectorMeta } from '@/connectors/confluence/meta'
import type { ConnectorConfig, ExternalDocument, ExternalDocumentList } from '@/connectors/types'
@@ -8,6 +9,155 @@ import { getConfluenceCloudId, normalizeConfluenceDomainHost } from '@/tools/con
const logger = createLogger('ConfluenceConnector')
/** Label prefixes for Confluence's built-in Info/Note/Warning/Tip macros, by their rendered CSS suffix. */
const CALLOUT_LABELS: Record<string, string> = {
information: '[INFO]',
note: '[NOTE]',
warning: '[WARNING]',
tip: '[TIP]',
error: '[ERROR]',
}
/**
* Inline formatting tags whose text flows directly into their surrounding
* sentence with no implied word break — e.g. `un<b>believe</b>able` must stay
* `unbelievable`, and `Hello<b>!</b>` must stay `Hello!`, not gain an
* artificial space. Anything not in this set (p, li, td, div, headings, br,
* etc.) is treated as a block boundary that always implies a break, even when
* the source HTML has no literal whitespace there.
*/
const INLINE_FORMATTING_TAGS = new Set([
'b',
'strong',
'i',
'em',
'u',
's',
'strike',
'del',
'ins',
'sup',
'sub',
'small',
'mark',
'code',
'span',
'a',
'abbr',
'cite',
'q',
'kbd',
'var',
'samp',
'time',
])
/**
* Cheerio's `.text()` concatenates every descendant text node with no
* separator at all, so pulling a macro body's text in one call fuses adjacent
* blocks together (e.g. a `<p>...for:</p>` immediately followed by
* `<li>GitLab</li>` becomes `for:GitLab`, corrupting the very word boundaries
* RAG chunking depends on). Simply joining every text node with a space isn't
* right either — that would corrupt genuinely inline-formatted text the same
* way. This walks the DOM, accumulating text through inline tags without a
* separator (preserving exact source adjacency) and flushing to a new segment
* at every other tag boundary (a block always implies a break, regardless of
* source whitespace) — matching how `html-parser.ts` already walks HTML for a
* related reason elsewhere in this codebase, extended with the inline/block
* distinction real Confluence rich text requires.
*/
function extractBlockJoinedText($: cheerio.CheerioAPI, $el: cheerio.Cheerio<any>): string {
const parts: string[] = []
let current = ''
const flush = () => {
const text = current.trim()
if (text) parts.push(text)
current = ''
}
const visit = ($node: cheerio.Cheerio<any>) => {
$node.contents().each((_, child) => {
if (child.type === 'text') {
current += $(child).text()
} else if (child.type === 'tag') {
const tag = child.tagName?.toLowerCase()
if (tag && INLINE_FORMATTING_TAGS.has(tag)) {
visit($(child))
} else {
flush()
visit($(child))
flush()
}
}
})
}
visit($el)
flush()
return parts.join(' ').trim()
}
/** Matches either flavor of panel/macro this function rewrites. */
const MACRO_SELECTOR = 'div.confluence-information-macro, div.panel'
/**
* Confluence's rendered `view` HTML wraps Info/Note/Warning/Tip macros in
* `confluence-information-macro confluence-information-macro-{type}` divs, and
* the customizable Panel macro in `.panel` > `.panelHeader` + `.panelContent`
* divs. `htmlToPlainText`'s blind tag-stripping discards the divs' classes along
* with the tags, so a red "do not use" warning panel becomes indistinguishable
* from a plain paragraph once flattened — and its trailing whitespace collapse
* would erase any newline-based separation too. Each detected panel is rewritten
* into a single bracketed label plus its own text so the callout semantic
* survives both the tag strip and the whitespace collapse.
*
* A panel can itself contain another panel or macro (e.g. a nested Note inside
* a Warning panel). Processing matches in document order — outermost first —
* would read a not-yet-converted nested macro as plain body text before it
* ever got its own label, silently dropping the inner callout's semantic, and
* `.find('.panelHeader')` would then risk pulling a nested panel's header up
* as if it were the outer panel's own title. Converting only "leaf" macros
* (ones with no remaining nested macro/panel inside them) and repeating until
* none are left processes innermost-first, so a nested macro is already a
* bracketed `<p>` by the time its parent's body/header text is read — at which
* point it correctly reads as plain text carrying its own label.
*/
export function preserveConfluenceCallouts(html: string): string {
if (!html) return html
const $ = cheerio.load(html)
let progressed = true
while (progressed) {
progressed = false
const leaves = $(MACRO_SELECTOR).filter((_, el) => $(el).find(MACRO_SELECTOR).length === 0)
if (leaves.length === 0) break
leaves.each((_, el) => {
const $el = $(el)
if ($el.hasClass('confluence-information-macro')) {
const type = ($el.attr('class') ?? '')
.match(/confluence-information-macro-(\w+)/)?.[1]
?.toLowerCase()
const label = (type && CALLOUT_LABELS[type]) || CALLOUT_LABELS.information
const macroBody = $el.find('.confluence-information-macro-body').first()
const body = extractBlockJoinedText($, macroBody.length > 0 ? macroBody : $el)
$el.replaceWith($('<p></p>').text(`${label} ${body}`))
} else {
const headerText = extractBlockJoinedText($, $el.find('.panelHeader').first())
const panelContent = $el.find('.panelContent').first()
const bodyText = extractBlockJoinedText($, panelContent.length > 0 ? panelContent : $el)
const label = headerText ? `[CALLOUT: ${headerText}]` : '[CALLOUT]'
$el.replaceWith($('<p></p>').text(`${label} ${bodyText}`))
}
progressed = true
})
}
return $.html()
}
/**
* Escapes a value for use inside CQL double-quoted strings.
*/
@@ -108,10 +258,12 @@ async function fetchLabelsForPages(
* invalidates every previously-synced Confluence document so a one-time
* re-hydration picks up content newly reachable by the current extraction
* (e.g. the switch from `storage` to rendered `view`, which expands Include
* Page / Excerpt macros). Without it, already-indexed pages whose version is
* unchanged classify as `unchanged` and keep their stale (empty) content.
* Page / Excerpt macros; or `preserveConfluenceCallouts`, which stops
* flattening panel/info/note/warning/tip macros into indistinguishable plain
* text). Without it, already-indexed pages whose version is unchanged
* classify as `unchanged` and keep their stale (pre-fix) content.
*/
const CONTENT_REPRESENTATION = 'view'
const CONTENT_REPRESENTATION = 'view-callouts'
/**
* Produces a canonical metadata stub with a deterministic contentHash that
@@ -288,7 +440,7 @@ export const confluenceConnector: ConnectorConfig = {
const body = page.body as Record<string, unknown> | undefined
const view = body?.view as Record<string, unknown> | undefined
const rawContent = (view?.value as string) || ''
const plainText = htmlToPlainText(rawContent)
const plainText = htmlToPlainText(preserveConfluenceCallouts(rawContent))
const labelMap = await fetchLabelsForPages(cloudId, accessToken, [String(page.id)])
const labels = labelMap.get(String(page.id)) ?? []