Skip to content

Fix CommonMark conformance bugs (631 → 652/652 spec examples) - #152

Closed
bobzhang wants to merge 2 commits into
mainfrom
fix-commonmark-conformance
Closed

bobzhang wants to merge 2 commits into
mainfrom
fix-commonmark-conformance

Conversation

@bobzhang

@bobzhang bobzhang commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Fixes CommonMark conformance bugs found while building a markdown-it compatible adapter on top of cmark 0.4.9. Example numbers below are the official CommonMark 0.31.2 ones (1-based).

Spec conformance

A new test, src/cmark_html/spec_examples_test.mbt, renders every example of src/data/test/spec.md with the XHTML renderer and compares it to the expected HTML. It fails if any example fails.

passing examples
before (main) 631 / 652
after 652 / 652

Failing before: 206, 209, 210, 354, 411, 412, 415, 429, 520, 540, 573–577, 585, 589, 606, 618, 621, 626.

Bugs fixed

# Bug Repro Fix
1 Lines silently dropped. When inline raw HTML ran past a line end and then failed to match, the rest of the paragraph disappeared. x <a\nb → <p>x &lt;a</p>. a <!-- b\nc\nd → <p>a &lt;!-- b</p> try_autolink_or_html records the tokens next_line consumes and puts them back on failure (tokens_next_line_consumed / tokens_restore).
1b A value-less attribute followed by a line break was rejected, and lines were skipped, because attribute returned the old line after next_line had already moved on. <b then\nc > was not raw HTML attribute now continues from the new line.
2 A lazy continuation line starting with a complete HTML tag became an HTML block. HTML blocks of type 7 can't interrupt a paragraph. - a\n</kbd>, > a\n</kbd> Treat HtmlBlockLine(EndBlank7) as a possible lazy continuation in block quotes, list items and footnotes.
3 Link labels were case-folded for ASCII only. [ΑΓΩ]: /φου\n\n[αγω] (206), [ẞ]\n\n[SS]: /url (540) link_label uses @casefold.casefold_char (full Unicode case folding).
4 Invalid raw HTML and autolinks were accepted. <33> (618), < a> (621), </1a>, <b 1=c>, <a b=<c>, <foo\+@bar.example.com> (606) Tag names must start with an ASCII letter. Attribute names must start with a letter, _ or :. Unquoted attribute values can't start with "'=<>` . The email atext set had \ where ' belongs.
5a An HTML block's location covered only its last line. <div>\nb\n</div> with locs=true block_struct_to_html_block starts the location at the first line.
5b An unclosed fenced code (or HTML) block at the end of the last list item or footnote got an extra empty line. - ```\n a\n → <pre><code>a\n\n</code></pre> end_doc also closes the blocks of the last list item and footnote.
6a Emphasis "rule of 3" was inverted ((marks.may_open || !opener.may_close)) and used the remaining counts instead of the delimiter run lengths. *foo**bar**baz* (411), *foo**bar* (412), *foo**bar*** (415), **foo*bar*baz** (429) Fixed the condition and added run_length to emphasis tokens.
6b Unicode symbols (category S) were not punctuation for flanking. *$*alpha. (354) New private is_unicode_punctuation (P or S). It uses charclass is_symbol, or \p{S} on JS.
6c <!--> and <!---> were not comments (CommonMark 0.31). 626 html_comment accepts them and ends at the first -->.
6d A link reference definition followed by text after its title was accepted. [foo]: /url "title" ok (209), [foo]: /url\n"title" ok (210) The after-title check used line.last + 1 instead of the title's end. A sequential let replaced OCaml's simultaneous let … and ….

Other bugs found along the way

Found by differential fuzzing against markdown-it-py 4.2 (commonmark preset), and checked against the spec:

  • HTML renderer: image alt text joined the segments of a line with newlines instead of concatenating them. ![foo *bar*](/u) gave alt="foo \nbar" (520, 573–577, 585, 589).
  • Emphasis algorithm: emphasis is now resolved with the spec's process emphasis procedure: a delimiter stack with openers_bottom, run over a linked list of the tokens. The previous approach scanned forward from openers and got cases wrong where a run can both open and close: *a **b*c* d must be <em>a **b</em>c* d. Strikethrough is processed in the same loop, and the existing strikethrough snapshots are unchanged. Multi-line emphasis locations now start on the opener's line. This is a separate commit.
  • Hard breaks: an escaped backslash at the end of a line was treated as a hard break. a\\ followed by a newline now renders a\ and a soft break.
  • Link text: blanks after [ were dropped: [ foo](/u) gave <a>foo</a>.
  • Code spans: a backtick run preceded by \ could never open a code span. \a`` now renders ``a ``.
  • Performance: without a guard, restoring tokens (fix 1) would make unterminated comments, PIs, CDATA sections and declarations quadratic. The first failing position of each kind is now remembered per token stream, so parsing stays linear. There is a regression test with 20,000 lines.

Tests

  • src/cmark_html/spec_examples_test.mbt: all 652 spec examples.
  • src/cmark_html/conformance_regression_test.mbt: regression tests for each bug above.
  • src/cmark/block_test.mbt: HTML block location.
  • Updated tests that encoded the old behaviour: raw_html_test.mbt (</1div>, <.div>, <!-->, <!--->), raw_html_wbtest.mbt (tag_name on 1div), inline_struct_wbtest.mbt (flanking after ~, closer index).

moon check, moon test (wasm-gc, wasm, js, native), moon fmt and moon info all pass. Public API is unchanged (no .mbti diff). The version is not bumped.

Known remaining differences (not changed here)

  • Leading spaces of continuation lines inside multi-line code spans and raw HTML are stripped. commonmark.js and cmark keep them. This is a spec ambiguity and no spec example covers it.
  • Trailing blank lines of an unterminated HTML block of types 1–5 at the end of a list item are kept. commonmark.js strips them.
  • Long runs of [ with a single ] at the end are still quadratic, as on main.

- Inline raw HTML that spans lines and then fails to match no longer drops
  the rest of the paragraph: the tokens consumed by `next_line` are
  restored. A value-less attribute followed by a line break (`<a b\nc>`)
  is now accepted.
- An HTML block start of type 7 can't interrupt a paragraph, so it is a
  lazy continuation line in block quotes, list items and footnotes.
- Link labels are matched with Unicode case folding (spec 206, 540).
- Tag names must start with an ASCII letter (`<33>`, `< a>`, spec 618,
  621); `\` is not allowed in email autolinks but `'` is (spec 606).
- `<!-->` and `<!--->` are HTML comments, comments end at the first `-->`
  (CommonMark 0.31, spec 626).
- A link reference definition with text after its title is rejected
  (spec 209, 210).
- Emphasis: implement the rule of 3 with the delimiter run lengths
  (spec 411, 412, 415, 429); Unicode symbols are punctuation for flanking
  (CommonMark 0.31, spec 354); marks that fail to open can still close an
  enclosing opener.
- An escaped backslash at the end of a line is not a hard break.
- Blanks after `[` are part of the link text.
- HTML block locations start on their first line; an unclosed fenced code
  (or HTML) block at the end of the last list item or footnote no longer
  gets an extra empty line.
- HTML renderer: image `alt` text concatenates the text of a line instead
  of joining it with newlines (spec 520, 573-577, 585, 589).

Adds a test running all the CommonMark spec examples (631 -> 652 of 652
pass) and regression tests.
The second pass resolved emphasis by scanning forward from each opener,
which differs from CommonMark when a delimiter run can both open and
close (it must first close the nearest possible opener, e.g.
`*a **b*c* d` is `<em>a **b</em>c* d`). Emphasis and strikethrough are
now processed with the "process emphasis" procedure of the spec
(delimiter stack with `openers_bottom`), on a linked list of the tokens.
Multi-line emphasis locations now start on the opener's line.

Also:
- Attribute names must start with a letter, `_` or `:`; unquoted
  attribute values can't start with `"'=<>` or a backtick.
- A backtick run preceded by `\` opens a code span after the escaped
  backtick.
- Remember unterminated comments, processing instructions, CDATA
  sections and declarations so that raw HTML parsing stays linear.
@bobzhang

bobzhang commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Closing in favour of smaller, focused PRs. Each one covers one bug and has its own regression tests:

I merged all of them into a local branch and checked it: 652/652 spec examples pass on native, js, wasm-gc and wasm, and all regression tests from this PR pass. A differential run over the spec examples plus about 11k random inputs, in strict and relaxed mode, found no output differences from this branch. The fix-commonmark-conformance branch is kept.

@bobzhang bobzhang closed this Oct 1, 2026
@bobzhang

bobzhang commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Update: after a Codex review of the split PRs, the linear-time guard for unterminated comments, PIs, CDATA sections and declarations moved from #156 into #155. With the 0.31 comment rule, unterminated comments on a single line were already quadratic without it. #156 now only restores the tokens and is still stacked on #155. #160 also gained a fix: a blank last line without a final newline, such as "- ```\n a\n ", is kept as code content. This deliberately differs from this PR's output. The combined branch was re-verified: 652/652 spec examples and all regression tests from this PR pass on native, js, wasm-gc and wasm.

@bobzhang

bobzhang commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

Correction to my previous update. After more Codex review rounds, #155 (0.31 HTML comments) was folded into #156, because the two depend on each other: the new comment rule needs token restoration, and the linear-time guard needs the new rule for comments. #156 is now based on main and contains the token restoration, the 0.31 comment rule and the guard for comments, PIs, CDATA sections and declarations. #160 also stops dropping the empty last line of a closed fence at the end of input. The replacement PRs are now #153, #154, #156, #157, #158, #159, #160, #161, #162, #163, #164, #165, #166, #167 (stacked on #161) and #168 (merge last). Re-verified on a local branch built from the pushed branches: 652/652 spec examples, and all regression tests from this PR pass on native, js, wasm-gc and wasm.

bobzhang added a commit that referenced this pull request Oct 2, 2026
Renders every example of src/data/test/spec.md (CommonMark 0.31.2) with
the XHTML renderer, compares it to the expected HTML and fails if any
example fails. All 652 examples pass once the conformance fixes split
out of #152 are merged; on main 631 pass.
bobzhang added a commit that referenced this pull request Oct 2, 2026
Renders every example of src/data/test/spec.md (CommonMark 0.31.2) with
the XHTML renderer, compares it to the expected HTML and fails if any
example fails. All 652 examples pass once the conformance fixes split
out of #152 are merged; on main 631 pass.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant