Repository navigation
Conversation
- Inline raw HTML that spans lines and then fails to match no longer drops the rest of the paragraph: the tokens consumed by `next_line` are restored. A value-less attribute followed by a line break (`<a b\nc>`) is now accepted. - An HTML block start of type 7 can't interrupt a paragraph, so it is a lazy continuation line in block quotes, list items and footnotes. - Link labels are matched with Unicode case folding (spec 206, 540). - Tag names must start with an ASCII letter (`<33>`, `< a>`, spec 618, 621); `\` is not allowed in email autolinks but `'` is (spec 606). - `<!-->` and `<!--->` are HTML comments, comments end at the first `-->` (CommonMark 0.31, spec 626). - A link reference definition with text after its title is rejected (spec 209, 210). - Emphasis: implement the rule of 3 with the delimiter run lengths (spec 411, 412, 415, 429); Unicode symbols are punctuation for flanking (CommonMark 0.31, spec 354); marks that fail to open can still close an enclosing opener. - An escaped backslash at the end of a line is not a hard break. - Blanks after `[` are part of the link text. - HTML block locations start on their first line; an unclosed fenced code (or HTML) block at the end of the last list item or footnote no longer gets an extra empty line. - HTML renderer: image `alt` text concatenates the text of a line instead of joining it with newlines (spec 520, 573-577, 585, 589). Adds a test running all the CommonMark spec examples (631 -> 652 of 652 pass) and regression tests.
The second pass resolved emphasis by scanning forward from each opener, which differs from CommonMark when a delimiter run can both open and close (it must first close the nearest possible opener, e.g. `*a **b*c* d` is `<em>a **b</em>c* d`). Emphasis and strikethrough are now processed with the "process emphasis" procedure of the spec (delimiter stack with `openers_bottom`), on a linked list of the tokens. Multi-line emphasis locations now start on the opener's line. Also: - Attribute names must start with a letter, `_` or `:`; unquoted attribute values can't start with `"'=<>` or a backtick. - A backtick run preceded by `\` opens a code span after the escaped backtick. - Remember unterminated comments, processing instructions, CDATA sections and declarations so that raw HTML parsing stays linear.
|
Closing in favour of smaller, focused PRs. Each one covers one bug and has its own regression tests:
I merged all of them into a local branch and checked it: 652/652 spec examples pass on native, js, wasm-gc and wasm, and all regression tests from this PR pass. A differential run over the spec examples plus about 11k random inputs, in strict and relaxed mode, found no output differences from this branch. The |
|
Update: after a Codex review of the split PRs, the linear-time guard for unterminated comments, PIs, CDATA sections and declarations moved from #156 into #155. With the 0.31 comment rule, unterminated comments on a single line were already quadratic without it. #156 now only restores the tokens and is still stacked on #155. #160 also gained a fix: a blank last line without a final newline, such as "- ```\n a\n ", is kept as code content. This deliberately differs from this PR's output. The combined branch was re-verified: 652/652 spec examples and all regression tests from this PR pass on native, js, wasm-gc and wasm. |
|
Correction to my previous update. After more Codex review rounds, #155 (0.31 HTML comments) was folded into #156, because the two depend on each other: the new comment rule needs token restoration, and the linear-time guard needs the new rule for comments. #156 is now based on |
Renders every example of src/data/test/spec.md (CommonMark 0.31.2) with the XHTML renderer, compares it to the expected HTML and fails if any example fails. All 652 examples pass once the conformance fixes split out of #152 are merged; on main 631 pass.
Renders every example of src/data/test/spec.md (CommonMark 0.31.2) with the XHTML renderer, compares it to the expected HTML and fails if any example fails. All 652 examples pass once the conformance fixes split out of #152 are merged; on main 631 pass.
Fixes CommonMark conformance bugs found while building a markdown-it compatible adapter on top of cmark 0.4.9. Example numbers below are the official CommonMark 0.31.2 ones (1-based).
Spec conformance
A new test,
src/cmark_html/spec_examples_test.mbt, renders every example ofsrc/data/test/spec.mdwith the XHTML renderer and compares it to the expected HTML. It fails if any example fails.main)Failing before: 206, 209, 210, 354, 411, 412, 415, 429, 520, 540, 573–577, 585, 589, 606, 618, 621, 626.
Bugs fixed
x <a\nb→<p>x <a</p>.a <!-- b\nc\nd→<p>a <!-- b</p>try_autolink_or_htmlrecords the tokensnext_lineconsumes and puts them back on failure (tokens_next_line_consumed/tokens_restore).attributereturned the old line afternext_linehad already moved on.<b then\nc >was not raw HTMLattributenow continues from the new line.- a\n</kbd>,> a\n</kbd>HtmlBlockLine(EndBlank7)as a possible lazy continuation in block quotes, list items and footnotes.[ΑΓΩ]: /φου\n\n[αγω](206),[ẞ]\n\n[SS]: /url(540)link_labeluses@casefold.casefold_char(full Unicode case folding).<33>(618),< a>(621),</1a>,<b 1=c>,<a b=<c>,<foo\+@bar.example.com>(606)_or:. Unquoted attribute values can't start with"'=<>`. The email atext set had\where'belongs.<div>\nb\n</div>withlocs=trueblock_struct_to_html_blockstarts the location at the first line.- ```\n a\n→<pre><code>a\n\n</code></pre>end_docalso closes the blocks of the last list item and footnote.(marks.may_open || !opener.may_close)) and used the remaining counts instead of the delimiter run lengths.*foo**bar**baz*(411),*foo**bar*(412),*foo**bar***(415),**foo*bar*baz**(429)run_lengthto emphasis tokens.*$*alpha.(354)is_unicode_punctuation(P or S). It uses charclassis_symbol, or\p{S}on JS.<!-->and<!--->were not comments (CommonMark 0.31).html_commentaccepts them and ends at the first-->.[foo]: /url "title" ok(209),[foo]: /url\n"title" ok(210)line.last + 1instead of the title's end. A sequentialletreplaced OCaml's simultaneouslet … and ….Other bugs found along the way
Found by differential fuzzing against markdown-it-py 4.2 (
commonmarkpreset), and checked against the spec:alttext joined the segments of a line with newlines instead of concatenating them.gavealt="foo \nbar"(520, 573–577, 585, 589).openers_bottom, run over a linked list of the tokens. The previous approach scanned forward from openers and got cases wrong where a run can both open and close:*a **b*c* dmust be<em>a **b</em>c* d. Strikethrough is processed in the same loop, and the existing strikethrough snapshots are unchanged. Multi-line emphasis locations now start on the opener's line. This is a separate commit.a\\followed by a newline now rendersa\and a soft break.[were dropped:[ foo](/u)gave<a>foo</a>.\could never open a code span.\a`` now renders ``a``.Tests
src/cmark_html/spec_examples_test.mbt: all 652 spec examples.src/cmark_html/conformance_regression_test.mbt: regression tests for each bug above.src/cmark/block_test.mbt: HTML block location.raw_html_test.mbt(</1div>,<.div>,<!-->,<!--->),raw_html_wbtest.mbt(tag_nameon1div),inline_struct_wbtest.mbt(flanking after~, closer index).moon check,moon test(wasm-gc, wasm, js, native),moon fmtandmoon infoall pass. Public API is unchanged (no.mbtidiff). The version is not bumped.Known remaining differences (not changed here)
[with a single]at the end are still quadratic, as onmain.