CommonMark flanking rules print ** on Japanese and Chinese pages

CommonMark only lets ** close emphasis when it is right-flanking. Japanese and Chinese have no space after a full stop, so bold ships as literal asterisks.

The Markdown mark, a black rounded rectangle enclosing an M and a downward arrow, on a white card
On this page

Introduction

I published an article in three languages, opened the Japanese page to read it, and found asterisks printed around a sentence that was meant to be bold.

**SIGTERM を受け取った Chrome は、終了する際に自分の SingletonLock を削除します。**これはテスト用に…

The cause is in CommonMark. A ** run closes emphasis only when it is right-flanking, and a run sitting between a 。 and a letter is neither left- nor right-flanking. English never meets the condition, because a space follows every full stop. CJK has no such space.

What decides whether a ** can close

A run of asterisks is not an instruction. It is a candidate, and the two characters on either side of it decide what the candidate is allowed to become.

Left-flanking and right-flanking

Bold is delimited by a pair of runs, the way brackets are. One run opens the emphasis and the other closes it, and the parser decides which job each run is eligible for before it can pair them at all.

A run that qualifies for neither job is not an error. It is left on the page as two asterisks.

Eligibility is decided per run, and it is decided by adjacency. A run can open only if it is touching the text it would open, and close only if it is touching the text it would close. A space on the inside breaks that contact:

This is **bold text** here both runs touch the text → renders
This is ** bold text** here opening run has a space → stays literal
This is **bold text ** here closing run has a space → stays literal

CommonMark gives those two eligibilities names, and states each as a condition on the characters either side of the run:

  • A run is left-flanking when the run is not followed by whitespace, and either not followed by punctuation, or else is preceded by whitespace or punctuation.
  • A run is right-flanking when the run is not preceded by whitespace, and either not preceded by punctuation, or else is followed by whitespace or punctuation.

The subject of every clause is the run itself, and “followed” means the single character sitting immediately after it. Left-flanking is the eligibility to open, right-flanking the eligibility to close.

Which run fails, and why, in a bold span that does not renderThe same sentence written three ways, with each asterisk run labelled. In the first, the opening run is left-flanking and opens, the closing run is right-flanking and closes, and the page shows bold text. In the second, a space follows the opening run, so that run is neither left- nor right-flanking and cannot open, while the closing run is still valid but has no partner. In the third the same happens to the closing run. Both broken sentences appear with their asterisks intact.what is writtenwhat the page showsThis is **bold text** hereopening run · left-flanking ✓ · opensclosing run · right-flanking ✓ · closesThis is bold text hereThis is **␣bold text** hereopening run · neither ✗ · cannot openclosing run · right-flanking ✓ · valid, but no partnerThis is ** bold text** hereThis is **bold text␣** hereopening run · left-flanking ✓ · valid, but no partnerclosing run · neither ✗ · cannot closeThis is **bold text ** hereEmphasis needs both jobs filled, so one disqualified run is enough to leave the asterisks on the page.
The same sentence three times, with each run labelled. A space on the inside disqualifies exactly one run; the other stays valid and simply has no partner, which is enough to leave both sets of asterisks on the page.

A run may open emphasis only if it is left-flanking, and close it only if it is right-flanking. Nothing else about the document matters: not the paragraph, not the matching run, only the character before and the character after.

The two characters that decide a delimiter runA ** run drawn between the character before it and the character after it, above a table of four combinations. Letter on both sides can open and close. Punctuation before and a letter after can open but not close, which is the case that breaks CJK. Letter before and punctuation after can close but not open. Punctuation before and a space after can close.What the parser reads?character before**?character afterbefore / aftercan opencan closeletter / letteryesyespunctuation / letteryesnoletter / punctuationnoyespunctuation / spacenoyesThe second row is the CJK case: a closer after 。 with a letter next to it.
The two characters either side of a run decide it. A closing run preceded by punctuation needs whitespace or punctuation on its right; a letter there disqualifies it.

Read the closing case slowly, because it is the one that bites. A run preceded by punctuation is right-flanking only if what follows is whitespace or punctuation. A letter on the right, with punctuation on the left, satisfies neither clause.

Why a full stop in English is always safe

English writes a space after a full stop, and that space is doing load-bearing work:

**It removes the lock.** This is a test.

The closing run is preceded by . and followed by a space, so the second clause of the right-flanking test is satisfied and the run closes. That space is the only reason English works. Every sentence-final bold span in English lands in that shape by default, which is why the rule is invisible to anyone writing in it.

Why the same sentence works in English and not in JapaneseThe same bold sentence in two languages. In English a space follows the closing run, which satisfies the right-flanking test and the emphasis closes. In Japanese the next sentence begins immediately after the closing run, so a letter sits against it and the run stays on the page as literal asterisks.English**It removes the lock.**␣This is…space after the closer → closesJapanese**ロックを削除します。**これは…letter after the closer → stays literal
The same sentence in both languages. The space after the English full stop is what makes the closing run right-flanking; Japanese has nothing in that position.

Japanese and Chinese put no space between sentences. The closing run therefore sits directly against the first character of the next sentence, and that character is almost always a letter.

The three shapes that fail in CJK

Three arrangements come up constantly in CJK technical prose, and all three produce a run that can do nothing:

…削除します。**これは closing run: 。 on the left, こ on the right
…対象を**「引用」**とする opening run: を on the left, 「 on the right
…ではなく**`client-id`**です opening run: く on the left, ` on the right

The first is a bold span ending on its own sentence. The second is a quoted phrase, where the Japanese bracket is punctuation and disqualifies the opener. The third is emphasis wrapped around an inline code span, where the backtick does the same.

Each repair moves the punctuation out of the emphasis rather than adding a space inside it:

…削除します**。これは
…対象を「**引用**」とする
…ではなく **`client-id`** です
The three failing shapes and their repairsThree rows. A bold span ending on a full stop, a bold span wrapped around a quoted phrase, and a bold span wrapped around an inline code span. Each is shown as it fails and as it is repaired: the first two move the punctuation outside the emphasis, and the third adds a space outside each delimiter because a backtick cannot be reordered.stays literalrenders…削除します。**これは…削除します**。これはsentence-final 。…対象を**「引用」**とする…対象を「**引用**」とするquote bracket…ではなく**`client-id`**です…ではなく **`client-id`** ですinline codeOnly the third needs a space; the first two move a character out of the emphasis.
The three shapes and what fixes each. The first two move a character out of the emphasis; only the third, where a backtick cannot be reordered, resorts to a space on the outside.

The third repair is the ugly one, because a space between Japanese text and a bold run is visible. It is the only option when the delimiter sits against a backtick that cannot move.

Four sweeps, and the one that broke three paragraphs

Repairing this by hand was not practical, so I swept it: move sentence-final punctuation outside a failing closer, move paired brackets outside a failing opener, and fall back to a space where neither applies.

The third sweep is where I broke things. I paired the delimiter runs per line, which is wrong whenever a paragraph wraps:

…把同一段流程分別跑在釋出版與修正版的打包檔上,**組字中的狀態完全相同,
結果卻不同**:

Read that second line alone and the ** on it is the first run on the line, so a per-line pass calls it an opener. It is a closer. The sweep tested it for left-flanking, found it wanting, and inserted a space in front of it — which is exactly the thing that stops a closer from being right-flanking.

Pairing per line reads a closing run as an openerA single paragraph wrapped across two source lines, with a bold run opening on the first line and closing on the second. Read line by line, the run on the second line is the first one seen there and is treated as an opener, so a repair pass puts a space in front of it and breaks it. Read per paragraph, it is the second run and is correctly treated as the closer.one paragraph, wrapped over two lines…打包檔上,**組字中的狀態完全相同,結果卻不同**:run 1 · openerpaired per linerun 1 · openerwrong: a space gets added in front of it結果卻不同␣**:paired per paragraphrun 2 · closerright: tested for closing, left alone結果卻不同**:
The same closing run, read two ways. Pairing per line makes it the first run and therefore an opener; pairing per paragraph keeps it the second, which is what it is.

A repair that can introduce the defect it repairs needs the same check pointed at its output. Running the renderer over every file afterwards is what found the three paragraphs I had broken, and a fourth sweep put them back.

The check that blocks it now

The rule is a function of two characters, so the check is short. It pairs runs per paragraph, applies both flanking tests, and reports any run that can neither open nor close where it is needed:

scripts/prose-check.ts
const canOpen = !isSpace(next) && (!isPunct(next) || isSpace(prev) || isPunct(prev));
const canClose = !isSpace(prev) && (!isPunct(prev) || isSpace(next) || isPunct(next));
const closing = i % 2 === 1;
if (closing ? canClose : canOpen) continue;

Two details are worth more than they look. A paragraph ends at a blank line, not at a gap in line numbers, or every stretch of prose between two code fences is treated as one unit.

And an odd number of runs is reported rather than skipped. A lone ** always renders literally, so it is itself the defect. Skipping the region would switch the check off for everything after it while still printing a clean count.

A gate nobody has seen fail is not a gate, so I broke a repaired file on purpose:

Terminal window
IN
node scripts/prose-check.ts <slug> ja
OUT
emphasis delimiters that cannot render: 1
line 184 closing …を削除します。**これはテスト用に起動した…

Restoring the file returned it to zero and exit 0. The check runs for ja and zh-tw only; English cannot produce the shape.

Summary

  1. A ** run opens emphasis only when left-flanking and closes only when right-flanking, and both tests read just the character before and the character after.
  2. A closing run preceded by punctuation needs whitespace or punctuation on its right, so 。**これは can neither open nor close and ships as literal asterisks.
  3. English never meets the condition because it writes a space after every full stop. The defect is invisible in a source locale and systematic in its translations.
  4. The repairs move punctuation out of the emphasis; a space outside the delimiter is the fallback for a run against a backtick.
  5. Pair delimiters per paragraph. A wrapped paragraph can begin a line with its closing run, and a per-line pass reads it as an opener.
  6. The rule is two characters wide, which makes it worth checking mechanically rather than reviewing for.

References

Share this article