Markdown as the source of truth in Git — PDF on demand built with WeasyPrint
Git as the source of truth for a specification: semantic line breaks for readable diffs, a check proving nothing changed, and WeasyPrint for PDF on demand.

On this page
Introduction
I kept losing track of which specification version I was implementing against. The documents lived in Google Drive, so git had no copy of them, and someone else’s edits landed without my working tree changing. When I diffed two exports a day apart, edits I had never seen had changed what the code had to do.
Committing the export would not have helped. A PDF is a binary git stores whole at every revision, and an HTML export re-wraps its own lines, so a one-word edit arrives as a rewritten document.
So the specification moved into the repository as Markdown, and the PDF became something built on demand with WeasyPrint.
Markdown in the repo, PDF on demand
Markdown was the only candidate a reviewer and an agent could both read line by line, and Claude Code can edit it in the same commit as the code that made it wrong.
I just need a way to convert it to PDF on demand so I can ship it to the client at any time, and WeasyPrint has made it easy.
Who owns the source of truth — the repository or a hosted editor
The question to settle before any documentation tooling is who edits the canonical version. Master in the repository means generating distributables from it. Master elsewhere means mirroring and tracing, which is a different job with a different payoff.
Here the master was the hosted editor, so committing a Markdown copy and generating a PDF from it would have produced a third rendering of a document whose source of truth was somewhere else, competing with both.
The two exports I diffed differed by a table and by a renumbered list of open questions. A renumbered list leaves every link working and silently repoints it at a different, plausible-looking entry, which is the kind of change no amount of careful reading catches.
Reflowing the prose so a diff is readable
Semantic line breaks, one sentence per line
The longest line in the exported specification was 519 characters. At that width a one-word correction produces a hunk containing the whole paragraph, and a reviewer has to read all of it to find what moved.
Semantic Line Breaks is the convention that fixes this: break after every sentence, and optionally after an independent clause. Markdown joins consecutive lines into one paragraph, so the rendered output is unchanged and the source becomes line-addressable.
The reflow ran as a script rather than by hand. Splitting after the Japanese full stop 。 is a rule a script can apply uniformly, and a script will not get bored halfway through a 500-line document and start tidying other things while it is in there.
Japanese soft breaks arrive as spaces in the PDF
One sentence per line has a side effect in Japanese that it does not have in English. A soft line break becomes a newline in HTML and a renderer collapses it to a space. English wants that space between two sentences.
Japanese does not put spaces between sentences at all, so every break added for the diff showed up in the PDF as a gap in the middle of a paragraph.
The CSS Text specification has segment-break rules that discard a break falling between two East Asian characters, and browsers implement them. Relying on that would mean trusting the PDF renderer to implement them too, so the build collapses those breaks itself before conversion:
CJK = ( "\u3000-\u303f" # punctuation "\u3040-\u309f" # hiragana "\u30a0-\u30ff" # katakana "\u4e00-\u9fff" # kanji "\uff00-\uffef" # full-width forms)CJK_SOFTBREAK_RE = re.compile(rf"(?<=[{CJK}])\n[ \t]*(?=[{CJK}])")The optional horizontal whitespace is easy to leave out and hard to notice. A continuation line inside a list item is indented, so without [ \t]* the pattern misses exactly the lines a specification full of bullet lists is made of.
The source keeps its one-sentence-per-line form. Only the copy handed to the renderer is collapsed.
The export cannot be committed as it comes. Backslash escapes, cross-document links and line breaks all need rewriting first, and that work touches nearly every line. That is exactly what makes “did any of the text change?” impossible to answer by reading.
Proving the migration changed nothing
Turning the export into something reviewable
The claim that needed proving was narrow: no character outside whitespace had changed.
The script that does the rewriting was built one component at a time. Strip the escapes, re-run. Rewrite the links, re-run. Convert the tables, re-run. Whatever the diff still reports is the part that is not settled yet, so the check doubles as the to-do list.
That is why the expected text is recomputed rather than stored. The script grows between runs, so a saved copy would be an older script’s output, and the comparison would be testing a version of the rewriting that no longer exists.
The hand stage is everything a script cannot decide. A metadata table, a revision history, stable identifiers for the list of open questions. Those are additions, and what the check has to establish is that they are only additions.
With whitespace removed from both sides, the two strings should differ only where something was deliberately added:
WS = re.compile(r"\s+")a, b = WS.sub("", expected), WS.sub("", current)Every deletion the diff then reports has to be accounted for.
This ran while the script was being written, not in CI. Once the Markdown is the source of truth it is supposed to change, so a permanent version of this check would fail on the first legitimate edit.
Remove deliberate deletions before the diff runs
Some deletions were deliberate: strikethrough markers, a column header, a link to a document that no longer existed. The check had to excuse those and fail on everything else, so the first version kept an allowlist of the exact strings permitted to disappear. It failed on strings that were on the list.
difflib does not align text the way a reader does. It looks for the longest common subsequences, and it finds them inside a deleted line.
The comparison runs over each document as one whitespace-stripped string, and that file linked the same hosted workspace a dozen times over. So a removed link had near-identical text elsewhere to align against.
Three scraps of it were classified as retained on that basis: oogle, oc, and a lone s. The line arrived as four disconnected deletions with those three sitting between them, and listing the full line did not help, because the full line never appeared as a single opcode.
So the deliberate deletions come out of the expected text by line, before the comparison runs. The check asserts the line was present before removing it, so a typo in the list fails loudly instead of quietly excusing something else.
The vertical tabs that survived every check
Fourteen U+000B vertical tabs were sitting in the migrated files, and nothing in the pipeline had reported one of them.
The exporter emits one wherever a table cell contains a line break. They are control characters, so no renderer draws them, grep does not surface them, and the whitespace-insensitive comparison above counted them as whitespace and skipped over them. str.rstrip() had already eaten one from the end of a line without anyone noticing.
Converting them to <br> is the fix, and doing that before the trailing-whitespace strip is what stops rstrip from taking one first. The durable part is the assertion:
found = { f"U+{ord(ch):04X}" for ch in set(text) if ch not in "\n\t" and unicodedata.category(ch) in ("Cc", "Cf", "Co", "Cs")}Building the PDF on demand
The build reads the Markdown and writes a PDF. Nothing in between is committed, and the source file is never rewritten:
One source for the version number
The project I ported the script from kept the document version in two places: a table at the top of the Markdown, and a constant in the build script that printed the cover page. Raising a revision meant editing four things, and the two copies were free to disagree.
The build reads the table instead. A parse_meta function splits the file at the first ---, takes the title, the subtitle and every two-column row above it, and hands them to the cover template. A missing version number aborts the build rather than producing a PDF with a blank field on the cover.
The PDF is gitignored
WeasyPrint stamps a generation timestamp into every file, so rebuilding an unchanged document produces different bytes and a committed PDF collects a multi-megabyte diff on every build. Nothing is committed, so the filename is free to come from the document’s own heading rather than its revision number.
WeasyPrint could not find its own libraries on macOS
brew install pango succeeded and the build still failed:
OSError: cannot load library 'libgobject-2.0-0': dlopen(libgobject-2.0-0, 0x0002): tried: 'libgobject-2.0-0' (no such file)WeasyPrint loads Pango and its dependencies through dlopen, which searches the dynamic loader’s paths. A Python installed by uv is not a Homebrew Python, so Homebrew’s lib directory is not among them. Pointing the loader at it fixes the build:
BREW_PREFIX := $(shell brew --prefix 2>/dev/null)DOCS_DYLD := $(if $(BREW_PREFIX),DYLD_FALLBACK_LIBRARY_PATH="$(BREW_PREFIX)/lib",)
docs-pdf: $(DOCS_DYLD) uv run --group docs python tools/build_docs_pdf.pyThe conditional is the part worth copying. Setting the variable unconditionally leaves "/lib" in it on a machine with no Homebrew, and DYLD_FALLBACK_LIBRARY_PATH replaces the default search list rather than extending it, so that breaks library loading for anyone who installed Pango another way.
Match trailing whitespace with [ \t]*, never \s*
The build strips each heading’s `<a id="…"></a>` anchor before conversion, so the table of contents does not inherit raw HTML. \s matches newlines, so \s*$ takes the blank line after the heading along with the anchor, and the collapsing rule then joins 針 to 本:
## 実装方針
本書では設計と実装の対応を示す。## 実装方針本書では設計と実装の対応を示す。// the heading swallowed its own first paragraphThe contents page carried that whole line, and [ \t]* is the fix. Neither rewrite is wrong on its own — Markdown ends a heading at the newline, so losing the blank line costs nothing until the collapsing rule removes the newline too.
Summary
Putting a document in git as the source of truth is a mastership decision before it is a tooling one. If the canonical copy lives somewhere else, generating a PDF only adds a third version of it. The payoff arrives once the repository owns the document and its diffs become reviewable.
Two things carried most of the cost:
- Prose has to be reflowed at sentence boundaries before a diff means anything, and in Japanese the renderer needs those breaks collapsed again on the way out.
- A migration touching every line has to prove it changed nothing. That needs a comparison ignoring whitespace, a separate assertion for the control characters such a comparison ignores, and deliberate deletions removed by line rather than matched as strings.
References
- Semantic Line Breaks, the convention of breaking a line after each substantial unit of thought
- WeasyPrint documentation, including the Pango and system library dependencies it loads at import time
- CSS Text Module Level 3, whose segment break transformation rules discard a break between two East Asian characters
- Python difflib, whose SequenceMatcher opcodes produce the longest-common-subsequence alignment described above

