Markdown as the source of truth in Git — PDF on demand built with WeasyPrint

Git as the source of truth for a specification: semantic line breaks for readable diffs, a check proving nothing changed, and WeasyPrint for PDF on demand.

The WeasyPrint logo, the lowercase word weasyprint in dark navy between a green chevron and a green rounded rectangle, on a white card
On this page

Introduction

I kept losing track of which specification version I was implementing against. The documents lived in Google Drive, so git had no copy of them, and someone else’s edits landed without my working tree changing. When I diffed two exports a day apart, edits I had never seen had changed what the code had to do.

Committing the export would not have helped. A PDF is a binary git stores whole at every revision, and an HTML export re-wraps its own lines, so a one-word edit arrives as a rewritten document.

So the specification moved into the repository as Markdown, and the PDF became something built on demand with WeasyPrint.

Markdown in the repo, PDF on demand

Markdown was the only candidate a reviewer and an agent could both read line by line, and Claude Code can edit it in the same commit as the code that made it wrong.

I just need a way to convert it to PDF on demand so I can ship it to the client at any time, and WeasyPrint has made it easy.

Who owns the source of truth — the repository or a hosted editor

The question to settle before any documentation tooling is who edits the canonical version. Master in the repository means generating distributables from it. Master elsewhere means mirroring and tracing, which is a different job with a different payoff.

Who owns the canonical copy decides whether a PDF helpsTwo cards side by side. On the left the repository owns the document: Markdown is the master and a PDF is generated from it, giving one canonical copy and reviewable diffs. On the right a hosted editor owns it: the document is exported into a Markdown mirror and a PDF is generated from that, giving three copies of one document and no diff finer than the whole file.Repository owns the documentMarkdownmastergeneratePDFhanded to readersOne canonical copy.Diffs are reviewable.Hosted editor owns the documentHosted editormasterexportMarkdown mirrorreflows on every exportgeneratePDFThree copies of one document.No diff finer than the file.
A Markdown source and a generated PDF support two different arrangements. Which one a project is in decides whether generating a PDF is useful or just a third copy.

Here the master was the hosted editor, so committing a Markdown copy and generating a PDF from it would have produced a third rendering of a document whose source of truth was somewhere else, competing with both.

The two exports I diffed differed by a table and by a renumbered list of open questions. A renumbered list leaves every link working and silently repoints it at a different, plausible-looking entry, which is the kind of change no amount of careful reading catches.

Reflowing the prose so a diff is readable

Semantic line breaks, one sentence per line

The longest line in the exported specification was 519 characters. At that width a one-word correction produces a hunk containing the whole paragraph, and a reviewer has to read all of it to find what moved.

Semantic Line Breaks is the convention that fixes this: break after every sentence, and optionally after an independent clause. Markdown joins consecutive lines into one paragraph, so the rendered output is unchanged and the source becomes line-addressable.

Hunk size for the same one-word correctionTwo panels. The upper panel is a paragraph stored as a single 519-character line, and every one of its four wrapped rows is highlighted as changed. The lower panel is the same paragraph stored one sentence per line, where only the second of four rows is highlighted. The correction is identical in both; only the size of the region a reviewer has to read differs.Source: one line of 519 charactersShown as changed:the whole paragraphSource: one sentence per lineShown as changed:one sentence
The same one-word correction, against one long line and against one sentence per line.

The reflow ran as a script rather than by hand. Splitting after the Japanese full stop 。 is a rule a script can apply uniformly, and a script will not get bored halfway through a 500-line document and start tidying other things while it is in there.

Japanese soft breaks arrive as spaces in the PDF

One sentence per line has a side effect in Japanese that it does not have in English. A soft line break becomes a newline in HTML and a renderer collapses it to a space. English wants that space between two sentences.

Japanese does not put spaces between sentences at all, so every break added for the diff showed up in the PDF as a gap in the middle of a paragraph.

The CSS Text specification has segment-break rules that discard a break falling between two East Asian characters, and browsers implement them. Relying on that would mean trusting the PDF renderer to implement them too, so the build collapses those breaks itself before conversion:

tools/build_docs_pdf.py
CJK = (
"\u3000-\u303f" # punctuation
"\u3040-\u309f" # hiragana
"\u30a0-\u30ff" # katakana
"\u4e00-\u9fff" # kanji
"\uff00-\uffef" # full-width forms
)
CJK_SOFTBREAK_RE = re.compile(rf"(?<=[{CJK}])\n[ \t]*(?=[{CJK}])")

The optional horizontal whitespace is easy to leave out and hard to notice. A continuation line inside a list item is indented, so without [ \t]* the pattern misses exactly the lines a specification full of bullet lists is made of.

The source keeps its one-sentence-per-line form. Only the copy handed to the renderer is collapsed.

The export cannot be committed as it comes. Backslash escapes, cross-document links and line breaks all need rewriting first, and that work touches nearly every line. That is exactly what makes “did any of the text change?” impossible to answer by reading.

Proving the migration changed nothing

Turning the export into something reviewable

The claim that needed proving was narrow: no character outside whitespace had changed.

The script that does the rewriting was built one component at a time. Strip the escapes, re-run. Rewrite the links, re-run. Convert the tables, re-run. Whatever the diff still reports is the part that is not settled yet, so the check doubles as the to-do list.

The loop that shrinks the diff until only deliberate additions remainThe original export feeds a mechanical script that gains one rule per round, producing an expected text recomputed on every run. That expected text and the file committed in the repository feed a comparison made with all whitespace removed. What the comparison reports is the remaining diff, the part not settled yet, which sends the author back to add the next rule to the script. The loop ends when only the deliberate additions are left.Original exportuntouchedMechanical scriptone rule added per roundExpected textrecomputed every runCurrent filecommitted in the repoCompare with all whitespace removedRemaining diffwhat is not settled yetadd the next ruleDone when only the deliberate additions are left
Each round adds one rule to the script, so the expected text is a different value every run. What the comparison reports is the work left, not a verdict.

That is why the expected text is recomputed rather than stored. The script grows between runs, so a saved copy would be an older script’s output, and the comparison would be testing a version of the rewriting that no longer exists.

The hand stage is everything a script cannot decide. A metadata table, a revision history, stable identifiers for the list of open questions. Those are additions, and what the check has to establish is that they are only additions.

With whitespace removed from both sides, the two strings should differ only where something was deliberately added:

WS = re.compile(r"\s+")
a, b = WS.sub("", expected), WS.sub("", current)

Every deletion the diff then reports has to be accounted for.

This ran while the script was being written, not in CI. Once the Markdown is the source of truth it is supposed to change, so a permanent version of this check would fail on the first legitimate edit.

Remove deliberate deletions before the diff runs

Some deletions were deliberate: strikethrough markers, a column header, a link to a document that no longer existed. The check had to excuse those and fail on everything else, so the first version kept an allowlist of the exact strings permitted to disappear. It failed on strings that were on the list.

difflib does not align text the way a reader does. It looks for the longest common subsequences, and it finds them inside a deleted line.

The comparison runs over each document as one whitespace-stripped string, and that file linked the same hosted workspace a dozen times over. So a removed link had near-identical text elsewhere to align against.

Three scraps of it were classified as retained on that basis: oogle, oc, and a lone s. The line arrived as four disconnected deletions with those three sitting between them, and listing the full line did not help, because the full line never appeared as a single opcode.

Why an exact-string allowlist never matches a deletionThree rows. The first is the link being deleted, shown whole. The second is what difflib reported: the same characters broken into seven pieces, because three scraps inside the link also occur in the links that stayed and were classified as kept, leaving four separate deletions around them. The third is the allowlist entry, which holds the whole line and therefore matches no single piece that difflib produced.The link being deletedhttps://docs.google.com/document/u/0/d/XXXXXXXXXXXXs-XXXXXXXX/editWhat difflib reportedhttps://docs.google.com/document/u/0/d/XXXXXXXXXXXXs-XXXXXXXX/editreported as deletedclassified as kept, because the same characters occur in links that stayedThe allowlist entrymatches no single piecehttps://docs.google.com/document/u/0/d/XXXXXXXXXXXXs-XXXXXXXX/edit
The allowlist holds the line. What the check reported was four deletions separated by three scraps it had matched against links elsewhere in the file, and the entry matches none of them. The document id is redacted; nothing else is.

So the deliberate deletions come out of the expected text by line, before the comparison runs. The check asserts the line was present before removing it, so a typo in the list fails loudly instead of quietly excusing something else.

The vertical tabs that survived every check

Fourteen U+000B vertical tabs were sitting in the migrated files, and nothing in the pipeline had reported one of them.

The exporter emits one wherever a table cell contains a line break. They are control characters, so no renderer draws them, grep does not surface them, and the whitespace-insensitive comparison above counted them as whitespace and skipped over them. str.rstrip() had already eaten one from the end of a line without anyone noticing.

Converting them to <br> is the fix, and doing that before the trailing-whitespace strip is what stops rstrip from taking one first. The durable part is the assertion:

found = {
f"U+{ord(ch):04X}"
for ch in set(text)
if ch not in "\n\t" and unicodedata.category(ch) in ("Cc", "Cf", "Co", "Cs")
}

Building the PDF on demand

The build reads the Markdown and writes a PDF. Nothing in between is committed, and the source file is never rewritten:

How one Markdown file becomes a PDFA pipeline running left to right. The Markdown in the repository is read and never written. Two rewrites are applied to a copy of it: heading anchors are lifted out, and newlines between Japanese characters are collapsed. The result becomes HTML, which WeasyPrint renders to PDF. A second branch reads the same Markdown with parse_meta to build the cover page, which also feeds WeasyPrint.Markdown in the reporead, never writtenlift heading anchorscollapse Japanese newlinesHTMLWeasyPrintPDFparse_metacover pagetitle, subtitle, version
Two rewrites run on a copy before conversion, and the cover page comes from a second, independent read of the same file. This is a different pipeline from the Google Docs migration the verification sections are about.

One source for the version number

The project I ported the script from kept the document version in two places: a table at the top of the Markdown, and a constant in the build script that printed the cover page. Raising a revision meant editing four things, and the two copies were free to disagree.

The build reads the table instead. A parse_meta function splits the file at the first ---, takes the title, the subtitle and every two-column row above it, and hands them to the cover template. A missing version number aborts the build rather than producing a PDF with a blank field on the cover.

The PDF is gitignored

WeasyPrint stamps a generation timestamp into every file, so rebuilding an unchanged document produces different bytes and a committed PDF collects a multi-megabyte diff on every build. Nothing is committed, so the filename is free to come from the document’s own heading rather than its revision number.

WeasyPrint could not find its own libraries on macOS

brew install pango succeeded and the build still failed:

OSError: cannot load library 'libgobject-2.0-0': dlopen(libgobject-2.0-0, 0x0002): tried: 'libgobject-2.0-0' (no such file)

WeasyPrint loads Pango and its dependencies through dlopen, which searches the dynamic loader’s paths. A Python installed by uv is not a Homebrew Python, so Homebrew’s lib directory is not among them. Pointing the loader at it fixes the build:

Makefile
BREW_PREFIX := $(shell brew --prefix 2>/dev/null)
DOCS_DYLD := $(if $(BREW_PREFIX),DYLD_FALLBACK_LIBRARY_PATH="$(BREW_PREFIX)/lib",)
docs-pdf:
$(DOCS_DYLD) uv run --group docs python tools/build_docs_pdf.py

The conditional is the part worth copying. Setting the variable unconditionally leaves "/lib" in it on a machine with no Homebrew, and DYLD_FALLBACK_LIBRARY_PATH replaces the default search list rather than extending it, so that breaks library loading for anyone who installed Pango another way.

Match trailing whitespace with [ \t]*, never \s*

The build strips each heading’s `<a id="…"></a>` anchor before conversion, so the table of contents does not inherit raw HTML. \s matches newlines, so \s*$ takes the blank line after the heading along with the anchor, and the collapsing rule then joins 針 to 本:

anchor lifted with [ \t]*$
## 実装方針
本書では設計と実装の対応を示す。
anchor lifted with \s*$
## 実装方針本書では設計と実装の対応を示す。
// the heading swallowed its own first paragraph

The contents page carried that whole line, and [ \t]* is the fix. Neither rewrite is wrong on its own — Markdown ends a heading at the newline, so losing the blank line costs nothing until the collapsing rule removes the newline too.

Summary

Putting a document in git as the source of truth is a mastership decision before it is a tooling one. If the canonical copy lives somewhere else, generating a PDF only adds a third version of it. The payoff arrives once the repository owns the document and its diffs become reviewable.

Two things carried most of the cost:

  1. Prose has to be reflowed at sentence boundaries before a diff means anything, and in Japanese the renderer needs those breaks collapsed again on the way out.
  2. A migration touching every line has to prove it changed nothing. That needs a comparison ignoring whitespace, a separate assertion for the control characters such a comparison ignores, and deliberate deletions removed by line rather than matched as strings.

References

Share this article