# Markdown as the source of truth in Git — PDF on demand built with WeasyPrint

> Git as the source of truth for a specification: semantic line breaks for readable diffs, a check proving nothing changed, and WeasyPrint for PDF on demand.

- Source: https://oharu121.com/blog/markdown-source-of-truth-git-weasyprint-pdf-pipeline/
- Published: 2026-09-19T14:26:43+09:00
- Tags: Markdown, PDF, WeasyPrint, Git, Developer Tooling

---
## Introduction

I kept losing track of which specification version I was implementing against. The documents lived in Google Drive, so `git` had no copy of them, and someone else's edits landed without my working tree changing. When I diffed two exports a day apart, **edits I had never seen had changed what the code had to do**.

Committing the export would not have helped. A PDF is a binary `git` stores whole at every revision, and an HTML export re-wraps its own lines, so a one-word edit arrives as a rewritten document.

So the specification moved into the repository as Markdown, and the PDF became something built on demand with WeasyPrint.

## Markdown in the repo, PDF on demand

**Markdown was the only candidate a reviewer and an agent could both read line by line**, and Claude Code can edit it in the same commit as the code that made it wrong.

I just need a way to convert it to PDF on demand so I can ship it to the client at any time, and WeasyPrint has made it easy.

## Who owns the source of truth — the repository or a hosted editor

**The question to settle before any documentation tooling is who edits the canonical version.** Master in the repository means generating distributables from it. Master elsewhere means mirroring and tracing, which is a different job with a different payoff.

*Figure — MastershipDecision: A Markdown source and a generated PDF support two different arrangements. Which one a project is in decides whether generating a PDF is useful or just a third copy.*

Here the master was the hosted editor, so committing a Markdown copy and generating a PDF from it would have produced a third rendering of a document whose source of truth was somewhere else, competing with both.

The two exports I diffed differed by a table and by a renumbered list of open questions. **A renumbered list leaves every link working and silently repoints it at a different, plausible-looking entry**, which is the kind of change no amount of careful reading catches.

## Reflowing the prose so a diff is readable

### Semantic line breaks, one sentence per line

The longest line in the exported specification was 519 characters. **At that width a one-word correction produces a hunk containing the whole paragraph**, and a reviewer has to read all of it to find what moved.

[Semantic Line Breaks](https://sembr.org/) is the convention that fixes this: break after every sentence, and optionally after an independent clause. Markdown joins consecutive lines into one paragraph, so **the rendered output is unchanged and the source becomes line-addressable**.

*Figure — DiffGranularity: The same one-word correction, against one long line and against one sentence per line.*

The reflow ran as a script rather than by hand. Splitting after the Japanese full stop `。` is a rule a script can apply uniformly, and a script will not get bored halfway through a 500-line document and start tidying other things while it is in there.

### Japanese soft breaks arrive as spaces in the PDF

One sentence per line has a side effect in Japanese that it does not have in English. A soft line break becomes a newline in HTML and a renderer collapses it to a space. English wants that space between two sentences.

**Japanese does not put spaces between sentences at all**, so every break added for the diff showed up in the PDF as a gap in the middle of a paragraph.

The CSS Text specification has segment-break rules that discard a break falling between two East Asian characters, and browsers implement them. Relying on that would mean trusting the PDF renderer to implement them too, so the build collapses those breaks itself before conversion:

```python title="tools/build_docs_pdf.py"
CJK = (
    "\u3000-\u303f"  # punctuation
    "\u3040-\u309f"  # hiragana
    "\u30a0-\u30ff"  # katakana
    "\u4e00-\u9fff"  # kanji
    "\uff00-\uffef"  # full-width forms
)
CJK_SOFTBREAK_RE = re.compile(rf"(?<=[{CJK}])\n[ \t]*(?=[{CJK}])")
```

The optional horizontal whitespace is easy to leave out and hard to notice. A continuation line inside a list item is indented, so without `[ \t]*` the pattern misses exactly the lines a specification full of bullet lists is made of.

The source keeps its one-sentence-per-line form. Only the copy handed to the renderer is collapsed.

**The export cannot be committed as it comes.** Backslash escapes, cross-document links and line breaks all need rewriting first, and that work touches nearly every line. That is exactly what makes "did any of the text change?" impossible to answer by reading.

## Proving the migration changed nothing

### Turning the export into something reviewable

The claim that needed proving was narrow: **no character outside whitespace had changed.**

The script that does the rewriting was built one component at a time. Strip the escapes, re-run. Rewrite the links, re-run. Convert the tables, re-run. **Whatever the diff still reports is the part that is not settled yet**, so the check doubles as the to-do list.

*Figure — VerificationStages: Each round adds one rule to the script, so the expected text is a different value every run. What the comparison reports is the work left, not a verdict.*

That is why the expected text is recomputed rather than stored. The script grows between runs, so a saved copy would be an older script's output, and the comparison would be testing a version of the rewriting that no longer exists.

The hand stage is everything a script cannot decide. A metadata table, a revision history, stable identifiers for the list of open questions. Those are additions, and what the check has to establish is that they are *only* additions.

With whitespace removed from both sides, the two strings should differ only where something was deliberately added:

```python
WS = re.compile(r"\s+")
a, b = WS.sub("", expected), WS.sub("", current)
```

Every deletion the diff then reports has to be accounted for.

**This ran while the script was being written, not in CI.** Once the Markdown is the source of truth it is supposed to change, so a permanent version of this check would fail on the first legitimate edit.

### Remove deliberate deletions before the diff runs

Some deletions were deliberate: strikethrough markers, a column header, a link to a document that no longer existed. The check had to excuse those and fail on everything else, so the first version kept an allowlist of the exact strings permitted to disappear. **It failed on strings that were on the list.**

`difflib` does not align text the way a reader does. It looks for the longest common subsequences, and **it finds them inside a deleted line**.

**The comparison runs over each document as one whitespace-stripped string**, and that file linked the same hosted workspace a dozen times over. So a removed link had near-identical text elsewhere to align against.

Three scraps of it were classified as retained on that basis: `oogle`, `oc`, and a lone `s`. The line arrived as four disconnected deletions with those three sitting between them, and listing the full line did not help, because the full line never appeared as a single opcode.

*Figure — DiffFragments: The allowlist holds the line. What the check reported was four deletions separated by three scraps it had matched against links elsewhere in the file, and the entry matches none of them. The document id is redacted; nothing else is.*

**So the deliberate deletions come out of the expected text by line, before the comparison runs.** The check asserts the line was present before removing it, so a typo in the list fails loudly instead of quietly excusing something else.

### The vertical tabs that survived every check

Fourteen `U+000B` vertical tabs were sitting in the migrated files, and **nothing in the pipeline had reported one of them.**

The exporter emits one wherever a table cell contains a line break. They are control characters, so no renderer draws them, `grep` does not surface them, and the whitespace-insensitive comparison above counted them as whitespace and skipped over them. `str.rstrip()` had already eaten one from the end of a line without anyone noticing.

Converting them to `<br>` is the fix, and doing that before the trailing-whitespace strip is what stops `rstrip` from taking one first. The durable part is the assertion:

```python
found = {
    f"U+{ord(ch):04X}"
    for ch in set(text)
    if ch not in "\n\t" and unicodedata.category(ch) in ("Cc", "Cf", "Co", "Cs")
}
```

> **CAREFUL**
>
> **A verifier that ignores whitespace needs a separate check for what it is ignoring.** Control characters fall in the gap between the two, which is why they survived a comparison specifically designed to catch changes.

## Building the PDF on demand

The build reads the Markdown and writes a PDF. Nothing in between is committed, and the source file is never rewritten:

*Figure — PdfPipeline: Two rewrites run on a copy before conversion, and the cover page comes from a second, independent read of the same file. This is a different pipeline from the Google Docs migration the verification sections are about.*

### One source for the version number

The project I ported the script from kept the document version in two places: a table at the top of the Markdown, and a constant in the build script that printed the cover page. Raising a revision meant editing four things, and the two copies were free to disagree.

**The build reads the table instead.** A `parse_meta` function splits the file at the first `---`, takes the title, the subtitle and every two-column row above it, and hands them to the cover template. A missing version number aborts the build rather than producing a PDF with a blank field on the cover.

### The PDF is gitignored

**WeasyPrint stamps a generation timestamp into every file**, so rebuilding an unchanged document produces different bytes and a committed PDF collects a multi-megabyte diff on every build. Nothing is committed, so the filename is free to come from the document's own heading rather than its revision number.

### WeasyPrint could not find its own libraries on macOS

`brew install pango` succeeded and the build still failed:

```text
OSError: cannot load library 'libgobject-2.0-0': dlopen(libgobject-2.0-0, 0x0002): tried: 'libgobject-2.0-0' (no such file)
```

WeasyPrint loads Pango and its dependencies through `dlopen`, which searches the dynamic loader's paths. A Python installed by `uv` is not a Homebrew Python, so Homebrew's `lib` directory is not among them. Pointing the loader at it fixes the build:

```make title="Makefile"
BREW_PREFIX := $(shell brew --prefix 2>/dev/null)
DOCS_DYLD := $(if $(BREW_PREFIX),DYLD_FALLBACK_LIBRARY_PATH="$(BREW_PREFIX)/lib",)

docs-pdf:
	$(DOCS_DYLD) uv run --group docs python tools/build_docs_pdf.py
```

The conditional is the part worth copying. Setting the variable unconditionally leaves `"/lib"` in it on a machine with no Homebrew, and **`DYLD_FALLBACK_LIBRARY_PATH` replaces the default search list rather than extending it**, so that breaks library loading for anyone who installed Pango another way.

## Match trailing whitespace with `[ \t]*`, never `\s*`

The build strips each heading's `` `<a id="…"></a>` `` anchor before conversion, so the table of contents does not inherit raw HTML. **`\s` matches newlines**, so `\s*$` takes the blank line after the heading along with the anchor, and the collapsing rule then joins 針 to 本:

```text title="anchor lifted with [ \t]*$"
## 実装方針

本書では設計と実装の対応を示す。
```

```text title="anchor lifted with \s*$"
## 実装方針本書では設計と実装の対応を示す。
// the heading swallowed its own first paragraph
```

The contents page carried that whole line, and `[ \t]*` is the fix. **Neither rewrite is wrong on its own** — Markdown ends a heading at the newline, so losing the blank line costs nothing until the collapsing rule removes the newline too.

## Summary

Putting a document in `git` as the source of truth is a mastership decision before it is a tooling one. **If the canonical copy lives somewhere else, generating a PDF only adds a third version of it.** The payoff arrives once the repository owns the document and its diffs become reviewable.

Two things carried most of the cost:

1. Prose has to be reflowed at sentence boundaries before a diff means anything, and in Japanese the renderer needs those breaks collapsed again on the way out.
2. A migration touching every line has to prove it changed nothing. That needs a comparison ignoring whitespace, a separate assertion for the control characters such a comparison ignores, and deliberate deletions removed by line rather than matched as strings.

## References

- [Semantic Line Breaks, the convention of breaking a line after each substantial unit of thought](https://sembr.org/)
- [WeasyPrint documentation, including the Pango and system library dependencies it loads at import time](https://doc.courtbouillon.org/weasyprint/stable/)
- [CSS Text Module Level 3, whose segment break transformation rules discard a break between two East Asian characters](https://www.w3.org/TR/css-text-3/#line-break-transform)
- [Python difflib, whose SequenceMatcher opcodes produce the longest-common-subsequence alignment described above](https://docs.python.org/3/library/difflib.html)
