A Claude Code translation pipeline adds a fluency gate — modeled on Andrew Ng's reflection pattern

An outside review caught defects in a zh-TW translation this Claude Code pipeline's QA never checked for; the fix adds a research-grounded Section C gate.

The Claude starburst logo in terracotta and the Claude wordmark in black on a white card
On this page

Introduction

An outside AI reviewer read the zh-TW translation of this blog’s own cloud-publish article and found two real defects: an intro sentence carrying English clause structure straight across, and the same source term rendered two different ways within one section: once in a figure caption, once in the paragraph right after it. This blog’s own translation review gate runs before every publish specifically to catch drift like that, so a defect an outside reviewer found first meant the gate had a blind spot nobody had named yet. The blind spot was narrower than the reviewer’s own suggestion of separate agents per locale. One of the two defects already had a rule written for it in this blog’s own house style, and the fix was a second, narrower review pass, modeled on published research into how AI translation QA gets built elsewhere. This article walks through what the review caught, why the existing gate couldn’t have caught it, the research behind the fix, and the gate now sitting in the pipeline. It is untested against a real translation as of this writing.

An outside review, checked line by line

The review came from a separate AI session grading a zh-TW translation this blog had already published. Most of what it flagged held up against the file itself: the Introduction’s second paragraph opened with a sentence dense enough to bury its own point, a modifier clause was stacked in front of 的 the way an English relative clause would be, and the same English term, “unattended,” came out as 無人看管 in a figure caption and 無人值守 in the paragraph right after it. One claim did not hold up. The review said the article still had no flow diagram. It did: a <Pipeline /> figure had shipped with the article from its first publish, showing exactly the local-versus-cloud split and the merge decision the reviewer said was missing.

That single wrong claim was worth keeping in mind for the rest of this: an outside review is useful evidence, not a verdict to act on unread.

Why the existing gate missed both

This blog’s publish pipeline already runs a translation check before every locale ships: Section B of its review gate, checking correctness against the source, completeness, the fixed heading strings, title re-casting, terminology against a shared glossary, and untranslated leakage. None of those items is built to catch a sentence that changed nothing in meaning but reads like a rendering of English word order, and none of them catches two different renderings of the same term when neither one is wrong and neither is in the glossary.

The first instinct was to add two more bullets to that same checklist. That held up for exactly as long as it took to be challenged: mixing an objective, checklist-driven pass with a subjective judgment call risks the subjective one getting the same shallow, box-ticking treatment as the items already ahead of it in line. Section B already had nine items on it; these two would only have been items ten and eleven.

What actually happens elsewhere

Settling that meant looking at how AI translation QA gets built outside this one blog, not guessing. Andrew Ng’s translation-agent project runs three prompts in sequence against the same model: an initial translation, a reflection pass that critiques the first draft against four named dimensions (accuracy, fluency, style, terminology), and a refinement pass that rewrites using that critique. Its refinement prompt names terminology “inappropriate for context, inconsistent use” almost verbatim as this blog’s own zh-TW incident. Translation vendors converge on the same shape from the other direction: Smartling and Lilt both run quality review as an architecturally distinct pass from the translation itself, and the Multidimensional Quality Metrics framework that WMT has used as its own human-evaluation standard since 2021 scores the same four categories Ng’s prompt independently converges on.

None of that argues for a team of specialized agents. Ng’s reflection loop is one model, two prompts, one session, not separate personas, and the multi-agent architectures the research turned up (a simulated translation company with a CEO, editors, and a proofreader) are validated for ultra-long literary translation, not a blog post. DeepL, for comparison, gets its fluency from training a stronger single model rather than running a second pass at all, which is not an option open to a pipeline calling a general-purpose API.

The evidence has real limits, worth stating plainly rather than skipped. Ng’s own README calls the project “not mature software,” and nothing found in the research measures a reflection pass’s detection rate for clause-stacking or self-consistency specifically. The case for this design rests on the prompt’s wording matching this blog’s own incident, not a benchmark.

A rule already on file

One thing the research didn’t need to justify: this blog’s own house-style guide already had a rule for the clause-stacking half of the problem, written months earlier for exactly this reason. Its target-language conventions state it plainly: “A source sentence with two or more subordinate clauses is split, not preserved.” That rule already existed when the zh-TW article shipped with the sentence it describes. Nothing in the review gate that actually runs at publish time ever pointed a reviewer at it for translated output specifically. It lived in the style guide translators are supposed to read, not in the checklist a review pass actually works through.

That is why the first instinct, adding two lines to an already-long checklist, would not have been enough on its own. A rule already existed in prose form and still went unchecked, which argues for a dedicated read with nothing else competing for attention, not one more line easy to skim past.

A second, narrower pass

The fix landed as a new section in the review gate: Section C, run right after Section B, on the same locale, reading the whole translated file once as a native reader rather than sentence-by-sentence against the source.

The translation review gate, now three stepsA flow diagram with four boxes in a row: Translate, Section B, Section C, and Publish, each connected by an arrow. Section B is labelled correctness, completeness, glossary drift. Section C is labelled fluency, self-consistency, register, and is marked as new.TranslateSection Bcorrectness, completeness, glossary driftnewSection Cfluency, self-consistency, registerPublish
Section C runs immediately after Section B, on the same locale, before the file ships.

It checks three things the old checklist never touched: whether a sentence carries the source’s clause structure across the language boundary against the rule already on file, whether the same recurring term was rendered the same way everywhere in the document regardless of what the glossary says, and whether the register holds steady from the Introduction to the Summary. Applying it to what the outside review actually flagged shows what it is meant to catch.

The two findings from the outside review, as Section C names themTwo rows. The first, labelled Fluency, shows a Chinese sentence before and after: before, one long clause stacking a chain of modifiers ahead of the possessive particle; after, the same meaning split into two shorter sentences. The second row, labelled Terminology, shows the same English term rendered two different ways before, and one consistent term used everywhere after.Fluencya clause carried straight across from English word orderBefore以一份完整、執行中沒有任何東西需要再問的規格出現After以一份完整的規格出現。執行到一半,不會有任何東西需要再問。Terminologythe same source term, "unattended," rendered two different waysBefore無人看管無人值守two renderings, a caption and the paragraph after itAfter無人值守one rendering, used throughout
Both findings from the outside review, matched against the two checks Section C runs.

Neither rendering of “unattended” is wrong on its own, and neither is in this blog’s glossary, which is exactly why Section B’s terminology check, built to catch drift from the glossary, had nothing to compare either one against. Section C is wired into both places translation actually ships from: the local /blog publish command, and the task list a cloud routine reads when it publishes an article unattended, so a routine running this without anyone watching runs the same check a human-triggered publish would.

Summary

This blog’s translation review gate now has a third section, reading a translated file whole and checking fluency, internal terminology consistency, and register, on top of the correctness and completeness checks it already ran. The two specific defects that started this, a clause-stacked sentence and a split rendering of one English word, are exactly the two categories Section C’s checklist names. Whether it actually catches either kind of defect in a real translation is still an open question. No source found in the research measures a reflection-style pass’s detection rate for either failure mode; the case for this design rests on matching wording between the incident and the published pattern it borrows from, not a measurement. The next article this blog translates through the full pipeline, starting with this one, is what tests it — not this write-up of the reasoning behind it.

References

Share this article