AI Workflows Beat Human Translators in 4 of 6 Content Types in China Localization Benchmark
Blind-scored benchmark of 774 English-to-Chinese outputs: human translators placed outside the top five on 4 of 6 content types, trailing post-edited Qwen by 22.2 points on marketing copy.

- Human translators finished 7th, 7th, 9th and 10th of 15 workflows on technical, product UI, UGC and marketing content; on marketing they scored 53.7 vs 75.9 for post-edited Qwen, a 22.2-point gap.
- Humans won informational (76.9) and SEO (74.1) content; the SEO lead over the best AI workflow (PE-Doubao/PE-Qwen at 71.3) was only 2.8 points.
- The study measured localization quality only across 774 outputs scored blind by Chinese native-speaking localizers; no ranking, traffic or conversion data was collected.
- Chinese models varied up to 22.3 points within a single content type (Qwen 70.4 vs Kimi 48.1 on technical), while post-editing lift averaged 5.6 points — model choice mattered more than the editing layer.
Professional human translators finished outside the top five in four of six content types in a blind-scored English-to-Chinese localization benchmark, placing 10th of 15 on marketing copy — 22.2 points behind the winning AI workflow. The study, a joint research project by localization provider EC Innovations and Jademond Digital, evaluated 774 localized outputs across six content types, seven workflow models and three task types (translation, transcreation, creation). It was published June 5, 2026, after outputs were produced in December 2025 and January 2026 and blind evaluation ran until early March.
The systems tested, all at their December 2025 versions accessed via web interfaces, were GPT-5.2 (ChatGPT), Gemini 3.0, Doubao 1.6, Qwen 3, Kimi K2, DeepSeek-V3.2, and Google Translate. "PE" denotes post-editing by human professionals. Chinese native-speaking professional localizers — different people from those who produced the human and post-edited versions — scored each output on three equally weighted dimensions: accuracy and consistency, fluency and language quality, and style and cultural adaptation.
Where humans won, and where they lost
Humans finished first on informational content (76.9) and SEO content (74.1). On the other four content types they placed 7th, 7th, 9th and 10th out of 15:
- Technical: human 64.8, 7th; PE-Qwen 79.6
- Product UI: human 63.0, 7th; PE-Qwen 73.1
- UGC: human 64.8, 9th; PE-Doubao and PE-Qwen 75.9
- Marketing: human 53.7, 10th; PE-Qwen 75.9
PE-Qwen won three categories outright, tied a fourth, and placed second on the remaining two.
On SEO content specifically — meta descriptions, headlines, keyword-carrying body copy — the human lead over the best AI workflow was only 2.8 points: humans scored 74.1, with PE-Doubao and PE-Qwen tied at 71.3. PE-DeepSeek, PE-Gemini and PE-Google MT each scored 66.7. Every raw model scored lower, with raw Google MT last at 55.6.
The study's authors characterize the pattern precisely: human expertise dominates where factual precision and terminological consistency matter and creative latitude is near zero — informational and SEO content. EC Innovations' reading of the marketing result: "Human translators appear to over-correct the language, smoothing copy toward formal correctness and stripping out the contemporary register that marketing content depends on. The same instinct that makes a linguist excellent at terminology discipline makes them a poor fit for writing that needs to sound like the internet."
What the study does not show
The study measured localization quality only. It collected no ranking, traffic or conversion data. The authors explicitly caution against presenting the quality gaps as ranking outcomes, citing Portent's 2021 crawl of 756,297 ranking pages, which found no correlation between readability and Google ranking position. The authors also disclose that both firms behind the study sell services the research evaluates, and they note that criticism of circularity in content-quality research "lands on research funded by localization companies too."
Category averages conceal a 22-point spread
On SEO content, raw Chinese LLM output and raw Western LLM output both averaged 60.7 — a 13.4-point gap behind humans. But the averages describe models that don't exist. The four Chinese models — Qwen (65.7), DeepSeek (61.1), Doubao (59.3), Kimi (56.5) — span 9.2 points on SEO, the tightest category in the study. On technical content the same four span 22.3 points, from Qwen at 70.4 to Kimi at 48.1; on marketing, 22.2 points. The two Western models, ChatGPT (62.0) and Gemini (59.3), never sit more than 8.4 points apart across all six content types.
Kimi finished last among the Chinese four in all six content types; removing it shifts the "Chinese LLM" figure by 1.3 to 4.9 points depending on category.
Post-editing is not a uniform quality layer
Post-editing lift averaged 5.6 points, but the effect varied from +30.6 (post-editing Google Translate output on UGC, from 33.3 raw to 63.9) to -3.7 (PE-ChatGPT on marketing). Three workflows got worse with post-editing: PE-ChatGPT on marketing (-3.7), PE-Kimi on marketing (-2.8), PE-ChatGPT on technical (-0.9). On technical content, post-editing lift ranged from -0.9 to +9.3 depending on the model.
The authors conclude that model choice matters more than the editing layer: the best-versus-worst Chinese model spread within one content type reaches 22.3 points, versus a 5.6-point average post-editing lift.
One finding relevant to legacy infrastructure: post-edited Google MT scored 66.7 on SEO, ahead of every raw model tested.
A note on Baidu
For Chinese SEO specifically, the authors argue content quality is rarely the binding constraint. Baidu weights site-level signals — domain history, ICP filing status, hosting geography — more heavily than page-level language quality, and treats slow sites served from outside the mainland unfavorably. An ICP filing with mainland or Hong Kong hosting, they write, can do more for visibility than the difference between a 60.7 and a 74.1 translation score.
Staleness as a finding
Every model was tested at its December 2025 version, and Qwen, Doubao, GPT and Gemini have all shipped since. The authors call the durable finding not "use Qwen" but that per-model differences on specific content types are large enough to be worth measuring yourself. Their recommended follow-up: tag localized pages by workflow and watch analytics for two quarters, and run a two-week internal bake-off on your own content — then run it again in six months.
via portent.com (Original)