RWS benchmark finds frontier AI models fail grammar in key markets

M-GATE tests 70 models across 30 languages, exposing blind spots that cost enterprises in multilingual AI deployments.

An open laptop displaying a geometric blue and grey pattern, three chrome abstract sculptures, a white mug, and a potted succulent rest on a light grey desk in a brightly lit office, with architectural models in the background.

RWS, the AIM-listed global AI solutions company, has published M-GATE, a linguist-designed benchmark that evaluates over 70 frontier AI models on grammar proficiency, translation accuracy and tokeniser efficiency across 30 languages. Developed by RWS's TrainAI data services unit, the tool is positioned as an independent signal for enterprises choosing which models to deploy across global markets, a decision that, the data suggests, most organisations are currently making on the strength of little more than a language-support list.

The headline finding is uncomfortable for several frontier labs: on a binary grammar test where random guessing yields roughly 50%, multiple models, including some considered top-tier, score below that threshold in certain languages. Meta's Muse Spark recorded just 23% on Fijian grammar, while OpenAI's GPT-5.5, which leads the benchmark on roundtrip translation overall, fell below random chance on the same language. Google's Gemini 3.1 Pro Preview tops the overall grammar leaderboard, but the broader pattern, according to RWS, is one of inconsistency: strength in one language rarely predicts strength in another.

The hidden cost of tokeniser inefficiency

Beyond accuracy, M-GATE surfaces a less-discussed enterprise risk: tokeniser cost variance. The benchmark found that the heaviest tokenisers consume more than ten times as many tokens to process some languages, Khmer is cited as an example, compared with English. That multiplier translates directly into inference costs at scale. Anthropic's Claude shows relatively balanced tokeniser usage overall, but records the fewest characters per token of any major lab in the benchmark, which drives per-query cost upward in absolute terms.

Speed dispersion is equally stark. The slowest models in the benchmark average roughly one hundred times the latency of the fastest, with worst-case response times measured in minutes. For enterprises running customer-facing multilingual pipelines, those tail latencies are an operational risk, not an edge case.

"Enterprise teams are being asked to choose between models on the strength of a language support list, '30 languages,' '50 languages', with no way to check whether any of the claims hold up," said Vasagi Kothandapani, CEO of TrainAI by RWS. "That's an expensive way to make a decision."

Cross-sector implications for AI procurement

The launch sits at a convergence point between AI capability marketing and enterprise procurement governance, a tension that is sharpening as agentic AI systems are delegated customer-facing and operational roles across retail, financial services, healthcare administration and media localisation. A model that performs well on English-language reasoning benchmarks but falls below chance on Swahili or Basque grammar is not a general-purpose multilingual system; it is a product with a significant deployment risk embedded in its specification.

This matters beyond the technology sector. For multinational retailers expanding AI-generated content into Southeast Asian or African markets, for global financial institutions deploying AI-summarised regulatory communications across multilingual jurisdictions, and for media groups localising content at scale, the cost of a misjudged model selection is not merely technical, it is reputational and, in regulated contexts, potentially legal.

RWS's own Content Unlocked 2026 study of 200 senior enterprise content leaders found that while 86% reported AI had accelerated content creation, 65% said it had simultaneously slowed localisation through additional rework. That finding points to a broader dynamic in the enterprise AI market: the productivity gains most often cited in frontier-model marketing are frequently offset by downstream quality remediation, the cost of which is rarely captured in benchmark comparisons.

The competitive context is also relevant for capital allocators tracking the AI infrastructure stack. Benchmark fragmentation, with most public leaderboards testing reasoning and mathematics rather than linguistic command, creates information asymmetry that advantages incumbents with strong English-centric reputations. Independent multilingual benchmarks like M-GATE represent an emerging layer of procurement infrastructure that could shift model selection away from brand recognition towards verifiable per-language performance data. That shift has implications for the relative positioning of US hyperscaler models against newer entrants from Alibaba, DeepSeek, Mistral and others that have built multilingual capability as a first-order design goal.