TransEvalnia Makes Translation Scores Explain Themselves

Sakana AI's evaluation framework asks an LLM to critique, compare, and rank translations before it ever emits a final verdict.

9 min read • View on GitHub • More from SakanaAI

A judge's bench made of stacked translation cards faces a set of scales, while a magnifying glass inspects a sentence fragment. The scene explains TransEvalnia's premise: the model must inspect, explain, and then rank translations instead of emitting a bare score.
TransEvalnia turns translation evaluation into a staged argument, not a single number.
Key Takeaways

Most translation metrics hand you a number and hide the argument. TransEvalnia does the opposite. It asks an LLM to critique each translation, line up those critiques, and then rank the outputs with a verdict that can be inspected by a human.

Why scorecards stopped being enough

Sakana AI built this repo for the point where translation quality gets too good for coarse metrics to be useful. BLEU can still tell you whether two strings overlap, but it cannot explain why a sentence sounds awkward, why a term is wrong, or why one version reads cleaner for a specific audience. TransEvalnia answers those questions with a prompting pipeline shaped by MQM, the Multidimensional Quality Metrics framework used in human evaluation.

Researchers at Sakana.ai have developed TransEvalnia, a translation evaluation and ranking system that uses prompting-based reasoning to assess translation quality.

Article Text, Content of MarkTechPost article · MarkTechPost on TransEvalnia

The middle layer is the product

The repo's core idea is not "ask an LLM which translation wins." It is more disciplined than that. First, evaluate_translations.py critiques each candidate on selected MQM dimensions. Then interleave_evaluations.py weaves those critiques together so the ranker can compare them theme by theme, and rank_translations.py turns the result into a final verdict.

That extra structure matters because it turns a judge model from a gut-feel grader into something closer to an analyst. The pipeline also reflects a practical concern that shows up all through the codebase: model calls fail, formats drift, and positional bias creeps in, so the system keeps adding guardrails instead of assuming the model will behave.

Two columns of translation critiques are fed through a loom and merged into one braided strip. The image explains interleaving, where separate evaluations are woven together so a ranker can compare them theme by theme.
Interleaving is the quiet trick that makes the final ranking easier to trust.

How the pipeline stays honest

The pipeline is built to reduce bias, preserve reasoning, and still end in a machine-readable decision.

The anti-bias story is easy to miss, but it is central. TransEvalnia uses permutations to rotate candidate order so the ranker does not simply favor translation one or translation last. After that, the repo still insists on a deterministic extraction step, which means even a chatty model has to collapse into a clean result before the system accepts it.

That final step is a tell. The codebase trusts LLMs for reasoning, but not for bookkeeping. rank_translations.py pulls the verdict back into a form software can use, which is exactly the kind of restraint you want in an evaluation tool.

What it beats, and what it trades away

SystemPrimary outputWhy people use itWhat it misses
BLEUN-gram overlap scoreFast baseline and easy to computeNo explanation, weak when outputs are already strong
COMET, BLEURT, MetricXLearned scalar scoreBetter correlation with human judgmentsStill just a number, not a critique
General LLM judgeScore or free-form verdictFlexible and promptableCan be inconsistent without scaffolding
TransEvalniaMQM critique, ranking, and Likert scoreTransparent, human-aligned evaluationHeavier pipeline and more moving parts

TransEvalnia has demonstrated competitive performance against leading models like MT-Ranker across various language pairs and tasks, including English-Japanese and Chinese-English.

Article Text, Content of itinai.com article · itinai on TransEvalnia

That comparison is the real message. TransEvalnia is not trying to kill modern metrics by being simpler. It is trying to outgrow them by making evaluation legible, which is what you want when translation quality is close enough that a single score no longer settles the argument.

The repo is more practical than it first looks

Under the hood, the stack is unusually flexible for a research repo. llm.py abstracts over AWS Bedrock, OpenAI, and local vLLM backends, so the same evaluation flow can run in cloud or on open models like Qwen and Llama without rewriting the pipeline. The prompt files are templated with Jinja2, which keeps the instructions and the language pair context separate from the orchestration code.

That makes the project feel less like a demo and more like a lab instrument. There is room for improvement, especially around test coverage and the global-flag style configuration, but the engineering bias is clear: keep the evaluator usable, keep the outputs parseable, and keep the model's reasoning visible.