What makes a translation good?
Ask a linguist, a project manager, a client and a marketing manager, and you may get very different answers. For one person, terminology will be the priority. For another, it will be whether the text sounds natural. Someone working on a legal document may care mostly about accuracy, while a marketing team may be much more concerned with tone and audience.
So when we talk about translation quality, what are we actually measuring?
Why frameworks like MQM help
This is one of the reasons frameworks such as Multidimensional Quality Metrics (MQM) are useful. MQM provides a common structure for evaluating translation issues. Accuracy, terminology, fluency, style, locale, audience appropriateness and markup are treated as different types of quality considerations, rather than being collapsed into a single impression of whether a translation “sounds good.”
That distinction matters. A terminology error and a stylistic issue are not necessarily equivalent. Neither are a wrong number in a technical document and an awkward sentence in a piece of internal communication.
MQM gives us a way to record those differences, classify errors and assign severity to them. It doesn’t tell us what a perfect translation looks like. It gives instead a more consistent way of discussing what is wrong when something isn’t working.
How we apply this in re[words]
This is also the principle behind the quality checks we use in re[words]. When we process an XLIFF file, the output is assessed across six dimensions: Accuracy, Fluency, Style, Terminology, Locale and Tags.
The system can then flag issues, make corrections where appropriate, and produce a report showing what was found in the file. For example, a glossary can be provided as part of the project context, so terminology isn’t left entirely to the model. The same quality dimensions can be applied to files coming from different CAT environments, and they can be used both when post-editing an existing translation and when generating a translation from scratch.
Why the report matters as much as the output
If we’re going to use AI to process large amounts of multilingual content, simply producing an output isn’t enough. We also need a way of looking at that output and deciding what deserves attention.
That’s where structured quality assessment becomes useful, as it can help separate the things that actually affect the quality of a deliverable from the things that are merely different from what someone would have written themselves. This distinction becomes increasingly important as AI-generated and AI-assisted translation become more common.
There will always be things that a metric cannot decide for us. Whether a particular tone is right for a brand, whether a sentence works for a specific audience, or whether a small deviation from the source is actually an improvement. These still require human judgment.
The goal isn’t to remove that judgment. It’s to make better use of it.
Want to see what this looks like in practice? Book a demo!
Fill out the form