Can two simple words cost 10 million dollars?
In 2012, HSBC proved that they could. The mistranslation of the payoff “Assume Nothing” triggered a domino effect across several international markets, leading to a costly emergency rebranding campaign. A major reputational setback for a bank whose identity was built on its ability to communicate effectively on a global scale. The case is cited by Al-Tarawneh and Al-Badawi in [a 2025 study] on how translation quality affects brand image and consumer perception.
This case highlights a fundamental truth: a linguistic mistake, even one that may appear minor, can undermine customer trust, affect brand perception and, in the worst-case scenarios, generate significant financial and reputational costs.
Today, this vulnerability is facing the speed of generative AI
The ability of LLMs to produce and translate huge volumes of content across dozens of languages and domains simultaneously has made workflows significantly faster and more scalable. However, a translation can sound fluent and convincing even when it contains omissions, terminology inconsistencies or hallucinations, meaning information that is not present in the source text. In addition, quality is not consistent: model performance varies depending on the language pair, domain and type of content.
As we already discussed in our article on the use of AI in localization, the question is no longer whether these tools should be integrated into translation workflows, but how to do so without compromising quality. And with content volumes continuing to grow, the real question becomes: how can translation quality be assessed in a systematic and controlled way?
Quality evaluation is not new, but traditional methods are no longer enough
Many traditional automatic metrics summarize quality through a single overall score. This is useful information, but it does not explain which errors are present, where they occur or what impact they have on the text. Errors with different levels of severity and different characteristics can, in fact, result in the same score, making it difficult to understand where action is needed.
Human evaluation, while still the benchmark for measuring quality, has two main limitations. On the one hand, it can be subject to variability: results may change depending on the linguist’s experience, the context and the framework used to classify errors. As a result, two professionals may assign different categories or severity levels to the same issue. On the other hand, a fully manual evaluation process is difficult to scale when facing the production volumes driven by generative AI.
Given these limitations, the need for an approach that makes quality evaluation more systematic, reproducible and scalable—while maintaining human oversight where needed—is clear.
Towards an Automated Linguistic Quality Assessment framework
This need was the starting point for our work on Automated LQA. Since 2024, we have been exploring how LLMs can be used to evaluate quality without giving up human expertise.
An Automated LQA system makes it possible to combine the speed of automation with structured linguistic criteria. The goal is to classify and describe errors systematically, generating results that can be compared over time and used to identify where improvements are needed.
A key role is played by MQM (Multidimensional Quality Metrics), a reference framework that enables errors to be classified by category—such as terminology, style and accuracy—and assigns each issue a severity level. This way, evaluation does not simply provide an overall score, but clearly highlights which aspects of a translation have affected the final outcome.
We covered this in more detail in this article.
Tailored evaluation and levels of automation
We have also considered the need to make evaluation flexible and adaptable to each brand’s specific requirements, including tone-of-voice guidelines, language style and proprietary terminology rules. This makes it possible to apply consistent criteria across very different needs.
Finally, we have explored different levels of automation and human involvement, from automated checks for high-volume workflows to hybrid processes that add human oversight. Quality assurance should not be a rigid choice between “fully manual” and “fully automated”: the right balance between scalability and control can vary depending on the content, volume and level of risk.
Measuring quality to improve the process
A structured evaluation process does more than identify errors: it helps detect recurring issues, understand their root causes and monitor quality over time. This deeper understanding becomes a valuable tool for improving translation processes and making data-driven decisions.
In a context where multilingual content is increasing in both volume and complexity, quality evaluation is therefore not just about checking the final output. It means building a reliable process where automation and linguistic expertise work together to reduce risks, maintain control over quality and support continuous improvement.
Sources
Al-Tarawneh, A. and Al-Badawi, M. (2025), “The Impact of Translation Quality on Brand Image and Consumer Perception”, in A. M. A. Musleh Al-Sartawi, M. Al-Okaily, A. A. Al-Qudah and F. Shihadeh (eds), From Machine Learning to Artificial Intelligence, Springer Nature Switzerland, pp. 1203–1212.
Vuoi contattarci? Compila il modulo.