TranslateGemma Quality Evaluation / Stress Test feat Alex Murauski

Episode 144 • March 03, 2026 • 01:00:52
TranslateGemma Quality Evaluation / Stress Test feat Alex Murauski
Nimdzi LIVE!
TranslateGemma Quality Evaluation / Stress Test feat Alex Murauski

Mar 03 2026 | 01:00:52

/

Show Notes

In this session, we will explore how we evaluated the translation quality of Google’s Gemma model using the MQM framework and a human-in-the-loop review process.

The case study walks through how LLM-generated translations were assessed using structured error typology, how linguistic quality was benchmarked, and how AI-enhanced workflows can combine automated generation with professional post-editing and evaluation.

We’ll discuss:

How MQM works in real-world AI evaluation

What kinds of errors LLMs produce across languages

Where AI performs well — and where it still struggles

How to design scalable human-in-the-loop evaluation workflows

What this means for localization vendors and enterprise buyers

The session is based on a real case study conducted by Alconost’s MT evaluation team using our MQM evaluation tool.

Full case:
https://alconost.mt/mqm-tool/case-studies/translategemma/

Other Episodes

Episode 3

April 18, 2021 • 00:50:53
Episode Cover

Scale Operations with OKR Methodology (feat. Daniela D'amato)

More info on this eLearning course and more here: https://www.nimdzi.com/courses        

Listen

Episode 126

November 15, 2024 • 00:56:35
Episode Cover

L10n in the Age of AI: Building Web 3.0 Products feat. Mohamed 'Mo' Ly

In today’s dizzyingly fast digital landscape, where AI-powered solutions and LLM-driven products are the norm, success hinges on mastering how to operationalize, productize, and...

Listen

Episode 50

August 05, 2022 • 01:00:06
Episode Cover

LSPs as growth enablers – A new paradigm? (feat. Olivier Marcheteau)

It’s a known fact that Translation & Localization go well beyond language and that the industry has long embraced a combined and ever-growing human+technology...

Listen