Nimdzi LIVE! · 2026-03-03 · 1h 1m
Key moments - from our scoring
Substance score
50 / 100
Five dimensions, 20 points each
This episode examines the practical realities of evaluating large language model translation quality through actual data rather than hype. Murauski walked through Alkonost's evaluation of Google's Translate Gemma using their MQM-based annotation tool, where 45 linguists spent 34 hours analyzing translations of scientific content across 16 languages. The key insight is that translation quality assessment isn't about setting an arbitrary threshold - it's about understanding what your threshold means in context. When they set a 99% pass threshold, all languages 'failed,' but lowering it to 95% (industry standard) showed many would pass. The study also exposed the fundamental weakness of single-annotator review: three independent linguists disagreed significantly on minor errors, style preferences, and terminology choices, with agreement varying by language structure. Notably, annotators spent anywhere from 30 minutes to two hours on identical segments, suggesting that unstructured review is inefficient. For smaller and mid-sized LSPs, the takeaway is that deploying and fine-tuning models like Gemma is now accessible with basic engineering skills and AI coding tools, not just enterprise-scale operations.
Benchmarking on production projects shows Translate Gemma performs comparably to Gemini and is one of the best translation models available, though performance varies significantly by language and content domain.
Yes, through prompt injection techniques officially documented in Google's documentation; Murauski successfully prompted the model to translate into unsupported languages like Hmong and Belarusian by instructing it to use special input formatting.
A 99% threshold is arbitrarily strict and meaningless without context; lowering it to industry-standard 95% changes which languages 'pass,' but the real issue is understanding how error types are weighted and what failure threshold aligns with business needs.
Linguists agreed on major errors (grammar, mistranslations) but disagreed on minor/stylistic issues, word order, and terminology - differences driven by language structure, training, expertise, and personal preference; Japanese showed higher agreement due to standardized syntax, while Polish showed more variance.
Without time limits, annotators ranged from 30 minutes to 2 hours on identical content; implementing a fixed time budget (e.g., 20 minutes) dramatically improves cost predictability while maintaining quality data.
Our reviewer’s read on each dimension, with quotes from the episode.
The episode contains genuine practitioner-level findings - annotator bias toward AI output, asymmetric automated metric reliability, and annotation time variance - but these are diluted heavily by host summaries, off-topic tangents about vibe coding, and repetitive conversational padding. The useful signal-to-noise ratio is mediocre for a 61-minute runtime.
I think the linguists, they just hate models. They somehow understood that that was a AI translation. So they mark every year because of. They just uh, don't like AI translations. I think they were biased. I think they were purely biased.
Some of them spent like half an hour for the, uh, whole annotation and some of them spent like around two hours. Uh, and it's a mystery for me why.
There is one genuinely novel idea - a future positive-reinforcement dataset that teaches models to improve translations rather than just flag errors - and the asymmetric reliability of automated metrics is an underappreciated point. The rest of the episode recycles standard localization-industry talking points: AI won't fully replace translators yet, agentic future, human in the loop.
I foresee and have an idea of uh, another kind of metric or the way uh, the language model is taught to translate... maybe there is a future world when we encourage machine to translate better and just give them positive feedback so that they not spot the euros but make the translation better.
going to production uh, without people is not yet, not yet there. So we are very excited sometimes about the quality of translation that AIs provide to us.
Alexander Murauski is a 20-year LSP founder with genuine engineering depth who personally ran the study under discussion, giving him direct practitioner authority. He falls short of the top tier - some answers are uncertain or speculative, and the study itself is explicitly framed as exploratory rather than definitive - but he is not a recycled thought-leader.
We are just running Lots of automated quality, uh benchmarking of uh different MTs between different language pairs and different domains. Every project, every production project starts with um every client just comes uh and asks well I need to translate everything by AI.
Our uh, benchmarking on production projects shows that it's, it's really good. So I would uh, place it uh, just uh, somewhere near Gemini.
The episode is anchored by real numbers from an actual study: 45 linguists, 34 hours, 6 days, 16 language pairs (10 supported + 6 unsupported), a Japanese inter-annotator kappa of 0.4, annotation times ranging from 15 minutes (Portuguese) to 90 minutes (Hmong), and named tools (Metric X, COMET Kiwi, Vertex AI, Hugging Face). The study is small and exploratory, and some claims - like the Vertex AI anomaly - remain unexplained, which limits the score.
45 different linguists spend a total of 34 different hours over six days to actually analyze the output
Japanese, uh the 0.4 is quite high agreement, it's moderate. So they are more or less agreeable.
The host occasionally demonstrates good instincts - probing the 99% pass threshold, asking about linguist resistance to AI - but too often fills airtime reading from the report aloud, making extended analogies (plumber, calculator), and letting speculative claims about agentic localization go entirely unchallenged. The conversation is collegial rather than disciplined.
I just want to read this inter annotator agreement, why multiple annotators matter and what low agreement actually tells us. With three evaluators per language working independently, we measured how often they agree and the answer is not very often.
Are linguists generally receptive to working on projects like this that are essentially like hybrids using technology but still wanting to enforce human quality evaluation?
Computed from the transcript - who did the talking, and the words that came up most.
In this session, we will explore how we evaluated the translation quality of Google’s Gemma model using the MQM framework and a human-in-the-loop review process. The case study walks through how LLM-generated translations were assessed using structured error typology, how linguistic quality was benchmarked, and how AI-enhanced workflows can combine automated generation with professional post-editing and evaluation. We’ll discuss: How MQM works in real-world AI evaluation What kinds of errors LLMs produce across languages Where AI performs well - and where it still struggles How to design scalable human-in-the-loop evaluation workflows What this means for localization vendors and enterprise buyers The session is based on a real case study conducted by Alconost’s MT evaluation team using our MQM evaluation tool. Full case:
Transcribed and scored by The B2B Podcast Index.
Speaker A: Hello, my name is Tucker Johnson and I am your host. Today as we experience NIMSY Live, uh, where we talk about the latest and greatest in translation, localization, internationalization, culturalization and all that fun stuff. Our global audience of international professionals needs to delight their customers on this program. We invite guests who like to have fun and have some value to add for our audience of globalization professionals. I'm always eager to provide a platform with to those with a good story or a good data set and today's guest has both. So if you have any ideas of topics you would like covered on future episodes of NIMZY Live or guests that we should put on this platform who have a good story to tell, let us know. You can send an email to livemsy.com or just reach out to me directly. Well, without further ado, I'm going to get into it today. Of course, make sure that you are subscribed subscribed to NIMSY Insights. Whether you are following us on LinkedIn, YouTube, Facebook, X any of those platforms are coming to you live on all. Of course, all of these episodes are archived on our YouTube channel where you can check those out at your leisure. We've got hours and hours of great content. If you're new to the industry or just want to stay up to date with all of the latest trends, it's a great place to start. So today we are going to be taking a practical look at large language model translation quality. Specifically, we're diving into the real world evaluation of Google's Gemma. I don't know. Our guest is going to tell me how to pronounce it Model using the MQM framework and human in the loop review process. What happens if you applied structured error typology to LLM output? Where does it perform well? Where does it fall short? That's what we're going to be talking about today. In today's episode we're moving beyond that hype and we're looking at the actual data today. Walking us through it, I'm joined by Alexander Murowski from from Alkonost. Uh, welcome Alexander. You were involved in this Translate Gemma quality evaluation using Alkonost MQM tool, giving you firsthand view of the LLM generated translations and whether they hold up under structured review. Alexander, welcome back to the show. It's good to have you, sir.
Speaker B: Uh, yeah, thank you Chakra for having me. A long time oc many years since last time almost exactly.
Speaker A: And well, I've been looking forward to this conversation because we're uh, I want to Kind of cut through the nitty gritty. Today. There's a lot of talk about large language models, what they do well, what they don't do well. And I feel that we're reaching that point where it's not just talk anymore. We're actually going to start seeing data, which is what you're going to talk about us with us today. Um, before we get started, I just want to set the scene a little bit, tell us a little bit about yourself, your company, and the study that we're talking about today.
Speaker B: So I run a company, been running it for almost 20 years. More than 20 years. This is a localization services provider. Uh, what we do is just translation, uh, localization of everything that goes global, like games, websites, e commerce, uh, educational things. Uh, mostly, mostly some technical stuff like we localize software. That's not translation, but localization mostly. And we are developing a lot in, inside the company because in the DNA, DNA of our company's software development. I'm a former software development, uh, engineer, uh, in the past and in the future in this, in this age of wipe coding, I write the code every day. So about this, that's my, that's my AIR era.
Speaker A: So you know, you, you've been coding since before vibe coding was even a term.
Speaker B: Absolutely. I remember times I coded using uh, assembler. That's so much time ago, so long ago. Uh, and through all 20 years we've been developing something at uh, Alchemist and uh, this engineering thing, uh, LLMs AIs is very, very our topic. I would say that we're happy that uh, from the business side of you, LLMs are disrupting industry very much. Now I must admit that, uh, it really, uh, just diverts some clients from um, traditional language services provider providers like ours. But I'm happy that uh, there are lots of opportunities, new opportunities coming, uh, through these agentic, um, I also already call it agentic localization. The new term like it comes after wipe coding into localization industry. Uh, soon we'll be talking about agentic localization. Everybody, every uh, who has this knowledge, some technical knowledge, is building something around localization, inside localization, outside localization. So we will be very soon see how agents, uh, will order translations. I think not people will be ordering, but some, um, Vibe coded code going to order some localization from people or from other agents. So this is coming, and that's just an interesting topic, but maybe different. Uh, now what I was going, um, uh, to show you, we've conducted an experiment around LLMs and automation and people. So AI and people and we just ask them to work together to see what happens. This the case is uh, about we're going to present today.
Speaker A: Well and this is kind of the question. This is useful for anybody out there to the data, the data we're going to review today. It's useful to anybody out there but I find it's particularly useful to smaller and medium sized LSPs. Um, you mentioned something earlier that I cut. I want to drill down a little bit. You said you know all of the advancements in AI and large language models, it's certainly disrupting the business side of operations and I don't think anyone's going to argue with that. I don't think that's a controversial statement. But what that leaves out is what about the language side? Because you can kind of bifurcate this into two different conversations. You can talk about like okay what is MT doing for the business side? Crazy stuff. I don't want to get into that right now. But what, what is it doing for the language side and how do we measure language but in terms of quality and you've, you've chosen in the study. I'm, I'm just going to pull it up here. Chosen in the studies, uh, take a look at it with human evaluation using MQM framework. Um, and that to me is super interesting. Going beyond the business and seeing. All right, well what can, what about equality essentially. Um, so you, you, you evaluate. Walk me through the study. You evaluated Google's Translate. First of all, is it Gemma or Gemma?
Speaker B: Well that's a good question. I never thought about it. I've never heard. I, I just read uh, it as Gemma but maybe it's Gamma like it like I don't know.
Speaker A: Here we are talking about the importance of quality and language and we don't even know how to pronounce the word. We're talking about real good examp. Geez. But anyways this is a model that can be fully you know downloaded and set up um, on a private server and you ran some different um, translation models through it or you ran some different content type through it. And for the content type you chose ah, scientific for something from a more or less scientific study, correct?
Speaker B: Uh, correct. Uh, and also in the beginning I would just make uh a certain disclaimer about this uh case study. It's called like uh, as we uh, assess the quality of uh translate Jammer. But I would not call this uh, uh it in this way because um, what we uh, this is this case study about the journey of making Such kind of experiment of this uh, kind of case study. If we were going to assess the quality of the LLM, uh the quality of the translation itself, that would be some, another approach because we would take longer text, different domains, different, different uh kind of kinds of text uh, and different languages perhaps. Uh, what we've done here is a kind of quality um assessment. In essence it's. Yes, but I cannot say that we assess the quality of translation. It's more like um, um like a journey, uh, like an example of uh, how it can be done, how we can run an LLM state of the art translation model on premises, how you can um organize uh professional linguist linguists to assess the quality using the uh, well known MKM M uh framework. So it's more like a story. Uh, I would not call it the quality assessment because it's two stories.
Speaker A: Uh well the challenge with, I mean the story is a good one and here's the story that I took away after reviewing everything is that deploying, testing and fine tuning LLMs is not something that smaller LSPs need to be afraid of. And by smaller I mean you're not Lionbridge, you're not translations.com or yes, transperfect or we localize. Right. And I, I believe that there's this. I believe the perceived barrier to entry is a lot higher than the actual barrier to entry and I love it to see that this technology and the benefits thereof is getting more accessible over time um, even as it's getting more complicated.
Speaker B: Yes, absolutely, absolutely. This uh case study shows the way. So uh, if you have some engineering background at least you know how to code a little bit. Uh with the modern agentic tools you definitely can set up a server, you can definitely run uh the model, you can select the hardware you want to run it and uh, with the agentic coding tools you can uh actually make it work. It won't be uh very optimal way, very so but you will get results. Uh this goes with the translation models uh uh and uh also general purpose models and also with metrics. If you want to run some bigger X commit uh Metric Automate or Google Metric X There is a way uh that you do it uh don't be afraid of doing things uh because you have a helping hand uh of agent uh tools nowadays and that resolves uh many things. So it's not impossible for just one engineer just to uh set up the model and translate through it. So any localization manager with a technical background could run Translate Gemma and and uh, just to translate everything uh with It. This is the. One of the best model models for translations in the world. Our uh, benchmarking on production projects shows that it's, it's really good. So I would uh, place it uh, just uh, somewhere near Gemini.
Speaker A: Well, well, I'm glad, I'm glad I have you here in person because I read through everything. It's just a Reader's Digest version for everybody. Um, and for our podcast listeners, I'm pulling the report up on screen. If you want a copy of the report, I put it into the chat of the LinkedIn Live session so you can follow along at home if you're watching us here or if you're listening to the recording. Uh, but you chose seven different segments from an academic paper and the academic paper was about multilingual speech recognition. Um, just to be meta. Uh, you translated it in the uh, using the, the LLM Google gemma, and then you exported it and analyzed it with human reviewers, right? Yes, under the MQM framework. So 45 different linguists spend a total of 34 different hours over six days to actually analyze the output. And to me this is always the best form of review of machine translation is what the actual humans say. Even better if it's a specific content type and those humans are specialists in those in that specific content. Content type. Um, but you provide a lot of data here on median annotation time per linguist, meaning how long did it take them to find the errors, correct it and everything, starting with Hmong, unsurprisingly going down to Portuguese at the bottom. What I wanted to take a look at though is I'm just going to scroll down here if you allow me. Please, please walk me through it.
Speaker B: A little story, uh, um, before um, of um. So how we, how did we choose uh um. Choose the languages so translate Gemma, Uh, just um. M. It supports uh, 54 languages, uh, if my memory says me right. So officially and uh, uh, what we selected just for fun, uh, we took 10 languages that are officially supported and uh, other six languages that are not officially supported. And if you um, translate with gemma, it expects the input in a special format. And if you ask uh, uh, for languages, for unsupported languages, then um, it gives you errors. But uh, as a prompt engineer, you can do a little prompt, uh, injection, um, to ask it to translate in the non supported languages and it's officially documented in their documentation. Uh, uh, so you can trick the
Speaker A: machine into translating languages that it tells you.
Speaker B: Okay, uh, uh, yes, it's, it's an LLM and you can trick it you can trick it to any language. Uh, so we tricked it to, to translate into six languages that are not officially supported. Uh, and including Hmong, including Belarusian and including some Arabic dialects.
Speaker A: You see these in the brown bars for those of you that are watching the screen. So Hmong, Belarusian, Arabic.
Speaker B: Yeah, yeah. So we, we got uh, we've experimented with um, with different dialects of Arabic. We, we are running a project currently in production around Arabic dialects. And so we were curious about uh, how it would be uh, translating differently into different Arabic dialects. So uh, we put that on test and um. M. So, and the trick, uh, the trick worked and uh, we asked Gemma to translate and it translated it its best and we got the translation. And regarding how do we choose the text? It was absolutely, I don't remember. It was some random thing. My assistant just uh, came up with this difficult one. If we were assessing the quality for true, um, we would be selecting different. We will be selecting like news, public domain, social, something very like uh, lightweight content. Because uh, for example wmt, uh, um, data set, they have only this kind of content. Uh, social, um, some general text, uh, uh, and usually the translation quality is assessed uh, using just general domain text. Um, not this kind of scientific paper. If I could return, if I return in time back, I would select different text. Because this text we selected, it's a really stress test for any model even for people. If you start to translate it into your native language, uh, I don't find the terms in my vocabulary that I use there. It's a scientific paper. So no, no Ukrainian, Russian will be uh, will be having those terms. So it's, it's tricky and that's something
Speaker A: that we sometimes forget when uh, we're evaluating the quality of machine translations. Like how good is the mt? Well, how good are humans? Like if it's a highly specialized source content, it's unreal, it's unrealistic to expect the quality to be better than humans, um, in a lot of cases because the training material is not there. And but that explains because you probably read my mind where my questions were going because not a lot of languages passed this test. Like if we. Yes, this is kind of an exploratory test. It's a stress test to push the limits and everything. But if this were an actual test with pass fail, not a lot of languages quote unquote passed it. Um, talk to us about that a little bit.
Speaker B: Um, uh, well, regarding the tests, uh the test pass if yours, if you have strict threshold of 99% uh then no model Passes. Because uh, even with the minor years, like in German there were a few minor years, but um, the text was very short and even minor years on this volume is unacceptable. Uh, but we're talking not about mine heroes here. There are lots of major ones M. Mistranslations. Uh, uh, mostly mistranslations. And in some languages they were like not using just different terms but some grammar things, some. Because the sentences uh, in this text are quite long. They are hard to translate for even so human. And here uh, is the model it make. It made lots of mistakes and we just wanted uh, to see how many. Uh. And there was a kind of. When I got those results, I saw um, I thought that it's kind of a disaster
Speaker A: until I read it further and I was like, oh, you set the pass bar at 99% so you set a really high pass threshold on this. So to say that all of the language just quote unquote failed would be. I mean, sure, but if you lowered that fail threshold or that pass threshold just a little bit. Um, I've typically seen pass thresholds at 95%, not 99% or anywhere in between.
Speaker B: Many of this would pass. Uh, I, I was misled a little bit. Uh, um, with our production, uh pass rate when we assess um, the quality for production, uh, things and um, when the human involved. When we assess human. So this 99 is. You can change it in, in the system. You just make it 95 and everything is fast.
Speaker A: So this is what I try to explain so many times I'm working with uh, typically client side localization, uh folks and who want to set up a quality framework or a quality program. And as part of that you have to decide what is pass, what is fail. And a common question that I get is well, what does. What do you. What do other client side organizations have as their pass rate? My.
Speaker B: It.
Speaker A: Yeah. It. Doesn't that mean anything? I could tell you 95, but what does that mean? How, how are they weighting their error types? Right? Because if you heavily weighting air types, then a lower or a lower uh, pass fail rate is going to be the same as a higher pace fast fail rate if you're not weighting those error types as heavily. So it's really kind of when you say 99, um, uh, threshold it's like okay, what does that really mean? But I can tell by the, for the sake of your study, you set the bar really high. Right? Like this is really high.
Speaker B: Absolutely. And also uh, in the real world, uh, when somebody translates and uh, there is A reviewer reviews and says well that's the. Here is a minus style, uh, some style, uh, or preferential year. There is a rebuttal stage when the linguist, they argue and then they come out with a, with some compromise with the solution. With some solution. Uh, and uh, we didn't give that chance to translate Gamma. So we just reported everything. Every minor, every preferential, every, every single era. Uh, uh, they just spotted. I um, think the linguists, they just hate models. They somehow understood that that was a AI translation. So they mark every year because of. They just uh, don't like AI translations. I think they were biased. I think they were purely biased. They made.
Speaker A: Which is good. You're stress testing it. You want, you want it to be, you want the deck to be stacked against the machine.
Speaker B: So when I was talking what I was saying about if we just the quality uh in the production environment so the uh, majority of those errors will be rejected by the author. Uh, like preferential or like uh. So, uh, it's um, many of the miners, the minor euros, they will be uh, they could be rejected if we gave uh, this chance to Gemma. But we didn't give it. So it's.
Speaker A: Well, and this comes back to. I want to talk about this idea of multiple reviewers and stuff because in my mind this is one of the strengths of traditional human LED review. And there's also one the weaknesses of traditional human read review is that you know, attention to detail and debating over should I use this term or this term, you know, these preferential changes that you're talking about that can be very time consuming. So I mean from a business perspective my brain says how can I eliminate that step? Right. From a linguist perspective I say no, that step is absolutely vital because that's, you know, that's where the magic happens. That's where we get those perfect translations that are going to surprise and delight our international customers. Right? In reality, the devil's probably the, the. The reality is probably somewhere in the middle. Right? Um, but you in this test use three evaluators per language. And I just want to read this inter annotator agreement, why multiple annotators matter and what low agreement actually tells us. With three evaluators per language working independently, we measured how often they agree and the answer is not very often. This isn't a flaw in our process. You know, like, like I was saying, like oh, that's an inefficiency. My, my project manager brain says oh, that's a flaw, it's an inefficiency. But your argument is it's a fundamental property of translation quality assessment. Different linguists notice different errors, weight them differently and bring different expertise. That's exactly why a single annotator is never enough for reliable MQM data.
Speaker B: Yeah, that was uh, one of our insights. Uh so we thought that uh two linguists is enough to assess. But we made three linguists and they don't agree with each other. Some of them agree, they agree on the number of the uh. On the quality of this of segment. Okay. They agree on major years. If it's a grammar. They. Everybody reports grammar but when it uh. When it's minor euros like style, some preferential things, uh word order uh which is correct but not be, might not be fluent. But for scientific uh. Uh article is fluency is not just that important like in marketing text for example. So uh. And they start to report many many minor years and they don't agree with each other. And I think in real life they would fight uh with each other uh to just uh. For their uh preferential variants. Uh anyway, uh, uh uh, it's interesting that um. Uh it depends on the language. I think it's somehow connected with uh the structure of the language. To take for example Japanese, uh the 0.4 is quite high agreement, it's moderate. So they are more or less agreeable. Uh um. And um, this is because the Japanese language is more like standard standardized regarding uh the, the. The uh. The word, the, the words order, the structures of the language. So if the language has structure they agree. But maybe Polish people, uh, they have another opinion. They agree, uh, they, they disagree more. Yeah and for uh Hmong language we only had one translator. So uh. If you scroll down let's, let's see what uh is uh. No, because yeah just one linguist there was uh. So our observation was that three linguists is not enough uh actually to assess the quality. So the multi dimensional thing first, uh, how they were trained first uh how they were pre selected. Uh who are they, who are the linguists. Um then uh. I would sample more like 5 to 10 linguists uh and to see how the core of them agree with, with each other. So the more data the better. Uh, and, and uh. One of the um objectives of our um this case study was to see how, how much money does it cost to uh make the year annotation. And uh. So if we knew that €1 costs $1, uh this gives us some, some, some information and we just could plan budgets and uh. What the tool uh we're using um has about it is it uh. Whenever A, uh, linguist marks an error, classifies it. The timestamp is written. So, uh, by, uh, just classifying €1,000, we've got 1,000 timestamps. So we could be able to measure, uh, the time they spent on the annotation. And that was quite another interesting thing. So some of them spent like half an hour for the, uh, whole annotation and some of them spent like around two hours. Uh, and it's a mystery for me why. So they're so different. Ah, so different. Uh, so whenever you conduct such a kind of production mkm, uh, annotation, you would be limiting that time. Uh, we gave freedom. Yeah, but next time, uh, we'll just, uh, ask. Uh, you have only 20 minutes to assess the quality. Mark as much as you. As you can, but you have 20 minutes. And this saves money. Given the freedom, you see 15 minutes to just one and a half hour. That's, uh, that's also quite interesting insight.
Speaker A: Yeah, well, I mean, from a business perspective, that makes a lot of sense. Like if this were a live project, I would absolutely limit the linguist time. I'd say you budgeted for one hour or I budgeted for 30 minutes. Please spend 30 minutes on this. Um, it's interesting though to see this data set where you did not limit their time because it shows how much time they would actually use if there were no constraints and. Yeah, absolutely. Interesting. So hmong linguist took 90 minutes, an hour and a half, um, going all the way down the list here, whereas the Portuguese linguist took an average of 15 minutes to review it. There's a big delta there. And I'd want to understand why on this. And also you brought it up that we're talking about, um, putting a monetary value to. Okay, how much does it actually cost to annotate this? I think the next step. Would that be okay, and how long will it take me to get a return on investment? How efficiently are those annotations being used? Um, and how effectively is the model being trained to improve, to create, I don't want to say better output, but output that it's more aligned to the preferences of the reviewers.
Speaker B: Yeah, looking at these numbers of times, I just. I can. We never asked translators why. Uh, we. But I assume that with common language it's not supported. There were so much m. So many errors. So it took time to. To mark them. Uh, the same with Belarusian and Ukrainian.
Speaker A: Yeah.
Speaker B: Japanese, uh, are very precise and they found every year. So it just, uh, because of, uh, the. This accuracy, this, um, this m. Kind of mentality. I Don't know, uh, what happens to Russian people. Uh, maybe, uh. I don't know why they, they, they. They um, needed more than an hour. Uh, but I see that the medium is, uh, is just, um, uh, 30 minutes. And I would, I would say that this is the target time. Well, and also it depends on the number of euros. So if, if we take German and French, there were not so many, not so many euros, uh, in. In the text. So they just, uh. It just took less time.
Speaker A: Well, and this is good data because as a project manager, remember I said I would always cap. Cap the amount. I would never send it to the linguist and say, use as much time as you want. I say I budgeted half an hour for this. Um, yes, let me know if you need more time. We can talk about it. But I'd always cap it. Right. Well, data like this allows me to kind of understand like. All right, what is a reasonable cap? Because that's a struggle I always had as a project manager. Okay, I'm gonna send this to linguist for review. I need to cap it at half an hour, but I don't know if that's going to be enough time for them. I don't know if it's going to be too much time. Am I losing money? Am I stressing out the linguist? Um, having some baseline data like this, kind of. It's a good sanity check to understand. Am I being unreasonable with my expectations towards a linguist? Because the linguist, as you mentioned, there's going to be a bias towards machine translation. They probably figured out it was machine translation, so did not go easy on that. Um, have you come across, you know, whether during the study or outside of the study, what is your experience working with linguists on projects like this that are essentially like hybrids using technology but still wanting to enforce human quality evaluation? Um, are linguists generally receptive to working on projects like this, or is it a struggle?
Speaker B: I think half a year ago, uh, there might be some struggle, uh, some resistance, uh, in linguists that were not willing to work with AI, uh, to proofread it, to. To just, uh, to edit. Uh, but now more or less, uh, this resistance is gone because everybody understands that this is not a new world coming. Uh, already came already come here, uh, with. With AI, And. And, uh, what we, what we have in production is when. Not the linguists, not the linguist editing the machine output. But it's called ape, um, m. Automated post editing.
Speaker A: Okay.
Speaker B: When the machine, the LLM, it just suggests the, uh, how to fix the translations performed by linguists and performed by other uh, other AIs. So uh, those uh, uh, those agents now using LLMs, they just leave comments, uh, say we're using a crowding platform, for example. But it, it's, it's possible in any, in any platform now when LLM, you just have your Gemini go and uh, read the translations and leave their honest opinion, the AI's honest opinion on the quality of the translation and suggest some corrections. Uh, and many times uh, it really works well because uh, uh, LLMs are just very good at following glossaries. Absolutely. They know the translation memory better than linguist. Uh, and they uh, spot very small like uh, some tags, some uh, wrong. Some very small things. Even this is this kind of spelling, especially spelling. Uh, they are excellent at sporting very subtle things that are not, might not be maybe overlooked by the linguist. And so now machines are editing the linguist's translation. So to this extent it just, it's just going to, towards this. So it's a mix, uh, agents and people. And we uh, see less and less resistance from uh, the translators, uh, towards AI. AI is just with us and we just work together helping each other. And AIs really help translators to sometimes uh, explain things, terminology to understand or get better wording. Because sometimes you are stuck with a phrase, you don't, you don't know how to render it and you're stuck. And uh, you know, just press uh, a button to generate some five or 10 variants of the translation and select from that. And that's very helpful.
Speaker A: Or use that as inspiration and write your own damn variation. Right. Like, and this is the thing, it's like when we're talking about using. Because um, let's take the conversation back a little bit because we're talking about, you know, how to LSPs leverage this at scale 64 languages, different content types. How can clients leverage this and make sure that they're not getting bad quality? Um, but let's take it back a little bit and talk about from the, the translator, the linguist perspective. And one thing that I've been talking, when I talk with translators is I always ask them are you using AI? And if not, why not? Right? Because this isn't something I, I think this idea from the translators, you said it was stronger six months ago. This idea that AI is bad, it's here to replace us. Um, that's AI that is being imposed upon us. But I've also talked to some translators that are using it themselves to do the translations. Now what I mean by that is not that they get a file from a client or an lsp and they use AI to translate it and send it back. But they have a window or a tab open next to their translation environment where they're saying oh, I'm not really sure how to translate this, can you give me some suggestion? And they're getting inspiration from it. Um, and I don't think that is anything to be ashamed of. I don't think that like perverts the craft of translation any more than using a dictionary would 20 years ago, you know, translate. If you were translating, you had a dictionary open next to you. Now if you're translating, you have ChatGPT open next to you.
Speaker B: Absolutely, absolutely. And uh, I foresee uh, with. And also some platforms, uh, already uh, has this implemented the translators editors, uh, they do the suggestions and uh, from uh, Google Translate for example, or from ChatGPT, whatever model you want, you get the suggestions. 5 suggestions, 6 suggestions, you can click uh, and get another suggestions and select what you like. And this is just a tool. Uh, and I would say that in the world not only translators are translating, uh, marketing persons translating into the native languages. Why not? They are sitting in crowding. They're not linguists, they're not translators. But they know what they want and they get it not from their hat, but uh, from an LLM inside the translation platform platforms editor. And that's uh, that's like uh, internal, internal translation. Um, houses of bigger companies, they just have their linguists inside and they are not linguists. They just, they uh, could be product managers, they could be project managers, whatever market or marketing specialist just translating some documentation, uh, or something without any linguists and AI, uh, since they are not professional translation, AI just uh, is a copilot, is a linguist copilot for them. Uh, and they are not super professional translators to uh, just infer the translation from their heads, but they can definitely select from three or four variants the suitable translation. And this is how the AI changing uh, the way of translations are done.
Speaker A: Well it's like I don't have to be a PhD professional mathematician, give me a calculator and I can do just, you know, I can do amazing things with the calculator even if I'm not a professional musician. So kind of, kind of the same concept on that. Yeah. Um, I saw a comment from Anne says yes, co translator AI is a great example of translators taking commands. It's just, yeah, one example, very fun tools for us. So translators are out there, they're, they're using AI. What you mentioned I think is an interesting phenomenon that we're seeing not just in translators, but across industries with AI, um, in general. But this idea of you don't have to be a professional translator, you have marketing people, people that are using AI to get the benefits of the knowledge that typically a translator would, would have. And it's this concept of leveling up existing team M members, whether they're project managers, translators, engineers, whatever is. You don't need to um. I hate kind of saying this but you don't need to hire rock stars anymore. You don't need to hire the very best translator. You can hire the middle mid range, cheaper translator and use AI. I've heard this comment or this sentiment floating around around there. What, what do you say around that I'm interested in your feedback?
Speaker B: Well, I, I think uh.
Speaker A: Right. That's kind of how I feel about it.
Speaker B: So at Alkonost we uh, we don't work with uh, translators that are not uh, satisfying certain level of uh. Our internal tests. Uh, there's a. It's very. So uh, you'd better work with stars.
Speaker A: Okay, you heard it here people.
Speaker B: So maybe someone uh, could be satisfied uh by um, uh, by not very professional translators. I, I don't.
Speaker A: I, I think it depends. It depends too on what kind of niche you're serving. What kind of clients are you're serving. Because I could foresee a situation. There are LSPs out there that they compete on price and we provide cheap, quick translations. That's the niche that they serve. There are some clients that they don't care about quality. They just want cheap. Fine, awesome. In those type of situations I think it makes sense to not have to worry about leverage AI to the extent possible.
Speaker B: But we had such kind of experiment, um back in time we had a nitro translate system, uh, which is uh, just uh, you can top up your balance, go and order uh, the translations. So sorry my son has come here.
Speaker A: It's all right. My minor at school. Otherwise they'd be pulling my beard right now.
Speaker B: Uh, so um, what we had in that system, it exists, uh, uh, just, just now. We, we had two uh, chairs of tiers of quality. Good, excellent and uh. Yeah, good and excellent. And we give choice uh, for people to select between good and excellent quality of translation. We measured the quality of our translators and. But uh, secretly they were all excellent translators translating in good and excellent quality. Um, uh. But uh. Uh, the good translation cost costed less. But everyone was buying Excel. Excellent. Yeah, it was expense more expensive. Nobody bought good. Some, some few people bought. But uh. Our statistics show that if it's translation, that it must be the best ever. Uh, the best. Uh, so I don't think that uh, somebody who is the, who will be just selecting between good translate and that translator plus LLM would select this bad translator plus. I don't believe. Uh, no, nobody will choose.
Speaker A: Yeah, I, I kind of equate it. I've in my mind, I've always equated to like hiring a plumber. If I hire a plumber and they ask me, do you want me to fix the leak with super high quality or do you want me to fix the leak just a little bit? Well, no, I want you to fix the leak like it doesn't matter. I want super high quality. I might not be thinking like high quality, but I want high quality. And that's what you're saying customers are, are expecting there. But I want to move the conversation on. Just move it, move it forward though, because there's a part in here, in the study that I wanted to look at which is the human versus automated Automatic Human versus automatic metrics. Right. And so you have a idea here is like, okay, how does the human, um, MQM scoring compare to the automated metrics? And this is. Well, this is particularly interesting to me. I'm working on some projects right now where we're doing a similar evaluation. But I think it's interesting to everybody who's wondering how do I evaluate the quality output of my machine translation. I don't want to be paying three human reviewers to review everything and annotate it. Um, is there an automated way? Um, and yes, there are, There are several. There are multiple automated ways. And you did it a comparison here of which one corresponded the closest to the human evaluation. Right. So which automated metric was the most human? Like and I'll just read it here. Um, human versus automatic metrics. How does state of the art automatic metrics compare to our linguist assessment? We ran two leading automatic MT evaluation metrics, Metric X from Google and Comment Kiwi from Unbabel on the same translations in quality estimation mode without reference translations. Both are neural metrics used in WMT evaluation campaigns. We compared their language rankings against our human MQM scores to measure how well machines can approximate expert judgment. And here we have the best is Metric X. Going down the list, Comet, uh, KIWI xl, Comet, Kiwi Base, Metric X, Excel, Comet KIWI XXL and Metric X. Um, walk me through this just a little bit or walk the audience, I should say through this.
Speaker B: Uh, yeah, this is my favorite, ah, topic. We are just running Lots of automated quality, uh benchmarking of uh different MTs between different language pairs and different domains. Every project, every production project starts with um every client just comes uh and asks well I need to translate everything by AI. And what happens next is we sit down and just uh classify the M, categorize the content, uh by its visibility, by its impact and then uh by its volume and then by domain and types of content. Like for example you have user generated user reviews, you have um support materials, you have, have website, you have marketing materials, you have uh your app or web UI strings. And those are different, uh different things. And uh, to select the right model for right languages we uh just do a lot of benchmarking. And so uh those metrics uh we use almost every day to measure everything. Uh so and that was a very fun part um of evaluating uh uh we took the biggest one. So those metrics are for um finding roughly saying they are for finding uh errors in translations. And they, they measure the, the number of roughly speaking they measure the number and the um uh and the severity of the errors found in translation. So they are trained to do this based on the data from uh WMG framework. Google studied it many years ago. And then this frame, this uh data set is growing and they have uh lots of data and they trained those LLM based metrics on the, on that data. Uh and that data is a human annotation, uh like this mkm, uh annotation. There are some, not not just MKM but also direct assessment, also Euro Spanish uh annotation. So uh, they have data um and they train the metrics and now the metrics are capable of um spotting euros in text and to somehow measure it uh and to give you um some number, some number telling this is good or this is bad. And uh, those metrics are available in different sizes. So the base model model can be smaller, uh can be bigger, it can be super heavy. And the, the bigger the base model, uh uh the more languages it supports, uh the more languages it knows it. Uh just so many data on the Internet uh but you need the uh hardware, um lots of hardware to run it uh to infer uh the value. So we just uh ran what we had everything and um, so according uh to the, to those metrics they pretty much agree with humans regarding uh that there are years, uh there are years and metrics tell that. Yes uh, they are but uh, we must be careful with um different side of this. If this metric says that there are no euros, it doesn't mean that there are no euros. So if it's 100% quality, you can go with your linguist and see that there might be years, but they are not spotted. But if the metrics shows that there are years, then there are just 100%. There are errors, they are told in this way. So, um, a good uh, benchmark to spot the bad quality of translation but not the measure of uh, high quality translation. You cannot rely on the, on the 100%. Uh, you the metric gives you. Uh, it might be not so but if it gives you 50% that it's 100 uh, percent that it's low quality of the segment or all the documents. So we're entering what we had uh, over the text. Over the text. And this is uh, this is animal Anomaly with uh, metric X served by Vertex AI. This is the. Yeah, uh, really API of Google.
Speaker A: It's really bad.
Speaker B: We should, and we should um. I think that's there is some mistake there. Uh, I think this is just a glitcher back we should investigate. But uh, maybe it's some misconfiguration or something else. Uh, but when served, uh, just uh, ad hoc from, from Google, uh, Google's API maybe, uh, um, I don't know why uh, but it's, it must be checked. Uh, but when served by hugging, uh, face inference endpoint, uh, the same model gave us what we expected. Uh, this is the same metric X. Uh but we uh, will be investigating why. Why it gave because maybe it's in our code. But maybe. Um. I don't know. I, I don't know. Uh, this is another finding. So if you go to Vertex AI and ask Google uh, Google Metric, uh, X, you will get this number and uh, and this number is low and why and so something wrong with that. Uh, this should be investigated. Yeah, maybe
Speaker A: right. So
Speaker B: to, to sum up, uh, uh, quality metrics of this kind reference free. Uh they. They agree, uh they agree with humans. If there are years, then metrics would say uh, there are years.
Speaker A: Yeah and I like what you said about, I mean the machines essentially I'm paraphrasing here, the machines are a lot better at finding errors or than they are about rating a good translation. I need to figure out a better way to say that. Right, but it's better at finding mistakes than finding awesomeness. Whatever the awesome, whatever the opposite of a mistake is. Right, Good, good translations. But I mean this brings into the, the debate the whole like what is a perfect translation, right? Is, is there, is there such a thing as 100 quality? Um, you bring three different reviewers to review the same thing. They're going to find three different things. As I'm sure you, you saw some examples of bats when doing this.
Speaker B: I foresee and have an idea of uh, another kind of metric or the way uh, the language model is taught to translate. So uh, now uh, with these metrics, uh, uh, the models, they are taught to find errors, so uh, they are rewarded uh, to find the errors. Um, but maybe there is a future world when we encourage machine to translate better and just give them positive feedback so that they not spot the euros but make the translation better. Yeah, so it's not a metric but uh, it's a, it's a. I um, would imagine a data set of uh, linguists uh, taking poor translations and then by some reasoning making them better. And this is, I, I think this is a future golden data set, uh, that, that data set, um, that could be used to train better models in translation because uh, everybody can come up with uh, bad translation. But how to make it better and, and why, uh, with explanation, with reasoning. Um, maybe somebody's building already this data set in the world.
Speaker A: I don't know, maybe it'll be you and 12 months from now you'll be back on here explaining. Well, I'm watching the clock here. We're running out of time today and I have a hard stop after this. Uh, any closing thoughts? I, I put the link really quickly. I put the link for this study into the LinkedIn. Um, but if anyone has any questions, make sure to hunt Alexander down. I'm sure he's here to answer them. But how can people find you? Um, what should people do with this information?
Speaker B: Uh, so type the Alkonost anywhere and you, and you get my LinkedIn and please just uh, add me on LinkedIn and that's it. How to find me. I'm very open. Uh, one of the final thoughts. So um, that yeah, this MT is not the official side of um, Alkonost. It's, it's our Laboratory. Laboratory. So alkanos.com is the uh, is the, the website.
Speaker A: I'm trying to promote you here. There we go.
Speaker B: Thank you very much. Uh, one final thought about this uh, case study. So uh, going to production uh, without people is not yet, not yet there. So we are very excited sometimes about the quality of translation that AIs provide to us. So uh, we see that brilliant examples of good translations and suddenly uh, 20% of the translations are not that good. So it's like almost brilliant everything but in, in very unexpected place you find the mistranslation or something that should have been checked by the linguists. So I see the, the future of translation in this. In this agendic world, uh, when the translations are really performed by LLMs, uh, but then checked, uh, with the. The real humans called each other, um, and sometimes retranslated. And if somebody is streaming of 100% AI translation. Maybe next year, but not today.
Speaker A: Well, today.
Speaker B: And it depends on the language.
Speaker A: It depends on the language. It depends upon the content type. It depends upon the
Speaker B: better build a pipeline with a human, uh, in the loop. They call it human in the loop. I don't like, I don't like the term, uh, because it's not human in the loop.
Speaker A: I've heard the term. I want to give credit to Marco Trombetti for this, but I've heard it's like human driving the bus or human at the wheel or something like that.
Speaker B: Yeah, that's a much better thing.
Speaker A: Yeah.
Speaker B: And with these cloth codes and cortexes, I uh, think it's. It. Um, the will. The wheel will be driven by um, code by some agent. Um, but this agent will be of course supervised by the bus is driven
Speaker A: by code, but the driver.
Speaker B: All the logistics, all the logistics. Uh, the localization service providers are providing with localization manager. It will disappear. No sending files, no even just uh, confirming deadlines. No this uh, and every. Any kind of logistics that can be automated will be automated. But there will be just the creator and just the translator and some agents in the middle. Even platforms. I don't believe that there are lots of future in platforms.
Speaker A: All right, I'm clipping it. I'm taking this as you're. I'm clipping this. I'm quoting you on it. Five years from now we'll see if you're right or if. Okay, completely Something completely different. So Alexander, uh, we are out of time today. I gotta run to my next one. Thank you so much for joining us today. Uh, I really appreciate it. Uh, guys, go check out the. The study over. Just search for Alkonost. I put the study link into the. The chat over on LinkedIn so you can link it directly. Ladies gentlemen, chat. We are out of time for. For today. If you enjoyed this Nimzy live experience, join us next time in the following weeks. We got some coming up. I don't think they're scheduled yet, but stay tuned. Follow NIMsy's LinkedIn page so that you'll be the first to get notified when we schedule new live streams. I appreciate our guest today, Alexander. I appreciate my colleagues here at Nimsy Insights doing all the hard work so that I can have these interesting conversations. And finally, I appreciate you, the audience join. Joining us live today, all the questions, comments and criticisms that we receive in the comments. I look forward to next time. Cheers.
Speaker B: Uh,
Speaker A: Sa.
Other episodes covering the same guests and topics, from across The B2B Podcast Index.