The message arrives in German at eight in the morning and the agent answers in German in eleven seconds. Flawless grammar, right facts, zero errors. And the customer still forwards the reply to their account manager with one line on top: "a person did not write this." They did not complain about the content. They complained that the company they have worked with for three years suddenly sounds like an instruction manual. Nobody opens a ticket for that, so the problem shows up on no dashboard.
The thesis in one line: the model translates well and transcreates badly. Switching on a language in your agent is a configuration toggle; keeping the brand voice inside that language is not. By default your agent speaks five languages like a translated manual — correct, flat and foreign — and in customer support that gets paid in trust, precisely in the markets where you have the least presence and the least room for error.
Multilingual AI agent and brand voice: the model translates well and transcreates badly
The reason this catches teams off guard is that the hard part solved itself. Three years ago, support in five languages meant five knowledge bases, five review cycles and five people. Today the model detects the customer’s language and answers in it without anybody maintaining anything on the side. That ease is real, and it is why the decision to open a language now gets made in a twenty-minute meeting that marketing is not in.
What did not solve itself is quality, and it has been measured. Oracle AI researchers published on arXiv, in the October 2025 revision, a study on multilingual consistency in enterprise applications with an uncomfortable result: even with advanced RAG systems they observed accuracy drops of up to 29% in non-English languages compared with English. The diagnosis matters more than the number: the model reasons internally in English even when you address it in another language, so the bias is not in the output translation, it sits upstream of it. Their own alignment technique recovers up to 23.9% of that accuracy — which confirms the gap is structural, not a one-off misconfiguration.
Scope of that figure, stated plainly: a lab benchmark over 500 business-domain documents human-translated into six non-English languages (Spanish, French and Portuguese among them), run by a team at an infrastructure vendor that sells AI solutions. It sizes the phenomenon at model level; it does not measure your agent and it is not a result of ours. And it is a universal datapoint, not a market one: it does not change depending on whether your company is in Spain or Italy.
Accuracy is also the visible half of the problem. An imprecise answer produces a second message and shows up in the metrics. A precise, lifeless answer produces nothing: the customer reads it, shrugs and drops a point of trust without telling you. It is the same mechanism that makes a polite, correct chatbot end up annoying customers, except here the mismatch is not about expectations. It is about accent.
What a badly shipped language costs has been measured since long before AI. CSA Research, in its "Can’t Read, Won’t Buy – B2C" report on consumer language preferences — a sample of 8,709 consumers across 29 markets, published July 2020 — asked what problems people run into on a site translated into their own language. The most cited answer was not missing translation: it was content quality, meaning the translation is bad or off target (34%), followed by no localized help (33%) and incomplete translation (26%). And nearly half — 48% — leave the site the moment they hit the problem instead of asking for help. Scope: a global 29-market survey, from 2020, about websites and apps, not about AI agents. It does not measure chatbots; it sets the baseline for what happens to a customer when a language is technically present and culturally absent.
The three things that break when a reply crosses languages
The failure is not random. When a reply crosses languages, three specific things break, and all three break in the same direction: toward neutral. When the model hesitates, it picks the most formal, most literal and most generic version of what it could have said. That is the safe bet in translation and the losing bet in brand.
| What breaks | How it shows up | Why the model gets it wrong by default |
|---|---|---|
| Level of formality | Your German site uses "du" and the agent replies with "Sie"; or in Portugal it answers in a Brazilian register that reads like another country | Formality is not in the source text, it is a brand decision. Without an explicit instruction the model picks the most formal register, because that is the one least likely to offend |
| Humour, indirectness and rhythm | The line with an edge comes out correct and defused: same meaning, zero gesture | Humour is the first thing lost in translation because it rides on shared references. The model does not strip it on purpose: it picks the most probable phrasing, and the most probable is never the sharpest |
| Legal and product references | The agent quotes a return window, a tax concept or a plan name that does not exist in that country | This is the expensive one. The model carries over content from the source language because that is what it has, and it has no way of knowing a fact is territorial unless you tell it |
The third deserves its own paragraph, because it is the only one that can end in an actual complaint. A support agent promising an Italian customer a return window that only applies in Spain is not making a tone error: it is making a false commercial claim on your company’s behalf. The operating rule that prevents it is boring and it works: every fact that changes by country — deadlines, prices, legal concepts, product names, payment methods — gets flagged as territorial in the knowledge base, and the agent is not allowed to resolve it by analogy. How to structure that base so the agent knows what it knows and what it does not is in training an agent on your own information.
The brand glossary ships per language, not translated
The most repeated process mistake: write the tone-of-voice document in your home language and have it translated into the other five. It looks like the logical move and it is exactly backwards. A translated glossary tells the model how the brand sounds in Spanish and leaves it to decide what the German equivalent is. That decision is the one you were trying not to delegate.
A per-language glossary is a different object: written directly in each language by somebody who speaks it in that market, with the decisions made rather than derived. Which pronoun you use. Which three words you never say. How a message opens and how it closes. How you say no without sounding like paperwork. Which terms stay in English because that industry keeps them in English. It is not a long document: one page per language is enough, and that page is the part of an AI agent’s instructions that returns the most per character written.
The cheap control almost nobody puts in
Here is the asymmetry that explains the neglect. Nobody ships an agent in their main language without reviewing replies. And almost nobody reviews the other languages, because the person who could judge them is not on the team that built the agent. The result is five secondary languages running unsupervised from day one, and the first signal that something is off arrives as a customer who stops replying.
The control that closes it is absurdly small next to the project: a weekly sample of real replies per language, reviewed by somebody native to that market, against that language’s glossary. Five or ten replies, not a hundred. Half an hour per language, not a day. And one person who answers for it, not a shared channel.
- Pull a random weekly sample per language, not the ones that generated complaints. You review those anyway; tone breaks in the ones nobody reported.
- Have it reviewed by somebody native to that market, not somebody who speaks the language. The difference matters: a Spaniard with good German catches errors, not register.
- Judge against that language’s written glossary, not the reviewer’s taste. With no glossary, the review produces a different opinion every week.
- Log the pattern, not the reply. A bad reply gets fixed by hand and changes nothing; a pattern — "in Italian it always opens with an apology" — gets fixed in the instruction and fixed for all of them.
- Check the following month that the pattern did not come back. Tone fixes decay on their own when the model changes or the knowledge base grows.
That loop — sample, score against a written criterion, catch the regression and fix it — is not a launch task: it is a function that runs. When it does not fit in the team, it is exactly what we build and operate in evaluating the quality of your AI agents, and with languages the only difference is that the criterion gets written five times instead of once.
When not to switch on the second language
The uncomfortable consequence of all of the above: if nobody can judge how your brand sounds in Italian, switching on Italian is not progress, it is delegating the company’s voice to a system that optimizes for probability. There are cases where that is fine and cases where it is not, and the line is reasonably sharp.
It is fine when the conversation is purely transactional and has no give: order status, opening hours, availability, a fact that either exists or does not. There, the correct flat answer is the correct answer, and language is logistics. It is not fine the moment there is negotiation, apology, retention, selling or bad news: in those four moments the brand is the message, and a neutral reply loses. So the practical criterion is not "which languages do we switch on" but "what kind of conversation do we let the agent run in a language we do not supervise".
And the mandatory companion to that line is the exit: in the languages you do not supervise, the threshold for handing off to a human has to be lower than in your main one. Not because the agent fails more often, but because when it does, nobody finds out. How that limit gets designed — what the agent does when it is not sure, and how the handoff happens without the customer noticing — is in what the agent does when it does not know.
None of this gets fixed by a better model, and that is the part that is hard to accept. Models will keep improving at multilingual accuracy, probably fast. But your brand’s tone in German is not information the model can derive from a Spanish text: it is a decision somebody at your company has to make and write down. If it is not written, the model makes it for you, and it makes it well — in the sense of correctly, and in no other sense.
The line for the next time somebody proposes switching on three more languages on Tuesday: speaking the language is not the same as sounding like you. The model does the first one for free. The second one has to be written, in every language, and checked every week.