AI Earbuds

How to Speak So Translation Gets It Right

By Mark Fulton · September 18, 2026

How to Speak So Translation Gets It Right

If you want better translation accuracy today, change how you speak before you change what you bought. The most common failures in live translation are not model failures. They are half-finished sentences, two questions stacked into one breath, a proper noun dropped into the middle of a clause, an idiom, or a pronoun whose referent was three sentences ago. Speak one complete idea per turn, finish it out loud, keep names and numbers in their own short sentence, say the subject instead of dropping it, and skip the sarcasm. That set of habits costs nothing, takes about ten minutes to learn, and improves results on every product in this category, including the free app already on your phone.

I sell a translation earbud kit and a browser translation app, so it is worth saying plainly that this article does not need you to buy either one. The habits below work on an expensive dedicated handheld and on a free phone app equally well, because they all operate on the same bottleneck: the sentence you actually handed the machine.

Why do most translation errors start with the speaker?

Every live translation product runs the same chain of jobs, and I broke that chain apart stage by stage in how translation earbuds actually work. The short version is that your audio gets captured, a system decides when your phrase ended, a recogniser turns the audio into text, a translation model turns that text into another language, and a voice reads it out.

The important property of that chain is that nothing at the back repairs a mistake made at the front. Machine translation is genuinely good at turning a sentence in one language into a sentence in another. It has no way to know that the sentence it received is not the sentence you meant. So a speaking habit that damages the text at stage three produces a fluent, confident, completely wrong result at stage five, delivered in the same calm synthetic voice as everything that worked.

This is also why the failure feels like the device malfunctioning. Nothing sounds broken. You hear a smooth sentence come out, the other person nods, and the misunderstanding only surfaces two minutes later.

Spontaneous speech is harder for machines than written text for reasons that have been studied formally for decades. Filled pauses, restarts, repairs and false starts are a recognised research problem in their own right, which is why NIST ran disfluency detection as a scored task inside its Rich Transcription evaluation series. When you say "so I was, uh, wondering if, actually no, can I just", you are producing exactly the input that a whole research programme exists to cope with.

The people who have solved this problem in practice are not engineers. They are the professions that have worked through human interpreters for a century. Their guidance is remarkably consistent, and it transfers almost unchanged to machines.

Which ten habits help most, in order?

Ranked by how much each one improves what the machine actually receives. The first three account for most of the gain. The tenth is the safety net that catches the other nine when they slip.

# Habit Instead of this Say this
1 One idea per turn, then stop. "I'm checking in, and I also wanted to ask whether the airport shuttle runs early, and if you can hold my bag after checkout." "I would like to check in." (wait) "Does the airport shuttle run at six in the morning?" (wait) "Can you hold my bag after checkout?"
2 Finish the sentence out loud. "I need something for, you know, my stomach, it's been a bit..." "I need medicine for stomach pain."
3 Say the subject and the object. "Can't eat that. Allergic." "I cannot eat peanuts. I am allergic to peanuts."
4 One question per turn. "Is breakfast included or do I pay separately, and what time does it start?" "Is breakfast included?" (wait) "What time does breakfast start?"
5 Give names their own sentence. "Can you take me to the Sathorn Grand on Soi 12, near the river?" "I want to go to a hotel." (wait) "The name is Sathorn Grand." (wait) "It is on Soi 12."
6 Plain verbs, not phrasal verbs. "Can you drop me off just up ahead by the corner?" "Please stop at the corner."
7 Say dates and times in full. "Let's push it to next Tuesday, end of day." "Let's move the meeting to Wednesday, the twenty-third of September, at five in the afternoon."
8 Cut sarcasm, jokes and padding. "I don't suppose you'd happen to have a room with a window, would you?" "I would like a room with a window."
9 Replace a stale pronoun with the noun. "And is it covered? Does it work for that too?" "Does my insurance pay for this medicine?"
10 Confirm back at the end. "Okay, great, thanks." "Please tell me the address you will drive to."

Habit 7 is worth a second look, because it contains two separate repairs. "Next Tuesday" is ambiguous even between two native English speakers standing in the same room, and "end of day" is an office idiom that does not survive a hop into most languages. Naming the weekday, the date and the clock time removes both.

Which sentence shapes survive translation best?

There is a shape that gets through almost intact, and it is boring on purpose: one subject, one verb, one object, stated in that order, in the present or simple past.

Four shapes reliably cause trouble.

Trailing clauses. "I was going to ask about the room, if that's alright, unless you're busy." The end-of-phrase detector has no idea where this stops, and the translation model has no main clause to anchor to.

Double negatives. "You don't have anything that isn't spicy?" This is the same warning the federal courts give attorneys. The New Jersey District Court's guide to the effective use of court interpreters tells lawyers to keep questions as straightforward as possible, and says plainly that lengthy questions and double negatives may confuse the witness. Machines inherit the problem and add one of their own: some languages answer a negative question with the opposite polarity to English, so "no" can come back meaning "correct, we have none" or "no, that's wrong, we do have some."

Dropped subjects. English speakers drop the subject constantly in casual speech. "Went yesterday." "Still waiting." A translation model must then guess a person and a number, and once it guesses wrong the rest of the paragraph inherits the error.

Stacked questions. Two questions in one breath usually produce one answer, and you will not know which question it answered.

The professional guidance says the same thing from the other direction. The HHS Office of Minority Health checklist for working effectively with an interpreter tells clinicians to use simple language, avoid jargon, and work sentence by sentence, because multiple sentences lead to information being left out. That was written about a trained human professional. The machine is considerably less forgiving.

What should you do with names, numbers and dates?

Proper nouns are the single most fragile thing you can say. The recogniser is trying to map your audio to words it has seen, and a hotel name, a street name, a surname or a medicine brand is often not in that vocabulary in the form you pronounce it. So it substitutes the nearest thing that is. "Sathorn" becomes "southern". "Ibuprofen" survives; a local brand name usually does not.

Four rules that fix most of it.

  • Isolate it. A name in its own short sentence gives the recogniser a clean start and a clean finish. A name buried mid-clause inherits the surrounding audio.
  • Show it rather than say it. For an address, a hotel, a medicine or a surname, hand over the written text. This is the one case where the best translation habit is to not use translation at all.
  • Say numbers in full words, one group at a time. "Room two, one, four" rather than "room two fourteen". Prices, platform numbers and phone numbers survive far better digit by digit.
  • Expand abbreviations. The same court guide tells attorneys to spell out jargon and abbreviations and to say the entire phrase where possible. "A and E", "ETA" and "ASAP" all translate badly, and "as soon as possible" translates perfectly.

Why do idioms and sarcasm fail?

An idiom is a phrase whose meaning is not the sum of its words. A translation model has two options: render the words, which produces nonsense, or recognise the phrase and substitute an equivalent, which only works for idioms common enough to appear in its training data. "It's raining cats and dogs" is famous enough to be handled. "Can you give me a ballpark?" often is not, and "let's touch base Thursday" almost never is.

Sarcasm fails for a deeper reason. The literal content of a sarcastic sentence is the opposite of what you mean, and the correction lives entirely in your tone of voice. The text handed to the translation model has no tone in it. Say "well, that's just perfect" to a hotel desk and the machine will tell them, sincerely, that everything is perfect.

The same goes for British-style politeness padding. "I don't suppose you could possibly..." is four words of hedging wrapped around one request, and the hedging is what survives while the request gets mangled. Directness reads as rude in English and as clear in translation, and clarity is what you are optimising for here. If you want warmth, put it in your face and your voice, not in your grammar.

Humour, rhyme, wordplay and film references belong in the same bucket. Save them for when you share a language.

How long should one turn be?

One complete idea, usually one sentence, rarely two. Then stop and let it run.

There is a practical mechanism behind this. Most systems decide you have finished talking by detecting a pause. Keep talking and either the system commits early and cuts your sentence in half, or it waits and the translation drifts further behind the conversation. In a loud room this gets significantly worse, which I covered in detail in what breaks in noisy places.

Short turns also limit the blast radius. When a fifteen-word turn fails, you lose fifteen words and you usually notice. When a ninety-word turn fails somewhere in the middle, you lose the thread and nobody notices at all.

This is what the courts do with witnesses, for exactly the same reason. The guide above tells attorneys to advise witnesses to pause regularly, because when witnesses do not pause it becomes very difficult for the interpreter to retain all the details of a long narrative. Machine or human, the working memory has a limit.

One more thing about pacing. Speak at a normal, even pace and at a normal volume. Slowing down to an exaggerated crawl actually hurts, because it distorts the rhythm and stress patterns a recogniser was trained on, and it inserts pauses mid-sentence that get read as end-of-turn. The HHS checklist makes the volume half of this point directly: speak clearly, and do not raise your voice or shout. The goal is clear, not loud, and not slow.

If you are new to the rhythm of a two-way exchange, what real-time translation actually means walks through the lag you should expect between turns.

How do you check the other person understood?

Assume nothing from a nod. A nod means "I hear a voice speaking my language", which is not the same as "the content arrived."

The healthcare profession solved this decades ago with the teach-back method, and the HHS checklist puts it near the end of every visit: rephrase and confirm that the person understands your directions. You are not testing them. You are testing the channel.

Three ways to do it that do not feel patronising:

  • Ask for the content back, not for a yes. "Please tell me the address you will drive to" beats "do you understand?" every time. A yes costs nothing to give.
  • Ask for a number back. Time, price, platform, room number, dose. Numbers are both the most fragile content and the easiest to verify.
  • Rephrase once yourself. If something matters, say it a second time in different words. Two independent phrasings rarely fail the same way, and the second one often repairs the first.

For anything with a real consequence, a medication, a medical history, a contract term or a safety instruction, add a written fallback. Type it, show it, and let both sides read it. I made the same argument at more length in do translation earbuds actually work, because knowing where a tool stops being appropriate is part of using it properly.

Try it twice

Here is the honest test, and you do not have to pay anything to run it. Go to the live demo on the homepage and speak a real sentence you would actually use, exactly the way you normally talk, with the trailing off and the stacked questions and the idiom in the middle. Listen to what comes back. Then say the same thing again using the ten habits above.

The gap between those two attempts is larger than the gap between most products in this category. That gap is free, and it belongs to you rather than to your hardware.

For what it is worth, our own answer is a $69 earbud kit with our browser translation app, AI Translation Live, as a $39 a year add-on. But the ten habits are the part that transfers everywhere, and they are the reason I would rather write this article than another setup guide.

FAQ

How do I make translation more accurate?

Change your input before you change your equipment. Speak one complete idea per turn, finish every sentence out loud, include the subject and the object instead of dropping them, put names and numbers in their own short sentence, and avoid idioms, sarcasm and stacked questions. These habits improve the text that reaches the translation model, and no model can recover meaning that never arrived. Reducing background noise and closing the distance to the microphone are the next two levers after that.

Should I speak slowly to a translator?

Speak clearly and evenly, not slowly. Exaggerated slowness distorts the rhythm and stress patterns that speech recognition was trained on, and the artificial pauses you insert mid-sentence often get read as the end of your turn, which chops the sentence in half. Normal pace, normal volume, complete sentences, and a real pause at the end of each one. Raising your voice does not help either, and the HHS interpreter guidance specifically warns against it.

Why does it mistranslate names?

Because a speech recogniser maps your audio onto words it already knows, and most proper nouns are either absent from that vocabulary or present in a different pronunciation than yours. When there is no match it substitutes the nearest ordinary word, so a hotel or street name becomes a common noun and the translation model then faithfully translates the wrong word. Give each name its own short sentence, or show it in writing, which is the most reliable fix by a wide margin.

How long should I speak before pausing?

One idea, which usually means one sentence and rarely more than two. Most systems use silence to detect that you have finished, so a long turn either gets cut early or falls further and further behind the conversation. Short turns also make failures visible: you notice a broken ten-word sentence immediately, while a broken ninety-word monologue just quietly loses the thread for both of you.