Translation Earbuds and Privacy: Where Audio Goes
By Mark Fulton · September 2, 2026

In almost every translation earbud setup on sale, your voice leaves your head and travels to a computer you do not own. The earbuds are a microphone and a speaker. The translating happens somewhere else, which means the audio has to get there, and every stop on that journey is a place where a copy could exist. Whether one does exist depends on the vendor's retention policy, not on the hardware. "Recording" and "processing" are separate questions, and most product pages answer neither. This page maps the actual path in four common setups, names who sits at each hop, and gives you the questions that force a straight answer out of any vendor, including us.
I sell a translation earbud kit and a browser translation app, so I have an interest in how this question gets answered. The way I have chosen to handle that is to publish our own data path in enough detail that you can check it, rather than print the word "private" on a product page and hope nobody asks.
Where does your voice actually go during translation?
Start by separating two things that both get sold as privacy.
The first is acoustic privacy. Nobody in the room hears the translation because it plays into your ear instead of out of a phone speaker. This is genuine, and it is the version most vendors advertise. It is also entirely about the room.
The second is data privacy. Who receives your speech, who stores it, for how long, and what they are allowed to do with it. This has nothing to do with where the sound comes out. A product can be excellent on the first and completely unexamined on the second, and in this category that combination is common.
The general path looks like this. Your voice hits the microphone in the earbud. It travels over Bluetooth to your phone, or to a handheld device. The app or device packages it and sends it over the internet to a backend. That backend usually hands the audio to a speech recognition model, then a translation model, then a speech synthesis model, which may all belong to a different company than the one whose name is on the box. The synthesised reply comes back down the same chain. I wrote the mechanical version of this in how translation earbuds actually work, and it matters here because you cannot reason about privacy in a system you believe is self-contained.
Here is the map across the four setups you will actually encounter.
| Setup | The path your voice takes | Who could hold a copy |
|---|---|---|
| Earbuds plus a vendor's phone app | Mic, to Bluetooth, to the app on your phone, then over the internet to the vendor's backend, then out to whichever speech and translation providers the vendor uses, then back | The app on your phone, the vendor, each sub-processor the vendor uses, and any cloud logging in between |
| Standalone device with its own connection | Mic, to the device, straight over cellular or Wi-Fi to the vendor's backend, then to the vendor's providers, then back | The device itself if it stores history, the vendor, each sub-processor, plus the carrier for connection metadata |
| On-device offline pack | Mic, to Bluetooth, to a small model running on your phone, then straight back to your ear | Only your phone, while it stays offline. The app can still upload history, diagnostics or crash reports later |
| Browser straight to the model | Mic, to the browser tab, then over an encrypted peer connection directly to the translation provider, then back | The translation provider, and your browser tab while it is open |
Read the third column as the list of parties who could keep something, not a list of parties who definitely do. That distinction is the whole subject. What turns "could" into "does" is a written retention policy, and the way to find out is to go looking for one.
Hop by hop, this is what each party is capable of holding:
- The earbuds. Nothing. Consumer translation earbuds have no meaningful storage and no independent connection. They are a transport.
- Your phone, or the browser tab. The app holds whatever it chooses to keep. Conversation history, saved transcripts, cached audio and an account you signed into are all decisions the developer made, not laws of physics.
- The vendor's backend. The hop people never picture. If the audio passes through the seller's own servers before reaching a model, the seller is holding your speech, however briefly, and their logging configuration decides what survives.
- The model provider. Speech recognition and translation are frequently bought in rather than built. The company on the box may not be the company processing the audio, and this is the hop least likely to be named anywhere on a product page.
- The other person's device, if they are running their own app on the other side of the conversation, with its own separate answer to all of the above.
What does "processed on device" usually mean?
It usually means a compressed model downloaded onto your phone, doing the translating locally, with the earbuds still acting as nothing more than a microphone and a speaker. Nothing is processed on the earbuds themselves, and vendors who offer an offline mode are generally clear about this once you read past the feature bullet.
For privacy this is the strongest position available, with two caveats worth more than the headline.
The first is that offline processing stops the audio leaving the phone in real time. It does not stop the app writing a transcript to disk, syncing history to an account when a connection returns, or uploading diagnostics. On-device processing and on-device storage are separate promises, and a product can make the first while quietly doing the second.
The second is that offline modes are limited. They cover a short list of language pairs at lower quality, and they usually have to be downloaded before you travel. The practical tradeoffs are in whether translation earbuds need Wi-Fi or data, and they are steep enough that offline is a fallback rather than a default for most people.
There is also a phrase to watch for. "Advanced encryption" is a claim about the audio in transit. It says nothing about what happens to that audio once it arrives and is decrypted at the other end, which is where retention decisions get made. A vendor can be telling the complete truth about encryption and still keep your conversations for a year.
What should you ask a vendor before buying?
Send these before you spend anything. A company with a defensible answer will have it written down already. A company without one will send you marketing copy, and that is itself an answer.
- Where is speech recognition and translation performed? On the device, on your servers, or on a third party's servers.
- Who are your sub-processors? Name them. A vendor that cannot name the companies processing customer audio has not thought about this.
- Is audio retained after a session ends, and for how long? Ask for a number, not a reassurance.
- Are transcripts stored, and where? Server side, on the device, or both.
- Is customer audio used to train or improve models? If so, is there an opt out, and is it on by default.
- Does the app keep conversation history, and does deleting it on the phone delete it on your servers.
- What happens to stored data if I close my account, or if the company is acquired or shuts down.
- What permissions does the app hold when it is not open? Background microphone access is a different risk profile from foreground only.
Two notes on grading the replies. If an answer is not in a dated written policy you can link to, treat it as unverified rather than false. And be equally sceptical of sweeping category-wide claims. When a page tells you that every product of this type is secure, it is describing a marketing position rather than an architecture.
What are the rules in a workplace or clinic?
This is where the question stops being about products. Nothing here is legal advice, the rules vary by country and often by state or province, and the only sound approach is to check your own jurisdiction and your own employer's policy.
What is useful is knowing the shape of the question, because it is the same shape almost everywhere.
The starting point is that speech is personal data. The European Data Protection Board's guidance for small organisations lists audio recordings containing the sounds of individuals as personal data, alongside names, email addresses and location data. Once something is personal data, an organisation processing it needs a lawful basis for doing so, has to inform the people involved, and has to keep it no longer than necessary.
Some conversations sit in a stricter category again. The GDPR treats data concerning health as a special category, and the regulation's own text sets out both the general prohibition and the narrow conditions under which such processing is permitted, including provisions tied to professional secrecy. In practice this is why a clinic cannot treat a consumer translation app the way a tourist treats it. A patient conversation is health data before it is anything else.
Three practical consequences:
- In an employer's meeting, the employer, not you, is usually the party carrying the compliance obligation. Ask before you introduce a tool that routes colleague and client speech to a company nobody approved. The etiquette and mechanics of doing that well are in translation earbuds for business meetings.
- In a clinical, legal or safeguarding setting, assume a professional interpreter is the correct answer and a consumer device is not, regardless of how well the device performs.
- On a managed work laptop or phone, your IT policy may already prohibit sending audio to unapproved processors. That is a policy question with a real answer, and someone in your organisation knows it.
Is anyone recording the other person?
The person in front of you did not agree to anything, and they may not have noticed a device is involved at all. That is the part of this subject the category avoids.
Two separate questions sit here. First, is a copy being made. That depends on the product, and the vendor questions above are how you find out. Second, is consent required before a copy is made. That depends on where you are standing.
In the United States the rules differ by state, and the split between one-party and all-party consent is exactly why a generic answer is useless. The Reporters Committee for Freedom of the Press maintains a state-by-state guide to the laws governing recording of calls and in-person conversations, which is a reasonable place to start reading before you rely on a rule of thumb. Outside the US, national data protection law and national wiretapping law both apply, and they do not always point the same way.
The behaviour that works in every jurisdiction is much simpler than the law. Say it out loud. "I'm using a translator, it listens while we talk, is that alright?" takes four seconds, and it is also better technique. People speak in cleaner, shorter sentences when they know a machine is in the loop, which improves what comes out the other end. Holding the phone screen where they can see it does the same job without words.
How does our setup differ?
I went back into the code before writing this section, because a privacy claim that has not been checked against the implementation is just a slogan.
AI Translation Live runs in the browser. When you press the button, the page asks our server for a short-lived credential. The server mints one from the translation provider and hands it back. The browser then opens its own encrypted peer connection straight to that provider and streams audio over it. Our server is not a hop on the audio path. It cannot be, because it is never part of that connection, and the long-lived key that would allow anything else never leaves the server.
The running transcript you see on screen arrives over the same direct connection and lives in the browser tab. It is not sent to us, and it is not written to your device. Close the tab and it is gone.
What our server does see, and I would rather list it than let you assume it is nothing:
- The output language you selected, so a session can be opened for it.
- Your licence key, if you have one, so time can be metered against it.
- Your IP address, as a counter that limits free demos. It expires after 24 hours.
- The duration of a session, for billing time against a licence.
- An analytics event recording that a demo started, and in which language, with no identifier attached.
And then the honest caveat, the part that would be missing if this were marketing copy. There is still a third party. The translation provider receives your audio while a session is open, because that is the company doing the translating. Our claim is narrower than "nobody hears you", and it is the one we can stand behind: we are not in the middle of it, we hold no audio, and we hold no transcripts. That is also what our privacy page says, in the same words, because the two should never drift apart.
The rest of the design follows from that. The demo on the home page needs no account and no email address, so there is nothing to collect in the first place. The kit is $69, and AI Translation Live is $39 a year with 5 hours of translation included.
Frequently asked questions
Are translation earbuds recording my conversations?
The earbuds themselves are not. They have no storage and no independent connection, so they cannot keep anything. Whether a recording exists depends on the app or service doing the translating, and that varies by vendor. Cloud-based products send your speech to a server by design, and what happens to it after arrival is a retention policy decision rather than a hardware fact. The only reliable way to know is to read the vendor's written policy and ask directly whether audio and transcripts are retained, for how long, and whether they are used for training.
Is it legal to translate someone without telling them?
It depends entirely on where you are, and nothing here is legal advice. Consent requirements for recording conversations vary by country and, in the United States, by state, with some jurisdictions requiring only one party to consent and others requiring everyone. Data protection law may apply on top of that, particularly in a workplace or a professional setting. Check your own jurisdiction against a primary source, and when in doubt, ask the person. Telling someone you are using a translator is quick, it removes the question entirely, and it tends to improve the conversation.
Does offline translation mean private translation?
It means the audio is not leaving your phone in real time, which is a real and significant difference. It does not automatically mean nothing is stored. An app translating locally can still write transcripts to the device, keep conversation history, sync that history to an account once a connection returns, or send diagnostic data. Offline processing and offline storage are two separate promises, so check that a vendor makes both rather than assuming the second follows from the first.
Can my employer use translation earbuds on calls?
Often yes, and the obligations usually sit with the employer rather than with you. If an organisation routes employee or client speech to an external processor, it generally needs a lawful basis, has to inform the people involved, and has to consider whether the vendor is an approved processor under its own policies. Some sectors and some conversations, particularly anything touching health or legal matters, carry stricter requirements. Ask whoever handles data protection or IT policy where you work before you introduce a tool onto meetings, and expect the answer to differ between a marketing call and a patient appointment.
If you would rather see the architecture than read about it, the demo on the home page runs browser to model. There is no signup, no account, and the audio does not route through us. Speak one sentence, listen to it come back, then go and ask every other vendor on your shortlist the eight questions above.