Voices

Text-to-Speech in Other Languages: What Actually Works

Voice quality varies enormously by language. Here's what to expect in Spanish, French, Portuguese, Italian, Japanese and beyond — and how to handle mixed-language documents.

Key takeaways

  • Voice quality tracks training data, so major European languages and Mandarin are excellent while smaller languages and regional variants lag behind.
  • Set the document's language explicitly — most mispronunciation in non-English text comes from the app applying English pronunciation rules.
  • Regional variety matters more than most people expect: European vs Latin American Spanish, and European vs Brazilian Portuguese, are different listening experiences.
  • Mixed-language documents are the hardest case; most systems don't switch mid-sentence, so expect quoted foreign phrases to be mangled.

Most writing about text-to-speech quietly assumes English. If you read in Spanish, Portuguese, Japanese or anything else, the experience is different — sometimes just as good, sometimes noticeably worse, and the reasons are predictable enough to plan around.

Quality follows training data

The blunt version: a language’s voice quality is roughly proportional to how much of it exists in digital form. Neural speech models learn from large corpora of recorded speech paired with text. Languages with enormous digital footprints get excellent voices; languages without them get whatever can be built from less.

That produces a rough tiering:

Excellent. English, Spanish, French, German, Italian, Portuguese, Mandarin Chinese, Japanese, Korean, Russian. Multiple natural neural voices, good prosody, several regional variants. Listening for an hour is comfortable.

Good. Dutch, Polish, Turkish, Arabic, Hindi, Swedish, Danish, Norwegian, Czech, Indonesian, Vietnamese, Thai. Solid neural voices, usually fewer of them, occasionally stiffer intonation on long sentences.

Variable. Smaller European languages, most African languages, regional South Asian languages, and minority or indigenous languages. You may find one voice, an older non-neural voice, or nothing.

Market forecasts consistently name multilingual capability as one of the strongest growth areas in speech synthesis, and coverage has improved fast — but the gap between tier one and tier three is still real, and no amount of marketing copy about “60+ languages” tells you which tier yours is in. Test with your own text before paying.

Set the language explicitly

This is the single most common fixable problem, and it accounts for most complaints about non-English text-to-speech being bad.

If Spanish text is being read with English vowels, the app is applying English grapheme-to-phoneme rules — the language setting is wrong, not the voice. Auto-detection works well on a long monolingual document and badly on short passages, proper nouns and anything mixed.

Find the language or voice setting and set it manually per document. It takes five seconds and it’s the difference between “this is unusable” and “this is fine.”

Regional variants matter more than you’d think

Choosing “Spanish” isn’t choosing a voice; it’s choosing a family. The differences are large enough to affect comfort over a long listen:

Spanish. Castilian (es-ES) uses the distinciónc before e/i and z pronounced as a soft th. Latin American variants (es-MX, es-AR, es-CO and others) don’t, and differ from each other in intonation and rhythm. Argentine Spanish has a distinctly different melody and the sh sound for ll and y. If you learned one, another will feel subtly wrong for hours before you work out why.

Portuguese. European (pt-PT) and Brazilian (pt-BR) Portuguese differ substantially in vowel reduction, rhythm and stress. Brazilian is more open and evenly stressed; European compresses unstressed vowels heavily. These are not interchangeable for comfortable listening.

French. Metropolitan (fr-FR) and Canadian (fr-CA) differ in vowels, intonation and some vocabulary.

English. The same applies — US, UK, Australian, Indian and Irish voices are all common, and it’s worth picking deliberately rather than defaulting.

Chinese. Mandarin and Cantonese are different languages, not accents. Simplified vs traditional characters is a separate axis again.

💡 Try the same paragraph in two or three variants of your language before settling. Ten minutes now saves a lot of low-grade irritation later.

Language-specific things that trip systems up

Japanese. The hard problem is that kanji have multiple readings depending on context — the same character is pronounced differently in different compounds, and proper nouns are notoriously unpredictable even for native speakers. Good Japanese TTS is genuinely impressive engineering; expect occasional wrong readings on names and unusual compounds. Pitch accent is another axis most systems handle approximately.

Arabic and Hebrew. Written without short vowels in normal text, so the system must infer vocalisation from context. Modern Standard Arabic is well supported; dialects much less so.

Tonal languages (Mandarin, Vietnamese, Thai). Tone is phonemic, so errors change meaning rather than just sounding odd. Quality here has improved a lot but is worth testing carefully.

German. Compound nouns can be arbitrarily long and novel, requiring the system to decompose words it has never seen. Usually handled well; occasionally not.

Languages with heavy inflection (Finnish, Hungarian, Polish, Turkish, Russian). Stress placement often depends on morphology, and errors sound distinctly foreign to a native ear.

The general mechanics of why any of this goes wrong are in why text-to-speech mispronounces words.

Mixed-language documents: the genuinely hard case

Academic papers with foreign-language quotations. Language-learning material with translations. Bilingual reports. Recipes with untranslated terms. CVs. Menus.

Most systems detect one language per document or per block and apply it throughout. A French phrase quoted inside an English paragraph gets English pronunciation rules, and vice versa. Some tools detect per paragraph, which helps if your document has cleanly separated sections and doesn’t help at all with inline switching.

Practical workarounds:

  • Split the document by language and import each part separately. Tedious, but it works and it’s the only reliable fix.
  • Accept the mangling for short inline phrases. A mispronounced quoted phrase in an otherwise clear paragraph is survivable.
  • Use a pronunciation dictionary if your app has one, for recurring foreign terms and names.
  • For language learning, this is actually less of a problem than it sounds, because you generally want each language read by its own voice anyway — so split by design.

For language learners specifically

Text-to-speech is a genuinely useful language-learning tool, with one caveat worth stating.

What it’s good for: hearing how written text sounds, building listening comprehension at a controllable speed, reading along with unfamiliar text so your eyes and ears learn the mapping together, and getting audio for material that has none.

The caveat: a synthesised voice is clear, consistent and neutral. Real speech is fast, sloppy, elided and full of regional variation. Synthetic audio is excellent scaffolding and a poor final destination — pair it with real recorded speech as you progress.

Settings that help: slow the speed to 0.8–0.9× when the language is new; turn on word highlighting so you’re reading and listening simultaneously (the bimodal reading effect is especially strong for learners); and use short passages repeatedly rather than long ones once.

Text-to-speech for language learning covers the full method.

Before you commit to a tool

  1. Take a real paragraph in your language — ideally one with names, numbers and a long sentence.
  2. Run it through each candidate’s free tier.
  3. Check the regional variants available, not just the language.
  4. Test a mixed-language passage if you read those.
  5. Listen for five minutes, not five seconds. Fatigue shows up over time, not in a demo clip.

Frateca reads documents, PDFs, articles and pasted text across languages with natural voices, and the site itself is available in Spanish, French, Italian, Portuguese and Japanese. Try it free with a paragraph in your language — free plan, no credit card required.

Stop reading. Start listening.

Frateca turns PDFs, articles, textbooks and web pages into natural audio you can play anywhere — on your commute, at the gym, or while you cook. Free plan included, no card required.

Try Frateca free

iOS · Android · Web · Free plan, no credit card required