Skip to main content

Text-to-Speech Models

AI-Public supports text-to-speech models that convert text into audio. These models are used in Text to Audio on the dashboard and in features that generate audio from a chat.

Current Catalog

ProviderModelNote
OpenAIGPT-4o mini TTSNaturally sounding speech with good control over tone and style.
GoogleGemini 3.1 Flash TTS PreviewNew Gemini speech model with precise control over style, pace, and tone.
European AIVoxtral Mini TTSEuropean text-to-speech based on Mistral Voxtral Mini.

Claude does not have its own text-to-speech model in the catalog. If Claude is enabled as a provider, speech models depend on the other configured providers.

Real-time Voice Chat

Voice chat uses a separate real-time model that combines listening and responding in one live conversation. For OpenAI, AI-Public uses GPT-Realtime 2.1 by default. Existing settings with GPT-Realtime 1.5 or GPT-Realtime 2 are automatically updated to this version. Gemini Live remains available as an alternative when Google is configured for real-time speech.

Primary Language and Support Language

For a voice assistant, the set language determines the main language the assistant speaks. The user's interface language remains available as a support language. For example, someone can practice a conversation in Spanish and briefly ask for explanations in their own language. Afterwards, the assistant switches back to Spanish.

A custom voice assistant loads its full configuration, including the system prompt, language, voice, files, and tool settings. You can use the system prompt to specify the conversation task, desired level, and feedback style.

Tools During Voice Chat

Voice assistants can, depending on the environment and their settings, automatically use read-only tools for internet information, the manual, previous chats, weather, and Wikipedia. In the assistant form, you can enable or disable a suitable tool. An enabled tool is set to Automatic: the assistant then decides when use is beneficial.

What a Speech Model Determines

A speech model determines how text is pronounced and which features are available. Consider:

  • the available voices;
  • the languages a voice supports;
  • the quality and naturalness of the pronunciation;
  • how instructions about pace, tone, accent, and pronunciation are followed.

Voices and Languages

Available voices differ per provider. AI-Public shows only voices tested for the chosen language or voices marked as multilingual by the provider in text to audio. If a voice is intended only for certain languages, that language is listed with the voice.

OpenAI and Google support most languages in the catalog. Voxtral Mini TTS can process multiple languages, but the current voice catalog contains tested voices for English and French. Therefore, this model starts by default with a suitable English combination and is not silently linked to Dutch text. If no tested voice is available for a chosen language, AI-Public shows a warning and you must choose another language or speech model.

A voice from another language can cause a foreign accent. Such a cross-lingual voice is therefore never automatically used as default or saved preference. Technical integrations can explicitly request this quality fallback.

System Prompt

In text to audio, the system prompt can be used to guide pronunciation and style. AI-Public fills in a language-appropriate base instruction. For Dutch, this requests Dutch vowels, stress, rhythm, and intonation without an English accent. Terms like AI, AI-Public, ChatGPT, and OpenAI may keep an English pronunciation, and Claude is pronounced as a French name. You can adjust the instruction for pace, tone, or a specific target audience.

Audio Quality and Fallback

Generated audio from OpenAI, Google, and Mistral is checked and stored as a standardized PCM16-WAV file. The technical metadata includes sample rate, number of channels, language, voice, and quality route. This keeps audio quality verifiable and prevents a differing provider response from being silently stored as a different audio format.

The read-aloud button for existing text can also use the system voice of the browser or device. This is visible as fallback quality. The application chooses a system voice matching the language where possible, but pronunciation and availability may vary per device.

Preferences

Users can save their text-to-audio settings as personal preferences. This way, model, language, voice, and pronunciation instructions do not have to be chosen repeatedly.

WhatsApp