Skip to main content

Text-to-Speech Models

AI-School supports text-to-speech models that convert text into audio. These models are used in Text to audio on the dashboard and in features that generate audio from a chat.

Current Catalog

ProviderModelNote
OpenAIGPT-4o mini TTSNaturally sounding speech with good control over tone and style.
GoogleGemini 3.1 Flash TTS PreviewNew Gemini speech model with precise control over style, pace, and tone.
European AIVoxtral Mini TTSEuropean text-to-speech based on Mistral Voxtral Mini.

Claude does not have its own text-to-speech model in the catalog. If Claude is enabled as a provider, speech models remain dependent on the other configured providers.

Real-time Voice Chat

Voice chat uses a separate real-time model that combines listening and responding in one live conversation. For OpenAI, AI-School uses GPT-Realtime 2.1 by default. Existing settings with GPT-Realtime 1.5 or GPT-Realtime 2 are automatically updated to this version. Gemini Live remains available as an alternative when Google is configured for real-time speech.

Language Training and Support Language

In a voice assistant, the set language determines the primary language the assistant speaks. The user's interface language remains available as a support language. For example, a learner can practice Spanish and briefly ask for explanations in Dutch. After the explanation, the language trainer switches back to Spanish.

The default language trainers for Dutch, English, German, Spanish, and French are designed for short conversations at A1 level. They ask one question at a time, adjust vocabulary and pace, and only correct the most important mistakes. A custom-made voice assistant uses its own system prompt, language, voice, files, and tool settings in the same way.

Tools During Voice Chat

Voice assistants can, depending on the environment and their settings, automatically use read-only tools for internet information, the manual, previous chats, weather, Wikipedia, and selected educational resources. In the assistant form, you can turn a suitable tool on or off. An enabled tool is set to Automatic: the assistant then decides when use is useful.

What a Speech Model Determines

A speech model determines how text is pronounced and which features are available. Think of:

  • the available voices;
  • the languages a voice supports;
  • the quality and naturalness of the pronunciation;
  • how instructions about pace, tone, accent, and pronunciation are followed.

Voices and Languages

The available voices differ per provider. AI-School shows only voices tested for the chosen language or voices marked as multilingual by the provider in text to audio. If a voice is intended only for certain languages, that language is listed with the voice.

OpenAI and Google support most languages in the catalog. Voxtral Mini TTS can process multiple languages, but the current voice catalog contains tested voices for English and French. Therefore, this model starts by default with a suitable English combination and is not silently linked to Dutch text. If no tested voice is available for a chosen language, AI-School shows a warning and you can choose another language or speech model.

A voice from another language can cause a foreign accent. Such a cross-lingual voice is therefore never automatically used as default or saved preference. Technical integrations can only explicitly request this quality fallback.

System Prompt

In text to audio, the system prompt can be used to guide pronunciation and style. AI-School fills in a language-appropriate base instruction for this. For Dutch, it requests Dutch vowels, stress, rhythm, and intonation without an English accent. Terms like AI, AI-School, ChatGPT, and OpenAI may keep an English pronunciation, and Claude is pronounced as a French name. You can adjust the instruction for pace, tone, or a specific target audience.

Audio Quality and Fallback

Generated audio from OpenAI, Google, and Mistral is checked and stored as a standardized PCM16-WAV file. The technical metadata includes sample rate, number of channels, language, voice, and quality route. This keeps audio quality verifiable and prevents a differing provider response from being silently stored as a different audio format.

The read-aloud button for existing text can also use the system voice of the browser or device. This is visible as fallback quality. The application chooses a system voice that matches the language where possible, but pronunciation and availability may vary per device.

Preferences

Users can save their text-to-audio settings as personal preferences. This way, model, language, voice, and pronunciation instructions do not have to be chosen repeatedly.

WhatsApp