Speech synthesis models
The text-to-speech model on Sawala Cloud, what it costs, and the languages it does — and does not — speak.
Speech synthesis turns text into spoken audio. On Sawala Cloud this happens inside a flow: a voice step takes the text you give it — a reply your assistant wrote, a notification, a summary — and returns an audio clip the flow can send as a voice message or store as a file.
Synthesis is billed by minutes of audio produced: a longer reply makes a longer clip and costs more.
MeloTTS
Made by MyShell, an AI studio that released MeloTTS as an open-source text-to-speech model.
It is the platform's voice-synthesis model, it runs on Sawala's infrastructure, and it produces natural, clearly-paced speech.
MeloTTS does not speak Bahasa Indonesia. Its languages are English, Spanish, French, Chinese, Japanese, and Korean. There is no Indonesian-speaking voice model on Sawala Cloud today, so an Indonesian voice reply is not something you can build on the platform yet. If your flow needs to speak Indonesian, send text instead.
Available in Flow only — chat surfaces never transcribe, draw, or speak, so this model does not appear in the Crew or Connect model pickers.
Back to the AI models overview.
Model list last verified: 2026-08-07.
Image generation models
The three text-to-image models on Sawala Cloud — who makes them, what drives their cost, and which one to start with.
API keys & access
How Sawala Cloud authenticates access — public vs secret API keys, project scope, the public API base URLs, and CLI tokens for the command-line tools.