Voices
What each backend can actually offer. The pipeline reads these capabilities directly — it never branches on a provider's name.
gttsdefault
Google Translate's speech endpoint. Free and dependable, but it exposes no voice selection and emits no word timings — so the pipeline runs speech recognition to get its clock.
Languages11
Max charactersunbounded
Cost / 1k chars$0.00
No voice selection
This backend declares an empty voice set — one voice per language, and no way to choose. The only lever is the accent, set with the tld option in Settings.
enesfrdeitptnlhijakozh
elevenlabsword timings
Returns character-level timings alongside the audio, so the whole speech-recognition stage is skipped. Costs money and needs an API key.
Languages11
Max characters5000
Cost / 1k chars$0.30