Self-hosted audiobook studio

Voices

Cloning voices for GPT-SoVITS, F5-TTS and Omni/IndexTTS-2 — a clip uploaded here is offered by all of them, whatever the engine label says. Upload a clean voice clip of any length — the app automatically extracts a ~9 second window (from the start time you choose), since the engines need a 3–10s reference. Provide the transcript for that window.

Getting good quality

  • Clean reference audio: one speaker, no music/background noise, no reverb. A studio/voiceover clip works far better than a noisy phone recording — noise is the main cause of a "tinny/mechanical" clone.
  • Transcript must exactly match the ~9s window (the part starting at your chosen start time), including punctuation. A mismatch makes it sound like a different/"off" person.
  • Pick a window where the speaker talks naturally and clearly with normal pacing.
  • Words randomly dragging out? That's a GPT-SoVITS timing artifact — the server defaults now reduce it; you can tune further via the SOVITS_* env vars (see README).

Add a voice

The ~9s reference window starts here. Use it to skip silence/noise at the beginning.

Existing voices

Achernar en · Omni
no transcript
ready
Aoede en · Omni
no transcript
ready
Autonoe en · Omni
no transcript
ready
Callirrhoe en · Omni
no transcript
ready
Charon en · Omni
no transcript
ready
Fenrir en · Omni
no transcript
ready
Kore en · Omni
no transcript
ready
Leda en · Omni
no transcript
ready
Orus en · Omni
no transcript
ready
Puck en · Omni
no transcript
ready
Sadie en · GPT-SoVITS
Hello, and thanks for listening. Today, I'd like to share a few thoughts about curiosity and the small habits that shape our everyday lives.
ready
sadie-excited en · GPT-SoVITS
Hello, and thanks for listening. Today, I'd like to share a few thoughts about curiosity and the small habits that shape our everyday lives.
ready
Zephyr en · Omni
no transcript
ready