Robustors

TTS (Text to Speech)

Robustors TTS is a self-hosted web studio that turns written scripts into natural-sounding speech.

It supports three ways to generate speech:

  • Auto: the studio picks a voice for you.
  • Design: you describe the voice you want, such as its gender, age, pitch, accent or a whisper, in English or Chinese.
  • Clone: it copies a voice from a saved voice in the shared library or a one-off uploaded clip.

You can fine-tune quality, speed and duration, and choose WAV, MP3, FLAC or OGG output.

Other features

  • Long scripts: very long text is generated piece by piece and joined into a single file, with a live progress bar and a cancel button.
  • Pronunciation dictionary: a shared list of custom pronunciations is applied to every request, so tricky words and acronyms come out right.
  • Voice Library: a shared, taggable set of reference voices.
  • History: each user sees their last 50 generations, with playback, full text and reference audio.
  • Job queue: a priority queue for batch work, with live position and time estimates.
  • Accounts and admin: secure sign-in for each team member. Admins can see and end user sessions and cancel queued jobs with a reason, and the job’s owner is notified.

API

Everything the Studio can do is also available through an HTTP API, so scripts, bots and other apps can generate speech, clone voices, manage the voice library and history, and submit large batches. Programs can check back for finished audio or have it delivered to them when it’s ready.

Output

0:00 / –:––
Generated speech sampleDownload WAV, 1.3 MB