Voiceover Studio is a self-hosted web app that turns a raw narration recording into a finished documentary-style mix.
You give it the voice and, if you want, a music track. It fixes the pacing of the pauses, keeps the music playing quietly under the speech, raises the music during meaningful silences, and exports a mastered MP3 or WAV. Your recordings stay on your own machine.
How it works
- Upload the voiceover (MP3, WAV, M4A, FLAC or OGG; up to 4 GB and 6 hours).
- Choose music: upload a track, reuse a saved one, generate one with AI, or pick No music.
- Pick how loud the music is while you speak: Soft, Slightly audible or More audible. Two sliders let you fine-tune the music level under speech and in pauses.
- Click Create documentary mix, compare Polished against Original, and download.
Jobs keep running in the background, so refreshing the browser doesn’t stop them. Projects save automatically.
Documentary pacing
- Finds the natural breaks in the narration and gives the important ones room to breathe, the way a documentary editor would.
- Leaves uncertain spots, lists and numbers exactly as recorded.
- Works in English, Urdu and Hindi.
Mixing and mastering
- Keeps a steady music bed under speech and lifts it smoothly in the longer pauses, without pumping on breaths.
- Lines up pauses with the music so the voice comes back in on the beat.
- Loops the music, adds fades, cleans up the voice and masters the result to a consistent, broadcast-ready loudness.
AI sound designer (optional)
- AI pause editor: reads the story, finds its key moments (hook, reveal, time or place jump, close) and places dramatic holds there. If it can’t, the automatic pauses are used and a note explains why.
- “Pauses in this mix” report: lists every changed pause with the words on either side, its old and new length, and why it changed. Click a row to hear it.
- AI music: composes an original instrumental track sized to your edited voiceover.
- Music styles: pick a direction for the score, such as a tense investigation or a classic documentary feel.
- Long-form mode for recordings over 10 minutes: splits the story into chapters and scores it with a set of recurring themes that crossfade between sections.
Music from a prompt
- A separate Music page for making tracks directly from a description.
- Settings: duration (5–600 s), BPM, key, time signature, 1–4 variants, and optional lyrics with a vocal language.
- Regenerate the same take or a variation of it, or press Surprise me.
- Results go to your personal music library, and you can put any of them into a project.
Render on Server or on Device
- Server: the normal render on the studio machine.
- Device: the final render runs in your browser instead. Both give the same result.
Advanced editor
- Edit or lock individual pauses, supply a script, and change the language.
- Waveform view, undo/redo, and a mode for recordings that already have voice and music combined (pause edits only).
- Export a project and summary as a ZIP.