Overview
VoiceStudio (previously OmniVoice-Studio) is a desktop application that bundles voice cloning, voice design, video dubbing, dictation and long-form audio production into one local workflow. Rather than shipping a single model, it acts as a front end over an engine catalogue — 16 text-to-speech engines and 11 speech-recognition engines, with a 646-language TTS catalogue whose real coverage and quality depend on which engine you pick. Engines are swapped from the Model Catalogue or with a keyboard shortcut, so a project can move between a fast small model and a heavier expressive one without leaving the app.
The project is local-first: voices, projects, settings and outputs stay on the machine by default, and the core creation path needs no account, API key or usage meter. Network-backed features are explicit opt-ins. Compute routing is automatic across CUDA, Apple Silicon MPS/MLX, ROCm on Linux and plain CPU, with optional remote workers for machines that cannot run the backend locally. Packages exist for macOS 13.3+ on Apple Silicon, Windows 10/11 x64, Linux x86_64 with glibc 2.39+, and Docker (CUDA, ROCm, CPU and worker-only profiles).
Beyond the desktop UI, VoiceStudio exposes a local REST/SSE/WebSocket API, an OpenAI-compatible audio API and an MCP server for synthesis and transcription, so it can be driven by scripts and agents as well as by hand. The application is AGPL-3.0; models you download keep their own upstream licence terms. The project describes itself as an active beta and points users at tagged releases for stable work.
What it does
- Zero-shot voice cloning from a short reference clip, plus voice design from age, accent, pitch, style and delivery instructions
- Video dubbing that transcribes, translates, preserves speakers, synthesizes and exports the finished video
- Stories and audiobooks: multi-voice scripts, EPUB/PDF import, chapter rendering and .m4b export
- A dictation widget bound to a system-wide shortcut, with live transcription and optional local-LLM cleanup
- Supporting audio tools — Demucs vocal isolation, Pyannote/WhisperX speaker diarization, a batch job queue and AudioSeal watermark embedding and detection
- Automation surfaces: a local REST/SSE/WebSocket API, an OpenAI-compatible audio API and an MCP server
- GPU auto-detect across CUDA, MPS, ROCm and CPU, with per-engine checks and optional remote workers
Getting started
The normal path is a packaged download rather than a source build. The first launch creates a managed Python environment and downloads the default model; later launches reuse both. On macOS the first launch needs a one-time right-click, then Open.
Install a packaged build
Grab the package for your platform from the latest release: an Apple Silicon DMG for macOS 13.3+, an x64 MSI for Windows 10/11 (pick the current-user build to install without admin rights), or an AppImage for Linux x86_64 with glibc 2.39+.
https://github.com/debpalash/VoiceStudio/releases/latestClone your first voice
Open Voice Cloning, add a clean voice sample — three seconds works, 5 to 15 seconds usually gives a better prompt — then enter text, choose a language and select Generate.
Switch engines when the default doesn't fit
Engines are installed, removed, selected and routed from the Model Catalogue, which shows engine, device and install state. Ctrl/Cmd+E switches engines straight from the status bar.
Run from source instead
After installing the development prerequisites from CONTRIBUTING.md, the repo builds with bun. Use `bun run dev` for the browser UI.
git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktopDiagnose a failed setup
The app ships self-checks; run them from Settings → About → Run self-check, or from the command line. A scrubbed diagnostic bundle can be saved from the app when opening an issue.
uv run python backend/main.py --diagnose --deepCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Produce cloned-voice narration for audiobooks or long-form video without sending scripts and reference audio to a hosted provider
- Dub an existing video into another language while keeping the original speaker layout
- Add system-wide local dictation with transcription that never leaves the machine
- Serve synthesis and transcription to your own scripts or agents through the local REST, OpenAI-compatible or MCP endpoints
How VoiceStudio compares
VoiceStudio alongside other open-source audio, music & voice tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Whisper | ★ 110k | OpenAI's speech recognition model that transcribes and translates audio across many languages. |
| GPT-SoVITS | ★ 62.2k | An open-source WebUI that clones a voice from a short audio sample and turns text into speech, with zero-shot and few-shot fine-tuning. |
| Voicebox | ★ 56.1k | Local-first voice studio that clones a voice from a short sample, generates speech across seven TTS engines and 23 languages, handles system-wide dictation, and speaks for agents over MCP. |
| VibeVoice | ★ 54.6k | Microsoft's text-to-speech model for generating long, expressive multi-speaker audio like podcasts. |
| whisper.cpp | ★ 54.1k | A dependency-free C/C++ port of Whisper built on ggml, running speech recognition on CPU, Metal, CUDA, Vulkan and NPUs from phones to servers. |
| VoiceStudio | ★ 51.5k | Clone voices, dub video and dictate entirely on your own hardware |
| Coqui TTS | ★ 46.1k | A library of text-to-speech models including the multilingual XTTS voice-cloning model. |
| ChatTTS | ★ 39.9k | ChatTTS is an open-source text-to-speech model tuned for dialogue, with multi-speaker support and fine-grained control over laughter, pauses, and prosody. |