█

AI/TLDR

VoiceStudio

Clone voices, dub video and dictate entirely on your own hardware

Audio, Music & VoiceOpen source
Language
Python
License
AGPL-3.0

Overview

VoiceStudio (previously OmniVoice-Studio) is a desktop application that bundles voice cloning, voice design, video dubbing, dictation and long-form audio production into one local workflow. Rather than shipping a single model, it acts as a front end over an engine catalogue — 16 text-to-speech engines and 11 speech-recognition engines, with a 646-language TTS catalogue whose real coverage and quality depend on which engine you pick. Engines are swapped from the Model Catalogue or with a keyboard shortcut, so a project can move between a fast small model and a heavier expressive one without leaving the app.

The project is local-first: voices, projects, settings and outputs stay on the machine by default, and the core creation path needs no account, API key or usage meter. Network-backed features are explicit opt-ins. Compute routing is automatic across CUDA, Apple Silicon MPS/MLX, ROCm on Linux and plain CPU, with optional remote workers for machines that cannot run the backend locally. Packages exist for macOS 13.3+ on Apple Silicon, Windows 10/11 x64, Linux x86_64 with glibc 2.39+, and Docker (CUDA, ROCm, CPU and worker-only profiles).

Beyond the desktop UI, VoiceStudio exposes a local REST/SSE/WebSocket API, an OpenAI-compatible audio API and an MCP server for synthesis and transcription, so it can be driven by scripts and agents as well as by hand. The application is AGPL-3.0; models you download keep their own upstream licence terms. The project describes itself as an active beta and points users at tagged releases for stable work.

What it does

  • Zero-shot voice cloning from a short reference clip, plus voice design from age, accent, pitch, style and delivery instructions
  • Video dubbing that transcribes, translates, preserves speakers, synthesizes and exports the finished video
  • Stories and audiobooks: multi-voice scripts, EPUB/PDF import, chapter rendering and .m4b export
  • A dictation widget bound to a system-wide shortcut, with live transcription and optional local-LLM cleanup
  • Supporting audio tools — Demucs vocal isolation, Pyannote/WhisperX speaker diarization, a batch job queue and AudioSeal watermark embedding and detection
  • Automation surfaces: a local REST/SSE/WebSocket API, an OpenAI-compatible audio API and an MCP server
  • GPU auto-detect across CUDA, MPS, ROCm and CPU, with per-engine checks and optional remote workers

Getting started

The normal path is a packaged download rather than a source build. The first launch creates a managed Python environment and downloads the default model; later launches reuse both. On macOS the first launch needs a one-time right-click, then Open.

Install a packaged build

Grab the package for your platform from the latest release: an Apple Silicon DMG for macOS 13.3+, an x64 MSI for Windows 10/11 (pick the current-user build to install without admin rights), or an AppImage for Linux x86_64 with glibc 2.39+.

texttext
https://github.com/debpalash/VoiceStudio/releases/latest

Clone your first voice

Open Voice Cloning, add a clean voice sample — three seconds works, 5 to 15 seconds usually gives a better prompt — then enter text, choose a language and select Generate.

Switch engines when the default doesn't fit

Engines are installed, removed, selected and routed from the Model Catalogue, which shows engine, device and install state. Ctrl/Cmd+E switches engines straight from the status bar.

Run from source instead

After installing the development prerequisites from CONTRIBUTING.md, the repo builds with bun. Use `bun run dev` for the browser UI.

bashbash
git clone https://github.com/debpalash/VoiceStudio.git
cd VoiceStudio
bun install
bun run desktop

Diagnose a failed setup

The app ships self-checks; run them from Settings → About → Run self-check, or from the command line. A scrubbed diagnostic bundle can be saved from the app when opening an issue.

bashbash
uv run python backend/main.py --diagnose --deep

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Produce cloned-voice narration for audiobooks or long-form video without sending scripts and reference audio to a hosted provider
  • Dub an existing video into another language while keeping the original speaker layout
  • Add system-wide local dictation with transcription that never leaves the machine
  • Serve synthesis and transcription to your own scripts or agents through the local REST, OpenAI-compatible or MCP endpoints

How VoiceStudio compares

VoiceStudio alongside other open-source audio, music & voice tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Whisper★ 110kOpenAI's speech recognition model that transcribes and translates audio across many languages.
GPT-SoVITS★ 62.2kAn open-source WebUI that clones a voice from a short audio sample and turns text into speech, with zero-shot and few-shot fine-tuning.
Voicebox★ 56.1kLocal-first voice studio that clones a voice from a short sample, generates speech across seven TTS engines and 23 languages, handles system-wide dictation, and speaks for agents over MCP.
VibeVoice★ 54.6kMicrosoft's text-to-speech model for generating long, expressive multi-speaker audio like podcasts.
whisper.cpp★ 54.1kA dependency-free C/C++ port of Whisper built on ggml, running speech recognition on CPU, Metal, CUDA, Vulkan and NPUs from phones to servers.
VoiceStudio★ 51.5kClone voices, dub video and dictate entirely on your own hardware
Coqui TTS★ 46.1kA library of text-to-speech models including the multilingual XTTS voice-cloning model.
ChatTTS★ 39.9kChatTTS is an open-source text-to-speech model tuned for dialogue, with multi-speaker support and fine-grained control over laughter, pauses, and prosody.