Overview
F5-TTS is an open-source text-to-speech model that generates speech in a target voice from a short reference audio clip plus its transcript. It is built on a Diffusion Transformer with ConvNeXt V2 and uses flow matching rather than a traditional autoregressive decoder.
It is aimed at developers and researchers who want voice cloning and zero-shot TTS they can run locally. You can use it through a pip package for inference, a command-line tool, or a Gradio web interface, and the repo also ships training and finetuning code for those who want to go further.
Within speech and audio tooling, F5-TTS sits in the voice-cloning and synthesis space. The project also includes E2 TTS (a reproduction of the E2 TTS paper) and a Sway Sampling inference strategy, plus a Triton/TensorRT-LLM runtime for higher-throughput serving.
What it does
- Clones a voice from a short reference audio clip and its transcript
- Diffusion Transformer (DiT) with ConvNeXt V2, trained with flow matching
- Sway Sampling, an inference-time flow-step strategy to improve output
- CLI inference tool (f5-tts_infer-cli) with reference audio and generation text flags
- Gradio web app with basic TTS, multi-style/multi-speaker, and voice chat
- Local-editable install for training and finetuning, plus a Triton/TensorRT-LLM runtime
Getting started
Set up a Python 3.10+ environment with PyTorch for your device, then install F5-TTS and run inference from the CLI or the Gradio web app.
Create an environment
Use Python 3.10 or newer (conda or virtualenv). Install FFmpeg as well.
conda create -n f5-tts python=3.11
conda activate f5-tts
conda install ffmpegInstall PyTorch and F5-TTS
Install PyTorch matched to your device first (the README lists NVIDIA, AMD, Intel, and Apple Silicon variants), then install the pip package for inference.
pip install f5-ttsRun CLI inference
Provide a reference clip, its transcript, and the text to generate. Leaving --ref_text empty lets an ASR model transcribe it (uses extra GPU memory).
f5-tts_infer-cli --model F5TTS_v1_Base \
--ref_audio "provide_prompt_wav_path_here.wav" \
--ref_text "The content, subtitle or transcription of reference audio." \
--gen_text "Some text you want TTS model generate for you."Or launch the web app
Start the Gradio interface for basic TTS, multi-speaker generation, and voice chat in the browser.
f5-tts_infer-gradioCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Clone a specific voice from a short sample to narrate scripts or articles
- Generate multi-speaker or multi-style dialogue for prototypes and demos
- Build a local voice chat assistant using the bundled Gradio app
- Finetune or train the model on your own data via the editable install
How F5-TTS compares
F5-TTS alongside other open-source audio, music & voice tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Whisper | ★ 110k | OpenAI's speech recognition model that transcribes and translates audio across many languages. |
| GPT-SoVITS | ★ 62.3k | An open-source WebUI that clones a voice from a short audio sample and turns text into speech, with zero-shot and few-shot fine-tuning. |
| Voicebox | ★ 56.3k | Local-first voice studio that clones a voice from a short sample, generates speech across seven TTS engines and 23 languages, handles system-wide dictation, and speaks for agents over MCP. |
| VibeVoice | ★ 54.6k | Microsoft's text-to-speech model for generating long, expressive multi-speaker audio like podcasts. |
| whisper.cpp | ★ 54.1k | A dependency-free C/C++ port of Whisper built on ggml, running speech recognition on CPU, Metal, CUDA, Vulkan and NPUs from phones to servers. |
| VoiceStudio | ★ 52.5k | A local-first desktop voice studio that clones and designs voices, dubs video, and handles dictation, running 16 text-to-speech and 11 speech-recognition engines on your own hardware. |
| Coqui TTS | ★ 46.1k | A library of text-to-speech models including the multilingual XTTS voice-cloning model. |
| F5-TTS | ★ 15.3k | Flow-matching text-to-speech that clones a voice from a short reference clip |
