█

AI/TLDR

F5-TTS

Flow-matching text-to-speech that clones a voice from a short reference clip

Audio, Music & VoiceOpen source
Language
Python
$conda create -n f5-tts python=3.11

Overview

F5-TTS is an open-source text-to-speech model that generates speech in a target voice from a short reference audio clip plus its transcript. It is built on a Diffusion Transformer with ConvNeXt V2 and uses flow matching rather than a traditional autoregressive decoder.

It is aimed at developers and researchers who want voice cloning and zero-shot TTS they can run locally. You can use it through a pip package for inference, a command-line tool, or a Gradio web interface, and the repo also ships training and finetuning code for those who want to go further.

Within speech and audio tooling, F5-TTS sits in the voice-cloning and synthesis space. The project also includes E2 TTS (a reproduction of the E2 TTS paper) and a Sway Sampling inference strategy, plus a Triton/TensorRT-LLM runtime for higher-throughput serving.

What it does

  • Clones a voice from a short reference audio clip and its transcript
  • Diffusion Transformer (DiT) with ConvNeXt V2, trained with flow matching
  • Sway Sampling, an inference-time flow-step strategy to improve output
  • CLI inference tool (f5-tts_infer-cli) with reference audio and generation text flags
  • Gradio web app with basic TTS, multi-style/multi-speaker, and voice chat
  • Local-editable install for training and finetuning, plus a Triton/TensorRT-LLM runtime

Getting started

Set up a Python 3.10+ environment with PyTorch for your device, then install F5-TTS and run inference from the CLI or the Gradio web app.

Create an environment

Use Python 3.10 or newer (conda or virtualenv). Install FFmpeg as well.

bashbash
conda create -n f5-tts python=3.11
conda activate f5-tts
conda install ffmpeg

Install PyTorch and F5-TTS

Install PyTorch matched to your device first (the README lists NVIDIA, AMD, Intel, and Apple Silicon variants), then install the pip package for inference.

bashbash
pip install f5-tts

Run CLI inference

Provide a reference clip, its transcript, and the text to generate. Leaving --ref_text empty lets an ASR model transcribe it (uses extra GPU memory).

bashbash
f5-tts_infer-cli --model F5TTS_v1_Base \
--ref_audio "provide_prompt_wav_path_here.wav" \
--ref_text "The content, subtitle or transcription of reference audio." \
--gen_text "Some text you want TTS model generate for you."

Or launch the web app

Start the Gradio interface for basic TTS, multi-speaker generation, and voice chat in the browser.

bashbash
f5-tts_infer-gradio

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Clone a specific voice from a short sample to narrate scripts or articles
  • Generate multi-speaker or multi-style dialogue for prototypes and demos
  • Build a local voice chat assistant using the bundled Gradio app
  • Finetune or train the model on your own data via the editable install

How F5-TTS compares

F5-TTS alongside other open-source audio, music & voice tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Whisper★ 110kOpenAI's speech recognition model that transcribes and translates audio across many languages.
GPT-SoVITS★ 62.3kAn open-source WebUI that clones a voice from a short audio sample and turns text into speech, with zero-shot and few-shot fine-tuning.
Voicebox★ 56.3kLocal-first voice studio that clones a voice from a short sample, generates speech across seven TTS engines and 23 languages, handles system-wide dictation, and speaks for agents over MCP.
VibeVoice★ 54.6kMicrosoft's text-to-speech model for generating long, expressive multi-speaker audio like podcasts.
whisper.cpp★ 54.1kA dependency-free C/C++ port of Whisper built on ggml, running speech recognition on CPU, Metal, CUDA, Vulkan and NPUs from phones to servers.
VoiceStudio★ 52.5kA local-first desktop voice studio that clones and designs voices, dubs video, and handles dictation, running 16 text-to-speech and 11 speech-recognition engines on your own hardware.
Coqui TTS★ 46.1kA library of text-to-speech models including the multilingual XTTS voice-cloning model.
F5-TTS★ 15.3kFlow-matching text-to-speech that clones a voice from a short reference clip