█

AI/TLDR

DiffSinger

Singing voice synthesis and text-to-speech via a shallow diffusion mechanism

Audio, Music & VoiceOpen source
Language
Python
License
MIT

Overview

DiffSinger is the official PyTorch implementation of the AAAI 2022 paper "DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism". The repository covers two systems from that paper: DiffSinger for singing voice synthesis (SVS), which turns lyrics plus pitch or MIDI information into a sung recording, and DiffSpeech for ordinary text-to-speech (TTS).

Stacked mel-spectrograms of the same utterance: DiffSpeech output on the lower half (bins 0-80) and FastSpeech 2 output on the upper half (bins 80-160)
DiffSpeech (bottom) against FastSpeech 2 (top) on the same sentence, from the project's mel visualisation.DiffSinger TTS docs ↗

Every pipeline follows the same shape: a frontend turns lyrics or text into a linguistic representation, a diffusion acoustic model produces a mel-spectrogram, and a vocoder turns that spectrogram into a waveform (HiFiGAN for speech, an NSF-HiFiGAN singing vocoder for SVS). In the shallow-diffusion variants, a pre-trained FastSpeech 2 or FFT-Singer checkpoint supplies the starting point the diffusion model refines. The repository ships several recipes: ground-truth-F0 singing on the PopCS dataset, and two MIDI-driven versions on the Opencpop dataset — version A predicts the F0 curve explicitly in a melody frontend, while version B drops explicit F0 prediction and lets the diffusion model predict pitch implicitly together with the spectrogram, which the docs say gives more natural pitch and a simpler pipeline.

A plug-in PNDM sampler accelerates inference: the model is trained with 1000 diffusion steps and, with the default `pndm_speedup` of 40, samples in 25 steps, with the speed-up adjustable from the command line. Pre-trained acoustic models and vocoders are published as GitHub release downloads, and the README links interactive SVS and TTS demos on Hugging Face Spaces. The README also thanks Team OpenVPI for their maintenance of DiffSinger at openvpi/DiffSinger.

What it does

  • Singing voice synthesis from lyrics plus ground-truth F0 (PopCS) or lyrics plus MIDI notes (Opencpop), and DiffSpeech text-to-speech on LJSpeech
  • Shallow diffusion: the diffusion acoustic model starts from a pre-trained FastSpeech 2 / FFT-Singer checkpoint rather than pure noise
  • MIDI version B predicts the F0 curve implicitly together with the mel-spectrogram, with a pitch extractor feeding the vocoder
  • Plug-in PNDM (PLMS) acceleration — 1000 training steps sampled in 25 steps by default, tunable via `--hparams="pndm_speedup=..."`
  • Pre-trained checkpoints for DiffSinger, DiffSpeech, FastSpeech 2 and an NSF HiFiGAN-Singing vocoder the docs describe as trained on ~70 hours of singing data
  • One entry point (`tasks/run.py --config ... --exp_name ...`) for binarized-data training, test-set inference and TensorBoard logging, plus a raw-input inference script for SVS

Getting started

This walkthrough follows the repository's MIDI SVS version B recipe on the Opencpop dataset. Opencpop is not redistributed by the project — request access by following the Opencpop team's own instructions first. The TTS (LJSpeech) and PopCS recipes use the same `tasks/run.py` entry point with different configs; see the docs folder.

Create a Python 3.8 environment

The README pins specific requirement files per GPU (`requirements_2080.txt` for 2080Ti/CUDA 10.2, `requirements_3090.txt` for 3090/CUDA 11.4); this is the plain venv route.

bashbash
python -m venv venv
source venv/bin/activate
pip install -U pip
pip install Cython numpy==1.19.1
pip install torch==1.9.0
pip install -r requirements.txt

Link and binarize the dataset

Point `data/raw/` at your extracted Opencpop folder, then pack it for training and inference. This generates `data/binary/opencpop-midi-dp`.

bashbash
ln -s /xxx/opencpop data/raw/
export PYTHONPATH=.
CUDA_VISIBLE_DEVICES=0 python data_gen/tts/bin/binarize.py --config usr/configs/midi/cascade/opencs/aux_rel.yaml

Put the pre-trained vocoder in place

Download the HifiGAN-Singing vocoder (`0109_hifigan_bigpopcs_hop128.zip`) and its pitch-extractor pendant (`0102_xiaoma_pe.zip`) from the repository's `pretrain-model` release and unzip both into `checkpoints/` before training the acoustic model.

Train the acoustic model

bashbash
export MY_DS_EXP_NAME=0228_opencpop_ds100_rel
CUDA_VISIBLE_DEVICES=0 python tasks/run.py --config usr/configs/midi/e2e/opencpop/ds100_adj_rel.yaml --exp_name $MY_DS_EXP_NAME --reset

# watch the run
tensorboard --logdir_spec exp_name
TensorBoard scalars from a DiffSinger training run showing curves for learning rate, mel loss, pitch, duration and voiced/unvoiced losses over about 160k steps
What `tensorboard --logdir_spec` shows while an acoustic model trains.DiffSinger README ↗

Run inference

Inference on the packed test set writes to `./checkpoints/MY_DS_EXP_NAME/generated_`; the raw-input script synthesizes from lyrics + notes + durations (defined as `inp` in the script) and writes to `./infer_out`. To skip training, unzip the published `0228_opencpop_ds100_rel.zip` checkpoint into `checkpoints/`.

bashbash
# from the packed test set
CUDA_VISIBLE_DEVICES=0 python tasks/run.py --config usr/configs/midi/e2e/opencpop/ds100_adj_rel.yaml --exp_name $MY_DS_EXP_NAME --reset --infer

# from raw inputs
python inference/svs/ds_e2e.py --config usr/configs/midi/e2e/opencpop/ds100_adj_rel.yaml --exp_name $MY_DS_EXP_NAME

Optional: sample faster with PNDM

The PNDM recipe uses the `ds1000.yaml` config; set the speed-up per run (1000 / 40 = 25 inference steps by default).

bashbash
CUDA_VISIBLE_DEVICES=0 python tasks/run.py --config usr/configs/midi/e2e/opencpop/ds1000.yaml --exp_name $MY_DS_EXP_NAME --reset --infer --hparams="pndm_speedup=40"

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Train a singing voice model on your own annotated lyrics + MIDI dataset using the Opencpop recipe as a template
  • Reproduce the AAAI 2022 DiffSinger and DiffSpeech results from the published configs and checkpoints
  • Generate sung audio from Chinese lyrics, note sequences and note durations with the raw-input inference script
  • Compare a diffusion acoustic model against a FastSpeech 2 / FFT-Singer baseline trained in the same framework

How DiffSinger compares

DiffSinger alongside other open-source audio, music & voice tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Whisper★ 110kOpenAI's speech recognition model that transcribes and translates audio across many languages.
GPT-SoVITS★ 62.2kAn open-source WebUI that clones a voice from a short audio sample and turns text into speech, with zero-shot and few-shot fine-tuning.
Voicebox★ 56.1kLocal-first voice studio that clones a voice from a short sample, generates speech across seven TTS engines and 23 languages, handles system-wide dictation, and speaks for agents over MCP.
VibeVoice★ 54.6kMicrosoft's text-to-speech model for generating long, expressive multi-speaker audio like podcasts.
whisper.cpp★ 54.1kA dependency-free C/C++ port of Whisper built on ggml, running speech recognition on CPU, Metal, CUDA, Vulkan and NPUs from phones to servers.
VoiceStudio★ 51.5kA local-first desktop voice studio that clones and designs voices, dubs video, and handles dictation, running 16 text-to-speech and 11 speech-recognition engines on your own hardware.
Coqui TTS★ 46.1kA library of text-to-speech models including the multilingual XTTS voice-cloning model.
DiffSinger★ 4.9kSinging voice synthesis and text-to-speech via a shallow diffusion mechanism