█

AI/TLDR

Whishper

A self-hosted transcription and subtitling suite that never sends your audio anywhere

Audio, Music & VoiceOpen source
Language
Svelte
License
AGPL-3.0
$docker --version

Overview

Whishper is an open-source audio transcription and subtitling suite that runs entirely on your own machine. You give it a media file or a URL, it transcribes the speech with Whisper, and you get back text you can download as TXT, JSON, VTT or SRT — or keep editing in the browser. Nothing is uploaded to a third-party service: the transcription, the translation and the subtitle editing all happen in containers you started, which is the whole point of the project.

It is built as a small stack of cooperating services rather than a single binary. A transcription API wraps Faster-Whisper, a Go backend coordinates the frontend, the job queue and the database, a SvelteKit frontend provides the web UI, and three third-party containers fill in the rest: LibreTranslate for subtitle translation, MongoDB for storing transcriptions, and Nginx so the whole thing is reachable on one domain. Faster-Whisper is what makes CPU-only transcription practical; if you have an NVIDIA GPU, Whishper can use it for a large speed-up.

The part that distinguishes it from a plain Whisper wrapper is the subtitle editor. Segments are listed with editable start and end timestamps beside the playing media, the line matching the current playback position is highlighted, and each segment is flagged with its characters-per-second reading rate so you can see which subtitles are too dense before you export them. You can split a segment, insert one, and switch the subtitle language, without leaving the page. Because Whishper accepts any source yt-dlp supports, the usual workflow is to paste a URL, wait for the job, then correct the result in the editor.

The Whishper subtitle editor: a video plays on the left while the right-hand pane lists numbered transcript segments with editable start and end timestamps, the text of each line, and a CPS reading-rate figure and duration beside it, with the segment matching the current playback position highlighted
The subtitle editor keeps the media beside the segment list, highlights the line at the playhead, and shows the characters-per-second rate of each segment so you can spot subtitles that read too fast.Whishper docs ↗

The project is AGPL-3.0 and distributed as Docker images. Its maintainer has frozen this branch while a full rewrite lands on the `v4` branch, so the documented v3 stack is what the published images and the install guide describe.

What it does

  • Transcribes audio or video from an uploaded file or from any URL that yt-dlp can fetch
  • Faster-Whisper as the Whisper backend, so CPU-only transcription is usable; NVIDIA GPU acceleration is supported for much faster runs
  • Exports to TXT, JSON, VTT and SRT, or copies the raw text to the clipboard
  • Subtitle editor in the browser: playback-synced highlighting, characters-per-second warnings, segment splitting and insertion, and subtitle-language selection
  • Translates finished transcriptions into any language LibreTranslate supports, in a container of your own
  • Runs 100% locally — transcription, translation and editing can all work offline once the images and models are pulled

Getting started

Whishper ships as a set of Docker containers driven by one Compose file and one `.env`. The quick-start script writes both for you; the manual path is the same files, edited by hand.

Check the prerequisites

You need Docker and the Compose plugin. For GPU transcription you additionally need an NVIDIA GPU with the NVIDIA Container Toolkit installed.

bashbash
docker --version
docker compose version

Run the quick-start script

The script fetches the Compose file, the example `.env` and the Nginx config, and walks you through the rest. Read the GPU-support page first if you intend to use an NVIDIA GPU.

bashbash
curl -fsSL -o get-whishper.sh https://raw.githubusercontent.com/pluja/whishper/main/get-whishper.sh
bash get-whishper.sh

Or install by hand

The manual path is the same three files plus the LibreTranslate data directories, which need to be owned by the container's user.

bashbash
curl -o docker-compose.yml https://raw.githubusercontent.com/pluja/whishper/main/docker-compose.yml
curl -o .env https://raw.githubusercontent.com/pluja/whishper/main/example.env
curl -o nginx.conf https://raw.githubusercontent.com/pluja/whishper/main/nginx.conf

mkdir -p ./whishper_data/libretranslate/{data,cache}
chown -R 1032:1032 whishper_data/libretranslate

Edit .env, then start the stack

`WHISHPER_HOST` is the URL you will actually open, port included — the frontend uses it to reach the backend, so `127.0.0.1` means Whishper only works from that machine. `WHISPER_MODELS` is the comma-separated list to download, and `LT_LOAD_ONLY` limits which LibreTranslate language pairs are fetched.

bashbash
# .env
LT_LOAD_ONLY=es,en,fr
WHISPER_MODELS=tiny,small
WHISHPER_HOST=http://127.0.0.1:8082
DB_USER=whishper
DB_PASS=whishper

docker compose up -d

Transcribe something and edit the subtitles

Open the host you configured, then paste a URL or upload a file, pick the language and model, and wait for the job. When it finishes, open it in the editor to fix the timings and wording, then download SRT or VTT.

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Reach for it when the recording is confidential — an interview, a medical or legal conversation, an internal meeting — and a hosted transcription API is not an option
  • Reach for it to subtitle a video end to end: transcribe, correct the timings in the editor, translate, export SRT
  • Reach for it to turn a backlog of podcast or lecture URLs into searchable text without paying per minute
  • Reach for it when you have a spare NVIDIA GPU and want a self-hosted alternative to a per-minute transcription service

How Whishper compares

Whishper alongside other open-source audio, music & voice tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Whisper★ 110kOpenAI's speech recognition model that transcribes and translates audio across many languages.
GPT-SoVITS★ 62.2kAn open-source WebUI that clones a voice from a short audio sample and turns text into speech, with zero-shot and few-shot fine-tuning.
Voicebox★ 56.1kLocal-first voice studio that clones a voice from a short sample, generates speech across seven TTS engines and 23 languages, handles system-wide dictation, and speaks for agents over MCP.
VibeVoice★ 54.6kMicrosoft's text-to-speech model for generating long, expressive multi-speaker audio like podcasts.
whisper.cpp★ 54.1kA dependency-free C/C++ port of Whisper built on ggml, running speech recognition on CPU, Metal, CUDA, Vulkan and NPUs from phones to servers.
VoiceStudio★ 51.5kA local-first desktop voice studio that clones and designs voices, dubs video, and handles dictation, running 16 text-to-speech and 11 speech-recognition engines on your own hardware.
Coqui TTS★ 46.1kA library of text-to-speech models including the multilingual XTTS voice-cloning model.
Whishper★ 3.1kA self-hosted transcription and subtitling suite that never sends your audio anywhere