Overview
Whisper transcribes well but times poorly: the models were trained to predict segment timestamps to roughly one-second accuracy and cannot produce word boundaries at all. whisper-timestamped, from LINAGORA's LinTO team, adds them. It applies dynamic time warping to the model's cross-attention weights — the approach demonstrated in Jong Wook Kim's Multilingual ASR notebook — and assigns a confidence score to every word and every segment along the way.
The implementation is deliberately a drop-in. It is an extension of the `openai-whisper` package and is meant to work with any version of it: import `whisper_timestamped` instead of `whisper`, call `transcribe(model, ...)` instead of `model.transcribe(...)`, and the output gains a `words` key on every segment with start, end and confidence. The CLI mirrors Whisper's, adding a CSV output format and writing word timestamps into the SRT, VTT and TSV files as well.
Two design choices matter in practice. When beam search is off, word alignment is computed on the fly as each speech segment is decoded, so no additional inference pass is needed — which is also why the defaults favour greedy decoding with no temperature fallback (`--accurate` restores Whisper's beam-search defaults). And memory use was treated as a first-class concern, so long files cost little more than a plain Whisper run.
The optional voice-activity detection step is the practical fix for a well-known Whisper failure: predicting phrases like "Thanks for watching!" over silence, a habit learned from the training data. Running silero (the default) or auditok before the model removes the non-speech regions that trigger it. The project's README is candid about the alternative approaches too — it explains why cross-attention DTW avoids the drawbacks of wav2vec-based alignment (one model per language, extra memory, awkward character normalisation, fragility around disfluencies) and why reading timestamp-token probabilities, as whisper.cpp and stable-ts do, can go badly out of sync. The maintainers flag the extension as experimental and note it can significantly affect performance.
What it does
- Word-level start/end timestamps from DTW over the model's own cross-attention weights — no second neural network
- A confidence score on every word and every segment
- No extra inference pass when beam search is off: alignment happens on the fly as each segment is decoded
- More accurate segment start/end estimation than stock Whisper, whose timestamps round to about a second
- Optional VAD (silero, auditok, auditok:v3.1) before transcription, to stop hallucinated text over silence
- Drop-in API and CLI compatibility with openai-whisper, plus a CSV output format and word timestamps inside SRT, VTT and TSV
- Works with fine-tuned Whisper checkpoints from the Hugging Face Hub or a local folder
- A `--plot` option that renders the word alignment, and helpers such as `remove_non_speech` for use on their own

Getting started
Install from PyPI on top of Python 3.9+ (3.7 is the floor) and ffmpeg. Optional extras: matplotlib for the alignment plot, onnxruntime and torchaudio for VAD, transformers for Hugging Face checkpoints.
Install
pip3 install whisper-timestampedTranscribe in Python
Import whisper_timestamped in place of whisper and call transcribe(model, ...). Every segment in the result gains a words list with start, end and confidence.
import whisper_timestamped as whisper
audio = whisper.load_audio("AUDIO.wav")
model = whisper.load_model("tiny", device="cpu")
result = whisper.transcribe(model, audio, language="fr")
import json
print(json.dumps(result, indent = 2, ensure_ascii = False))Or use the CLI
The command line mirrors Whisper's. Note two differences in the defaults: no output directory is set (pass --output_dir .) and there is no verbose output (pass --verbose True). --accurate restores Whisper's beam-search defaults.
whisper_timestamped audio1.flac audio2.mp3 audio3.wav --model tiny --output_dir .Run VAD first on noisy or padded audio
Stripping non-speech regions before the model is the documented remedy for hallucinated phrases over silence. remove_non_speech is also exposed on its own.
from whisper_timestamped import remove_non_speech
audio_speech, segments, convert_timestamps = remove_non_speech(audio, vad="silero")Use a fine-tuned checkpoint
load_model accepts a Hugging Face repo id or a local path, so a language-specific fine-tune works the same way.
whisper_timestamped --model NbAiLab/whisper-large-v2-nob <...>Commands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Reach for it when subtitles need word-level timing — karaoke-style highlighting, word-synced captions, clip extraction on an exact word
- Reach for it when a transcript has to be searchable down to the word, with a confidence score to flag what a human should check
- Reach for it when Whisper keeps inventing text over silence and you need VAD in front of it
- Reach for it when whisperX's wav2vec alignment is awkward for your language set and you would rather not carry a second model per language
How whisper-timestamped compares
whisper-timestamped alongside other open-source audio, music & voice tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Whisper | ★ 110k | OpenAI's speech recognition model that transcribes and translates audio across many languages. |
| GPT-SoVITS | ★ 62.2k | An open-source WebUI that clones a voice from a short audio sample and turns text into speech, with zero-shot and few-shot fine-tuning. |
| Voicebox | ★ 56.1k | Local-first voice studio that clones a voice from a short sample, generates speech across seven TTS engines and 23 languages, handles system-wide dictation, and speaks for agents over MCP. |
| VibeVoice | ★ 54.6k | Microsoft's text-to-speech model for generating long, expressive multi-speaker audio like podcasts. |
| whisper.cpp | ★ 54.1k | A dependency-free C/C++ port of Whisper built on ggml, running speech recognition on CPU, Metal, CUDA, Vulkan and NPUs from phones to servers. |
| VoiceStudio | ★ 51.5k | A local-first desktop voice studio that clones and designs voices, dubs video, and handles dictation, running 16 text-to-speech and 11 speech-recognition engines on your own hardware. |
| Coqui TTS | ★ 46.1k | A library of text-to-speech models including the multilingual XTTS voice-cloning model. |
| whisper-timestamped | ★ 2.9k | Word-level timestamps and per-word confidence for Whisper, recovered from the model's own cross-attention with dynamic time warping — no second model, no extra inference pass |