WhisperX
WhisperX: Time-Accurate Transcription with Speaker Diarization for Long-Form Audio
WhisperX is a fast, feature-rich automatic speech recognition (ASR) system designed to deliver word-level timestamps and robust speaker diarization for long-form audio. Built on OpenAI’s Whisper and augmented with a suite of optimizations, WhisperX targets real-time or near-real-time transcription while maintaining high accuracy through advanced alignment and speaker tagging. The project emphasizes batched inference, effective VAD (voice activity detection) preprocessing, and flexible deployment, enabling users to transcribe meetings, lectures, podcasts, and other multi-speaker content with precision.
[whisperx-arch image] This repository provides fast automatic speech recognition (70x realtime with large-v2) with word-level timestamps and speaker diarization.
- 70x realtime transcription using whisper large-v2 with batched inference
- fastex counting via faster-whisper backend, demanding under 8 GB GPU memory for large-v2 with beam_size=5
- Accurate word-level timestamps using wav2vec2 alignment
- Multispeaker ASR using speaker diarization from pyannote-audio
- VAD preprocessing, reducing hallucination and batching without degrading WER
Note: WhisperX exists in a landscape with related services such as Recall.ai’s Meeting Transcription API, which emphasizes diarization by pulling speaker data and separate audio streams from meeting platforms, aiming for 100% accurate speaker labeling with actual names. WhisperX differentiates itself through a strong emphasis on speed (70x realtime for large-v2), open-source flexibility, and an end-to-end pipeline that combines Whisper with wav2vec2 alignment and pyannote-based diarization.
Images and figures in the Input illustrate the architecture and pipeline that WhisperX uses to combine transcription, alignment, and diarization. The visual helps readers understand how raw audio becomes segmented by speaker, then aligned at the word level for accurate timing.
Section 1: What WhisperX Delivers
WhisperX is not merely a faster wrapper around Whisper; it is an integrated pipeline that addresses key pain points in transcription of long-form, multi-speaker audio:
Speed without sacrificing fidelity
Batched inference accelerates transcription by processing multiple frames or segments concurrently.
Uses the faster-whisper backend to deliver substantial speedups on large-v2 models while maintaining alignment accuracy.
Word-level precision
Alignment with wav2vec2-based models enables word-level timing, improving subtitling and downstream processing.
Addresses a common issue with vanilla Whisper, whose timestamps are often utterance-level and can drift by seconds.
Speaker diarization
Multispeaker ASR is achieved by integrating speaker labeling from pyannote-audio diarization pipelines.
The system can tag words with speaker IDs, enabling clear, speaker-by-speaker transcripts.
Robust preprocessing
VAD-based preprocessing reduces hallucinations and allows for more efficient batching.
The pipeline is designed to minimize WER degradation while enhancing real-time performance.
Language and phoneme support
WhisperX supports multiple languages and includes language-specific alignment models.
It can automatically select default alignment models for common languages and guide users in using phoneme-based ASR models when needed.
Training and benchmarking context
WhisperX is anchored to a preprint that details the approach to batching, alignment, and diarization, including benchmarking on diverse datasets.
The project’s v3 release emphasizes per-sentence transcript segmentation and improved diarization, with a substantial speed-up using the faster-whisper backend.
Section 2: A Visual Overview
The WhisperX architecture image (whisperx-arch) illustrates the end-to-end process: audio input is consumed by a fast ASR inference engine, then aligned with a robust phoneme-aware model, and finally diarized into speaker-labeled segments. The pipeline highlights how word-level timing is anchored to phoneme-based alignment, while a diarization module assigns speaker identities to segments. This orchestration is key to producing transcripts that are both temporally precise and easy to read for subtitle generation, meeting notes, and accessibility needs.
Section 3: New Developments and Achievements
WhisperX has seen notable progress and recognition in the research community and practical deployments:
Competition success
Won 1st place at the Ego4D transcription challenge, signaling strong performance in real-world transcription tasks.
Conference acceptance
WhisperX was accepted at INTERSPEECH 2023, reflecting its contribution to the field of speech processing and diarization.
Version progress and performance improvements
v3 introduced segment-per-sentence formatting, leveraging nltk sent_tokenize for better subtitling and diarization.
v3 also delivered a 70x speed-up by enabling batched whisper with the faster-whisper backend.
v2 brought code cleanup and default VAD filtering enabled, aligning with the paper’s recommendations.
Preprint and benchmarking
The ArXiv preprint provides benchmark results and details of the WhisperX pipeline, including efficient batch inference for large-v2 with significantly reduced latency.
Section 4: Setup and Installation
WhisperX aims to be approachable for both casual users and developers who want to modify or extend the pipeline. The setup guidance emphasizes GPU acceleration, though CPU usage is supported.
CUDA installation (GPU acceleration)
Linux: Install CUDA toolkit 12.8
Windows: Download and install CUDA toolkit 12.8
If you are not using a GPU, skip CUDA installation and use CPU mode
Simple installation (recommended)
PyPI: pip install whisperx
uvx: uvx whisperx (for those using the uvx toolkit)
Advanced installation options
Option A: Install from GitHub
- uvx git+https://github.com/m-bain/whisperX.git
Option B: Developer installation (for contributing)
- git clone https://github.com/m-bain/whisperX.git
- cd whisperX
- uv sync --all-extras --dev
Note: The development version may include experimental features and bugs; use the stable PyPI release for production.
Dependencies and notes
You may also need to install ffmpeg, Rust, and related tools according to WhisperX and OpenAI Whisper setup instructions.
Language diarization requires a Hugging Face access token to enable speaker diarization:
- Provide token with --hf_token and accept the speaker-diarization-community-1 model terms if you intend to label speakers.
Speaker diarization setup
To enable speaker diarization, supply a Hugging Face access token and select the appropriate diarization model (e.g., speaker-diarization-community-1).
Section 5: Usage and Practical Commands
WhisperX provides a command-line interface (CLI) that makes it straightforward to transcribe audio files and obtain Word-Level timestamps with speaker labels.
English transcription (default)
whisperx path/to/audio.wav
To visualize word timings in the subtitle file, add --highlight_words True
Alignment and timing improvements
The default path performs transcription with forced alignment to wav2vec2.0 large for better timing.
Subtitles and timing with speaker IDs
To label transcripts with speaker IDs, specify known speaker counts if possible:
- whisperx path/to/audio.wav --model large-v2 --diarize --highlight_words True
- If you know there are 2 speakers: whisperx path/to/audio.wav --model large-v2 --diarize --minspeakers 2 --maxspeakers 2
Model and alignment choices
For higher timestamp accuracy, use a larger alignment model and batch size settings (beware of GPU memory constraints):
- whisperx path/to/audio.wav --model large-v2 --alignmodel WAV2VEC2ASRLARGELV60K960H --batchsize 4
CPU mode
To run on CPU instead of GPU (e.g., on macOS), use: whisperx path/to/audio.wav --compute_type int8 --device cpu
Language-specific usage
WhisperX supports multiple languages and automatically selects default alignment models for common languages (en, fr, de, es, it) via torchaudio or Hugging Face. If the detected language isn’t in the default list, users can choose a phoneme-based ASR model from Hugging Face.
Example: German transcription
whisperx --model large-v2 --language de path/to/audio.wav
For German, a phoneme-based ASR alignment model may be required for best results.
Section 6: Python Usage (Snippet Overview)
WhisperX provides a Python API to integrate transcription into custom pipelines. A typical workflow includes loading the WhisperX model, transcribing, performing alignment, and assigning speaker labels.
Basic steps
Load WhisperX model (large-v2) on a specified device
Load and preprocess audio
Transcribe with batch processing
Align transcripts with a wav2vec2 alignment model
Load a diarization model and assign speaker IDs to segments
Key commands (illustrative)
Import modules, load models, and run transcription
Align outputs with alignment model
Run diarization to obtain speaker segments
Combine diarization with transcription to produce speaker-tagged segments
Practical notes
You can reduce memory usage by lowering batchsize, using a smaller ASR model (e.g., base), or selecting a lighter computetype (int8).
If memory is tight, consider clearing caches and unloading models after use.
Section 7: Demos and Cloud Access
Demos illustrate WhisperX’s capabilities and provide cloud-based access for testing if local GPUs are unavailable.
Replicate demos
WhisperX large-v3: Demonstrations on Replicate (cloud API)
WhisperX large-v2: Demonstrations on Replicate
WhisperX medium: Demonstrations on Replicate
Try before you run
If you don’t have access to GPUs, these cloud demos offer a convenient way to experience WhisperX performance and diarization capabilities.
Section 8: Technical Details and Considerations
WhisperX’s technical design centers on batching, alignment, VAD, and diarization, with attention to GPU memory and transcription quality.
Batching and alignment
The paper behind WhisperX explains how batching interacts with the wav2vec2 alignment model to produce accurate word-level timestamps efficiently.
A key trade-off exists between batch size and GPU memory usage; larger batch sizes can improve throughput but require more memory.
VAD and segmentation
VAD preprocessing reduces hallucinations and makes batching more stable, contributing to better overall transcription quality.
OpenAI Whisper and limitations
WhisperX builds on OpenAI Whisper, but its timestamps are not inherently word-level. WhisperX addresses this by using alignment with wav2vec2 and segment-level processing.
The limitation notes include:
- Some transcript words may not align if characters aren’t in the alignment model’s dictionary (e.g., “2014.” or “£13.60”).
- Overlapping speech remains challenging for both Whisper and WhisperX.
- Diarization is not perfect and depends on the quality of the diarization model and audio.
Phoneme-based ASR and language models
WhisperX supports phoneme-based ASR models, particularly when language-specific alignment models are not readily available.
Users can search the Hugging Face model hub for language-specific phoneme models and test them in WhisperX to improve alignment for non-English languages.
Section 9: Contribute and Community
WhisperX invites contributions from multilingual developers and researchers who want to experiment with phoneme models and language-specific alignment. A few pathways exist:
Add new language models
Train or test phoneme-based models for languages not in the default alignment model set.
Share results with examples showing accuracy improvements on multi-speaker audio.
Bug fixes and feature improvements
Submit pull requests for improvements in diarization accuracy, VAD performance, and model flushing for low-GPU-memory scenarios.
Documentation and tutorials
Improve usage examples, add more end-to-end pipelines, and create tutorials for streaming transcription.
Section 10: TODO and Roadmap
WhisperX maintains a running list of tasks and improvements:
Completed items
Multilingual initialization
Automatic align model selection based on language detection
Python usage
Incorporating speaker diarization
Model flush for low memory
Faster-whisper backend
Sentence-level segments (NLTK toolbox)
Improved alignment logic
Silero VAD as alternative VAD
Speaker diarization enhancements
In progress / future goals
Update examples with diarization and word highlighting
Restore subtitle .ass output functionality
Add benchmarking code (e.g., TEDLIUM for SPD/WER and word segmentation)
Section 11: Contact and Support
If you have questions or want to discuss WhisperX, you can reach out to the maintainer:
- Email: [email protected]
Support is also encouraged via the project’s Buy Me a Coffee link, which helps sustain ongoing development.
Section 12: Acknowledgements
WhisperX builds on a broad ecosystem of tools and research:
Core architecture and theory
OpenAI’s Whisper as the foundational ASR model
PyTorch-based forced alignment and wav2vec2 alignment concepts
Pyannote audio for VAD and diarization
Silero-VAD and other VAD backends as optional components
Faster-whisper as a backend for speed improvements
CTranslate2 for model efficiency and deployment
People and institutions
VGG (Visual Geometry Group) at the University of Oxford
The broader open-source community contributing to WhisperX, including the OS contributors and those who have supported this work financially
Acknowledgements to collaborators and models
Special thanks to the authors and maintainers of pyannote-audio, pyannote speaker diarization models, Silero VAD, and Hugging Face community models for diarization and phoneme-based ASR
Section 13: Citation
If you use WhisperX in your research, please cite the associated paper:
- Bain, Max; Huh, Jaesung; Han, Tengda; Zisserman, Andrew. WhisperX: Time-Accurate Speech Transcription of Long-Form Audio. INTERSPEECH 2023.
This citation provides a formal acknowledgment of the approach, methodology, and results presented in the WhisperX study, enabling others to reproduce and build upon the work.
Closing Thoughts
WhisperX stands out as a practical, high-performance solution for transcribing long-form, multi-speaker audio. By combining OpenAI Whisper with wav2vec2 alignment and pyannote-based diarization, WhisperX delivers word-level timestamps and speaker labels at impressive speeds. The project’s emphasis on batching, VAD, and language-aware alignment makes it a versatile tool for researchers, developers, and content producers who require accurate transcripts with intelligible speaker separation. Whether you’re processing corporate meetings, academic lectures, or podcasts, WhisperX provides a robust framework to produce high-quality, richly annotated transcripts suitable for subtitling, indexing, and downstream analysis.
Images from the Input:
- whisperx-arch: This figure visually encapsulates WhisperX’s pipeline and the interaction between transcription, alignment, and diarization stages. It helps readers grasp how words are timestamped and mapped to speakers across long audio streams.
If you’d like additional visuals or a step-by-step walkthrough with screenshots of running WhisperX on a sample dataset, I can create a companion guide that blends narrative explanations with annotated figures to further illuminate the process.
Enjoying this project?
Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.
Repository:https://github.com/m-bain/whisperX
GitHub - m-bain/whisperX: WhisperX
WhisperX is a fast, feature-rich automatic speech recognition (ASR) system designed to deliver word-level timestamps and robust speaker diarization for long-for...
github - m-bain/whisperx


