WhisperX
GitHub Repo
MIT
June 30, 2026 at 02:22 AM
0 views

WhisperX

@m-bainProject Author

WhisperX: Time-Accurate Transcription with Speaker Diarization for Long-Form Audio

WhisperX is a fast, feature-rich automatic speech recognition (ASR) system designed to deliver word-level timestamps and robust speaker diarization for long-form audio. Built on OpenAI’s Whisper and augmented with a suite of optimizations, WhisperX targets real-time or near-real-time transcription while maintaining high accuracy through advanced alignment and speaker tagging. The project emphasizes batched inference, effective VAD (voice activity detection) preprocessing, and flexible deployment, enabling users to transcribe meetings, lectures, podcasts, and other multi-speaker content with precision.

[whisperx-arch image] This repository provides fast automatic speech recognition (70x realtime with large-v2) with word-level timestamps and speaker diarization.

  • 70x realtime transcription using whisper large-v2 with batched inference
  • fastex counting via faster-whisper backend, demanding under 8 GB GPU memory for large-v2 with beam_size=5
  • Accurate word-level timestamps using wav2vec2 alignment
  • Multispeaker ASR using speaker diarization from pyannote-audio
  • VAD preprocessing, reducing hallucination and batching without degrading WER

Note: WhisperX exists in a landscape with related services such as Recall.ai’s Meeting Transcription API, which emphasizes diarization by pulling speaker data and separate audio streams from meeting platforms, aiming for 100% accurate speaker labeling with actual names. WhisperX differentiates itself through a strong emphasis on speed (70x realtime for large-v2), open-source flexibility, and an end-to-end pipeline that combines Whisper with wav2vec2 alignment and pyannote-based diarization.

Images and figures in the Input illustrate the architecture and pipeline that WhisperX uses to combine transcription, alignment, and diarization. The visual helps readers understand how raw audio becomes segmented by speaker, then aligned at the word level for accurate timing.

Section 1: What WhisperX Delivers

WhisperX is not merely a faster wrapper around Whisper; it is an integrated pipeline that addresses key pain points in transcription of long-form, multi-speaker audio:

  • Speed without sacrificing fidelity

  • Batched inference accelerates transcription by processing multiple frames or segments concurrently.

  • Uses the faster-whisper backend to deliver substantial speedups on large-v2 models while maintaining alignment accuracy.

  • Word-level precision

  • Alignment with wav2vec2-based models enables word-level timing, improving subtitling and downstream processing.

  • Addresses a common issue with vanilla Whisper, whose timestamps are often utterance-level and can drift by seconds.

  • Speaker diarization

  • Multispeaker ASR is achieved by integrating speaker labeling from pyannote-audio diarization pipelines.

  • The system can tag words with speaker IDs, enabling clear, speaker-by-speaker transcripts.

  • Robust preprocessing

  • VAD-based preprocessing reduces hallucinations and allows for more efficient batching.

  • The pipeline is designed to minimize WER degradation while enhancing real-time performance.

  • Language and phoneme support

  • WhisperX supports multiple languages and includes language-specific alignment models.

  • It can automatically select default alignment models for common languages and guide users in using phoneme-based ASR models when needed.

  • Training and benchmarking context

  • WhisperX is anchored to a preprint that details the approach to batching, alignment, and diarization, including benchmarking on diverse datasets.

  • The project’s v3 release emphasizes per-sentence transcript segmentation and improved diarization, with a substantial speed-up using the faster-whisper backend.

Section 2: A Visual Overview

The WhisperX architecture image (whisperx-arch) illustrates the end-to-end process: audio input is consumed by a fast ASR inference engine, then aligned with a robust phoneme-aware model, and finally diarized into speaker-labeled segments. The pipeline highlights how word-level timing is anchored to phoneme-based alignment, while a diarization module assigns speaker identities to segments. This orchestration is key to producing transcripts that are both temporally precise and easy to read for subtitle generation, meeting notes, and accessibility needs.

Section 3: New Developments and Achievements

WhisperX has seen notable progress and recognition in the research community and practical deployments:

  • Competition success

  • Won 1st place at the Ego4D transcription challenge, signaling strong performance in real-world transcription tasks.

  • Conference acceptance

  • WhisperX was accepted at INTERSPEECH 2023, reflecting its contribution to the field of speech processing and diarization.

  • Version progress and performance improvements

  • v3 introduced segment-per-sentence formatting, leveraging nltk sent_tokenize for better subtitling and diarization.

  • v3 also delivered a 70x speed-up by enabling batched whisper with the faster-whisper backend.

  • v2 brought code cleanup and default VAD filtering enabled, aligning with the paper’s recommendations.

  • Preprint and benchmarking

  • The ArXiv preprint provides benchmark results and details of the WhisperX pipeline, including efficient batch inference for large-v2 with significantly reduced latency.

Section 4: Setup and Installation

WhisperX aims to be approachable for both casual users and developers who want to modify or extend the pipeline. The setup guidance emphasizes GPU acceleration, though CPU usage is supported.

  • CUDA installation (GPU acceleration)

  • Linux: Install CUDA toolkit 12.8

  • Windows: Download and install CUDA toolkit 12.8

  • If you are not using a GPU, skip CUDA installation and use CPU mode

  • PyPI: pip install whisperx

  • uvx: uvx whisperx (for those using the uvx toolkit)

  • Advanced installation options

  • Option A: Install from GitHub

    • uvx git+https://github.com/m-bain/whisperX.git
  • Option B: Developer installation (for contributing)

    • git clone https://github.com/m-bain/whisperX.git
    • cd whisperX
    • uv sync --all-extras --dev
  • Note: The development version may include experimental features and bugs; use the stable PyPI release for production.

  • Dependencies and notes

  • You may also need to install ffmpeg, Rust, and related tools according to WhisperX and OpenAI Whisper setup instructions.

  • Language diarization requires a Hugging Face access token to enable speaker diarization:

    • Provide token with --hf_token and accept the speaker-diarization-community-1 model terms if you intend to label speakers.
  • Speaker diarization setup

  • To enable speaker diarization, supply a Hugging Face access token and select the appropriate diarization model (e.g., speaker-diarization-community-1).

Section 5: Usage and Practical Commands

WhisperX provides a command-line interface (CLI) that makes it straightforward to transcribe audio files and obtain Word-Level timestamps with speaker labels.

  • English transcription (default)

  • whisperx path/to/audio.wav

  • To visualize word timings in the subtitle file, add --highlight_words True

  • Alignment and timing improvements

  • The default path performs transcription with forced alignment to wav2vec2.0 large for better timing.

  • Subtitles and timing with speaker IDs

  • To label transcripts with speaker IDs, specify known speaker counts if possible:

    • whisperx path/to/audio.wav --model large-v2 --diarize --highlight_words True
    • If you know there are 2 speakers: whisperx path/to/audio.wav --model large-v2 --diarize --minspeakers 2 --maxspeakers 2
  • Model and alignment choices

  • For higher timestamp accuracy, use a larger alignment model and batch size settings (beware of GPU memory constraints):

    • whisperx path/to/audio.wav --model large-v2 --alignmodel WAV2VEC2ASRLARGELV60K960H --batchsize 4
  • CPU mode

  • To run on CPU instead of GPU (e.g., on macOS), use: whisperx path/to/audio.wav --compute_type int8 --device cpu

  • Language-specific usage

  • WhisperX supports multiple languages and automatically selects default alignment models for common languages (en, fr, de, es, it) via torchaudio or Hugging Face. If the detected language isn’t in the default list, users can choose a phoneme-based ASR model from Hugging Face.

  • Example: German transcription

  • whisperx --model large-v2 --language de path/to/audio.wav

  • For German, a phoneme-based ASR alignment model may be required for best results.

Section 6: Python Usage (Snippet Overview)

WhisperX provides a Python API to integrate transcription into custom pipelines. A typical workflow includes loading the WhisperX model, transcribing, performing alignment, and assigning speaker labels.

  • Basic steps

  • Load WhisperX model (large-v2) on a specified device

  • Load and preprocess audio

  • Transcribe with batch processing

  • Align transcripts with a wav2vec2 alignment model

  • Load a diarization model and assign speaker IDs to segments

  • Key commands (illustrative)

  • Import modules, load models, and run transcription

  • Align outputs with alignment model

  • Run diarization to obtain speaker segments

  • Combine diarization with transcription to produce speaker-tagged segments

  • Practical notes

  • You can reduce memory usage by lowering batchsize, using a smaller ASR model (e.g., base), or selecting a lighter computetype (int8).

  • If memory is tight, consider clearing caches and unloading models after use.

Section 7: Demos and Cloud Access

Demos illustrate WhisperX’s capabilities and provide cloud-based access for testing if local GPUs are unavailable.

  • Replicate demos

  • WhisperX large-v3: Demonstrations on Replicate (cloud API)

  • WhisperX large-v2: Demonstrations on Replicate

  • WhisperX medium: Demonstrations on Replicate

  • Try before you run

  • If you don’t have access to GPUs, these cloud demos offer a convenient way to experience WhisperX performance and diarization capabilities.

Section 8: Technical Details and Considerations

WhisperX’s technical design centers on batching, alignment, VAD, and diarization, with attention to GPU memory and transcription quality.

  • Batching and alignment

  • The paper behind WhisperX explains how batching interacts with the wav2vec2 alignment model to produce accurate word-level timestamps efficiently.

  • A key trade-off exists between batch size and GPU memory usage; larger batch sizes can improve throughput but require more memory.

  • VAD and segmentation

  • VAD preprocessing reduces hallucinations and makes batching more stable, contributing to better overall transcription quality.

  • OpenAI Whisper and limitations

  • WhisperX builds on OpenAI Whisper, but its timestamps are not inherently word-level. WhisperX addresses this by using alignment with wav2vec2 and segment-level processing.

  • The limitation notes include:

    • Some transcript words may not align if characters aren’t in the alignment model’s dictionary (e.g., “2014.” or “£13.60”).
    • Overlapping speech remains challenging for both Whisper and WhisperX.
    • Diarization is not perfect and depends on the quality of the diarization model and audio.
  • Phoneme-based ASR and language models

  • WhisperX supports phoneme-based ASR models, particularly when language-specific alignment models are not readily available.

  • Users can search the Hugging Face model hub for language-specific phoneme models and test them in WhisperX to improve alignment for non-English languages.

Section 9: Contribute and Community

WhisperX invites contributions from multilingual developers and researchers who want to experiment with phoneme models and language-specific alignment. A few pathways exist:

  • Add new language models

  • Train or test phoneme-based models for languages not in the default alignment model set.

  • Share results with examples showing accuracy improvements on multi-speaker audio.

  • Bug fixes and feature improvements

  • Submit pull requests for improvements in diarization accuracy, VAD performance, and model flushing for low-GPU-memory scenarios.

  • Documentation and tutorials

  • Improve usage examples, add more end-to-end pipelines, and create tutorials for streaming transcription.

Section 10: TODO and Roadmap

WhisperX maintains a running list of tasks and improvements:

  • Completed items

  • Multilingual initialization

  • Automatic align model selection based on language detection

  • Python usage

  • Incorporating speaker diarization

  • Model flush for low memory

  • Faster-whisper backend

  • Sentence-level segments (NLTK toolbox)

  • Improved alignment logic

  • Silero VAD as alternative VAD

  • Speaker diarization enhancements

  • In progress / future goals

  • Update examples with diarization and word highlighting

  • Restore subtitle .ass output functionality

  • Add benchmarking code (e.g., TEDLIUM for SPD/WER and word segmentation)

Section 11: Contact and Support

If you have questions or want to discuss WhisperX, you can reach out to the maintainer:

Support is also encouraged via the project’s Buy Me a Coffee link, which helps sustain ongoing development.

Section 12: Acknowledgements

WhisperX builds on a broad ecosystem of tools and research:

  • Core architecture and theory

  • OpenAI’s Whisper as the foundational ASR model

  • PyTorch-based forced alignment and wav2vec2 alignment concepts

  • Pyannote audio for VAD and diarization

  • Silero-VAD and other VAD backends as optional components

  • Faster-whisper as a backend for speed improvements

  • CTranslate2 for model efficiency and deployment

  • People and institutions

  • VGG (Visual Geometry Group) at the University of Oxford

  • The broader open-source community contributing to WhisperX, including the OS contributors and those who have supported this work financially

  • Acknowledgements to collaborators and models

  • Special thanks to the authors and maintainers of pyannote-audio, pyannote speaker diarization models, Silero VAD, and Hugging Face community models for diarization and phoneme-based ASR

Section 13: Citation

If you use WhisperX in your research, please cite the associated paper:

  • Bain, Max; Huh, Jaesung; Han, Tengda; Zisserman, Andrew. WhisperX: Time-Accurate Speech Transcription of Long-Form Audio. INTERSPEECH 2023.

This citation provides a formal acknowledgment of the approach, methodology, and results presented in the WhisperX study, enabling others to reproduce and build upon the work.

Closing Thoughts

WhisperX stands out as a practical, high-performance solution for transcribing long-form, multi-speaker audio. By combining OpenAI Whisper with wav2vec2 alignment and pyannote-based diarization, WhisperX delivers word-level timestamps and speaker labels at impressive speeds. The project’s emphasis on batching, VAD, and language-aware alignment makes it a versatile tool for researchers, developers, and content producers who require accurate transcripts with intelligible speaker separation. Whether you’re processing corporate meetings, academic lectures, or podcasts, WhisperX provides a robust framework to produce high-quality, richly annotated transcripts suitable for subtitling, indexing, and downstream analysis.

Images from the Input:

  • whisperx-arch: This figure visually encapsulates WhisperX’s pipeline and the interaction between transcription, alignment, and diarization stages. It helps readers grasp how words are timestamped and mapped to speakers across long audio streams.

If you’d like additional visuals or a step-by-step walkthrough with screenshots of running WhisperX on a sample dataset, I can create a companion guide that blends narrative explanations with annotated figures to further illuminate the process.

Enjoying this project?

Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.

Project
whisperx
Created
June 30
Last Updated
June 30, 2026 at 02:22 AM

Find more projects like this

One email a week: new and trending developer tools, fresh comparisons, and what shipped. Unsubscribe in one click.