Denoising, speaker diarization and transcription in a single streamlined process. It's perfect for transcribing podcasts, interviews, or any multi-speaker audio content, as long as they have clear audio. In the output, you'll get a JSON file with the transcript, speaker labels, and timestamps.
- Audio Source Separation: Extract vocals from background music/noise
- Speaker Diarization: Identify and separate different speakers
- Transcription: Convert speech to text with timestamps
- Post-processing: Consolidate transcripts for readability
- Clean Display: Real-time progress updates without cluttering the console
- Step Skipping: Start the pipeline from any step
- Chunked Transcription:
--choptranscribes long audio in fixed-length pieces with globally consistent speaker labels - Cross-platform: Supports Windows, macOS, and Linux
- Python 3.10+
- FFmpeg (for audio processing)
- CUDA-compatible GPU (recommended, but CPU mode available)
- Hugging Face token (optional, for enhanced speaker diarization accuracy)
# Clone the repository
git clone https://github.com/nullwiz/audiopipe.git
cd audiopipe
# Install as a package (gives you the `audiopipe` command)
pip install -e .
# or just the runtime dependencies for running the scripts directly
pip install -r requirements.txt
# For macOS (with Homebrew)
brew install ffmpegSet a Hugging Face token so pyannote can download its models (either name works):
export HUGGING_FACE_TOKEN=hf_... # or HF_TOKEN=hf_...# Basic usage - runs all steps in sequence
python pipeline.py input.mp3
# Resume from a specific step:
python pipeline.py input.mp3 --start-step 2 # Skip separation, start from diarization
python pipeline.py input.mp3 --start-step 3 # Skip to transcription step
# Optional parameters:
python pipeline.py input.mp3 --num-speakers 3 --language en
# Write everything somewhere else and use a faster model on CPU:
python pipeline.py input.mp3 -o runs/ep01 --model openai/whisper-large-v3-turbo
# For very long audio files (>1 hour), transcribe in chunks:
python pipeline.py input.mp3 --chop --chunk-minutes 10audiopipe ... is equivalent to python pipeline.py ... after pip install -e ..
The process consists of three main steps that can be run together or separately:
-
Separation (Step 1): Extracts vocals from background using Demucs
- Input: Any audio/video file
- Output:
output/combined_vocals.wav - Note: Files under 60MB are processed as a single unit; larger files are chunked automatically
-
Diarization (Step 2): Identifies different speakers
- Input:
output/combined_vocals.wav - Output:
output/combined_vocals_diarized.json - Tip: Use
--num-speakersfor better results when speaker count is known
- Input:
-
Transcription (Step 3): Converts complete audio to text, then maps speakers
- Input:
output/combined_vocals.wavand diarization data - Output:
output/final_transcription.json - Architecture: Whisper on the whole file (or on
--chunk-minutespieces with--chop) → speaker mapping → consolidation of consecutive same-speaker segments (disable with--no-consolidate) - Tip:
--languageskips Whisper's auto-detection;--model openai/whisper-large-v3-turbois much faster on CPU
- Input:
With --chop, separation and diarization still run once on the full audio, so speaker labels stay consistent across chunks; only Whisper is run per chunk and the timestamps are offset back into the global timeline.
The pipeline creates several files during processing, all stored in the output/ directory:
combined_vocals.wav: Extracted voices/speech from the inputcombined_background.wav: Background music/noise separated from the inputspeakers/SPEAKER_XX/*.wav: Individual audio segments for each speaker
-
combined_vocals_diarized.json: Speaker diarization results showing who speaks when{ "speakers": ["SPEAKER_01", "SPEAKER_02", ...], "segments": [ {"speaker": "SPEAKER_01", "start": 0.0, "end": 2.5}, {"speaker": "SPEAKER_02", "start": 2.7, "end": 5.1}, ... ] } -
final_transcription.json: Complete transcription with speaker attribution in chronological order{ "segments": [ {"text": "Complete sentence or phrase", "start": 0.1, "end": 2.5, "speaker": "SPEAKER_01"}, {"text": "Another speaker's response", "start": 2.7, "end": 5.1, "speaker": "SPEAKER_02"}, {"text": "Continuing conversation", "start": 5.3, "end": 8.0, "speaker": "SPEAKER_01"}, ... ] }
separated/: Intermediate stems from Demucsdemucs_chunks/: Input pieces when a large file is chunked before separationchunks/: Vocals pieces when using--chop
The presence of these files allows the pipeline to resume from different steps:
- If
combined_vocals.wavexists, audio separation can be skipped (step 1) - If
combined_vocals_diarized.jsonexists, diarization can be skipped (step 2)
- Audio:
.mp3,.wav,.m4a,.flac,.ogg - Video (extracts audio):
.mp4,.mov,.avi,.mkv
python pipeline.py INPUT_AUDIO [OPTIONS]
Arguments:
INPUT_AUDIO Path to input audio/video file
Options:
--num-speakers, -n INT Number of speakers (optional, auto-detected if not specified)
--language, -l STRING Language code for transcription (auto-detected if omitted)
--start-step, -s [1-3] Start from step: 1=separation, 2=diarization, 3=transcription
--device, -d [cpu|cuda|mps] Device to use for processing (auto-detected if not specified)
--output-dir, -o DIR Directory for all outputs (default: output)
--model, -m NAME Whisper model (default: openai/whisper-large-v3)
--chop, -c Transcribe in fixed-length chunks
--chunk-minutes INT Chunk length for --chop (default: 15)
--no-consolidate Keep raw Whisper chunks; don't merge same-speaker runs
--help Show this help message
python dem.py INPUT_FILE [--device cpu|cuda|mps] [--vocals-only] [-o DIR]
Arguments:
INPUT_FILE Path to input audio/video file
python diarize.py INPUT_AUDIO [OPTIONS]
Arguments:
INPUT_AUDIO Path to vocals audio file (usually output/combined_vocals.wav)
Options:
--num-speakers, -n INT Exact number of speakers
--min-speakers INT Lower bound (optional)
--max-speakers INT Upper bound (optional)
Without any of these, pyannote picks the speaker count itself.
For macOS users, there are two operation modes:
For Macs without dedicated NVIDIA GPUs:
# Add this to your .bashrc or .zshrc
export PYTORCH_ENABLE_MPS_FALLBACK=1
# Run with CPU-only flag
python pipeline.py input.mp3 --device cpuFor M1/M2/M3 Macs, you can utilize Metal Performance Shaders:
# Install PyTorch with MPS support
pip install torch torchvision torchaudio
# Run with MPS device
python pipeline.py input.mp3 --device mpsThe final output is a JSON file with chronological segments:
{
"segments": [
{
"text": "Transcript text for this segment",
"start": 0.5,
"end": 4.2,
"speaker": "SPEAKER_01"
},
{
"text": "Response from another speaker",
"start": 4.5,
"end": 7.8,
"speaker": "SPEAKER_02"
},
...
]
}-
Audio Processing:
- Standard mode processes complete audio files for best quality
- For very long files (>1 hour), use
--chopto split into 15-minute chunks - If you get memory errors, try using
--device cpuwhich uses less memory
-
Transcription Accuracy:
- Specify the language with
--languagefor better results - Complete audio transcription provides better context than chunking
- Improved accuracy for clear audio with minimal background noise
- Specify the language with
-
Speaker Identification:
- If speakers are not correctly identified, try setting
--num-speakers - Better results when speakers have distinct voices and don't talk over each other
- Hugging Face token improves diarization accuracy but is not required
- If speakers are not correctly identified, try setting
pip install -r requirements-dev.txt
# Fast unit tests (no models, no torch needed)
pytest
# Integration tests: real Demucs / pyannote / Whisper runs on test/data/
pytest test/test_integration.py -v --integration
# Full pipeline end to end (slow)
pytest test/test_integration.py::test_full_pipeline -v --integration --runslowIntegration tests need HUGGING_FACE_TOKEN (or --hf-token). See README.test.md.