diarOpen is an open, modular toolkit for building speaker-diarization and transcription pipelines from independently usable stages.
It combines:
- Demucs vocal separation
- faster-whisper speech recognition
- torchaudio MMS word-level forced alignment
- MSDD, Sortformer, and pyannote diarization
- Overlap-based word-to-speaker assignment
- Diarization repair and boundary refinement
- Speaker enrollment and identity matching
- Punctuation restoration
- Per-stage caching
- Partial pipeline execution
- Batch processing
- Third-party pipeline plugins
The base package is lightweight. Heavy ML dependencies are installed through optional extras.
- Features
- Installation
- Quick Start
- Pipeline
- Usage
- Plugins
- Development
- Project Structure
- Output
- License
Select the diarization backend at runtime:
--diarizer msdd
--diarizer sortformer
--diarizer pyannote| Backend | Description |
|---|---|
msdd |
NeMo MSDD with clustering and neural refinement |
sortformer |
NeMo Sortformer end-to-end diarization |
pyannote |
pyannote/speaker-diarization-community-1 |
Note: The bundled Sortformer implementation currently supports up to four speakers.
| Capability | Description |
|---|---|
| Caching | Stage-specific disk caching |
| Mapping | Temporal word-to-speaker assignment |
| Repair | Turn merging and boundary refinement |
| Enrollment | Match speakers against voice profiles |
| Review | Manual speaker corrections |
| Batch | Independent processing of multiple files |
| Plugins | Third-party pipeline stages |
| Alignment | Word-level MMS/CTC alignment |
A cached pipeline looks like:
audio
│
├── resolve
├── demucs ← cached
├── diarization ← cached
├── whisper ← cached
├── alignment ← cached
└── mapping
Changing an unrelated parameter, such as the Whisper model, does not invalidate diarization results.
- Python 3.10+
- FFmpeg
git clone https://github.com/serptail/diarOpen-framework.git
cd diarOpen-framework
pip install -e .pip install -e ".[full]"pip install -e ".[dev]"pip install -e ".[full,dev]"| Extra | Purpose |
|---|---|
full |
Complete bundled pipeline |
ingest |
yt-dlp downloading |
vad-silero |
Silero VAD |
ecapa-hdbscan |
ECAPA + HDBSCAN |
enhance |
Demucs + Silero preprocessing |
captions |
YouTube transcripts + WhisperX |
dev |
pytest + Ruff |
Combine extras as needed:
pip install -e ".[full,dev,ingest]"Verify FFmpeg:
ffmpeg -versionThe pyannote backend requires access to the appropriate Hugging Face model and an authentication token:
python -m diaropen.diarize \
-a clip.wav \
--diarizer pyannote \
--hf-token "$HF_TOKEN" \
--num-speakers 2Some non-Latin alignment workflows require uroman:
pip install uromanRun the full pipeline:
python -m diaropen.diarize \
-a clip.wav \
--diarizer pyannote \
--hf-token "$HF_TOKEN" \
--num-speakers 2The installed pipeline entry point is also available:
diaropen-pipeline --helpTypical output:
clip.srt
clip.txt
Example:
[00:00:01.240 --> 00:00:03.810] SPEAKER_00: Welcome back.
[00:00:03.810 --> 00:00:05.920] SPEAKER_01: Thanks for having me.
The reference pipeline is built from small, independently usable stages:
┌──────────────► msdd
│
resolve → demucs → decode
│
└──────────────► whisper
│
▼
ctc_align
│
▼
mapping
│
▼
write
Stages live under:
diaropen/pipeline/steps/
The generic runner and cache are:
diaropen/pipeline/core.py
diaropen/pipeline/cache.py
Shared pipeline configuration:
diaropen/pipelines_default.py
| Stage | Depends on | Produces | Cache |
|---|---|---|---|
resolve |
— | Audio path + content hash | No |
demucs |
resolve |
Vocal/source audio | Yes |
decode |
demucs |
Waveform + language | No |
msdd |
decode |
Diarization turns | Yes |
whisper |
decode |
Transcript + language | Yes |
ctc_align |
whisper |
Word timestamps | Yes |
mapping |
Alignment + diarization | Speaker-labelled transcript | No |
write |
mapping |
Output files | No |
Stages return their results through the pipeline context:
def step(ctx):
...
return resultThe result becomes:
ctx.data[step_name]The runner resolves dependencies automatically, allowing complete or partial execution.
Whisper transcript
│
▼
CTC / MMS
│
▼
word timestamps
│
▼
speaker mapping
ctc_align.py uses torchaudio MMS forced alignment. uroman can be used for scripts requiring romanization.
GPU-heavy stages run sequentially:
Demucs ──► Diarization ──► Whisper ──► Forced alignment
There is no --gpu-parallel option. Sequential execution reduces VRAM contention and allows stages to release resources before the next model runs.
python -m diaropen.diarize \
-a clip.wav \
--diarizer msdd \
--num-speakers 2python -m diaropen.diarize \
-a clip.wav \
--diarizer sortformerpython -m diaropen.diarize \
-a clip.wav \
--diarizer pyannote \
--hf-token "$HF_TOKEN" \
--num-speakers 2msdd is the default backend in the bundled pipeline.
List available stages:
python -m diaropen.diarize --list-stepsDiarization only:
python -m diaropen.diarize \
-a clip.wav \
--steps resolve,demucs,decode,msdd \
--diarizer pyannote \
--hf-token "$HF_TOKEN" \
--num-speakers 2Transcription + alignment:
python -m diaropen.diarize \
-a clip.wav \
--no-stem \
--steps resolve,demucs,decode,whisper,ctc_alignThe runner automatically resolves required dependencies.
Use a custom cache directory:
python -m diaropen.diarize \
-a clip.wav \
--cache-dir .diarize_cacheDisable caching:
python -m diaropen.diarize \
-a clip.wav \
--no-cacheClear the cache:
python -m diaropen.diarize \
-a clip.wav \
--clear-cacheCache keys contain the input audio content and parameters relevant to each stage.
Enroll known speakers:
python -m diaropen.diarize \
-a interview.wav \
--enroll-profile speakers/alice.wav \
--enroll-profile speakers/bob.wavCreate a manual-review package:
python -m diaropen.diarize \
-a clip.wav \
--review-out reviewApply edited results:
python -m diaropen.diarize \
-a clip.wav \
--review-edits review/edits.jsonpython -m diaropen.batch \
--input-dir ./clips \
--pattern "*.wav" \
-- \
--diarizer pyannote \
--hf-token "$HF_TOKEN" \
--num-speakers 9Each file runs independently with its own pipeline state and cache.
The base installation also provides ML-independent utilities:
diaropen fetch <url> [-o out_dir]
diaropen convert <in_path> <out_path> [--sr 16000]
diaropen encode <in_path> <out_path> --preset mp3_voice
diaropen list-stages [category]| Option | Purpose |
|---|---|
--diarizer |
msdd, sortformer, or pyannote |
--num-speakers N |
Speaker count |
--whisper-model MODEL |
faster-whisper model |
--steps ... |
Selected pipeline stages |
--no-stem |
Disable Demucs |
--vad-sensitivity |
MSDD VAD sensitivity |
--vad-backend |
marblenet or silero |
--merge-gap-ms N |
Short-turn merge gap |
--merge-short-ms N |
Short-turn threshold |
--pitch-semitone-tol N |
Pitch tolerance |
--no-boundary-refinement |
Disable boundary refinement |
--enroll-profile PATH |
Speaker profile |
--review-out DIR |
Review package |
--review-edits FILE |
Apply review edits |
--similarity-report FILE |
Speaker diagnostics |
--cache-dir DIR |
Cache directory |
--no-cache |
Disable caching |
--clear-cache |
Clear cache |
--dump-json FILE |
Dump intermediate data |
--output-formats ... |
Output formats |
--quiet |
Reduce output |
--no-color |
Disable terminal colours |
Complete reference:
python -m diaropen.diarize --helpThird-party stages can be registered through the diaropen.stages entry-point group.
Plugin discovery is handled by:
diaropen/pipeline/registry.py
diaropen/pipeline/bootstrap.py
List registered stages:
python -m diaropen.diarize --list-stepsInstall development dependencies:
pip install -e ".[dev]"For the complete development environment:
pip install -e ".[full,dev]"Run tests:
pytestRun Ruff:
ruff check .Packaging and dependency configuration is defined in pyproject.toml.
diaropen/
├── diarize.py
├── batch.py
├── cli.py
├── pipelines_default.py
│
├── pipeline/
│ ├── core.py
│ ├── cache.py
│ ├── console.py
│ ├── registry.py
│ ├── bootstrap.py
│ └── steps/
│ ├── resolve_step.py
│ ├── demucs_step.py
│ ├── decode_step.py
│ ├── diarize_step.py
│ ├── whisper_step.py
│ ├── ctc_step.py
│ ├── mapping_step.py
│ └── write_step.py
│
└── stages/
├── helpers.py
├── mapping.py
├── ctc_align.py
├── similarity.py
├── enrollment.py
├── review.py
└── encode.py
Supported output formats include:
.srt.txt.vtt.json
A typical run can produce:
clip.wav
clip.srt
clip.txt
clip.vtt
diarization.rttm
Intermediate JSON and review files are available through the corresponding CLI options.
MIT