Skip to content

About

Open-source Python framework for building modular speaker diarization and transcription pipelines with caching, plugins, multi-backend AI support, forced alignment, and speaker-aware audio processing.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

diarOpen

diarOpen is an open, modular toolkit for building speaker-diarization and transcription pipelines from independently usable stages.

It combines:

  • Demucs vocal separation
  • faster-whisper speech recognition
  • torchaudio MMS word-level forced alignment
  • MSDD, Sortformer, and pyannote diarization
  • Overlap-based word-to-speaker assignment
  • Diarization repair and boundary refinement
  • Speaker enrollment and identity matching
  • Punctuation restoration
  • Per-stage caching
  • Partial pipeline execution
  • Batch processing
  • Third-party pipeline plugins

The base package is lightweight. Heavy ML dependencies are installed through optional extras.

Contents

Features

Diarization backends

Select the diarization backend at runtime:

--diarizer msdd
--diarizer sortformer
--diarizer pyannote
Backend Description
msdd NeMo MSDD with clustering and neural refinement
sortformer NeMo Sortformer end-to-end diarization
pyannote pyannote/speaker-diarization-community-1

Note: The bundled Sortformer implementation currently supports up to four speakers.

Pipeline capabilities

Capability Description
Caching Stage-specific disk caching
Mapping Temporal word-to-speaker assignment
Repair Turn merging and boundary refinement
Enrollment Match speakers against voice profiles
Review Manual speaker corrections
Batch Independent processing of multiple files
Plugins Third-party pipeline stages
Alignment Word-level MMS/CTC alignment

A cached pipeline looks like:

audio
  │
  ├── resolve
  ├── demucs       ← cached
  ├── diarization  ← cached
  ├── whisper      ← cached
  ├── alignment    ← cached
  └── mapping

Changing an unrelated parameter, such as the Whisper model, does not invalidate diarization results.

Installation

Requirements

  • Python 3.10+
  • FFmpeg

Clone and install

git clone https://github.com/serptail/diarOpen-framework.git
cd diarOpen-framework
pip install -e .

Install the complete pipeline

pip install -e ".[full]"

Development

pip install -e ".[dev]"

Development with the complete pipeline

pip install -e ".[full,dev]"

Optional extras

Extra Purpose
full Complete bundled pipeline
ingest yt-dlp downloading
vad-silero Silero VAD
ecapa-hdbscan ECAPA + HDBSCAN
enhance Demucs + Silero preprocessing
captions YouTube transcripts + WhisperX
dev pytest + Ruff

Combine extras as needed:

pip install -e ".[full,dev,ingest]"

External dependencies

Verify FFmpeg:

ffmpeg -version

The pyannote backend requires access to the appropriate Hugging Face model and an authentication token:

python -m diaropen.diarize \
  -a clip.wav \
  --diarizer pyannote \
  --hf-token "$HF_TOKEN" \
  --num-speakers 2

Some non-Latin alignment workflows require uroman:

pip install uroman

Quick Start

Run the full pipeline:

python -m diaropen.diarize \
  -a clip.wav \
  --diarizer pyannote \
  --hf-token "$HF_TOKEN" \
  --num-speakers 2

The installed pipeline entry point is also available:

diaropen-pipeline --help

Typical output:

clip.srt
clip.txt

Example:

[00:00:01.240 --> 00:00:03.810] SPEAKER_00: Welcome back.
[00:00:03.810 --> 00:00:05.920] SPEAKER_01: Thanks for having me.

Pipeline

The reference pipeline is built from small, independently usable stages:

                    ┌──────────────► msdd
                    │
resolve → demucs → decode
                    │
                    └──────────────► whisper
                                         │
                                         ▼
                                     ctc_align
                                         │
                                         ▼
                                       mapping
                                         │
                                         ▼
                                       write

Code organization

Stages live under:

diaropen/pipeline/steps/

The generic runner and cache are:

diaropen/pipeline/core.py
diaropen/pipeline/cache.py

Shared pipeline configuration:

diaropen/pipelines_default.py

Stages

Stage Depends on Produces Cache
resolve — Audio path + content hash No
demucs resolve Vocal/source audio Yes
decode demucs Waveform + language No
msdd decode Diarization turns Yes
whisper decode Transcript + language Yes
ctc_align whisper Word timestamps Yes
mapping Alignment + diarization Speaker-labelled transcript No
write mapping Output files No

Stages return their results through the pipeline context:

def step(ctx):
    ...
    return result

The result becomes:

ctx.data[step_name]

The runner resolves dependencies automatically, allowing complete or partial execution.

Alignment

Whisper transcript
       │
       ▼
   CTC / MMS
       │
       ▼
word timestamps
       │
       ▼
speaker mapping

ctc_align.py uses torchaudio MMS forced alignment. uroman can be used for scripts requiring romanization.

GPU execution

GPU-heavy stages run sequentially:

Demucs ──► Diarization ──► Whisper ──► Forced alignment

There is no --gpu-parallel option. Sequential execution reduces VRAM contention and allows stages to release resources before the next model runs.

Usage

Diarization

MSDD

python -m diaropen.diarize \
  -a clip.wav \
  --diarizer msdd \
  --num-speakers 2

Sortformer

python -m diaropen.diarize \
  -a clip.wav \
  --diarizer sortformer

pyannote

python -m diaropen.diarize \
  -a clip.wav \
  --diarizer pyannote \
  --hf-token "$HF_TOKEN" \
  --num-speakers 2

msdd is the default backend in the bundled pipeline.

Partial execution

List available stages:

python -m diaropen.diarize --list-steps

Diarization only:

python -m diaropen.diarize \
  -a clip.wav \
  --steps resolve,demucs,decode,msdd \
  --diarizer pyannote \
  --hf-token "$HF_TOKEN" \
  --num-speakers 2

Transcription + alignment:

python -m diaropen.diarize \
  -a clip.wav \
  --no-stem \
  --steps resolve,demucs,decode,whisper,ctc_align

The runner automatically resolves required dependencies.

Caching

Use a custom cache directory:

python -m diaropen.diarize \
  -a clip.wav \
  --cache-dir .diarize_cache

Disable caching:

python -m diaropen.diarize \
  -a clip.wav \
  --no-cache

Clear the cache:

python -m diaropen.diarize \
  -a clip.wav \
  --clear-cache

Cache keys contain the input audio content and parameters relevant to each stage.

Enrollment and review

Enroll known speakers:

python -m diaropen.diarize \
  -a interview.wav \
  --enroll-profile speakers/alice.wav \
  --enroll-profile speakers/bob.wav

Create a manual-review package:

python -m diaropen.diarize \
  -a clip.wav \
  --review-out review

Apply edited results:

python -m diaropen.diarize \
  -a clip.wav \
  --review-edits review/edits.json

Batch processing

python -m diaropen.batch \
  --input-dir ./clips \
  --pattern "*.wav" \
  -- \
  --diarizer pyannote \
  --hf-token "$HF_TOKEN" \
  --num-speakers 9

Each file runs independently with its own pipeline state and cache.

Lightweight CLI

The base installation also provides ML-independent utilities:

diaropen fetch <url> [-o out_dir]
diaropen convert <in_path> <out_path> [--sr 16000]
diaropen encode <in_path> <out_path> --preset mp3_voice
diaropen list-stages [category]

Key options

Option Purpose
--diarizer msdd, sortformer, or pyannote
--num-speakers N Speaker count
--whisper-model MODEL faster-whisper model
--steps ... Selected pipeline stages
--no-stem Disable Demucs
--vad-sensitivity MSDD VAD sensitivity
--vad-backend marblenet or silero
--merge-gap-ms N Short-turn merge gap
--merge-short-ms N Short-turn threshold
--pitch-semitone-tol N Pitch tolerance
--no-boundary-refinement Disable boundary refinement
--enroll-profile PATH Speaker profile
--review-out DIR Review package
--review-edits FILE Apply review edits
--similarity-report FILE Speaker diagnostics
--cache-dir DIR Cache directory
--no-cache Disable caching
--clear-cache Clear cache
--dump-json FILE Dump intermediate data
--output-formats ... Output formats
--quiet Reduce output
--no-color Disable terminal colours

Complete reference:

python -m diaropen.diarize --help

Plugins

Third-party stages can be registered through the diaropen.stages entry-point group.

Plugin discovery is handled by:

diaropen/pipeline/registry.py
diaropen/pipeline/bootstrap.py

List registered stages:

python -m diaropen.diarize --list-steps

Development

Install development dependencies:

pip install -e ".[dev]"

For the complete development environment:

pip install -e ".[full,dev]"

Run tests:

pytest

Run Ruff:

ruff check .

Packaging and dependency configuration is defined in pyproject.toml.

Project Structure

diaropen/
├── diarize.py
├── batch.py
├── cli.py
├── pipelines_default.py
│
├── pipeline/
│   ├── core.py
│   ├── cache.py
│   ├── console.py
│   ├── registry.py
│   ├── bootstrap.py
│   └── steps/
│       ├── resolve_step.py
│       ├── demucs_step.py
│       ├── decode_step.py
│       ├── diarize_step.py
│       ├── whisper_step.py
│       ├── ctc_step.py
│       ├── mapping_step.py
│       └── write_step.py
│
└── stages/
    ├── helpers.py
    ├── mapping.py
    ├── ctc_align.py
    ├── similarity.py
    ├── enrollment.py
    ├── review.py
    └── encode.py

Output

Supported output formats include:

  • .srt
  • .txt
  • .vtt
  • .json

A typical run can produce:

clip.wav
clip.srt
clip.txt
clip.vtt
diarization.rttm

Intermediate JSON and review files are available through the corresponding CLI options.

License

MIT

About

Open-source Python framework for building modular speaker diarization and transcription pipelines with caching, plugins, multi-backend AI support, forced alignment, and speaker-aware audio processing.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages