WhisperX Tiny is a fast and accurate speech recognition model with speaker diarization capabilities. Built on OpenAI's Whisper with additional features for alignment and speaker segmentation.
Links
Tags

VibeVoice ASR 7B (C++ / GGML, Q4_K) - long-form speech-to-text with speaker diarization. Returns per-speaker JSON segments with start/end timestamps. English-only. ~10 GB download.
Links
Tags
Repository: localaiLicense: apache-2.0
MOSS-Transcribe-Diarize 0.9B, an end-to-end audio understanding model for joint multi-speaker transcription, speaker diarization and timestamping in a single pass. Q5_K GGUF for the moss-transcribe-cpp backend (C++/ggml port of OpenMOSS MOSS-Transcribe-Diarize), byte-identical to the reference at ~1/6 the size. Faster than the reference PyTorch runtime on CPU (1.6 to 2.2x).
Links
Tags
Repository: localaiLicense: cc-by-nc-4.0
Sortformer Diarization 4-speaker v1 (audio.cpp, Q8_0) - speaker diarization for up to four speakers, served by the audio-cpp backend through /v1/audio/diarization. Returns per-segment start, end and speaker label; it does not transcribe, so pair it with an ASR model for text. Q8_0 because upstream records it as a clean Pass, the same as 16-bit, at two thirds of the size. Licensing: the base checkpoint nvidia/diar_sortformer_4spk-v1 is CC BY-NC 4.0, so commercial use is not permitted.
Links
Tags
Repository: localaiLicense: cc-by-4.0
Streaming Sortformer Diarization 4-speaker v2 (nemo-speech-cpp, Q8_0) - speaker diarization for up to four speakers, served by the nemo-speech-cpp backend through /v1/audio/diarization. Returns per-segment start, end and speaker label; it does not transcribe, so pair it with an ASR model for text. This is the streaming variant: use it when you cannot wait for the full recording. For offline batch work the non-streaming v1 is faster and more accurate.
Links
Tags
Repository: localaiLicense: openmdw-1.1
Nemotron 3.5 ASR Streaming 0.6B (nemo-speech-cpp, Q8_0) - multilingual streaming speech to text for 40 language-locales, with native punctuation and capitalization. Served by the nemo-speech-cpp backend through /v1/audio/transcriptions and the streaming transcription path. A standalone ASR entry. For per-word speaker tags, attach the sortformer diarization model with the diar_model option.
Links
Tags
Repository: localaiLicense: openmdw-1.1
Nemotron 3.5 ASR Streaming 0.6B with streaming Sortformer Diarization 4-speaker v2 (nemo-speech-cpp, Q8_0). Multilingual streaming ASR with per-word speaker tags: the ASR model transcribes and the attached sortformer model labels each word with its speaker. Served through /v1/audio/transcriptions; segments are cut at each speaker change. Q8_0 for both models. The diarization model is the streaming v2 variant, for use when you cannot wait for the full recording.
Links
Tags
Repository: localaiLicense: cc-by-4.0
Parakeet TDT 0.6B v3 (nemo-speech-cpp, Q8_0) - multilingual streaming speech to text for 25 languages. Served by the nemo-speech-cpp backend through /v1/audio/transcriptions and the streaming transcription path. A standalone ASR entry. For per-word speaker tags, install the nemo-speech-cpp-parakeet-tdt-0.6b-v3-diarized entry instead, which attaches the streaming Sortformer diarization model.
Links
Tags
Repository: localaiLicense: cc-by-4.0
Parakeet TDT 0.6B v3 with streaming Sortformer Diarization 4-speaker v2 (nemo-speech-cpp, Q8_0). Multilingual streaming ASR with per-word speaker tags: the ASR model transcribes and the attached sortformer model labels each word with its speaker. Served through /v1/audio/transcriptions; segments are cut at each speaker change. Q8_0 for both models. The diarization model is the streaming v2 variant, for use when you cannot wait for the full recording.
Links
Tags