adventists-ai/MOSS-Transcribe-Diarize-Whisper-Encoder
MOSS-Transcribe-Diarize-Whisper-Encoder
The audio encoder of OpenMOSS-Team/MOSS-Transcribe-Diarize, repackaged as a standard transformers Whisper encoder (WhisperModel config: dmodel 1024, 24 layers, 80 mel bins, 307 M parameters) so that it can feed a frozen LLM through a DuplexJev connector. The weights are those of the original model's `whisperencoder, unchanged (model.encoder.*); the tokenizer files are copied from openai/whisper-small` only so that the Whisper processor loads. Reloading gives outputs identical to the original encoder (max difference 0).
In a linear probe on mean-pooled last-layer states (3,000 train / 800 test clips) it separates speaker gender at 99.0% and four-way emotion at 88.2%, close to Qwen3-ASR-0.6B (99.2 / 90.9) and well above Whisper-large-v3-turbo (88.4 / 82.8).
Used by the DuplexJev-B[-Para]-MOSS-Transcribe-* connectors; Decider.from_pretrained("adventists-ai/DuplexJev-...") fetches it automatically. Source: github.com/adventists-ai/duplexjev.
License and attribution
Apache-2.0, as the original model. The weights are MOSS-Transcribe-Diarize by MOSI.AI / OpenMOSS. Please cite:
@misc{moss_transcribe_diarize_2026,
title={MOSS Transcribe Diarize Technical Report},
author={{MOSI.AI}},
year={2026},
eprint={2601.01554},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2601.01554}
}