Team Ai
Modelpublic

adventists-ai/MOSS-Transcribe-Diarize-Whisper-Encoder

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
0likes22downloads
Model Card

MOSS-Transcribe-Diarize-Whisper-Encoder

The audio encoder of OpenMOSS-Team/MOSS-Transcribe-Diarize, repackaged as a standard transformers Whisper encoder (WhisperModel config: dmodel 1024, 24 layers, 80 mel bins, 307 M parameters) so that it can feed a frozen LLM through a DuplexJev connector. The weights are those of the original model's `whisperencoder, unchanged (model.encoder.*); the tokenizer files are copied from openai/whisper-small` only so that the Whisper processor loads. Reloading gives outputs identical to the original encoder (max difference 0).

In a linear probe on mean-pooled last-layer states (3,000 train / 800 test clips) it separates speaker gender at 99.0% and four-way emotion at 88.2%, close to Qwen3-ASR-0.6B (99.2 / 90.9) and well above Whisper-large-v3-turbo (88.4 / 82.8).

Used by the DuplexJev-B[-Para]-MOSS-Transcribe-* connectors; Decider.from_pretrained("adventists-ai/DuplexJev-...") fetches it automatically. Source: github.com/adventists-ai/duplexjev.

License and attribution

Apache-2.0, as the original model. The weights are MOSS-Transcribe-Diarize by MOSI.AI / OpenMOSS. Please cite:

bibtex
@misc{moss_transcribe_diarize_2026,
  title={MOSS Transcribe Diarize Technical Report},
  author={{MOSI.AI}},
  year={2026},
  eprint={2601.01554},
  archivePrefix={arXiv},
  primaryClass={cs.SD},
  url={https://arxiv.org/abs/2601.01554}
}