Team Ai
Modelpublic

moonshine-ai/moonshine-streaming-tiny-ja

sourceHugging Facemitupdated 2mo agoView on Hugging Face
2likes124downloads
Model Card

Moonshine Streaming Tiny — Japanese

Japanese streaming speech recognition, 27.0M parameters. Same architecture as moonshine-ai/moonshine-streaming-tiny, trained for Japanese with a 12,288-entry Japanese tokenizer.

Moonshine Streaming pairs a 50 Hz time-domain audio frontend with a sliding-window Transformer encoder, so it transcribes incrementally rather than waiting for an utterance to finish. It is intended for on-device use on edge-class hardware.

Checkpoint identity

This repository is a conversion of a specific training checkpoint, recorded here because the training run that produced it was still in progress when this snapshot was taken and a better one may replace it:

Checkpointja12k_tiny_stageC_best.safetensors
StageC (read-speech mix)
Architectureslinkier_prime_adapted
Tokenizertokenizer_ja12k.json, vocab 12,288
Snapshot taken2026-08-23
Tensors / parameters163 / 27.0M

If you need reproducibility, pin the revision of this repository rather than tracking main.

Usage

bash
pip install --upgrade transformers datasets[audio]
python
from transformers import MoonshineStreamingForConditionalGeneration, AutoProcessor
import torch

model = MoonshineStreamingForConditionalGeneration.from_pretrained(
    "moonshine-ai/moonshine-streaming-tiny-ja"
).eval()
processor = AutoProcessor.from_pretrained("moonshine-ai/moonshine-streaming-tiny-ja")

inputs = processor(audio, return_tensors="pt", sampling_rate=16000)

# Cap the output length. Like other seq2seq ASR models this one can fall into a
# repetition loop, and short or noisy clips are where it happens.
seq_lens = inputs.attention_mask.sum(dim=-1)
max_new_tokens = int((seq_lens * 6.5 / 16000).max().item()) + 2

generated = model.generate(**inputs, max_new_tokens=max_new_tokens)
print(processor.batch_decode(generated, skip_special_tokens=True)[0])

Pass the `attention_mask`. The encoder applies its per-layer sliding windows only when it is given one; called without a mask it attends over the whole utterance instead, which is a different model from the one that was trained. The processor returns the mask, so the snippet above is the safe form. The processor also pads audio to a whole number of 80-sample frames, which the frontend requires.

Architecture

Encoder6 layers, width 320, 8 heads, sliding windows (16, 4) on the first two and last two layers and (16, 0) between
Decoder6 layers, width 320, 8 heads, RoPE over 32 of each head's 40 dimensions
Frontend50 Hz features, CMVN, asinh compression, two causal stride-2 convolutions
Adapterlearned absolute positional embeddings before the decoder

The lookahead layers give roughly 80 ms of lookahead; the intermediate layers have none.

Training data

Trained on a large-scale automatically labeled Japanese corpus, plus a read-speech mix in the final stage:

  • —Podcast crawl, roughly 109,000 hours.
  • —YouTube crawl, roughly 50,000 hours.
  • —Stage C read-speech mix, including Common Voice Japanese.

The podcast and YouTube transcripts are pseudo-labels: they were produced by running a Whisper-family teacher model over crawled audio, not by human transcription. The model therefore inherits the teacher's error modes, including its handling of proper nouns, numerals and code-switching, and its transcription conventions for a language written without spaces. No human-verified transcript was used for the bulk of training.

Evaluation

Japanese is scored on character error rate with spaces removed (cer_nospace), never WER. Japanese is written without spaces, so tokenization differences alone can read as several hundred percent WER while the characters are correct.

suite_ja is FLEURS Japanese (650 utterances) and ReazonSpeech Japanese (5,263).

Full panels, batch 8

PanelCER
fleurs_ja11.50
reazonspeech_ja26.73
macro19.115

Seeded 400-utterance sample, batch 1

Batch 1 is the honest number for deployment. Batched evaluation zero-pads short clips up to the longest in the batch, and that trailing silence flatters the model by more than a point on spontaneous speech.

PanelCER
fleurs_ja11.62
reazonspeech_ja27.77
macro19.70

This repository against the training checkpoint

These weights were converted from the neo training checkpoint, and the conversion was checked by measurement rather than inspection: same seeded sample, same batch size, same normalizer.

`fleurs_ja``reazonspeech_ja`macro
Training checkpoint11.6227.7719.699
This repository11.3528.0819.712

397/400 and 385/400 transcripts are byte-identical. The residual comes from the frontend's 80-sample frame alignment, which this path pads and the training path does not.

Limitations

  • —Machine-labeled training data. See above; the model reproduces its teacher's mistakes as well as its strengths.
  • —Repetition loops on short clips. About 1.25% of ReazonSpeech utterances in the batch-1 sample run away, costing 0.34 CER. Cap the output length.
  • —Short-utterance sensitivity. Utterances with references under ~15 characters are far harder than the macro number suggests (CER above 60% on that bucket of spontaneous speech) and are where numerically small changes produce large per-utterance swings.
  • —Evaluated only on read speech (FLEURS) and spontaneous speech (ReazonSpeech). No evaluation of telephony, children's speech, heavy dialect, or noisy far-field conditions.
  • —Snapshot of an in-progress run. See the checkpoint identity above.

Out-of-scope use

Not intended for non-consensual surveillance, speaker identification, or high-stakes decisions.

License

MIT.

moonshine-ai/moonshine-streaming-tiny-ja · Team Ai