ZihanLiummyycc/diffusion-asr-dbank-interface-lora
Diffusion-ASR interface and LoRA for DementiaBank
Adapted parameters accompanying Diffusion and Flow Matching ASR for Elderly Speech. Code and experiment recipes.
What this repository contains
- Updated scope:
acoustic_plus_lora. - Independent parameter elements: 63,827,002.
- Stored tensor elements: 87,298,420; shared-parameter aliases can make this larger.
- 270 named tensors in
adaptation.safetensors. - Exact base-asset revisions and SHA256 hashes in
base_assets.json.
These are final parameter values, not additive deltas. Frozen base weights, optimizers, training logs, private manifests, raw predictions and clinical data are not included. This is a native component checkpoint, not a standalone Transformers model or an automatically loadable standard PEFT adapter.
Loading
Clone the accompanying code and use the model-specific environment recorded there:
git clone https://github.com/ZihanLiummyycc/diffusion-asr-elderly-speech.git
cd diffusion-asr-elderly-speech
git checkout 29709f075399be63ed6d1e536a0a0ca3bba0a072Download the component repository to a separate directory. Add that directory to Python's import path for load_component.py and load_adaptation.py. The examples below assume component_dir is that directory and code_root is the cloned code repository. Resolve base files using the immutable revisions and hashes in `base_assets.json`, not an arbitrary latest model.
import sys
sys.path.insert(0, str(code_root / "benchmarks/repos/Diffusion-ASR"))
sys.path.insert(0, str(component_dir))
from models.WhisperLLaDA import WhisperLLaDA
from load_component import load_component
model = WhisperLLaDA(
whisper_model=str(whisper_large_v3_dir), llada_model=str(llada_dir),
gen_len=128, lora=True, lora_rank=8, lora_alpha=32, lora_dropout=0.1,
second_per_window=0.333333, second_stride=0.333333,
task_prompt="Transcribe the audio:",
).eval()
load_component(model, component_dir)The constructor also reads the BERT configuration recorded under qformer_config. This package contains the adapted Q-Former, speech query tokens, projection, and LLaDA LoRA parameters. It does not require reloading the original ASR adaptation initialization over these final values. For native inference, use the repository's generate method and Whisper feature extractor. The historical eval used mode=decoding, gen_len=128, block_length=32, and 128 total steps; the short integrity probe used 8 steps, not this evaluation budget.
Evaluation and verification
The original recorded evaluation on 928 DementiaBank utterances reported WER 24.94% for this configuration. This is a historical aggregate checked against the original experiment, not a result recomputed during publication. Consult the paper/code for the data splits, normalization, model-specific decoding settings and exploratory-comparison limitations. No clinical examples are included here.
File/tensor integrity checks, strict native loading, and original-versus-exported output comparisons passed on two synthetic inputs using an NVIDIA A40. verification.json records the short probe budget and runtime. Those probes do not measure WER, latency, robustness or privacy leakage. A fresh installation from this public repository has not been evaluated as a corpus benchmark.
Intended use and limitations
Research on ASR adaptation and decoding for elderly speech. This is not a diagnostic model and has not been validated for clinical decisions. Transcription errors and dataset bias remain possible. Absence of raw clinical files does not establish absence of memorization. Use only data you are authorized to process.
Attribution and terms
See LICENSE.md, retained upstream notices in licenses/, and the base model repositories in base_assets.json. This upload does not change upstream terms or authorize redistribution of DementiaBank data. No new blanket MIT/Apache license is assigned to the complete derivative bundle.
