datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Qwen3.8-Flash-Next-GGUF-metricsQwen3.8-Flash-Next-expert-activation-map
Qwen3.8-Flash-Next Expert Activation Map
Per-(layer, expert) routing and output-importance statistics for
Qwen3.8-Flash-Next (Qwen4Exp architecture, 48 MoE layers x 512 routed
experts, top-k 10), measured on the unquantized bf16 checkpoint over a
2751-prompt, 28-domain calibration corpus.
The purpose is to answer, per layer, which experts carry the model's routed
output so that expert-level decisions (bf16 protection under quantization,
offload residency, pruning, warm-start… See the full description on the dataset page: https://huggingface.co/datasets/tcclaviger/Qwen3.8-Flash-Next-expert-activation-map.
