Team Ai
20 results

flash-next

aswinkumar99 /qwen3.8-flash-next-expert-traces Qwen3.8-Flash-Next expert routing traces Token-level routing traces of a deployed MoE model: for every token and every one of the 48 MoE layers, which experts the router chose, the top-32 router logits behind that choice, and the exact hidden state the router read — plus, in v3, the state at many layers per token, the post-final-norm state the LM head consumes, and the LM head's top-8 next-token candidates. The corpus exists to answer one question: how well can the next tokens'… See the full description on the dataset page: https://huggingface.co/datasets/aswinkumar99/qwen3.8-flash-next-expert-traces.text-generation2 likes804 downloads1mo agoHugging FaceAtomicChat /Qwen3.8-Flash-Next-GGUF-metricstabularn<1K0 likes283 downloads2mo agoHugging FaceJerrybro /qwen38-flash-next-int2-rtx5090-research Qwen3.8-Flash-Next Mixed INT2 AutoRound on a Single RTX 5090 This benchmark and reproducibility artifact documents SGLang inference for the mixed-INT2 AutoRound Qwen3.8-Flash-Next checkpoint on one NVIDIA RTX 5090 Blackwell GPU. It covers a validated 256K / 262,144-token long context, MoE autotuning, tiered KV cache, CPU offload, and the device-local evidence showing why another Blackwell GPU's tuning configuration should not be copied blindly. Headline inference… See the full description on the dataset page: https://huggingface.co/datasets/Jerrybro/qwen38-flash-next-int2-rtx5090-research.n<1K1 likes184 downloads20d agoHugging FaceAtomicChat /Qwen3.8-Flash-Next-Abliterated-Uncensored-GGUF-metrics Qwen3.8-Flash-Next-Abliterated-Uncensored-GGUF: measurements Everything behind the numbers on the model card, from one run on 2026-10-07/08: Qwen/Qwen3.8-Flash-Next@de4b8e4d, llama.cpp 980aef8c, 8x RTX PRO 6000 (sm_120), CUDA 13. Path What kld/ The original BF16 model's logits over the held-out neutral and code sets (87 chunks at 4096 context), the reference for every KLD. ablit/data/manifest.json Prompt sources with revisions, the split sizes and the sha256 of… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/Qwen3.8-Flash-Next-Abliterated-Uncensored-GGUF-metrics.0 likes167 downloads2d agoHugging Faceyyyang /gameworld-qwen38-flash-next-trajectory GameWorld: Qwen3.8-Flash-Next Trajectories This dataset is organized for inspecting trajectories. It contains 1,020 completed trajectories from Qwen/Qwen3.8-Flash-Next (both General and Computer-Use (CUA) interfaces) on the GameWorld 400-step benchmark: 170 tasks, and 3 runs per task. Leaderboard Each model was evaluated on 170 tasks with a 400-step limit and three runs per task (510 trajectories per interface). SR is the success rate, and PG is mean task… See the full description on the dataset page: https://huggingface.co/datasets/yyyang/gameworld-qwen38-flash-next-trajectory.0 likes113 downloads3d agoHugging Facetcclaviger /Qwen3.8-Flash-Next-expert-activation-map Qwen3.8-Flash-Next Expert Activation Map Per-(layer, expert) routing and output-importance statistics for Qwen3.8-Flash-Next (Qwen4Exp architecture, 48 MoE layers x 512 routed experts, top-k 10), measured on the unquantized bf16 checkpoint over a 2751-prompt, 28-domain calibration corpus. The purpose is to answer, per layer, which experts carry the model's routed output so that expert-level decisions (bf16 protection under quantization, offload residency, pruning, warm-start… See the full description on the dataset page: https://huggingface.co/datasets/tcclaviger/Qwen3.8-Flash-Next-expert-activation-map.textother1K<n<10K0 likes83 downloads25d agoHugging Face