datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwopus-dflash-swe20-runtime-results
Qwopus / DFlash SWE20 Runtime Results
Local RTX 3090 Ti benchmark artifacts for 20 long SWE-bench Lite prompts. The run compares Qwopus 3.6 GGUF variants, llama.cpp MTP speculative decoding, QuinsZouls, and DFlash DDTree configurations at 64K context with q8/q8 KV unless noted.
The quality score is a reproducible proxy rubric over gold-patch signals, not official SWE-bench pass/fail. It checks touched-file matches, identifier overlap, patch-like concreteness, test signal, length… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwopus-dflash-swe20-runtime-results.osmqwopus-dflash-article-assetsbonsai2-dflash2-bench
Ternary Bonsai 2 27B + DFlash2 on one NVIDIA L4 — benchmark, 2026-09-24
Engine: llama-server from PrismML-Eng/llama.cpp prism @ ee8ad0ef6 + the DFlash2 cherry-pick (PrismML-Eng/llama.cpp#261),
CUDA 12.8 (sm_89). GPU: NVIDIA L4 24 GB (AWS g6.12xlarge, one server per GPU), driver 535.
Greedy (temperature 0), one request at a time, -np 1. Target GGUF prism-ml/Ternary-Bonsai-2-27B-gguf.
Drafter r3 = naklitechie/Qwen3.8-27B-DFlash2-ternary-bonsai2 Q4_K_M; stock =… See the full description on the dataset page: https://huggingface.co/datasets/naklitechie/bonsai2-dflash2-bench.dflash-code-multilingual-teacher-responses-qwen235b
Code + Multilingual Teacher Responses (Qwen3-235B-A22B-Instruct-2507)
This repo now contains 302,800 total samples across the main blended
data.jsonl / .parquet file plus a second Nemotron-only file
(nemotron_code_teacher_responses.jsonl / .parquet). All responses were
generated by Qwen3-235B-A22B-Instruct-2507 in non-thinking mode
(enable_thinking=false) to match downstream speculator training and eval.
Built in two batches: an initial 59,506-row batch (50K code + 9.5K… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/dflash-code-multilingual-teacher-responses-qwen235b.gemma4-mtp-dflash-speed-bench-results
Gemma 4 MTP vs DFlash SPEED-Bench Results
This dataset contains the raw JSON result files from a Gemma 4 speculative decoding benchmark on a single H100 80GB.
Companion GitHub repository:
https://github.com/Gladiator07/gemma4_mtp_dflash
Contents
The data/ directory contains 440 JSON files:
2 target models:
google/gemma-4-31B-it
google/gemma-4-26B-A4B-it
4 serving variants per target:
baseline decoding
MTP with num_speculative_tokens=8
MTP with num_speculative_tokens=16… See the full description on the dataset page: https://huggingface.co/datasets/Gladiator/gemma4-mtp-dflash-speed-bench-results.dflash-flm-regen-qwen3-8bqwen3-4b-dflash-official100k-prepared
Qwen3-4B DFlash Official-100K Prepared Dataset
This is the prepared datasets.load_from_disk() artifact used by the
Qwen3-4B official-route DFlash recipe:
100,000 examples
columns: input_ids, loss_mask, seq_len
max training sequence length used by the recipe: 3072
route: Qwen3 no-thinking / enable_thinking=false
Use it with:
from huggingface_hub import snapshot_download
from datasets import load_from_disk
path = snapshot_download(… See the full description on the dataset page: https://huggingface.co/datasets/jiamingshan/qwen3-4b-dflash-official100k-prepared.MoS-DFlash-Evidence
MoS-DFlash aggregate experiment evidence
This dataset repository contains aggregate, reviewer-facing evidence for the
MoS-DFlash experiments. It does not contain prompts, per-prompt generations,
training data, credentials, internal paths, or raw training logs.
B5 Qwen3-4B fixed-budget replication
releases/b5-qwen3-4b-fixed-budget/ contains:
the frozen result summary;
the matched-training-volume aggregate trajectory;
plot-ready aggregate trajectories;
run… See the full description on the dataset page: https://huggingface.co/datasets/ryan-0608/MoS-DFlash-Evidence.dflash-qwen3-8b-nemotron-evol-rollout-ctx16k
Qwen3-8B Regenerated DFlash Rollout, ctx16k filtered
This dataset contains target-regenerated assistant responses for DFlash/speculative
draft-model training.
Generation Setup
Target model: Qwen/Qwen3-8B
Serving stack: vLLM
Sampling: temperature=0.6, top_p=0.95, top_k=20
max_tokens=3072
Context filter: rows that exceeded max_model_len=16384 during rollout are
excluded from this split, even if later recovered with a 32k refill.
Qwen3 thinking: disabled via chat… See the full description on the dataset page: https://huggingface.co/datasets/jingwut/dflash-qwen3-8b-nemotron-evol-rollout-ctx16k.qwen3.8-27b-uncensored-dflash2-m4-pro-benchmark
Qwen3.8-27B (MLX 4-bit) + DFlash2 speculative decoding on M4 Pro — benchmark recipe
This is a benchmark recipe, not redistributed weights. It records the exact
hardware, software, and commands used to measure a 2.06x generation-throughput
speedup with DFlash speculative decoding, and how to rerun it.
Result
HumanEval, 20 samples, max 256 new tokens, temperature 0 (greedy), reasoning
off, block size 5, paired baseline and DFlash under identical settings. Other… See the full description on the dataset page: https://huggingface.co/datasets/hamiejuice/qwen3.8-27b-uncensored-dflash2-m4-pro-benchmark.dflash-qwen3-8b-qwen235b-instruct-bs16-prepared-datak3-dflash-sequence-compaction-validation-7cf713b
K3 DFlash Sequence-Compaction GB300 Validation
This repository is the evidence bundle for an independent correctness review of the
TensorRT-LLM K3 DFlash sequence-compaction patch at commit
7cf713b08b6892aae44a12d41b1a67d029d4b234.
The implementation was evaluated against the design contract at
f7789542915749fc9e6cd9b165b3a271cbe92184.
Bottom line
No patch correctness defect was found within the implemented and supported envelope. The patch's… See the full description on the dataset page: https://huggingface.co/datasets/srivy-together/k3-dflash-sequence-compaction-validation-7cf713b.qwen2.5-vl-3b-dflashouro_14_dflash_regen_openr1math220kqwen36-27b-dflash-stagebDflash-small-pretrain-dataouro_26_dflash_regen_openr1math220kdflash-hidden-states-cachellama-cpp-dflash2-cuda-binchimere-dflash-datadflashkv_builtglm52-dflash-traces-v2
