dflash
Datasets
All datasets matching “dflash”qwopus-dflash-swe20-runtime-results
Qwopus / DFlash SWE20 Runtime Results
Local RTX 3090 Ti benchmark artifacts for 20 long SWE-bench Lite prompts. The run compares Qwopus 3.6 GGUF variants, llama.cpp MTP speculative decoding, QuinsZouls, and DFlash DDTree configurations at 64K context with q8/q8 KV unless noted.
The quality score is a reproducible proxy rubric over gold-patch signals, not official SWE-bench pass/fail. It checks touched-file matches, identifier overlap, patch-like concreteness, test signal, length… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwopus-dflash-swe20-runtime-results.osmqwopus-dflash-article-assetsbonsai2-dflash2-bench
Ternary Bonsai 2 27B + DFlash2 on one NVIDIA L4 — benchmark, 2026-09-24
Engine: llama-server from PrismML-Eng/llama.cpp prism @ ee8ad0ef6 + the DFlash2 cherry-pick (PrismML-Eng/llama.cpp#261),
CUDA 12.8 (sm_89). GPU: NVIDIA L4 24 GB (AWS g6.12xlarge, one server per GPU), driver 535.
Greedy (temperature 0), one request at a time, -np 1. Target GGUF prism-ml/Ternary-Bonsai-2-27B-gguf.
Drafter r3 = naklitechie/Qwen3.8-27B-DFlash2-ternary-bonsai2 Q4_K_M; stock =… See the full description on the dataset page: https://huggingface.co/datasets/naklitechie/bonsai2-dflash2-bench.dflash-code-multilingual-teacher-responses-qwen235b
Code + Multilingual Teacher Responses (Qwen3-235B-A22B-Instruct-2507)
This repo now contains 302,800 total samples across the main blended
data.jsonl / .parquet file plus a second Nemotron-only file
(nemotron_code_teacher_responses.jsonl / .parquet). All responses were
generated by Qwen3-235B-A22B-Instruct-2507 in non-thinking mode
(enable_thinking=false) to match downstream speculator training and eval.
Built in two batches: an initial 59,506-row batch (50K code + 9.5K… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/dflash-code-multilingual-teacher-responses-qwen235b.gemma4-mtp-dflash-speed-bench-results
Gemma 4 MTP vs DFlash SPEED-Bench Results
This dataset contains the raw JSON result files from a Gemma 4 speculative decoding benchmark on a single H100 80GB.
Companion GitHub repository:
https://github.com/Gladiator07/gemma4_mtp_dflash
Contents
The data/ directory contains 440 JSON files:
2 target models:
google/gemma-4-31B-it
google/gemma-4-26B-A4B-it
4 serving variants per target:
baseline decoding
MTP with num_speculative_tokens=8
MTP with num_speculative_tokens=16… See the full description on the dataset page: https://huggingface.co/datasets/Gladiator/gemma4-mtp-dflash-speed-bench-results.qwen3-4b-dflash-official100k-prepared
Qwen3-4B DFlash Official-100K Prepared Dataset
This is the prepared datasets.load_from_disk() artifact used by the
Qwen3-4B official-route DFlash recipe:
100,000 examples
columns: input_ids, loss_mask, seq_len
max training sequence length used by the recipe: 3072
route: Qwen3 no-thinking / enable_thinking=false
Use it with:
from huggingface_hub import snapshot_download
from datasets import load_from_disk
path = snapshot_download(… See the full description on the dataset page: https://huggingface.co/datasets/jiamingshan/qwen3-4b-dflash-official100k-prepared.
