Team Ai
20 results

dflash

jakeatx /qwopus-dflash-swe20-runtime-results Qwopus / DFlash SWE20 Runtime Results Local RTX 3090 Ti benchmark artifacts for 20 long SWE-bench Lite prompts. The run compares Qwopus 3.6 GGUF variants, llama.cpp MTP speculative decoding, QuinsZouls, and DFlash DDTree configurations at 64K context with q8/q8 KV unless noted. The quality score is a reproducible proxy rubric over gold-patch signals, not official SWE-bench pass/fail. It checks touched-file matches, identifier overlap, patch-like concreteness, test signal, length… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwopus-dflash-swe20-runtime-results.texttext-generation10K<n<100K0 likes323 downloads4mo agoHugging Facejunafinity /osmqwopus-dflash-article-assetsimagen<1K0 likes288 downloads4mo agoHugging Facenaklitechie /bonsai2-dflash2-bench Ternary Bonsai 2 27B + DFlash2 on one NVIDIA L4 — benchmark, 2026-09-24 Engine: llama-server from PrismML-Eng/llama.cpp prism @ ee8ad0ef6 + the DFlash2 cherry-pick (PrismML-Eng/llama.cpp#261), CUDA 12.8 (sm_89). GPU: NVIDIA L4 24 GB (AWS g6.12xlarge, one server per GPU), driver 535. Greedy (temperature 0), one request at a time, -np 1. Target GGUF prism-ml/Ternary-Bonsai-2-27B-gguf. Drafter r3 = naklitechie/Qwen3.8-27B-DFlash2-ternary-bonsai2 Q4_K_M; stock =… See the full description on the dataset page: https://huggingface.co/datasets/naklitechie/bonsai2-dflash2-bench.0 likes147 downloads12d agoHugging Faceinference-optimization /dflash-code-multilingual-teacher-responses-qwen235b Code + Multilingual Teacher Responses (Qwen3-235B-A22B-Instruct-2507) This repo now contains 302,800 total samples across the main blended data.jsonl / .parquet file plus a second Nemotron-only file (nemotron_code_teacher_responses.jsonl / .parquet). All responses were generated by Qwen3-235B-A22B-Instruct-2507 in non-thinking mode (enable_thinking=false) to match downstream speculator training and eval. Built in two batches: an initial 59,506-row batch (50K code + 9.5K… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/dflash-code-multilingual-teacher-responses-qwen235b.texttext-generation100K<n<1M2 likes138 downloads1mo agoHugging FaceGladiator /gemma4-mtp-dflash-speed-bench-results Gemma 4 MTP vs DFlash SPEED-Bench Results This dataset contains the raw JSON result files from a Gemma 4 speculative decoding benchmark on a single H100 80GB. Companion GitHub repository: https://github.com/Gladiator07/gemma4_mtp_dflash Contents The data/ directory contains 440 JSON files: 2 target models: google/gemma-4-31B-it google/gemma-4-26B-A4B-it 4 serving variants per target: baseline decoding MTP with num_speculative_tokens=8 MTP with num_speculative_tokens=16… See the full description on the dataset page: https://huggingface.co/datasets/Gladiator/gemma4-mtp-dflash-speed-bench-results.text-generationn<1K0 likes110 downloads5mo agoHugging Facejiamingshan /qwen3-4b-dflash-official100k-prepared Qwen3-4B DFlash Official-100K Prepared Dataset This is the prepared datasets.load_from_disk() artifact used by the Qwen3-4B official-route DFlash recipe: 100,000 examples columns: input_ids, loss_mask, seq_len max training sequence length used by the recipe: 3072 route: Qwen3 no-thinking / enable_thinking=false Use it with: from huggingface_hub import snapshot_download from datasets import load_from_disk path = snapshot_download(… See the full description on the dataset page: https://huggingface.co/datasets/jiamingshan/qwen3-4b-dflash-official100k-prepared.100K<n<1M0 likes92 downloads4mo agoHugging Face