datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
terminal-bench-2.1-deepseek-v4-flash-trajectories
DeepSeek-V4-Flash on Terminal-Bench 2.1 — full agent trajectories
Every agent trajectory from a controlled scaffold comparison: the same model, the
same machine, the same 89 tasks, only the agent harness changed.
Scaffold
Solved
terminus-2 (Terminal-Bench's own agent)
53 / 89
dsh sdk-minimal (DeepSeek Harness)
61 / 89
Paired: dsh solved 18 tasks terminus-2 missed, terminus-2 solved 8 that dsh
missed, 43 both, 18 neither, 2 not scorable (see Caveats).… See the full description on the dataset page: https://huggingface.co/datasets/openguardrails/terminal-bench-2.1-deepseek-v4-flash-trajectories.DeepSeek-V4-Flash-0731-REAM-calibration-stats
DeepSeek-V4-Flash-0731 — expert calibration statistics (REAM line)
Layerwise routed-expert statistics of
deepseek-ai/DeepSeek-V4-Flash-0731
(43 MoE layers × 256 experts), collected by running the full model over a
~4.9M-token multi-domain calibration mix (multi-turn dialogs, thinking and
direct modes, rendered with the model's own chat encoder). These are the
statistics behind the REAM144/96 release line — published so that expert
selection, pruning, merging and routing research… See the full description on the dataset page: https://huggingface.co/datasets/WaveCut/DeepSeek-V4-Flash-0731-REAM-calibration-stats.swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.deepseek-v4-flash-rocm-vllm-repro
Reproducing DeepSeek-V4-Flash on AMD ROCm with vLLM: 32K Correctness and TopK Sweep
This article summarizes an engineering reproduction of
deepseek-ai/DeepSeek-V4-Flash on an AMD ROCm ModelScope DSW instance. The work
focuses on a practical question: can a complex, fast-moving DeepSeek-V4-Flash
serving path be turned into a reproducible ROCm baseline with explicit
correctness gates?
The answer from this run is yes, with an important boundary: the current setup
is a fallback-heavy… See the full description on the dataset page: https://huggingface.co/datasets/lyydfys/deepseek-v4-flash-rocm-vllm-repro.deepseek-v4-flash-0731-m3-ultra
DeepSeek-V4-Flash-0731 on M3 Ultra 512 GB — benchmark dataset
Independent performance characterization of
Vontra/DeepSeek-V4-Flash-0731-MXFP4-MLX
on a single Mac Studio M3 Ultra (80-core GPU, 512 GB unified memory).
Engine: omlx 0.5.7 · OS: macOS 26.6 (25G72) · MLX: 0.32.0
Recommended configuration
omlx serve --model-dir /opt/models --port 8033 \
--hot-cache-max-size 256GB --initial-cache-blocks 512
// ~/.omlx/model_settings.json
{"version": 1, "models":… See the full description on the dataset page: https://huggingface.co/datasets/guruswami-ai/deepseek-v4-flash-0731-m3-ultra.Wikipedia-FA-EN-DeepSeek-V4-Flash-0731
Wikipedia Persian to English — DeepSeek V4 Flash 0731
Rolling, machine-generated English translations of Persian Wikipedia articles
from Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge, configuration
full_articles_fa_without_en. 129,816 translations are
currently published in 26 immutable Parquet shards.
The target release contains 129,816 translations;
five source rows have empty plain_text and are not translated. Shards are
published only after 5,000 complete, validated records… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.DeepSeek-v4-Flash-ChatThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Teich Test
This directory contains newline-delimited JSON training examples generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-flash.
Rows: 6313
Format
Each file is newline-delimited JSON where every line is already a training example.
Chat-only datasets include messages… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Flash-Chat.scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928
Scale-SWE DeepSeek V4 Flash 0731 Think Rollouts
Successful AweAgent trajectories generated with deepseek-v4-flash-0731 in think mode.
Dataset summary
Source task instances: 3,393
Rollouts per source instance: 4
Total attempted rollouts: 13,572
Successful exported trajectories: 7,928
Unique instances represented by successful trajectories: 2,250
Scaffold: aweagent
Tool-call format: openai_function
The export retains assistant reasoning_content, function tool… See the full description on the dataset page: https://huggingface.co/datasets/wjn922-01/scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928.DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x
DeepSeek V4 Flash 0731 Teacher Distillation — 40,513 Retained Rows
Teacher-distillation corpus generated with
deepseek-ai/DeepSeek-V4-Flash-0731.
The original manifest contained 45,000 unique seeds.
Following generation, QC, retry-based repair, quarantine auditing,
and recovery adjudication, 40,513 rows were retained.
Composition
Bucket
Rows
Coding
5,601
Agentic
9,982
Cyber blue
13,000
Controlled cyber red
6,999
Tool use
4,931
Total
40,513… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x.deepseek-v4-flash-swebench-replay
deepseek-v4-flash-swebench-replay
中文
这是一个 DeepSeek V4 Flash 在 SWE-bench 上的 agentic replay 数据集仓库。
它的目标是让使用者不需要部署 SWE-bench,也不需要复现 Docker/benchmark 环境,就可以直接查看和重放模型的多轮推理与工具调用轨迹。
当前包含的数据
verified_agentic
lite_agentic
当前不包含的数据
单轮 single-turn trace
verified_mini_agentic(当前本地仅完成 31/50,因此不纳入首版)
分数汇总
verified_agentic: 354 / 500, Acc/Pass@1 = 70.8
lite_agentic: 182 / 300, Acc/Pass@1 = 60.67
数据来源
这些轨迹由… See the full description on the dataset page: https://huggingface.co/datasets/fxiao0369/deepseek-v4-flash-swebench-replay.Deepseek-V4-Flash-11000x
Sherlock Thinking Alpha DeepSeek V4 Flash Distillation
Seed Prompt Dataset
Prompts are sourced from TeichAI/sherlock-thinking-alpha-11000x.
Model
Solutions and reasoning traces were generated with deepseek-ai/DeepSeek-V4-Flash.
DeepSeek-V4-Flash is part of the DeepSeek-V4 preview series. Its model card describes it as a Mixture-of-Experts language model with 284B total parameters, 13B activated parameters, and a 1M-token context length. The model repository is… See the full description on the dataset page: https://huggingface.co/datasets/SLoonker/Deepseek-V4-Flash-11000x.llm_timeline_deepseek_v4_flash-pi
Coding agent session traces
This dataset contains coding agent session traces collected while working on LLM Timeline web app using the prompt from coding-agent-bench-prompts
DeepSeek-V4-Flash-0731-REAM-Healing-Mix-2048
DeepSeek-V4 Flash REAM Healing Mix 2048
Private, deterministic healing mixtures for the 104-expert REAM-pruned
DeepSeek-V4-Flash-0731-120B checkpoint. It was prepared after direct A/B
generation tests found post-pruning degradation in factual accuracy, language
control, thinking delimiters, and long-code repetition.
seq1024 base config
Rows: 2,048
Maximum sequence length: 1,024 DeepSeek-V4 tokens
Non-padding tokens: 1,581,071
Supervised assistant tokens: 1,092… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/DeepSeek-V4-Flash-0731-REAM-Healing-Mix-2048.denovoswe-distill2767-deepseek-v4-flash-0731-rollout8-instance674-trajectories3985
DeNovoSWE Distill 2767 — DeepSeek V4 Flash NL2Repo Trajectories
This dataset contains 3,985 difficulty-filtered NL2Repo SFT trajectories from 674 repository
instances. Each instance was sampled with eight rollouts using deepseek-v4-flash-0731.
Selection
Tasks receive a static difficulty score derived only from Stage 1–4 artifacts. The 1,141
successful tasks are split into five equal-count difficulty bins. A rollout is retained when its
evaluator score is greater… See the full description on the dataset page: https://huggingface.co/datasets/wjn922-01/denovoswe-distill2767-deepseek-v4-flash-0731-rollout8-instance674-trajectories3985.deepseek-v4-flash-instruct-308xTrace of DeepSeek V4 Flash LLM.
WARNING: This trace was made WITHOUT reasoning. Use it to finetune only instruct models.
Data count (Total: 308):
English - 198
Russian - 110
Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline.
This model was NOT free, and I had to use OpenRouter for it. Crypto donations for future projects like this are available on my personal page
DeepSeek_V4_Flash_distilled_dataset_5k
DeepSeek V4 Flash — Distilled Reasoning Dataset
A synthetic dataset of 5,099 unique reasoning traces designed to mirror the step-by-step thinking style of DeepSeek V4 Flash. Generated entirely with template-based parameterized generation (no LLM API calls).
Format
JSONL (one JSON object per line):
{
"id": "ds4f_math_000042",
"domain": "mathematics",
"subdomain": "algebra",
"difficulty": "easy",
"prompt": "Solve 3x + 7 = 22.",
"reasoning_trace":… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/DeepSeek_V4_Flash_distilled_dataset_5k.s1K-DeepSeek-V4-Flash-Max-Thinking
DeepSeek V4 Flash s1K Distillation
A curated collection of high quality distillation datasets, reasoning traces, and fine-tuning pipelines generated using the DeepSeek V4 Flash (Max Thinking) teacher model.
DeepSeek V4 Flash (Max Thinking) was selected as the teacher model for this distillation pipeline. Although the Max Thinking mode generates a high volume of output tokens per prompt compared to other models, its token pricing remains unmatched. The model offers exceptional… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/s1K-DeepSeek-V4-Flash-Max-Thinking.deepseek-v4-flash-filler-lens-demo
DeepSeek-V4-Flash filler-token lens captures — demo subset
Per-position logit-lens activations and top-k attention recorded from
deepseek-ai/DeepSeek-V4-Flash on a three-product arithmetic task, with and without
filler tokens.
This is the public demo subset (7 captures) of a larger private collection. It exists
so the attention viewer in the accompanying repo runs without special access.
Code, full results and write-up: https://github.com/safwanalbeshti/filler-effect-writeup… See the full description on the dataset page: https://huggingface.co/datasets/SafwanAlbeshti/deepseek-v4-flash-filler-lens-demo.med-synth-questions-gemma-3-27b-deepseek-v4-flash
Med Synth Questions (Gemma-3 + DeepSeek V4 Flash)
Synthetic reasoning traces and answers for medical questions from openmed-community/med-synth-questions-gemma-3-27b-it. Each record contains a medical question with SYNTH-style reasoning and a generated answer by DeepSeek V4 Flash.
Dataset Summary
29,148 records (2 dupes + 3,410 incomplete/truncated removed from 32,560 source)
29,148 reasoning turns (99.2% format compliance)
Average 1,591 chars per reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/med-synth-questions-gemma-3-27b-deepseek-v4-flash.deepseek-v4-flash-0731-frontierscience-research-trajectories
DeepSeek-V4-Flash-0731 FrontierScience Research Trajectories
This public dataset contains a completed sci-eval evaluation of
DeepSeek-V4-Flash-0731 on the public FrontierScience Research split.
Protocol
Candidate request model: api_deepseek_deepseek-v4-flash
Returned model identities: deepseek/deepseek-v4-flash and, for passthrough retries, deepseek-v4-flash
Dataset: openai/frontierscience, Research test split
Dataset revision:… See the full description on the dataset page: https://huggingface.co/datasets/Siyuc/deepseek-v4-flash-0731-frontierscience-research-trajectories.deepseek-v4-flash-reap-observations-v1
deepseek-v4-flash-reap-observations-v1
Observations/calibration artifacts collected from the deepseek-v4-flash lineage during REAP/quantization runs.
Provenance
Dataset provenance: contains model prompts, traces, and/or generated outputs collected during REAP/quantization runs. Output content follows the upstream model's usage terms, but no dataset license is asserted here. Contact the uploader before redistribution.
Acknowledgements
Cerebras… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/deepseek-v4-flash-reap-observations-v1.deepseek-v4-flash-ga-hs3-rawswebench-verified-deepseek-v4-flashscale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance1000-trajectories3379v3-2k-traj-deepseek-v4-flashswebench-multilingual-deepseek-v4-flashr2egym_deepseek_v4_flash_no_track_1390ir2egym_deepseek_v4_flash_1390ideepseek-v4-flash-swe-cot
DeepSeek-V4-Flash SWE Agent Trajectories (with raw chain-of-thought)
795 multi-turn software-engineering agent trajectories generated by
DeepSeek-V4-Flash-0731 at reasoning_effort=max, each one executed in a real
repository inside an isolated container and verified by running the repository's own
tests. 469 are verified-correct.
Every assistant turn preserves reasoning_content — the model's raw chain-of-thought,
not a summary. That is the point of this dataset: the DeepSeek API… See the full description on the dataset page: https://huggingface.co/datasets/blythet/deepseek-v4-flash-swe-cot.s1K-DeepSeek-V4-Flash-Max-Thinking
DeepSeek V4 Flash s1K Distillation
A curated collection of high quality distillation datasets, reasoning traces, and fine-tuning pipelines generated using the DeepSeek V4 Flash (Max Thinking) teacher model.
DeepSeek V4 Flash (Max Thinking) was selected as the teacher model for this distillation pipeline. Although the Max Thinking mode generates a high volume of output tokens per prompt compared to other models, its token pricing remains unmatched. The model offers exceptional… See the full description on the dataset page: https://huggingface.co/datasets/OnlyDanial/s1K-DeepSeek-V4-Flash-Max-Thinking.
