deepseek-v4-flash
terminal-bench-2.1-deepseek-v4-flash-trajectories
DeepSeek-V4-Flash on Terminal-Bench 2.1 — full agent trajectories
Every agent trajectory from a controlled scaffold comparison: the same model, the
same machine, the same 89 tasks, only the agent harness changed.
Scaffold
Solved
terminus-2 (Terminal-Bench's own agent)
53 / 89
dsh sdk-minimal (DeepSeek Harness)
61 / 89
Paired: dsh solved 18 tasks terminus-2 missed, terminus-2 solved 8 that dsh
missed, 43 both, 18 neither, 2 not scorable (see Caveats).… See the full description on the dataset page: https://huggingface.co/datasets/openguardrails/terminal-bench-2.1-deepseek-v4-flash-trajectories.DeepSeek-V4-Flash-0731-REAM-calibration-stats
DeepSeek-V4-Flash-0731 — expert calibration statistics (REAM line)
Layerwise routed-expert statistics of
deepseek-ai/DeepSeek-V4-Flash-0731
(43 MoE layers × 256 experts), collected by running the full model over a
~4.9M-token multi-domain calibration mix (multi-turn dialogs, thinking and
direct modes, rendered with the model's own chat encoder). These are the
statistics behind the REAM144/96 release line — published so that expert
selection, pruning, merging and routing research… See the full description on the dataset page: https://huggingface.co/datasets/WaveCut/DeepSeek-V4-Flash-0731-REAM-calibration-stats.swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.deepseek-v4-flash-rocm-vllm-repro
Reproducing DeepSeek-V4-Flash on AMD ROCm with vLLM: 32K Correctness and TopK Sweep
This article summarizes an engineering reproduction of
deepseek-ai/DeepSeek-V4-Flash on an AMD ROCm ModelScope DSW instance. The work
focuses on a practical question: can a complex, fast-moving DeepSeek-V4-Flash
serving path be turned into a reproducible ROCm baseline with explicit
correctness gates?
The answer from this run is yes, with an important boundary: the current setup
is a fallback-heavy… See the full description on the dataset page: https://huggingface.co/datasets/lyydfys/deepseek-v4-flash-rocm-vllm-repro.deepseek-v4-flash-0731-m3-ultra
DeepSeek-V4-Flash-0731 on M3 Ultra 512 GB — benchmark dataset
Independent performance characterization of
Vontra/DeepSeek-V4-Flash-0731-MXFP4-MLX
on a single Mac Studio M3 Ultra (80-core GPU, 512 GB unified memory).
Engine: omlx 0.5.7 · OS: macOS 26.6 (25G72) · MLX: 0.32.0
Recommended configuration
omlx serve --model-dir /opt/models --port 8033 \
--hot-cache-max-size 256GB --initial-cache-blocks 512
// ~/.omlx/model_settings.json
{"version": 1, "models":… See the full description on the dataset page: https://huggingface.co/datasets/guruswami-ai/deepseek-v4-flash-0731-m3-ultra.Wikipedia-FA-EN-DeepSeek-V4-Flash-0731
Wikipedia Persian to English — DeepSeek V4 Flash 0731
Rolling, machine-generated English translations of Persian Wikipedia articles
from Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge, configuration
full_articles_fa_without_en. 129,816 translations are
currently published in 26 immutable Parquet shards.
The target release contains 129,816 translations;
five source rows have empty plain_text and are not translated. Shards are
published only after 5,000 complete, validated records… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.
