FINAL-Bench/Darwin-180B-RSI
Darwin-180B-RSI
180B Mixture-of-Experts · vision-language · #1 on ten Hugging Face official leaderboards: nine by this model (MDPBench 83.65 · IFStruct 98.95 · AIME 2026 100 · HMMT Feb 2026 100 · GPQA Diamond 94.44 · MMLU-Pro 88.12 · MMMU-Pro 79.48 · LEXam 68.94 · LEXam-hard 45.72) and ExtractBench 90.29 by its next round R3 · self-improving
💻 Run it on your own machine — [POCKET-Darwin-180B-GGUF](https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF): the 4-bit GGUF of R3 (111 GB) runs on a laptop with an 8 GB GPU and 32 GB RAM, CPU only at 18–21 tok/s, a 128 GB mini PC or one DGX Spark — MMLU-Pro identical to BF16 (87.65%).
🥇 NEW: #1 on [MDPBench](https://huggingface.co/datasets/Delores-Lin/MDPBench) (83.65), multilingual document parsing across 17 languages, digital and photographed pages, ahead of Kimi-K3 (83.6) and the dedicated OCR models.
🥇 NEW: #1 on [IFStruct](https://huggingface.co/datasets/LiquidAI/ifstruct-v1.0) (98.95), the structured-output compliance benchmark (valid JSON/YAML that follows a requested schema), ahead of Agents-A1 (93.25) and gpt-oss-20b (91.95).
🥇 NEW: [Darwin-180B-RSI-R3](https://huggingface.co/FINAL-Bench/Darwin-180B-RSI-R3), the next self-improvement round of this model, is #1 on the [ExtractBench](https://huggingface.co/datasets/llamaindex/ExtractBench) leaderboard (90.29) and #3 on EvasionBench (77.83).
reasoning · MoE 512 experts · 262K long context · image + text · Korean + English · self-improvement · ZTC
<p align="center"> <a href="https://vidraft.net"><img src="https://img.shields.io/badge/🌐VIDRAFT-vidraft.net-111827?style=for-the-badge"></a> <a href="https://huggingface.co/datasets/Idavidrein/gpqa"><img src="https://img.shields.io/badge/GPQADiamond-94.44%25%231-gold?style=for-the-badge"></a> <a href="https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro"><img src="https://img.shields.io/badge/MMLU--Pro-88.12%25%231-2563eb?style=for-the-badge"></a> <a href="https://huggingface.co/datasets/MMMU/MMMUPro"><img src="https://img.shields.io/badge/MMMU--Pro-79.48%25%231-0891b2?style=for-the-badge"></a> <a href="https://huggingface.co/datasets/MathArena/aime2026"><img src="https://img.shields.io/badge/AIME2026-100%25%231-dc2626?style=for-the-badge"></a> <a href="https://huggingface.co/datasets/MathArena/hmmtfeb2026"><img src="https://img.shields.io/badge/HMMTFeb2026-100%25%231-ea580c?style=for-the-badge"></a> <a href="https://huggingface.co/datasets/LEXam-Benchmark/LEXam"><img src="https://img.shields.io/badge/LEXam(Law)-68.94%25%231-4f46e5?style=for-the-badge"></a> <a href="https://huggingface.co/datasets/joelniklaus/LEXam-hard"><img src="https://img.shields.io/badge/LEXam--hard(Law)-45.72%231-6d28d9?style=for-the-badge"></a> <a href="https://huggingface.co/datasets/LiquidAI/ifstruct-v1.0"><img src="https://img.shields.io/badge/IFStruct-98.95%25%231-b45309?style=for-the-badge"></a> <a href="https://huggingface.co/datasets/Delores-Lin/MDPBench"><img src="https://img.shields.io/badge/MDPBench-83.65%231-0f766e?style=for-the-badge"></a> <a href="https://huggingface.co/datasets/llamaindex/ExtractBench"><img src="https://img.shields.io/badge/ExtractBench(R3)-90.29%231-059669?style=for-the-badge"></a> <a href="https://huggingface.co/FINAL-Bench/Darwin-27B-RSI"><img src="https://img.shields.io/badge/Self--Improving-RSI-e11d48?style=for-the-badge"></a> <a href="https://huggingface.co/collections/FINAL-Bench/ztc-models-jev-ecosystems"><img src="https://img.shields.io/badge/ZTC-Zero--TokenConfidence-7c3aed?style=for-the-badge"></a> </p> <p align="center"> <a href="https://arxiv.org/abs/2605.14386"><img src="https://img.shields.io/badge/arXiv-2605.14386DarwinFamily-b31b1b?style=for-the-badge"></a> <a href="https://huggingface.co/papers/2609.20269"><img src="https://img.shields.io/badge/Paper-2609.20269LatinSquare-b31b1b?style=for-the-badge"></a> <a href="https://huggingface.co/collections/FINAL-Bench/darwin-family"><img src="https://img.shields.io/badge/🧬Collection-DarwinFamily-16a34a?style=for-the-badge"></a> <a href="https://huggingface.co/collections/FINAL-Bench/ztc-models-jev-ecosystems"><img src="https://img.shields.io/badge/🏛️Collection-ZTCModels-7c3aed?style=for-the-badge"></a> <a href="https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF"><img src="https://img.shields.io/badge/💻POCKET4--bit-Laptop·CPUonly·DGX_Spark-0f766e?style=for-the-badge"></a> </p>
The newest flagship of the Darwin family: #1 on MDPBench, IFStruct, AIME 2026, HMMT Feb 2026, GPQA Diamond, MMLU-Pro, MMMU-Pro, LEXam and LEXam-hard, and a model that gets better by learning from its own verified work.
🏆 Head-to-head with Chinese frontier models


Scores as listed on the Hugging Face official benchmark leaderboards (self-reported by each model's publisher). 🥇 = #1 on that leaderboard. "—" = not reported.
<sub>This comparison covers open-weight models listed on the Hugging Face official benchmark leaderboards; closed API models are not included. Leaderboard values are each publisher's own reported numbers; settings (samples, voting, thinking budget) differ across models. Darwin-180B-RSI settings are listed in the evaluation protocol below.</sub>
🧬 The Darwin Family
<p align="center"> <a href="https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC"><img src="https://img.shields.io/badge/Darwin--397B--ZTC-GPQA93.43-16a34a"></a> <a href="https://huggingface.co/FINAL-Bench/Darwin-28B-REASON"><img src="https://img.shields.io/badge/Darwin--28B--REASON-GPQA89.39-16a34a"></a> <a href="https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus"><img src="https://img.shields.io/badge/Darwin--35B--A3B--Opus-♥98-e11d48"></a> <a href="https://huggingface.co/FINAL-Bench/Darwin-36B-Opus"><img src="https://img.shields.io/badge/Darwin--36B--Opus-♥97-e11d48"></a> <a href="https://huggingface.co/FINAL-Bench/Darwin-4B-Genesis"><img src="https://img.shields.io/badge/Darwin--4B--Genesis-♥63-e11d48"></a> </p> <p align="center"> <a href="https://huggingface.co/FINAL-Bench/Darwin-9B-NEG"><img src="https://img.shields.io/badge/Darwin--9B--NEG-♥57-e11d48"></a> <a href="https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF"><img src="https://img.shields.io/badge/POCKET--35B-824K↓-1f6feb"></a> <a href="https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF"><img src="https://img.shields.io/badge/POCKET--26B-365K↓-1f6feb"></a> <a href="https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF"><img src="https://img.shields.io/badge/POCKET--EN-♥43-1f6feb"></a> <a href="https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF"><img src="https://img.shields.io/badge/POCKET--KR-♥36-1f6feb"></a> <a href="https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF"><img src="https://img.shields.io/badge/POCKET--Darwin--180B-NEW·4--bit·laptop-1f6feb"></a> </p>
Darwin is VIDRAFT's measurement-driven reasoning model family — 50+ official models, 400+ community derivatives, and now two places in the GPQA Diamond top 3 (Darwin-180B-RSI #1 · Darwin-397B-ZTC #3).
🧬 Darwin — evolve the parent, keep what works
Darwin treats a strong open model as a parent. It measures where the parent is weak, and strengthens exactly those parts — instead of re-training everything and risking what already works.
- Diagnose before you change. Every Darwin generation starts from a measured weakness map of the parent.
- Change little, precisely. Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved.
- Proven capability over new guesses. Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; Darwin-180B-RSI adds a new ingredient — the model's own verified work.
- Measured, not claimed. Every change must beat the parent on held-out tests before it ships.
Lineage
📄 Darwin Platform & Research
- Darwin Family — MRI trust-weighted evolutionary merging for training-free scaling of language-model reasoning (arXiv:2605.14386)
- Placement Is Free, Composition Is Not — the Latin square as a provably-balanced construction for heterogeneous sequence-mixer stacks (2609.20269) — the AETHER architecture line
- FINAL Bench — VIDRAFT's measurement-driven evaluation framework (SSRN)
- Four-layer Pre-AGI roadmap — Darwin → AETHER → PROMETHEUS → HEPHAESTUS
- Collections: Darwin Family · ZTC Models — JEV ecosystems
🔁 RSI — a model that improves from its own work
Recursive self-improvement (RSI) is the core of this generation. Instead of distilling a bigger teacher, the model improves by learning from itself:
- Solve — the model works through practice problems it has never seen in evaluation.
- Verify — its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned.
- Learn — it is re-trained on the reasoning that turned out to be correct.
- Repeat — the improved model becomes the next solver.
What it bought in this release:
Same or better accuracy with shorter reasoning — cheaper and faster to serve. Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter).
Next round: Darwin-180B-RSI-R3
Darwin-180B-RSI-R3 is the next RSI round, trained from this model with the same recipe. Its own leaderboard entries (submitted 2026-10-05):
Model-level RSI vs. harness-level RSI
Darwin-180B-RSI is Model-level RSI: the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically. Harness-level RSI (e.g., Google's RRSI) improves the prompts, tools and workflow around a fixed model. It's like rewriting an employee's manual, while Model-level RSI is the employee getting smarter. The two are complementary.
🏛️ ZTC — it knows before it answers
Zero-Token Confidence (ZTC) reads the model's own internal state once, before generation, and returns the probability that the answer it is about to give is correct — no extra tokens, no second model.
{"answer": "...", "confidence": 0.93, "ztc_score": 1.84, "truncated": false}Use it to gate actions: when confidence is low, do not call the tool, escalate, or answer "I don't know".
What ships with this model
from handler import EndpointHandler
h = EndpointHandler("FINAL-Bench/Darwin-180B-RSI") # local snapshot path
print(h({"inputs": "What is 17 * 23?"}))
# [{"answer": "...391...", "confidence": 0.97, "ztc_score": 2.1, "truncated": false}]Same format as Darwin-397B-ZTC. The readout is fitted only on practice data that is disjoint from every benchmark reported here.
Status: early release. This probe reaches an AUROC of 0.64 on our held-out validation split: a coarse signal for routing and review, not a correctness guarantee. A retrained probe with more data will replace it. Details in `ztc/README.md`.
🏆 Results
Evaluation protocol
Common to every benchmark
Per benchmark
All numbers are self-measured and reproducible with the settings above. Majority-vote scores are system scores (several samples per question) and are labeled as such.
MMLU-Pro by category (single sample) — strongest in math 95.0 · biology 94.6 · physics 92.5; room to grow in law and history.
⚙️ Specifications
🚀 Quickstart
Serving with vLLM (8 × B200 or equivalent)
vllm serve FINAL-Bench/Darwin-180B-RSI \
--tensor-parallel-size 8 --enable-expert-parallel \
--max-model-len 135168 --trust-remote-codeChat Completions (OpenAI-compatible)
from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI",
messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}],
temperature=1.0, top_p=0.95, extra_body={"top_k": 20})
print(r.choices[0].message.content)Transformers
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "FINAL-Bench/Darwin-180B-RSI"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")Tip: this is a thinking model. Give it room — a thinking budget of 32K–131K tokens is recommended for hard reasoning. Short budgets truncate the reasoning and cost accuracy.
⚠️ Limitations and disclosure
- Scores are self-measured with the settings stated in the Results table; majority-vote numbers use several samples per question.
- Very long reasoning is normal for hard problems; a short thinking budget will truncate answers and lower accuracy.
- Like every LLM, the model can be confidently wrong — use the ZTC confidence readout to gate high-stakes actions.
🔗 Related Darwin Models
- [Darwin-180B-RSI-R3](https://huggingface.co/FINAL-Bench/Darwin-180B-RSI-R3): the next RSI round of this model, ExtractBench #1 and EvasionBench #3
- [Darwin-397B-ZTC](https://huggingface.co/FINAL-Bench/Darwin-397B-ZTC) — 397B MoE (FP8), GPQA Diamond 93.43 %, ZTC on board
- [Darwin-28B-REASON](https://huggingface.co/FINAL-Bench/Darwin-28B-REASON) — 28B, GPQA Diamond 89.39 %
- [Darwin-36B-Opus](https://huggingface.co/FINAL-Bench/Darwin-36B-Opus) — 36B MoE, GPQA Diamond 88.4 %
- [Darwin-27B-RSI](https://huggingface.co/FINAL-Bench/Darwin-27B-RSI) — 27B, the first Darwin RSI model
- [Darwin-9B-NEG](https://huggingface.co/FINAL-Bench/Darwin-9B-NEG) — 9B with Negentropy distillation, GPQA Diamond 84.3 %
- [ZTC-Judge-27B](https://huggingface.co/FINAL-Bench/ZTC-Judge-27B) — standalone ZTC judge
📚 Citation
@misc{darwin180b_rsi_2026,
title = {Darwin-180B-RSI: Recursive Self-Improvement on Verified Answers for a 180B Mixture-of-Experts Reasoning Model},
author = {FINAL-Bench / Darwin Research Team},
year = {2026},
howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-180B-RSI}},
note = {GPQA Diamond 94.44 \% · MMLU-Pro 88.12 \%}
}
@misc{darwin_family_2026,
title = {Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning},
author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Shin, Junghoon},
year = {2026},
eprint = {2605.14386},
archivePrefix = {arXiv}
}
@misc{latin_square_2026,
title = {Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks},
author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Kim, Minseo},
year = {2026},
eprint = {2609.20269},
archivePrefix = {arXiv}
}📜 License
Darwin-180B-RSI is a derivative of Qwen3.8-Flash-Next and is distributed under the Qwen Community License 1.0 (see LICENSE).
🏢 About
Built by [VIDRAFT](https://vidraft.net) · evaluated with FINAL-Bench.
This model is part of the Darwin Family.
