FINAL-Bench/Darwin-4B-David
### ๐ฑ Run it on your phone or a GPU-less PC โ POCKET ยท ๐ [Try it live (CPU chat)](https://huggingface.co/spaces/FINAL-Bench/POCKET-35B-CPU) VIDRAFT's on-device family: a 35B model that runs on iPhone and on CPU with no GPU โ stock llama.cpp, no fork.     Darwin-4B-David โ The First Second-Generation Darwin Model
<p align="center"> <a href="https://huggingface.co/FINAL-Bench/Darwin-4B-Opus"><img src="https://img.shields.io/badge/๐งฌGen1-Darwin--4B--Opus-blue?style=for-the-badge" alt="Gen1"></a> <a href="https://huggingface.co/FINAL-Bench/Darwin-4B-David"><img src="https://img.shields.io/badge/๐งฌGen2-Darwin--4B--David-blue?style=for-the-badge" alt="Gen2"></a> <a href="https://huggingface.co/FINAL-Bench/Darwin-4B-Genesis"><img src="https://img.shields.io/badge/โญ_Gen3-Darwin--4B--Genesis-gold?style=for-the-badge" alt="Gen3"></a> </p>
<p align="center"> <a href="https://huggingface.co/FINAL-Bench/Darwin-9B-Opus"><img src="https://img.shields.io/badge/๐งฌModel-Darwin--9B--Opus-blue?style=for-the-badge" alt="9B"></a> <a href="https://huggingface.co/spaces/FINAL-Bench/Darwin-9B-Opus"><img src="https://img.shields.io/badge/๐Space-9BDemo-purple?style=for-the-badge" alt="9B Space"></a> <a href="https://huggingface.co/FINAL-Bench/Darwin-31B-Opus"><img src="https://img.shields.io/badge/๐งฌModel-Darwin--31B--Opus-blue?style=for-the-badge" alt="31B"></a> <a href="https://huggingface.co/spaces/FINAL-Bench/Darwin-31B-Opus"><img src="https://img.shields.io/badge/๐Space-31BDemo-purple?style=for-the-badge" alt="31B Space"></a> </p>
<p align="center"> <a href="https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus"><img src="https://img.shields.io/badge/๐งฌModel-Darwin--35B--A3B--Opus-blue?style=for-the-badge" alt="35B"></a> <a href="https://huggingface.co/spaces/FINAL-Bench/Darwin-35B-A3B-Opus"><img src="https://img.shields.io/badge/๐Space-35BDemo-purple?style=for-the-badge" alt="35B Space"></a> <a href="https://huggingface.co/FINAL-Bench/Darwin-35B-A3B-Opus-Q8-GGUF"><img src="https://img.shields.io/badge/๐ฆGGUF-Q8--Official-yellow?style=for-the-badge" alt="Q8 GGUF"></a> <a href="https://huggingface.co/bartowski/FINAL-BenchDarwin-35B-A3B-Opus-GGUF"><img src="https://img.shields.io/badge/๐ฆGGUF-bartowski-yellow?style=for-the-badge" alt="bartowski GGUF"></a> </p>
<p align="center"> <a href="https://huggingface.co/spaces/FINAL-Bench/Leaderboard"><img src="https://img.shields.io/badge/๐FINALBench-Leaderboard-green?style=for-the-badge" alt="FINAL Bench"></a> <a href="https://huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard"><img src="https://img.shields.io/badge/๐ALLBench-Leaderboard-orange?style=for-the-badge" alt="ALL Bench"></a> </p>
Gemma 4 E4B Dense | 4.5B Params | Thinking Mode | 128K Context | 140+ Languages | BF16 | Apache 2.0 The first-ever second-generation Darwin model โ "Evolution of Evolution"
Overview
Darwin-4B-David is the first second-generation (Generation 2) model in Darwin history โ a model evolved from an already-evolved model.
The first-generation Darwin-4B-Opus (Father) was evolved from the original gemma-4-E4B-it using the Darwin V6 engine. Darwin-4B-David was born by crossbreeding this first-generation evolved model with DavidAU's DECKARD-Expresso-Universe (Mother). This is the first realization of Darwin's core concept: "Merge = Evolve" applied recursively.
The name "David" pays tribute to the Mother model's creator DavidAU, while evoking the biblical David who defeated Goliath โ symbolizing how a 4.5B small model challenges models many times its size.
Family Tree
<p align="center"> <img src="family.png" alt="Darwin-4B-David" width="100%"> </p>
Generation Comparison
Parent Models
Model Diagnostic Scan (MDS)
<p align="center"> <img src="s1.png" alt="Father (Darwin-4B-Opus) MDS Scan" width="48%"> <img src="s2.png" alt="Mother (DECKARD-Expresso-Universe) MDS Scan" width="48%"> </p>
Left: Father (Darwin-4B-Opus) โ REASONING concentration in later layers (dist 0.4), MATH activation throughout. Already optimized through Gen-1 evolution. Right: Mother (DECKARD-Expresso-Universe) โ Strong KOREAN hotspot (dist 1.5), signature of Unsloth deep tuning. Remaining regions show uniform distribution.
Benchmarks
Key Results
GPQA Diamond Evaluation Details
GPQA Diamond (graduate-level scientific reasoning) was evaluated using generative (thinking mode) evaluation.
Why maj@8:
- Single-sample (greedy/pass@1) is vulnerable to stochastic variation with do_sample
- 8 independent generations with majority voting reflects the model's stable reasoning capability
- maj@k is standard practice in frontier model benchmarks (AIME, MATH, etc.)
Note on 50-question sampling:
- GPQA Diamond contains 198 questions total; 50 questions represent 25.3% of the full set
- 50 questions ร 8 samples = 400 total generations, providing sufficient statistical confidence
- Full 198-question evaluation is planned
Note on lm-eval Loglikelihood Results
ARC-Challenge and KMMLU show identical scores to the original model. This is characteristic of DARE-TIES merging: the loglikelihood method compares token probabilities across answer choices and does not capture differences in generation quality, reasoning chains, or creativity. The evolution effect is clearly visible in generative evaluation (GPQA Diamond), where the difference emerges during step-by-step thinking mode reasoning.
MRI-Guided Evolution Recipe
Darwin V6's Model MRI scanned weight divergence across all 42 layers and automatically assigned independent weight ratios to each layer.
Parent Comparison
<p align="center"> <img src="parent_comparison.png" alt="Father vs Mother layer-wise importance comparison" width="100%"> </p>
Evolution Parameters
Darwin V6 vs Conventional Merging
Significance of Second-Generation Evolution
- Proof of "Evolution of Evolution": The first systematic case of recursive evolution (2+ generations) in the open-source model merging community. Darwin V6 + MRI automates the entire process.
- 85% GPQA Diamond at 4.5B parameters: +26.4%p over the original 58.6%. This surpasses the 31B-class gemma-4-31B (84.3%) with only 4.5B parameters โ an exceptional result in parameter efficiency.
- Apache 2.0 + Edge deployment: Preserves the Gemma 4 E4B architecture, enabling deployment on Jetson Orin NX 16GB and consumer GPUs with no commercial restrictions.
- Multimodal preservation: Father's vision encoder (~150M) and audio encoder (~300M) are frozen during evolution, maintaining image/video/audio input capabilities.
- Community synergy: Mother model creator DavidAU is an active contributor on HuggingFace. Darwin-4B-David symbolizes collaborative evolution within the open-source ecosystem.
Model Specifications
Usage
Transformers
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
tokenizer = AutoTokenizer.from_pretrained("FINAL-Bench/Darwin-4B-David", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
"FINAL-Bench/Darwin-4B-David",
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
)
messages = [{"role": "user", "content": "Prove that sqrt(2) is irrational."}]
text = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=4096, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))Disable Thinking Mode
text = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)VRAM Requirements
Darwin Opus Family
Roadmap
- Full 198-question GPQA Diamond evaluation (maj@8)
- MTI (Minimal Test-Time Intervention) serving โ expected additional +9-11% reasoning accuracy
- GRPO + TinyLoRA reinforcement learning
- SSD self-distillation
- Cross-architecture breeding research (Transformer ร Mamba FFN transplantation)
References
- DARE-TIES: Yadav et al., 2023 (https://arxiv.org/abs/2311.03099) โ re-implemented, not library-dependent
- Darwin V6 Engine: https://huggingface.co/spaces/ginigen-ai/DARWIN-V5-BACKUP
- FINAL Bench: https://huggingface.co/spaces/FINAL-Bench/Leaderboard
- DavidAU DECKARD Series: https://huggingface.co/DavidAU
- MTI: Minimal Test-Time Intervention (arXiv:2510.13940)
Built By
Citation
@misc{vidraft_darwin_4b_david_2026,
title = {Darwin-4B-David: First Second-Generation Evolutionary Merge Model},
subtitle = {Recursive Evolution Achieves 85\% GPQA Diamond with 4.5B Parameters},
author = {VIDRAFT},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-4B-David}}
}This model is introduced in Darwin Family.
