final-bench
OSC-Superconductor
OSC — Open Superconductor Challenge
Screen thousands of 2D materials for unconventional (d-wave) superconductivity — from your own laptop — and climb an open, verified leaderboard.
🏆 Leaderboard, live challenge & submission: https://huggingface.co/spaces/FINAL-Bench/OSC-Leaderboard
TL;DR (the short answer)
OSC is a free, open-science competition on Hugging Face to find the next candidate twisted-2D superconductor. We provide a computed effective model (t, U… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/OSC-Superconductor.ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.MDIT-Bench-Triple-FinalWorld-Model
🌍 World Model Bench (WM Bench) v1.0
Beyond FID — Measuring Intelligence, Not Just Motion
WM Bench is the world's first benchmark for evaluating the cognitive capabilities of World Models and Embodied AI systems.
🎯 Why WM Bench?
Existing world model evaluations focus on:
FID / FVD — image and video quality ("Does it look real?")
Atari scores — performance in fixed game environments
WM Bench measures something different: Does the model think correctly?… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/World-Model.Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
---
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/Metacognitive.Darwin-27B-JEV-decision-index
Darwin-27B-JEV — Decision Index run
Complete, untouched results of Darwin-27B-JEV on the frozen Decision Index 0.2 suite, produced with the official reproduction kit (apolinario/decision-index, commit 19ad28e) and its http engine.
Run
runs/darwin-27b-jev/
Rows / questions
158,927 / 593,768
Status
complete: true · errors 0 · unsupported 0
Decision Index (kit 0.2 scoring)
54.63
Decision Index (kit 0.2.1, edition-0.2.1/scores.json)
61.17
Engine… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/Darwin-27B-JEV-decision-index.
