Team Ai
20 results

final-bench

FINAL-Bench /OSC-Superconductor OSC — Open Superconductor Challenge Screen thousands of 2D materials for unconventional (d-wave) superconductivity — from your own laptop — and climb an open, verified leaderboard. 🏆 Leaderboard, live challenge & submission: https://huggingface.co/spaces/FINAL-Bench/OSC-Leaderboard TL;DR (the short answer) OSC is a free, open-science competition on Hugging Face to find the next candidate twisted-2D superconductor. We provide a computed effective model (t, U… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/OSC-Superconductor.tabular1K<n<10K39 likes8.1k downloads1d agoHugging FaceFINAL-Bench /ALL-Bench-Leaderboard 🏆 ALL Bench Leaderboard 2026 The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file. Dataset Summary ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.imagetext-generationn<1K26 likes797 downloads7mo agoHugging FaceAlbertmade /MDIT-Bench-Triple-Finalimage100K<n<1M0 likes651 downloads1y agoHugging FaceFINAL-Bench /World-Model 🌍 World Model Bench (WM Bench) v1.0 Beyond FID — Measuring Intelligence, Not Just Motion WM Bench is the world's first benchmark for evaluating the cognitive capabilities of World Models and Embodied AI systems. 🎯 Why WM Bench? Existing world model evaluations focus on: FID / FVD — image and video quality ("Does it look real?") Atari scores — performance in fixed game environments WM Bench measures something different: Does the model think correctly?… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/World-Model.textothern<1K48 likes460 downloads5mo agoHugging FaceFINAL-Bench /Metacognitive FINAL Bench: Functional Metacognitive Reasoning Benchmark "Not how much AI knows — but whether it knows what it doesn't know, and can fix it." --- Overview FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs). Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/Metacognitive.documenttext-generationn<1K108 likes385 downloads8mo agoHugging FaceFINAL-Bench /Darwin-27B-JEV-decision-index Darwin-27B-JEV — Decision Index run Complete, untouched results of Darwin-27B-JEV on the frozen Decision Index 0.2 suite, produced with the official reproduction kit (apolinario/decision-index, commit 19ad28e) and its http engine. Run runs/darwin-27b-jev/ Rows / questions 158,927 / 593,768 Status complete: true · errors 0 · unsupported 0 Decision Index (kit 0.2 scoring) 54.63 Decision Index (kit 0.2.1, edition-0.2.1/scores.json) 61.17 Engine… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/Darwin-27B-JEV-decision-index.1 likes249 downloads14d agoHugging Face