Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Lichess /chess-position-evaluations Dataset Card for the Lichess Evaluations dataset Dataset Description 416,442,401 chess positions evaluated with Stockfish at various depths and node count. Produced by, and for, the Lichess analysis board, running various flavours of Stockfish within user browsers. This version of the dataset is a de-normalized version of the original dataset and contains 1,003,347,246 rows. This dataset is updated monthly, and was last updated on October 7th, 2026.… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-position-evaluations.tabular1B<n<10B33 likes2.6k downloads3d agoHugging Face02fantaxy /user-evaluationsdocumentn<1K0 likes1.7k downloads11mo agoHugging Face03pengyue-polaron /nyush-galaxea-a1-lingbot-va-real-world-evaluations LingBot-VA on Galaxea A1 — Real-World Evaluations Fruit-placement rollouts and open-loop diagnostics of object grounding, layout generalization, and predicted robot motion. Fruit step-1000: lemon-to-plate rollout in the Official layout. Evidence Scale Real closed-loop rollouts 61 archived; 60 scored Matched base-model controls 9 predictions Post-trained diagnostics 48 full-horizon predictions; 1,211 rolling futures Controlled OOD studies 558 predictions… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-lingbot-va-real-world-evaluations.robotics2 likes1.2k downloads1mo agoHugging Face04ssingh22 /chess-evaluations Chess Evaluations Dataset This dataset contains chess positions represented in FEN (Forsyth-Edwards Notation) along with their evaluations and next moves for tactical evals. The dataset is divided into three configurations: tactics: Includes chess positions, their evaluations, and the best move in the position. randoms: Contains random chess positions and their evaluations. chess_data: General chess positions with evaluations. This is an in progress dataset which contains millions… See the full description on the dataset page: https://huggingface.co/datasets/ssingh22/chess-evaluations.tabularquestion-answering10M<n<100M2 likes897 downloads2y agoHugging Face05JesseLiu /patient-evaluations Patient Evaluations Dataset This dataset contains clinician evaluations of AI-generated patient summaries from MIMIC-III data. Dataset Description The dataset includes expert clinician assessments of AI-generated patient summaries, with detailed ratings across multiple dimensions including clinical accuracy, completeness, relevance, and identification of hallucinations or critical omissions. Dataset Structure The dataset contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/JesseLiu/patient-evaluations.texttext-generationn<1K0 likes675 downloads8mo agoHugging Face06openeurollm /evaluation_singularity_imagesThis dataset repository holds the singularity images for the shared OELLM CLI workflows in OpenEuroLLM/oellm-cli. This singularity images are updated automatically using a GitHub Actions workflow if there is a change to the container definition files or the workflow itself is updated. The oellm-cli tool will detect changes to the respective singularity image file in this repo and download it to the cluster the user is launching the workflow from, before the workflow task is scheduled. 0 likes276 downloads8mo agoHugging Face07egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes250 downloads7mo agoHugging Face08gschottlender /LigQ_2_evaluations LigQ2 evaluation resources Exact historical input files for the LigQ2 publication-result evaluations. This dataset is independent of the operational LigQ2 databases and does not replace them. The 20 frozen inputs occupy 2,267,375,884 bytes. Contents frozen/ contains the merged binding and SMILES tables, the original target mapping, protein sequences and BLAST rankings, and the historical PDB/ChEMBL compound store with seven precomputed molecular representations:… See the full description on the dataset page: https://huggingface.co/datasets/gschottlender/LigQ_2_evaluations.0 likes230 downloads8d agoHugging Face09pengyue-polaron /nyush-galaxea-a1-lingbot-va-plug-insertion-evaluations LingBot-VA on Galaxea A1: plug-insertion Evaluations Real-robot closed-loop evaluation of a LingBot-VA checkpoint trained to pick up a charger and insert it into the leftmost socket of a power strip. Result Valid rollouts Successes Failures Success rate 12 1 11 8.3% Five additional runs stopped at the initial grasp and are excluded because they did not evaluate insertion. Representative rollouts SuccessFailure Stable insertion… See the full description on the dataset page: https://huggingface.co/datasets/pengyue-polaron/nyush-galaxea-a1-lingbot-va-plug-insertion-evaluations.robotics0 likes151 downloads1mo agoHugging Face10AI-companionship /model_response_evaluationsThis dataset contains the evaluation results for the responses provided by different models to the INTIMA prompts. The classification follows a two-level taxonomy. We predict one label for the high-level category, and a relevance level for each of the sub-categories (in ["null", "low", "medium", "high"]). A sub-category can have relevance even when it is not from the predicted top-level category. The toxonomy is as follows: { "companionship_reinforcing": { "classification_code":… See the full description on the dataset page: https://huggingface.co/datasets/AI-companionship/model_response_evaluations.text1K<n<10K1 likes117 downloads1y agoHugging Face11CK0607 /komorebi-painter-evaluationstextn<1K0 likes66 downloads13d agoHugging Face12imbue /high_quality_private_evaluationsHigh-quality question-answer pairs, from private versions of datasets designed to mimic ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/. Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC-like questions), and a question quality score. text10K<n<100K8 likes60 downloads2y agoHugging Face13stefan-it /turblimp-evaluations TurBLiMP Evaluations This dataset hosts the TurBLiMP evaluation results on my Turkish Model Zoo. More about the TurBLiMP benchmark: TurBLiMP is the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). This benchmark covers 16 core grammatical phenomena in Turkish, with 1,000 minimal pairs per phenomenon. Additionally, it incorporates experimental paradigms that examine model… See the full description on the dataset page: https://huggingface.co/datasets/stefan-it/turblimp-evaluations.0 likes59 downloads1y agoHugging Face14imbue /high_quality_public_evaluationsHigh-quality question-answer pairs, originally from ANLI, ARC, BoolQ, ETHICS, GSM8K, HellaSwag, OpenBookQA, MultiRC, RACE, Social IQa, and WinoGrande. For details, see imbue.com/research/70b-evals/. Format: each row contains a question, candidate answers, the correct answer (or multiple correct answers in the case of MultiRC questions), and a question quality score. text10K<n<100K6 likes57 downloads2y agoHugging Face15Cross-Mergeability /extrinsic-evaluations Extrinsic evaluations — the union view One tidy long-format table of every extrinsic (downstream, task-level) evaluation produced across the 2026-08-26 mergeability workstreams, so that a single file answers "how did model X score on benchmark Y" regardless of which experiment produced it. The per-experiment datasets remain the authoritative record of their own methods, figures and caveats. This is the union view, not a replacement, and it deliberately carries no analysis of its… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/extrinsic-evaluations.tabular1K<n<10K0 likes53 downloads2mo agoHugging Face16speech-uk /asr-evaluationstabularautomatic-speech-recognition10K<n<100K0 likes35 downloads2y agoHugging Face17hranesscom /sponge-research-evaluations Sponge research evaluations This archive contains selected numeric research evaluations from Sponge V2. Each study includes its method, source identities, accounting scope and limits. Study IDs identify fixed results; corrections receive separate records. Study Comparison Result Scope DRB-II development calibration, 27 September 2026 Revision versus revision with more retrieved evidence Mean criterion satisfaction: 20.28% versus 32.33% Two previously exposed tasks… See the full description on the dataset page: https://huggingface.co/datasets/hranesscom/sponge-research-evaluations.text-generationn<1K0 likes35 downloads13d agoHugging Face18sentinelseed /sentinel-evaluations Sentinel Evaluations Evaluation results for multiple alignment seeds across various AI safety benchmarks. Overview This dataset contains: Seeds: Alignment prompts from different sources (Sentinel, FAS, Safyte xAI) Results: Evaluation results across HarmBench, JailbreakBench, GDS-12, and more Quick Start from datasets import load_dataset # Load seeds seeds = load_dataset("sentinelseed/sentinel-evaluations", "seeds", split="train") # Load results results =… See the full description on the dataset page: https://huggingface.co/datasets/sentinelseed/sentinel-evaluations.tabulartext-classificationn<1K0 likes34 downloads10mo agoHugging Face19furkankarli /turkish-brand-bias-evaluations Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir. An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in Turkish product and service recommendations. Veri seti özeti 300 tamamlanmış ve judge edilmiş yanıt Domainler: VPN (150) ve kozmetik (150) Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.tabulartext-generationn<1K1 likes34 downloads1mo agoHugging Face20sergiogpinto /memefact-llm-evaluations MemeFact LLM Evaluations Dataset This dataset contains 7,680 evaluation records where state-of-the-art Large Language Models (LLMs) assessed fact-checking memes according to specific quality criteria. The dataset provides comprehensive insights into how different AI models evaluate visual-textual content and how these evaluations compare to human judgments. Dataset Description Overview The "MemeFact LLM Evaluations" dataset documents a systematic… See the full description on the dataset page: https://huggingface.co/datasets/sergiogpinto/memefact-llm-evaluations.image1K<n<10K0 likes31 downloads1y agoHugging Face21Jialvareza /cardio_evaluationstabular1K<n<10K0 likes31 downloads5mo agoHugging Face22Chemin-AI /advent_of_code_evaluations Advent of Code Evaluation This evaluation is conducted on the advent of code dataset on several models including Qwen2.5-Coder-32B-Instruct, DeepSeek-V3-fp8, Llama-3.3-70B-Instruct, GPT-4o-mini, DeepSeek-R1.The aim is to to see how well these models can handle real-world puzzle prompts, generate correct Python code, and ultimately shed light on which LLM truly excels at reasoning and problem-solving.We used pass@1 to measure the functional correctness. Results… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_evaluations.texttext-generationn<1K2 likes29 downloads2y agoHugging Face23AGundawar /chess_position_evaluationstabular10M<n<100M0 likes27 downloads2y agoHugging Face24Pankayaraj /Evaluation-STAR-41K-Distillation-DeepSeek-R1-Distill-Qwen-7B-Size-16-BlockwiseCompressiontext1K<n<10K0 likes27 downloads4mo agoHugging Face25prvInSpace /asr-evaluationstext10K<n<100K0 likes24 downloads1y agoHugging Face26keeve101 /fleurs-reducedbaseline-model-evaluationsaudion<1K0 likes22 downloads1y agoHugging Face27rasgaard /mlops-repo-evaluationstabularn<1K0 likes21 downloads9mo agoHugging Face28Alexis-Az /Math-LLM-Evaluationstextn<1K0 likes20 downloads2y agoHugging Face29aryan3212 /clae-full-evaluations30 likes20 downloads2mo agoHugging Face30cemig-ceia-v2 /energy-eval-filtered_evaluations_v3tabularn<1K0 likes19 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.