Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /frames-benchmark FRAMES: Factuality, Retrieval, And reasoning MEasurement Set FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning. Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941. Dataset Overview 824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.texttext-classificationn<1K267 likes15k downloads2y agoHugging Face02NoeFlandre /benchmark-llms-landuse-relevance Land-use relevance benchmark v3-multilingual · 85 languages x 300 items/language · 25,500 items · binary yes/no labels. Code Package version recorded in run metadata: 0.2.0 (some runs lack version metadata). Task and prompt Does a sentence describe a place's land or environment in ways visible to satellites? English prompt · greedy decoding · seed 0 · max_new_tokens=4096 · bfloat16 · batch varies by model. unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.tabulartext-classification10K<n<100K0 likes3.2k downloads8d agoHugging Face03NoeFlandre /geoparser-benchmark-results Geoparser benchmark results Geoparsing pipelines from geoparser scored on English and multilingual corpora, run on Grid'5000 (one Tesla T4). Benchmarks Benchmark Languages Docs Toponyms Source GeoVirus en 229 2167 WikiNews articles on epidemics (Gritta et al., 2018). HIPE-2020 de, en, fr 129 1516 Historical Swiss, Luxembourgish and American newspapers, OCR. NewsEye de, fi, fr, sv 77 1772 Historical European newspapers, OCR (HIPE-2022). TopRes19th… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/geoparser-benchmark-results.tabularn<1K0 likes2.8k downloads2d agoHugging Face04L-FAME-Dataset-Benchmark /L-FAME L-FAME: Longitudinal Focused Attention Meditation EEG Dataset and Benchmark A longitudinal 64-channel EEG dataset and benchmark for studying focused attention meditation (FAM) and how its neural signatures evolve across a six-week training period. 74 healthy adults were recorded at a pre-intervention baseline; 44 of them returned for a post-intervention follow-up. Three FAM techniques are systematically compared: Hare Krishna mantra (HK), SA-TA-NA-MA mantra (SA), and Breath… See the full description on the dataset page: https://huggingface.co/datasets/L-FAME-Dataset-Benchmark/L-FAME.textothern<1K0 likes2.4k downloads4mo agoHugging Face05inria-soda /tabular-benchmark Tabular Benchmark Dataset Description This dataset is a curation of various datasets from openML and is curated to benchmark performance of various machine learning algorithms. Repository: https://github.com/LeoGrin/tabular-benchmark/community Paper: https://hal.archives-ouvertes.fr/hal-03723551v2/document Dataset Summary Benchmark made of curation of various tabular data learning tasks, including: Regression from Numerical and Categorical Features… See the full description on the dataset page: https://huggingface.co/datasets/inria-soda/tabular-benchmark.tabulartabular-classification10M<n<100M51 likes2.3k downloads3y agoHugging Face06vals-ai /finance_agent_benchmark Finance Agent Benchmark Dataset We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings. We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.textn<1K10 likes1.7k downloads1y agoHugging Face07witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.5k downloads16d agoHugging Face08scaledown /vllm-inference-benchmarkstextn<1K0 likes1.1k downloads8d agoHugging Face09AbBibench /Antibody_Binding_Benchmark_DatasetWe introduce AbBiBench (Antibody Binding Benchmarking), a benchmarking framework for optimizing antibody binding affinity. This dataset contains the sequences of mutants, experimentally measured affinity values, and the structures of antigen-antibody complexes. Mutant structure files for AbBiBench dataset is available at https://zenodo.org/records/16557372 text100K<n<1M8 likes994 downloads5mo agoHugging Face10Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Gandalf Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K7 likes984 downloads11mo agoHugging Face11redmadrobot-rnd /pii_benchmark Russian PII NER Evaluation Dataset Dataset Description This dataset is designed for evaluating PII (Personally Identifiable Information) detection and Named Entity Recognition (NER) systems on Russian-language text. It targets guardrail and anonymization pipelines that must reliably find personal data (names, addresses, contacts) and Russian identity-document numbers (passport, SNILS, INN, OMS, etc.) in text. The test set combines real, manually annotated examples… See the full description on the dataset page: https://huggingface.co/datasets/redmadrobot-rnd/pii_benchmark.texttoken-classification1K<n<10K21 likes764 downloads4mo agoHugging Face12ibm-research /Wikipedia_contradict_benchmark Wikipedia contradict benchmark Wikipedia contradict benchmark is a dataset consisting of 253 high-quality, human-annotated instances designed to assess LLM performance when augmented with retrieved passages containing real-world knowledge conflicts. The dataset was created intentionally with that task in mind, focusing on a benchmark consisting of high-quality, human-annotated instances. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Wikipedia_contradict_benchmark.textquestion-answeringn<1K28 likes664 downloads2y agoHugging Face13gschottlender /LigQ_2_benchmark Ligand-Based Target Benchmark Splits This dataset contains target-level ligand benchmark splits generated from the same pipeline used for the enrichment-factor evaluations. Each target has one folder per random_state. Inside each split: known_actives.csv: active ligands available to the search method. These are the query/reference ligands. evaluation_actives.csv: held-out active ligands for the same target. These are the positives to recover during evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/gschottlender/LigQ_2_benchmark.texttabular-classification100K<n<1M0 likes565 downloads4mo agoHugging Face14nanimani /local-llm-benchmark Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB) English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe Manual evaluation results of local GGUF model variants on a single consumer machine, combining two fully independent benchmarks: technical/ uncensored/ Measures capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.tabulartext-generation1K<n<10K2 likes518 downloads20d agoHugging Face15MothMalone /data-preprocessing-automl-benchmarks Data Preprocessing AutoML Benchmarks This repository contains text classification datasets with known data quality issues for preprocessing research in AutoML. Usage Load a specific dataset configuration like this: from datasets import load_dataset # Example for loading the TREC dataset dataset = load_dataset("MothMalone/data-preprocessing-automl-benchmarks", "trec") Available Datasets Below are the details for each dataset configuration available in this… See the full description on the dataset page: https://huggingface.co/datasets/MothMalone/data-preprocessing-automl-benchmarks.texttext-classification100K<n<1M0 likes483 downloads1y agoHugging Face16LindertLab /DMS-Fold2-Benchmark-Dataset DMS-Fold2 Benchmark Dataset This repository contains the benchmark datasets used to evaluate DMS-Fold2. The files are organized into two benchmark sets: CASP/CAMEO targets (casp_cameo_targets/) Mega-scale targets (megascale_targets/) Each target includes the sequence and experimental data required to reproduce the benchmark inputs. Directory Structure paper_benchmarks/ ├── casp_cameo_targets/ │ ├── alignments.tar.gz │ ├── fastas/ │ ├── pdbs/ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/LindertLab/DMS-Fold2-Benchmark-Dataset.text100M<n<1B0 likes452 downloads3mo agoHugging Face17RajvardhanPatil07 /halluguard-med-300-question-benchmark HalluGuard-Med: 300-Question Medical QA Benchmark Published by Rajvardhan Patil. Release 1.0, 4 October 2026. A frozen English-language evaluation dataset containing 300 MedQuAD questions, released reference answers, 900 MedGemma responses, and claim-level automatic assessments. The benchmark compares Raw MedGemma, RAG only, and HalluGuard-Med under recorded generation settings. The questions and references originate from MedQuAD by Asma Ben Abacha and Dina Demner-Fushman. This… See the full description on the dataset page: https://huggingface.co/datasets/RajvardhanPatil07/halluguard-med-300-question-benchmark.tabularquestion-answering1K<n<10K1 likes418 downloads2d agoHugging Face18microsoft /delulu-fim-benchmarkDelulu — Fill-in-the-Middle Code Hallucination Benchmark A verified multilingual benchmark for code-completion hallucinations. Every golden completion compiles. Every hallucination provably doesn't. 📄 Read the preprint on arXiv → Every Delulu sample ships as a self-contained Docker image. The viewer above lets you browse the dataset, pull a sample's verifier, and re-run verify golden / verify hallucinated / verify patch <your-completion> with one… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/delulu-fim-benchmark.texttext-generation1K<n<10K3 likes397 downloads5mo agoHugging Face19tum-nlp /neural-news-benchmark AI-generated News Detection Benchmark neural-news is a benchmark dataset designed for human/AI news authorship classification in English, Turkish, Hungarian, and Persian. Presented in Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian @ NLP for Positive Impact Workshop @ EMNLP2024. Dataset Details The dataset includes equal parts human-written and AI-generated news articles, raw and pre-processed. Curated… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/neural-news-benchmark.texttext-classification10K<n<100K4 likes376 downloads2y agoHugging Face20namanvats /harbor-goose-openhands-benchmark Same Model, Opposite Results: Goose vs OpenHands Turn Budget Study on Harbor Terminal-Bench-Pro Trial-level results from a small controlled study comparing two agent harnesses — Goose and OpenHands-SDK — on a frozen 40-task Harbor Terminal-Bench-Pro slice. All runs used minimax/minimax-m2.5 via OpenRouter with Daytona as the sandbox backend. Key Findings Reducing the turn budget from 100 to 60 pushed the two harnesses in opposite directions under the base setup:… See the full description on the dataset page: https://huggingface.co/datasets/namanvats/harbor-goose-openhands-benchmark.tabularn<1K3 likes364 downloads6mo agoHugging Face21ysdede /asr_benchmark_storetabularn<1K1 likes359 downloads3mo agoHugging Face22latam-gpt /Trueque-Benchmark-beta-0.1 🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture 🌐 Language versions: Español | Português ⚠️ Official Disclaimer: Beta Release (v0.1) Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America. Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.textquestion-answeringn<1K9 likes356 downloads3mo agoHugging Face23theResearchNinja /benchmarkResults_violentUTF_cybersecurityBehavior Overview Interdependent cybersecurity addresses the complexities and interconnectedness of various systems, emphasizing the need for collaborative and holistic approaches to mitigate risks. This field focuses on how different components, from technology to human factors, influence each other, creating a web of dependencies that must be managed to ensure robust security. Despite significant investments in cybersecurity, many organizations struggle to effectively manage cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/benchmarkResults_violentUTF_cybersecurityBehavior.tabular100K<n<1M1 likes335 downloads11mo agoHugging Face24polinaeterna /tabular-benchmark Tabular Benchmark Dataset Description This dataset is a curation of various datasets from openML and is curated to benchmark performance of various machine learning algorithms. Repository: https://github.com/LeoGrin/tabular-benchmark/community Paper: https://hal.archives-ouvertes.fr/hal-03723551v2/document Dataset Summary Benchmark made of curation of various tabular data learning tasks, including: Regression from Numerical and Categorical Features… See the full description on the dataset page: https://huggingface.co/datasets/polinaeterna/tabular-benchmark.tabulartabular-classification1M<n<10M0 likes318 downloads3y agoHugging Face25imodels /tabular-benchmark-797-classificationtabular1K<n<10K0 likes281 downloads3y agoHugging Face26dima0000 /ninfer-benchmarks NInfer on one RTX 5090 — recorded benchmark evidence Historical results from 19–20 September 2026, published by dima0000. This is a collection of benchmark evidence, not a model checkpoint or a live inference service. It preserves successful measurements, failed checks and incomplete experiments. Hardware: one NVIDIA RTX 5090 with 32 GB VRAM per trial. Models: regular and uncensored Qwen3.8-27B with NVFP4 weights. Engine: NInfer commit 9e163eee4b8acec21ab0ac765107b6a3f287b217… See the full description on the dataset page: https://huggingface.co/datasets/dima0000/ninfer-benchmarks.tabularn<1K0 likes281 downloads13d agoHugging Face27omegaprime669 /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.imagetext-generationn<1K0 likes279 downloads2mo agoHugging Face28py-feat /benchmarks py-feat benchmarks Live benchmark data for py-feat and a cross-tool comparison against OpenFace 3.0, LibreFace, and PyAFAR. Powers the py-feat live dashboard. Updated by scheduled benchmark runs. Files File What accuracy.csv Tidy long table: one row per (tool, dataset, modality, metric). Covers AU F1 (DISFA+), 7-class emotion (AffectNet-val, RAF-DB), valence/arousal CCC (AffectNet-val), and gaze angular error (Columbia). throughput.csv py-feat… See the full description on the dataset page: https://huggingface.co/datasets/py-feat/benchmarks.tabularn<1K0 likes273 downloads4mo agoHugging Face29enlatics /Enlatics_benchmarking GAIA-style Evaluation Results (Public) This dataset contains GAIA-inspired benchmark question results for LLM evaluation. What is inside grok_answers.csv: question-by-question outputs, model used, response time (seconds), and run status. Notes These tasks are designed in a GAIA-style (multi-hop, web-grounded questions). Official GAIA leaderboard submission access was restricted for our account, so results are published here for transparency and… See the full description on the dataset page: https://huggingface.co/datasets/enlatics/Enlatics_benchmarking.textn<1K0 likes263 downloads8mo agoHugging Face30zachz /prompt-injection-benchmark Prompt Injection Benchmark A curated dataset of labeled prompt injection attacks and benign prompts for testing and benchmarking injection detection systems. Dataset Description This dataset contains 200 examples across 7 attack categories, plus 100 benign prompts. Each example is labeled with: text: The prompt text label: injection or benign category: Attack category (e.g., instruction_override, role_hijack) severity: low, medium, high, or critical Attack… See the full description on the dataset page: https://huggingface.co/datasets/zachz/prompt-injection-benchmark.texttext-classificationn<1K1 likes252 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.