datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.b3-agent-security-benchmark-weak[paper] [blogpost] [game]
b3 AI Security Benchmark: Breaking Agent Backbones
Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge.
This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security
of Backbone LLMs in AI Agents.
The high quality dataset was used to evaluate the security of more than 30 LLMs.
Dataset Summary
Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.delulu-fim-benchmarkDelulu — Fill-in-the-Middle Code Hallucination Benchmark
A verified multilingual benchmark for code-completion hallucinations.
Every golden completion compiles. Every hallucination provably doesn't.
📄 Read the preprint on arXiv →
Every Delulu sample ships as a self-contained Docker image. The viewer above lets you browse the dataset, pull a sample's verifier, and re-run verify golden / verify hallucinated / verify patch <your-completion> with one… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/delulu-fim-benchmark.Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.local-llm-benchmark
Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB)
English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe
Manual evaluation results of local GGUF model variants on a single consumer machine,
combining two fully independent benchmarks:
technical/
uncensored/
Measures
capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.story_writing_benchmark
Story Evaluation Dataset
This dataset contains stories generated by Large Language Models (LLMs) across multiple languages, with comprehensive quality evaluations. It was created to train and benchmark models specifically on creative writing tasks.
This benchmark evaluates an LLM's ability to generate high-quality short stories based on simple prompts like "write a story about X with n words." It is similar to TinyStories but targets longer-form and more complex content, focusing… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/story_writing_benchmark.Cognitive_Atrophy_Benchmark
Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets
This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four mental-health conversation datasets. It is a building block for a forthcoming evaluation framework that quantifies cognitive atrophy — the gradual erosion of users' own reasoning, recall, and decisional autonomy when an LLM… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/Cognitive_Atrophy_Benchmark.C3-BenchMark
C^3-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking
Paper: C^3-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking
GitHub: https://github.com/Tencent-Hunyuan/C3-Benchmark
📖 Overview
Agents based on large language models leverage tools to modify environments, revolutionizing how AI interacts with the physical world. Unlike traditional NLP tasks that rely solely on historical dialogue for responses, these agents must consider more… See the full description on the dataset page: https://huggingface.co/datasets/tencent/C3-BenchMark.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.auditable-fincrime-benchmark
Auditable Financial-Crime Compliance Benchmark
A matched-twin benchmark of auditable explanations for financial-crime compliance, spanning two domains — anti-money-laundering (AML) and market abuse. Every scenario is judged not by a yes/no label but against a four-element rubric: the mechanism of the abuse, the rule that is defeated (loophole in AML, breach in market abuse), the distinguishing facts that separate a violation from its innocent twin, and the decision (REPORT /… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousbabygoat/auditable-fincrime-benchmark.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.singapore-legal-ai-benchmark
Singapore Legal AI Benchmark
Public research release of 102 Singapore legal research questions, model
responses from 6 systems, and overlapping grades on five dimensions.
Headline metrics are overlapping binary flags, not a ranking and not a
partition of 100%.
Interactive explorer
Open the explorer →
— comparison table, category heatmap, per-question comparison, and every answer
with its sources and grades.
(Space page)
Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.TARA_Turkish_LLM_Benchmark
TARA: Turkish Advanced Reasoning Assessment Veri Seti
*Img Credit: Open AI ChatGPT
**English version is given below.**
Evaluation Notebook / Değerlendirme Not Defteri
Dataset Summary
TARA (Turkish Advanced Reasoning Assessment), Türkçe dilindeki Büyük Dil Modellerinin (LLM'ler) gelişmiş akıl yürütme yeteneklerini çoklu alanlarda ölçmek için tasarlanmış, zorluk derecesine göre sınıflandırılmış bir benchmark veri setidir. Bu veri seti, LLM'lerin sadece bilgi… See the full description on the dataset page: https://huggingface.co/datasets/emre/TARA_Turkish_LLM_Benchmark.l4-gpu-llm-benchmark-leaderboard
🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB)
An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU.
📊 Executive Summary & Key Takeaways
⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.qwen3.8-27b-medical-benchmarks
Qwen3.8-27B — Medical Benchmark Evaluation
Evaluation report for Qwen/Qwen3.8-27B across 32 medical benchmark configurations
(27,312 scored items), under a single fixed inference configuration, with two grading regimes:
deterministic and model-graded.
The model was evaluated as released. No weights were modified.
Field
Value
Model under evaluation
Qwen/Qwen3.8-27B
Evaluation dates
2026-09-27 to 2026-09-30 (UTC)
Benchmark configurations
32 (24 deterministic, 8… See the full description on the dataset page: https://huggingface.co/datasets/himajahealth/qwen3.8-27b-medical-benchmarks.tam-benchmarks
Tasks over Application Manuals (TAM)
TAM is a benchmark for evaluating long-horizon procedural reasoning: the ability of a language-model system to follow a large application manual, resolve cross-references, apply interdependent constraints, and produce an exact answer. Unlike short-horizon multi-hop tasks, TAM requires systems to maintain consistency across dozens of decisions drawn from manuals containing tens of thousands of rules. An early missed exception or incorrect… See the full description on the dataset page: https://huggingface.co/datasets/manulife/tam-benchmarks.agentic-reasoning-benchmark
Agentic & Reasoning Benchmark (ARB) – Expanded
Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning.
Überblick
Eigenschaft
Wert
Anzahl Beispiele
2.550
Kategorien
8
Schwierigkeitsgrade
easy / medium / hard
Formate
CSV + JSON
Reproduzierbarkeit
Generator-Skript (seed=42) enthalten
Lizenz
CC-BY-4.0
Kategorien
Kategorie
Anzahl
Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.trilingual-cultural-bias-redteaming-benchmark
Trilingual Cultural Bias Red-Teaming Benchmark (HR–SR–HU)
Overview
This is a small qualitative benchmark for red-teaming large language models in Croatian (HR), Serbian (SR), and Hungarian (HU).
The benchmark tests how models respond to provocative, culturally and historically loaded questions, when they are asked to role-play a patriotic citizen of a given country and answer in their own native language.
The goal is not factual QA accuracy, but to observe reasoning… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/trilingual-cultural-bias-redteaming-benchmark.open-models-benchmark-results
⚡ Local LLM Evaluation Leaderboard
Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs.
💻 Hardware & System Specifications
All evaluations are executed under standardized local cluster environments:
Specification
Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.hungarian-riddles-benchmark
Hungarian Riddles Benchmark
Overview
This dataset is a cultural and reasoning benchmark based on 100 metaphorical, trivia-style Hungarian riddles.
The riddles are intentionally tricky and culturally grounded. They are designed to test answer correctness and reasoning quality, not only surface-level language fluency.
Dataset structure
Each row contains one riddle with reference material for evaluation.
Fields
ID – unique identifier
topic – general… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/hungarian-riddles-benchmark.MedPRESS_Benchmark
MedPRESS
MedPRESS is a multi-turn benchmark for evaluating patient-pressure-induced medical sycophancy in large language models. The dataset tests whether a model maintains a safe medical stance when a user repeatedly pressures it toward an unsafe or false health belief.
The benchmark is designed around five-turn conversations. Each row contains one medical scenario, the unsafe or false belief being pressured, the expected safe stance, and five progressively stronger user turns.… See the full description on the dataset page: https://huggingface.co/datasets/samanjoy2/MedPRESS_Benchmark.ponys-multilingual-ai-character-consistency-benchmark
Ponys Multilingual AI Character Consistency Benchmark
This repository contains a preregistered test instrument, not collected product
results and not an independent product ranking.
140 fixed test cases across seven locales
four dimensions: persona, register, relationship state, and visual identity
three planned clean-session runs per case
result state: not_collected
publisher: Ponys.ai Research (official first-party research)
official source: https://ponys.ai/
research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.llm-iso-benchmark
LLM ISO Tax-Optimization Benchmark
Frontier AI models were given one incentive stock option (ISO) exercise-optimization
problem and asked for the schedule that maximizes after-tax net final value (NFV) at a
four-year horizon. The headline finding holds across two rounds of five frontier models
each: every model overshoots the achievable after-tax outcome, by roughly 2x to 20x. The
provable optimum is computed by a deterministic optimizer and is reproducible against a
public… See the full description on the dataset page: https://huggingface.co/datasets/AlphaLatitude/llm-iso-benchmark.Cognitive_Atrophy_Benchmark
Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets
Status: Anonymous submission to the NeurIPS 2026 Evaluations & Datasets Track.
Author identities, affiliations, and acknowledgements are intentionally omitted during double-blind review and will be added upon acceptance.
This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four… See the full description on the dataset page: https://huggingface.co/datasets/CABenchmark/Cognitive_Atrophy_Benchmark.turkish-pii-masking-benchmark
Turkish PII Masking Benchmark (1,000 test cases)
A hand-built, fully synthetic benchmark for evaluating instruction-conditional PII masking
in Turkish. Designed to stress-test small LLMs (0.3B-1B) fine-tuned for privacy filtering in
banking/ERP user messages. All values are synthetic (no real persons); all sentence patterns
were written specifically for this benchmark (no training-set overlap).
Task
Given instruction (the masking policy) and input, the model must… See the full description on the dataset page: https://huggingface.co/datasets/cagrigungor/turkish-pii-masking-benchmark.agent-evaluation-benchmark
Agent Evaluation Benchmark
A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories.
Overview
This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more.
Categories
Category
Test Cases
Description
Weather
5
Forecasts, UV index, climate history
Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.El-TARA_Spanish_LLM_Benchmark
El-Tara: Evaluación de Razonamiento Avanzado en Español
Dataset Summary
El-Tara (Evaluación de Razonamiento Avanzado en Español) is a benchmark dataset designed to assess the advanced reasoning capabilities of Large Language Models (LLMs) in Spanish. It is adapted from the original TARA (Turkish Advanced Reasoning Assessment) dataset.
Similar to TARA, El-Tara aims to test higher-order cognitive skills across multiple domains, using synthetically generated questions… See the full description on the dataset page: https://huggingface.co/datasets/emre/El-TARA_Spanish_LLM_Benchmark.IndoBloom-AQG-Benchmark-Corpus
📚 Indo-Bloom AQG Benchmark Corpus (All Models)
⚠️ RESEARCH ARTIFACT STATUS: BENCHMARK / SILVER CORPUS (Stage 1)
This dataset serves as the comparative benchmark corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM).
Current State: LLM-Generated QA Pairs Evaluated via Rule-Based Evaluator
Next Stage: Expert Annotation (Stage 2) → Gold Standard
🔒 FROZEN — Benchmark v1.0
This version is permanently frozen to ensure reproducibility of the experimental… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/IndoBloom-AQG-Benchmark-Corpus.llm-smartrouter-benchmark
LLM SmartRouter & Agent Highway Latency & Cost Benchmark (v1.4.0)
Empirical performance benchmark dataset comparing direct model endpoints (OpenAI, Anthropic Claude, Google Gemini) against the PixelRouter / BLUN SmartRouter proxy layer and Autonomous Agent Web Highway (https://api.pixeloffice.eu/v1).
v1.4.0 Benchmark Highlights
Anthropic Claude Messages API: Sub-35ms proxy routing for native /v1/messages payloads with 94%+ cost savings.
Machine Web Highway… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-smartrouter-benchmark.agent-sandbox-negotiation-benchmark
Agent Sandbox Negotiation Benchmark v1
Overview
A dataset of simulated multi-agent negotiations generated using the open-source Agent Sandbox framework.
This dataset captures the final negotiation outcomes, turn depths, strategy alignments, and agreed prices of local LLMs (Llama-3 and Mistral) engaged in intense, adversarial price negotiations at massive scale.
Dataset Statistics
Simulations: 24,122
Strategies: 4 (Balanced, Aggressive, Conservative, Adaptive)… See the full description on the dataset page: https://huggingface.co/datasets/ScareRezume/agent-sandbox-negotiation-benchmark.
