Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01witcheer /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.imagetext-generationn<1K1 likes1.5k downloads3d agoHugging Face02Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K9 likes1.3k downloads4d agoHugging Face03microsoft /delulu-fim-benchmarkDelulu — Fill-in-the-Middle Code Hallucination Benchmark A verified multilingual benchmark for code-completion hallucinations. Every golden completion compiles. Every hallucination provably doesn't. 📄 Read the preprint on arXiv → Every Delulu sample ships as a self-contained Docker image. The viewer above lets you browse the dataset, pull a sample's verifier, and re-run verify golden / verify hallucinated / verify patch <your-completion> with one… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/delulu-fim-benchmark.texttext-generation1K<n<10K3 likes396 downloads5mo agoHugging Face04latam-gpt /Trueque-Benchmark-beta-0.1 🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture 🌐 Language versions: Español | Português ⚠️ Official Disclaimer: Beta Release (v0.1) Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America. Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.textquestion-answeringn<1K9 likes340 downloads3mo agoHugging Face05nanimani /local-llm-benchmark Local LLM Benchmark — Technical and Uncensored Behavior (NVIDIA RTX 5070 Ti 16GB) English | 简体中文 | 繁體中文 | 한국어 | Español | 日本語 | हिन्दी | Русский | Português | తెలుగు | Français | Deutsch | Italiano | Tiếng Việt | العربية | اردو | বাংলা | فارسی | Română | Türkçe Manual evaluation results of local GGUF model variants on a single consumer machine, combining two fully independent benchmarks: technical/ uncensored/ Measures capability: coding, systems, networking, DB, agents… See the full description on the dataset page: https://huggingface.co/datasets/nanimani/local-llm-benchmark.tabulartext-generation1K<n<10K2 likes334 downloads24d agoHugging Face06lars1234 /story_writing_benchmark Story Evaluation Dataset This dataset contains stories generated by Large Language Models (LLMs) across multiple languages, with comprehensive quality evaluations. It was created to train and benchmark models specifically on creative writing tasks. This benchmark evaluates an LLM's ability to generate high-quality short stories based on simple prompts like "write a story about X with n words." It is similar to TinyStories but targets longer-form and more complex content, focusing… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/story_writing_benchmark.tabulartext-generation10K<n<100K6 likes254 downloads2y agoHugging Face07abadawi /Cognitive_Atrophy_Benchmark Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four mental-health conversation datasets. It is a building block for a forthcoming evaluation framework that quantifies cognitive atrophy — the gradual erosion of users' own reasoning, recall, and decisional autonomy when an LLM… See the full description on the dataset page: https://huggingface.co/datasets/abadawi/Cognitive_Atrophy_Benchmark.tabulartext-generation10K<n<100K2 likes253 downloads3mo agoHugging Face08tencent /C3-BenchMark C^3-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking Paper: C^3-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking GitHub: https://github.com/Tencent-Hunyuan/C3-Benchmark 📖 Overview Agents based on large language models leverage tools to modify environments, revolutionizing how AI interacts with the physical world. Unlike traditional NLP tasks that rely solely on historical dialogue for responses, these agents must consider more… See the full description on the dataset page: https://huggingface.co/datasets/tencent/C3-BenchMark.texttext-generationn<1K7 likes241 downloads1y agoHugging Face09omegaprime669 /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.imagetext-generationn<1K0 likes218 downloads3mo agoHugging Face10Anonymousbabygoat /auditable-fincrime-benchmark Auditable Financial-Crime Compliance Benchmark A matched-twin benchmark of auditable explanations for financial-crime compliance, spanning two domains — anti-money-laundering (AML) and market abuse. Every scenario is judged not by a yes/no label but against a four-element rubric: the mechanism of the abuse, the rule that is defeated (loophole in AML, breach in market abuse), the distinguishing facts that separate a violation from its innocent twin, and the decision (REPORT /… See the full description on the dataset page: https://huggingface.co/datasets/Anonymousbabygoat/auditable-fincrime-benchmark.tabulartext-classification10K<n<100K0 likes217 downloads5d agoHugging Face11humanlong /emotion-negotiation-benchmarks Emotion-Aware LLM Negotiation Benchmarks Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency. The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.tabulartext-generationn<1K0 likes151 downloads4mo agoHugging Face12JonathanSu /singapore-legal-ai-benchmark Singapore Legal AI Benchmark Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%. Interactive explorer Open the explorer → — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page) Overall (n = 612)… See the full description on the dataset page: https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark.tabulartext-generationn<1K0 likes148 downloads8d agoHugging Face13emre /TARA_Turkish_LLM_Benchmark TARA: Turkish Advanced Reasoning Assessment Veri Seti *Img Credit: Open AI ChatGPT **English version is given below.** Evaluation Notebook / Değerlendirme Not Defteri Dataset Summary TARA (Turkish Advanced Reasoning Assessment), Türkçe dilindeki Büyük Dil Modellerinin (LLM'ler) gelişmiş akıl yürütme yeteneklerini çoklu alanlarda ölçmek için tasarlanmış, zorluk derecesine göre sınıflandırılmış bir benchmark veri setidir. Bu veri seti, LLM'lerin sadece bilgi… See the full description on the dataset page: https://huggingface.co/datasets/emre/TARA_Turkish_LLM_Benchmark.textquestion-answeringn<1K28 likes138 downloads1y agoHugging Face14mayank-dubey-ai /l4-gpu-llm-benchmark-leaderboard 🚀 Local LLM Serving & Quality Benchmark Leaderboard (NVIDIA L4 24GB) An exhaustive, reproducible benchmark study measuring real-world serving performance (TTFT, TPOT, throughput, peak VRAM, energy consumption, and cost) alongside rigorous task quality gates (HumanEval+, MMLU-Pro, BFCL v4 tool calling, and RULER needle retrieval) for open-weight LLMs on a single NVIDIA L4 24GB GPU. 📊 Executive Summary & Key Takeaways ⚡ Best Throughput & Coding Workhorse:… See the full description on the dataset page: https://huggingface.co/datasets/mayank-dubey-ai/l4-gpu-llm-benchmark-leaderboard.tabulartext-generationn<1K0 likes138 downloads2mo agoHugging Face15himajahealth /qwen3.8-27b-medical-benchmarks Qwen3.8-27B — Medical Benchmark Evaluation Evaluation report for Qwen/Qwen3.8-27B across 32 medical benchmark configurations (27,312 scored items), under a single fixed inference configuration, with two grading regimes: deterministic and model-graded. The model was evaluated as released. No weights were modified. Field Value Model under evaluation Qwen/Qwen3.8-27B Evaluation dates 2026-09-27 to 2026-09-30 (UTC) Benchmark configurations 32 (24 deterministic, 8… See the full description on the dataset page: https://huggingface.co/datasets/himajahealth/qwen3.8-27b-medical-benchmarks.tabularquestion-answeringn<1K0 likes129 downloads9d agoHugging Face16manulife /tam-benchmarks Tasks over Application Manuals (TAM) TAM is a benchmark for evaluating long-horizon procedural reasoning: the ability of a language-model system to follow a large application manual, resolve cross-references, apply interdependent constraints, and produce an exact answer. Unlike short-horizon multi-hop tasks, TAM requires systems to maintain consistency across dozens of decisions drawn from manuals containing tens of thousands of rules. An early missed exception or incorrect… See the full description on the dataset page: https://huggingface.co/datasets/manulife/tam-benchmarks.documenttext-generationn<1K0 likes110 downloads29d agoHugging Face17roskosmos19 /agentic-reasoning-benchmark Agentic & Reasoning Benchmark (ARB) – Expanded Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning. Überblick Eigenschaft Wert Anzahl Beispiele 2.550 Kategorien 8 Schwierigkeitsgrade easy / medium / hard Formate CSV + JSON Reproduzierbarkeit Generator-Skript (seed=42) enthalten Lizenz CC-BY-4.0 Kategorien Kategorie Anzahl Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.textquestion-answering1K<n<10K1 likes78 downloads1mo agoHugging Face18boczkakaroly /trilingual-cultural-bias-redteaming-benchmark Trilingual Cultural Bias Red-Teaming Benchmark (HR–SR–HU) Overview This is a small qualitative benchmark for red-teaming large language models in Croatian (HR), Serbian (SR), and Hungarian (HU). The benchmark tests how models respond to provocative, culturally and historically loaded questions, when they are asked to role-play a patriotic citizen of a given country and answer in their own native language. The goal is not factual QA accuracy, but to observe reasoning… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/trilingual-cultural-bias-redteaming-benchmark.texttext-generationn<1K0 likes69 downloads9mo agoHugging Face19ahmedBargady /open-models-benchmark-results ⚡ Local LLM Evaluation Leaderboard Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs. 💻 Hardware & System Specifications All evaluations are executed under standardized local cluster environments: Specification Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.tabulartext-generationn<1K1 likes60 downloads2mo agoHugging Face20boczkakaroly /hungarian-riddles-benchmark Hungarian Riddles Benchmark Overview This dataset is a cultural and reasoning benchmark based on 100 metaphorical, trivia-style Hungarian riddles. The riddles are intentionally tricky and culturally grounded. They are designed to test answer correctness and reasoning quality, not only surface-level language fluency. Dataset structure Each row contains one riddle with reference material for evaluation. Fields ID – unique identifier topic – general… See the full description on the dataset page: https://huggingface.co/datasets/boczkakaroly/hungarian-riddles-benchmark.imagequestion-answeringn<1K0 likes59 downloads9mo agoHugging Face21samanjoy2 /MedPRESS_Benchmarkgated MedPRESS MedPRESS is a multi-turn benchmark for evaluating patient-pressure-induced medical sycophancy in large language models. The dataset tests whether a model maintains a safe medical stance when a user repeatedly pressures it toward an unsafe or false health belief. The benchmark is designed around five-turn conversations. Each row contains one medical scenario, the unsafe or false belief being pressured, the expected safe stance, and five progressively stronger user turns.… See the full description on the dataset page: https://huggingface.co/datasets/samanjoy2/MedPRESS_Benchmark.texttext-generationn<1K0 likes53 downloads2mo agoHugging Face22wujoe132 /ponys-multilingual-ai-character-consistency-benchmark Ponys Multilingual AI Character Consistency Benchmark This repository contains a preregistered test instrument, not collected product results and not an independent product ranking. 140 fixed test cases across seven locales four dimensions: persona, register, relationship state, and visual identity three planned clean-session runs per case result state: not_collected publisher: Ponys.ai Research (official first-party research) official source: https://ponys.ai/ research feeds:… See the full description on the dataset page: https://huggingface.co/datasets/wujoe132/ponys-multilingual-ai-character-consistency-benchmark.tabulartext-generationn<1K0 likes51 downloads1mo agoHugging Face23AlphaLatitude /llm-iso-benchmark LLM ISO Tax-Optimization Benchmark Frontier AI models were given one incentive stock option (ISO) exercise-optimization problem and asked for the schedule that maximizes after-tax net final value (NFV) at a four-year horizon. The headline finding holds across two rounds of five frontier models each: every model overshoots the achievable after-tax outcome, by roughly 2x to 20x. The provable optimum is computed by a deterministic optimizer and is reproducible against a public… See the full description on the dataset page: https://huggingface.co/datasets/AlphaLatitude/llm-iso-benchmark.tabulartext-generationn<1K0 likes49 downloads4mo agoHugging Face24CABenchmark /Cognitive_Atrophy_Benchmark Cognitive Atrophy Benchmark — LLM Responses Across Four Mental-Health Conversation Datasets Status: Anonymous submission to the NeurIPS 2026 Evaluations & Datasets Track. Author identities, affiliations, and acknowledgements are intentionally omitted during double-blind review and will be added upon acceptance. This dataset releases the LLM-response component of the Cognitive Atrophy Benchmark: five large language models prompted under identical conditions across four… See the full description on the dataset page: https://huggingface.co/datasets/CABenchmark/Cognitive_Atrophy_Benchmark.tabulartext-generation10K<n<100K0 likes46 downloads5mo agoHugging Face25cagrigungor /turkish-pii-masking-benchmark Turkish PII Masking Benchmark (1,000 test cases) A hand-built, fully synthetic benchmark for evaluating instruction-conditional PII masking in Turkish. Designed to stress-test small LLMs (0.3B-1B) fine-tuned for privacy filtering in banking/ERP user messages. All values are synthetic (no real persons); all sentence patterns were written specifically for this benchmark (no training-set overlap). Task Given instruction (the masking policy) and input, the model must… See the full description on the dataset page: https://huggingface.co/datasets/cagrigungor/turkish-pii-masking-benchmark.texttext-generation1K<n<10K0 likes44 downloads2mo agoHugging Face26aiagentkarl /agent-evaluation-benchmark Agent Evaluation Benchmark A benchmark dataset for evaluating AI agent tool-use capabilities across 55+ test cases spanning 14 categories. Overview This benchmark tests whether AI agents can correctly select and use the right MCP tools for real-world tasks. It covers data retrieval, blockchain queries, security analysis, academic research, and more. Categories Category Test Cases Description Weather 5 Forecasts, UV index, climate history Blockchain… See the full description on the dataset page: https://huggingface.co/datasets/aiagentkarl/agent-evaluation-benchmark.texttext-generationn<1K0 likes34 downloads7mo agoHugging Face27emre /El-TARA_Spanish_LLM_Benchmark El-Tara: Evaluación de Razonamiento Avanzado en Español Dataset Summary El-Tara (Evaluación de Razonamiento Avanzado en Español) is a benchmark dataset designed to assess the advanced reasoning capabilities of Large Language Models (LLMs) in Spanish. It is adapted from the original TARA (Turkish Advanced Reasoning Assessment) dataset. Similar to TARA, El-Tara aims to test higher-order cognitive skills across multiple domains, using synthetically generated questions… See the full description on the dataset page: https://huggingface.co/datasets/emre/El-TARA_Spanish_LLM_Benchmark.textquestion-answeringn<1K1 likes31 downloads2y agoHugging Face28Firmansyah-Ibrahim /IndoBloom-AQG-Benchmark-Corpus 📚 Indo-Bloom AQG Benchmark Corpus (All Models) ⚠️ RESEARCH ARTIFACT STATUS: BENCHMARK / SILVER CORPUS (Stage 1) This dataset serves as the comparative benchmark corpus for the Indo-Bloom research project at Universitas Negeri Malang (UM). Current State: LLM-Generated QA Pairs Evaluated via Rule-Based Evaluator Next Stage: Expert Annotation (Stage 2) → Gold Standard 🔒 FROZEN — Benchmark v1.0 This version is permanently frozen to ensure reproducibility of the experimental… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/IndoBloom-AQG-Benchmark-Corpus.tabulartext-generation10K<n<100K0 likes30 downloads7mo agoHugging Face29pixeloffice /llm-smartrouter-benchmark LLM SmartRouter & Agent Highway Latency & Cost Benchmark (v1.4.0) Empirical performance benchmark dataset comparing direct model endpoints (OpenAI, Anthropic Claude, Google Gemini) against the PixelRouter / BLUN SmartRouter proxy layer and Autonomous Agent Web Highway (https://api.pixeloffice.eu/v1). v1.4.0 Benchmark Highlights Anthropic Claude Messages API: Sub-35ms proxy routing for native /v1/messages payloads with 94%+ cost savings. Machine Web Highway… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/llm-smartrouter-benchmark.tabulartext-generationn<1K0 likes30 downloads1mo agoHugging Face30ScareRezume /agent-sandbox-negotiation-benchmark Agent Sandbox Negotiation Benchmark v1 Overview A dataset of simulated multi-agent negotiations generated using the open-source Agent Sandbox framework. This dataset captures the final negotiation outcomes, turn depths, strategy alignments, and agreed prices of local LLMs (Llama-3 and Mistral) engaged in intense, adversarial price negotiations at massive scale. Dataset Statistics Simulations: 24,122 Strategies: 4 (Balanced, Aggressive, Conservative, Adaptive)… See the full description on the dataset page: https://huggingface.co/datasets/ScareRezume/agent-sandbox-negotiation-benchmark.tabulartext-generation10K<n<100K1 likes20 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.