Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01McGill-NLP /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.imagerobotics1K<n<10K4 likes23k downloads1y agoHugging Face02vals-ai /finance_agent_benchmark Finance Agent Benchmark Dataset We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings. We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/vals-ai/finance_agent_benchmark.textn<1K11 likes1.9k downloads1y agoHugging Face03Lakera /b3-agent-security-benchmark-weak[paper] [blogpost] [game] b3 AI Security Benchmark: Breaking Agent Backbones Highly contextalized prompt injections crowd-sourced during the Agent Breaker Challenge. This is a low-quality version of the data behind Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents. The high quality dataset was used to evaluate the security of more than 30 LLMs. Dataset Summary Purpose: This dataset contains crowdsourced adversarial attacks… See the full description on the dataset page: https://huggingface.co/datasets/Lakera/b3-agent-security-benchmark-weak.tabulartext-classificationn<1K9 likes1.3k downloads4d agoHugging Face04FLARE25-Agent-Xray /CBIS-DDSM-SEGimage1K<n<10K1 likes858 downloads1y agoHugging Face05yuanyyaa /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/yuanyyaa/agent-reward-bench.imagerobotics1K<n<10K0 likes775 downloads6mo agoHugging Face06DeusHorizon /agent-web-index Agent Web Index — how much of the web can AI assistants actually read? 50,413 domains measured live. 26% of them cannot be read by at least one of ChatGPT, Claude, Perplexity or Gemini. Updated daily. Live index: https://shop.lumnika.com/ai-readiness/ Every row here is the result of real HTTP requests, not an estimate and not a re-publication of someone else's crawl: each domain's homepage is requested once as a browser and once as each of the published AI crawler user-agents… See the full description on the dataset page: https://huggingface.co/datasets/DeusHorizon/agent-web-index.tabular10K<n<100K0 likes557 downloads18h agoHugging Face07BothBosu /multi-agent-scam-conversation Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset with Agentic Personalities Dataset Description The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset with Agentic Personalities is an enhanced collection of simulated phone conversations between two AI agents, one acting as a scammer or non-scammer and the other as an innocent receiver. Each dialogue is labeled as either a scam or non-scam interaction. This dataset is designed to help develop… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/multi-agent-scam-conversation.texttext-classification1K<n<10K12 likes358 downloads2y agoHugging Face08HaomingLuo /AgentFEM-Bracket-Elasticity From finite elements to physical AI 176 three-dimensional FEM geometries. Real CAD, meshes and displacement fields. An open, reproducible starting point for geometry-conditioned physical learning. Interactive lab · Trained model · AgentFEM · AgentFEM-Learning What is inside A symmetric double-arm support bracket under a uniform downward top-seat load, with fixed foot undersides. Four parameters vary the arm half-width, half-thickness, waist and bow. Material:… See the full description on the dataset page: https://huggingface.co/datasets/HaomingLuo/AgentFEM-Bracket-Elasticity.tabularothern<1K0 likes322 downloads20d agoHugging Face09FatimahEmadEldin /agent-intrusion-escalation-forensics Both Sides Detected It, Neither Escalated: Concurrency and Escalation Failure in the July 2026 Autonomous Agent Intrusion This repository contains the corpus, ingestion pipeline and report for a forensic reconstruction of the July 2026 autonomous agent intrusion, submitted to the Apart Research & CeSIA AI Incident Response Sprint, Track 2 (Forensics and Forecasting). By: Fatimah Mohamed Emad Elden Trouve Labs Detection was not the binding… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/agent-intrusion-escalation-forensics.documentn<1K0 likes275 downloads27d agoHugging Face10AdityaaXD /Multi-Agent_Reinforcement_Learning_Trading_System_Data 📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024). GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024). 📝… See the full description on the dataset page: https://huggingface.co/datasets/AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data.tabulartime-series-forecasting1K<n<10K14 likes260 downloads8mo agoHugging Face11AgenticFinLab /PortBench-Market PortBench Market Base Dataset Dataset Description A ten-year (Jan 2015–Dec 2025) daily financial dataset covering 183 instruments across six heterogeneous asset classes, designed for multi-asset portfolio management research and LLM evaluation. Asset Coverage Asset Class Instruments Data Fields Sources Equities 126 OHLCV + return Yahoo Finance (ETFs: broad market, sector, factor, international) Bonds 16 Close + return (ETFs); yield… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-Market.tabulartime-series-forecasting1K<n<10K3 likes243 downloads4mo agoHugging Face12agentic-moral-alignment /matrix-game-evaltabular10K<n<100K0 likes225 downloads5mo agoHugging Face13sanjaydoss /Multi-Agent_Reinforcement_Learning_Trading_System_Data 📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024). GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024).… See the full description on the dataset page: https://huggingface.co/datasets/sanjaydoss/Multi-Agent_Reinforcement_Learning_Trading_System_Data.tabulartime-series-forecasting1K<n<10K20 likes221 downloads1mo agoHugging Face14cx-cmu /deepresearchgym-agentic-search-logs DeepResearchGym Agentic Search Logs This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617). The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253. All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.tabulartext-retrieval10M<n<100M16 likes214 downloads8mo agoHugging Face15Yunhao-Feng /AgentHazard AgentHazard A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents 🌐 Website | 📊 Dataset | 📄 Paper | 📖 Appendix 🎯 Overview AgentHazard is a comprehensive benchmark for evaluating harmful behavior in computer-use agents. Unlike traditional prompt-level safety benchmarks, AgentHazard focuses on execution-level failures that emerge through the composition of locally plausible steps across multi-turn, tool-mediated trajectories. Key Features… See the full description on the dataset page: https://huggingface.co/datasets/Yunhao-Feng/AgentHazard.tabular1K<n<10K0 likes211 downloads6mo agoHugging Face16AgentsSci /EMNLP_Cost-Aware-Protocol-Routing Cost-Aware Protocol Routing: Matched Protocol Outcomes The short version. We ran the same 6,803 reasoning problems through four different LLM collaboration setups — from a single direct answer up to a four-agent deliberation — and recorded, for every problem, which ones got it right. Then we asked whether a model can look at a problem beforehand and predict which setup is worth paying for. It can predict whether it will fail. It cannot predict which collaboration protocol will… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/EMNLP_Cost-Aware-Protocol-Routing.tabulartabular-classification10K<n<100K0 likes206 downloads17d agoHugging Face17AgentNativeResearchLab /rl-research-envs RL Research Envs Four self-contained research tasks for coding agents, plus 48 full agent runs on them (3 agents × 4 tasks × 4 trials), with trajectories, submitted code, and verifier output. Each task gives the agent a starter solution.py, a small labelled dev split, and a self-scoring script. The agent has 60 minutes to improve the file. A hidden test split then scores it. The tasks are small statistics / signal-processing problems where the obvious approach scores 0 and a… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/rl-research-envs.tabularn<1K0 likes189 downloads13d agoHugging Face18AgenticFinLab /PyFi-600K Dataset Card for PyFi-600K This dataset card aims to be a introduction for PyFi-600K, A financial VLM dataset containing 600K question-answer pairs generated via Adversarial agents. AgenticFinLab/PyFi-600K/ ├── README.md # Dataset documentation and description ├── images.zip # Compressed image files ├── PyFi-600K-dataset.csv # Q&A pairs in CSV format ├── PyFi-600K-dataset.json # Q&A pairs in JSON format ├── PyFi-600K-chain-dataset.json # Chain of Thought Q&A pairs dataset └──… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PyFi-600K.imagequestion-answering100K<n<1M1 likes167 downloads10mo agoHugging Face19agentic-moral-alignment /persona-and-other-evals Qwen3.5-9B AMA adapters — persona evals Inference code, the data it produced, and the tools that turn that data into tables and an HTML viewer. The evals are Anthropic's persona set, scored in three regimes: teacher-forced logprob of the answer literal, greedy answer with the reasoning block pre-closed, and a full 16k-budget reasoning trace. Pinned models base unsloth/Qwen3.5-9B @ 005429cee5cb648998cf2b70eebdd83175989c9a util… See the full description on the dataset page: https://huggingface.co/datasets/agentic-moral-alignment/persona-and-other-evals.tabularn<1K0 likes164 downloads19d agoHugging Face20skandotai /agentic-ontology-of-work Agentic Ontology of Work (AOW) The Agentic Ontology of Work (AOW) is an open model for describing work done by AI agents, people, and systems in an enterprise. It defines each part of the work, the information each part carries, and how the parts connect. This dataset contains version 2.0.0 of the ontology, its validation files, and validated example data. Website: https://skandotai.github.io/agentic-ontology-of-work/ Source repository:… See the full description on the dataset page: https://huggingface.co/datasets/skandotai/agentic-ontology-of-work.textn<1K0 likes162 downloads16d agoHugging Face21agentlans /readabilityDescription: This dataset comprises approximately 200,000 paragraphs and readability metrics from each of four sources: HuggingFace's Fineweb-Edu Ronen Eldan's TinyStories Wikipedia-2023-11-embed-multilingual-v3 (English only) ArXiv Abstracts-2021. Each paragraph falls within the character range of 50 to 2000. Format: JSON, with each row representing a paragraph and containing both the text and its corresponding readability grade. Features: Text: A paragraph of text from one of the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/readability.texttext-classification100K<n<1M1 likes155 downloads2y agoHugging Face22Sepideh2027 /AgentYear: 2025License: MITAuthor: Sepideh Moafi PathogenAgentAI Instruction Dataset Dataset Description A ClinVar-derived dataset developed as part of the PathogenAgentAI research software project. The dataset is released in two parallel formats: Tabular version (train.csv, valid.csv, test.csv) — structured genomic-variant data for classical ML and analysis. BioGPT instruction version (biogpt_train.csv, biogpt_valid.csv, biogpt_test.csv) — instruction-style data… See the full description on the dataset page: https://huggingface.co/datasets/Sepideh2027/Agent.texttext-generation1M<n<10M0 likes152 downloads22d agoHugging Face23BothBosu /single-agent-scam-conversations Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset Dataset Description The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams. Dataset Structure The dataset consists of three columns: dialogue: The transcribed conversation between the caller and receiver. type: The specific type of scam or non-scam interaction. labels: A binary label indicating whether the conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/single-agent-scam-conversations.texttext-classification1K<n<10K2 likes149 downloads2y agoHugging Face24sumitaidev /agent-memory-resilience-benchmark Agent Memory Resilience & Poisoning Benchmark Dataset Summary This benchmark dataset evaluates resilience, negative transfer, and memory poisoning mitigation in autonomous LLM agent architectures (such as LangGraph, AutoGen, and CrewAI). When autonomous agents record distilled self-reflections after attempting tasks, external stochastic failures or subtle API deprecations often cause agents to commit defective strategies into episodic memory. Under standard… See the full description on the dataset page: https://huggingface.co/datasets/sumitaidev/agent-memory-resilience-benchmark.tabularreinforcement-learning1K<n<10K1 likes137 downloads17d agoHugging Face25Xieji-Li /Derm1M-AgentAuggated Derm1M-AgentAug Knowledge-enriched captions for 413,369 dermatological images, generated by MAGEN (Multi-Agent data GENeration) and used to pretrain O-MAKE. MAGEN rewrites part of the corpus through a foundation-model-assisted captioning agent with a diagnostic tool, verifying each result by retrieval; captions it did not improve on keep the original Derm1M text, and the agent_generated column records which is which. Every caption is additionally decomposed into distinct… See the full description on the dataset page: https://huggingface.co/datasets/Xieji-Li/Derm1M-AgentAug.textzero-shot-image-classification100K<n<1M4 likes131 downloads2mo agoHugging Face26agentic-learning-ai-lab /daily-oracle Daily Oracle 📰 Project Website📝 Paper - Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle Daily Oracle is a continuous evaluation benchmark using automatically generated QA pairs from daily news to assess how the future prediction capabilities of LLMs evolve over time. Dataset Details Question Type: True/False (TF) & Multiple Choice (MC) Current Version* Time Span: 2020.01.01 - 2026.07.18 Size: 20,376 TF questions and 18,557 MC… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/daily-oracle.textquestion-answering10K<n<100K4 likes130 downloads3mo agoHugging Face27witcheer /agentic-score-leaderboard 🛠️ Agentic Score Leaderboard — one RTX 5090 How well do local models actually drive a tool-using agent loop? Not single-call function-calling benchmarks — a real loop: native OpenAI tool-calling through llama-server, multi-step deterministic tasks, programmatic verification. Everything runs on a single RTX 5090 32GB. Updated 2026-06-17 · llama.cpp b9562 · --jinja native tool-calling · temp 0. Leaderboard # model params Agentic Score success tool-eff… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/agentic-score-leaderboard.tabularn<1K3 likes126 downloads3mo agoHugging Face28Koplos /finance_agent_benchmark Finance Agent Benchmark Dataset We present the Finance Agent Benchmark, featuring challenging and diverse real-world finance research problems which require LLMs to perform complex analysis with the use of of recent SEC filings. We construct the benchmark using a taxonomy of nine financial task categories, developed in consultation with experts from banks, hedge funds, and private equity firms. The dataset includes 537 expert-authored questions, covering tasks from information… See the full description on the dataset page: https://huggingface.co/datasets/Koplos/finance_agent_benchmark.textn<1K1 likes126 downloads2mo agoHugging Face29agentlans /tatoeba-english-translations Tatoeba English Translation Dataset Dataset Summary This dataset is derived from the Tatoeba database, focusing on English sentences and their translations. It includes assessments of English sentences using text quality, sentiment, and readability models. The dataset is designed for tasks related to multilingual text quality, readability, and sentiment analysis. Supported Tasks and Leaderboards Quality Assessment Readability Prediction Sentiment Analysis… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/tatoeba-english-translations.tabulartext-classification1M<n<10M2 likes118 downloads2y agoHugging Face30replit /agent-challenge Replit Agent Challenge For comprehensive details about the challenge, visit our GitHub repository. Dataset Overview This dataset comprises a collection of instructions and file states specifically curated for the agent challenge. It is derived from a subset of SWE-Bench-Lite. Schema Structure The dataset follows this schema: - File_before: [Initial state of the file] - Instructions: [Steps to transform the file to its final state] - File_after: [Resulting state… See the full description on the dataset page: https://huggingface.co/datasets/replit/agent-challenge.textn<1K3 likes114 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.