Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01taekbae /llm-agent-backup-20260930 Convention belief in LLM agents An interpretability harness for zero-shot coordination in Hanabi. A language model watches a partner that follows one of two signalling conventions and has to work out which one. The interest is not in whether it succeeds but in where in the network the failure lives: whether the convention is represented at all, whether that representation drives the answer, and whether feeding it back repairs behaviour. Measurements and the design record are… See the full description on the dataset page: https://huggingface.co/datasets/taekbae/llm-agent-backup-20260930.0 likes2.8k downloads10d agoHugging Face02Exgentic /agent-llm-traces-v2 Exgentic Agent LLM Traces v2 — Agent Chat Only OpenTelemetry-shaped execution traces for 10,057 agent runs across 6 benchmarks (AppWorld, SWE-bench, BrowseCompPlus, τ²-bench Airline/Retail/Telecom), filtered to the agent under test's chat-only LLM calls. This is the dataset for replay testing, behavioral analysis, or any task where you care about what the benchmarked model actually did — not the eval scaffolding around it. This v2 release expands upon Exgentic/agent-llm-traces… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces-v2.tabulartext-generation10K<n<100K2 likes2.6k downloads3mo agoHugging Face03Exgentic /agent-llm-traces Multi-Benchmark LLM Agent Traces A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization. Collected by Exgentic - A platform for LLM observability and performance optimization. Dataset Overview This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/Exgentic/agent-llm-traces.tabulartext-generation1K<n<10K23 likes1.1k downloads4mo agoHugging Face04open-llm-leaderboard-old /details_llm-agents__tora-code-7b-v1.0 Dataset Card for Evaluation run of llm-agents/tora-code-7b-v1.0 Dataset Summary Dataset automatically created during the evaluation run of model llm-agents/tora-code-7b-v1.0 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-7b-v1.0.7 likes938 downloads3y agoHugging Face05GloriaaaM /LLM-Agent-Harness-Survey English | 中文 Agent Harness for Large Language Model Agents: A Survey ⭐ This repo is actively maintained. If you find it useful, please star the repo to stay updated and help others find it. The agent execution harness — not the model — is the primary determinant of agent reliability at scale.This survey formalizes the harness as a first-class architectural object H = (E, T, C, S, L, V), surveys 110+ papers, blogs and reports across 23 systems, and maps 9 open… See the full description on the dataset page: https://huggingface.co/datasets/GloriaaaM/LLM-Agent-Harness-Survey.documentn<1K10 likes695 downloads5mo agoHugging Face06open-llm-leaderboard-old /details_llm-agents__tora-code-34b-v1.0 Dataset Card for Evaluation run of llm-agents/tora-code-34b-v1.0 Dataset automatically created during the evaluation run of model llm-agents/tora-code-34b-v1.0 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-34b-v1.0.2 likes684 downloads3y agoHugging Face07open-llm-leaderboard-old /details_llm-agents__tora-13b-v1.0 Dataset Card for Evaluation run of llm-agents/tora-13b-v1.0 Dataset automatically created during the evaluation run of model llm-agents/tora-13b-v1.0 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-13b-v1.0.4 likes547 downloads3y agoHugging Face08tecnologiactc /automl_llm_agent_m2 AutoML-LLM Agent Module 2 Benchmark This dataset contains the Module 2 benchmark for evaluating an assistant that converts a Module 1 recipe, a user request, and a processed tabular dataset into an auditable autogluon.cloud.TabularCloudPredictor configuration. The repository is scoped to Module 2 only. Tables module2_cases: one row per Module 2 evaluation case. module2_queries: user requests for each case. module2_reference_configs: legacy reference JSON files… See the full description on the dataset page: https://huggingface.co/datasets/tecnologiactc/automl_llm_agent_m2.texttabular-classificationn<1K0 likes533 downloads3mo agoHugging Face09llm-agents /CriticBench Dataset Card for Dataset Name CriticBench is a comprehensive benchmark designed to assess LLMs' abilities to generate, critique/discriminate and correct reasoning across a variety of tasks. CriticBench encompasses five reasoning domains: mathematical, commonsense, symbolic, coding, and algorithmic. It compiles 15 datasets and incorporates responses from three LLM families. Dataset Details Dataset Description Curated by: THU Funded by [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/llm-agents/CriticBench.textquestion-answering1K<n<10K16 likes298 downloads3y agoHugging Face10visionscaper /agentic-llm-pretraining-1.7b Agentic LLM Pretraining Dataset A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.texttext-generation1M<n<10M3 likes232 downloads9mo agoHugging Face11RobinChen2001 /A-Survey-for-LLM-Agent-Trajectory-Analysis A Survey for LLM Agent Trajectory Analysis This dataset repository hosts the survey paper A Survey for LLM Agent Trajectory Analysis: From Failure Attribution to Enhancement and a structured metadata snapshot of the companion paper collection from Awesome-LLM-Agent-Trajectory-Analysis. The repository is intended for discovery, citation, and lightweight analysis of the literature around LLM agent trajectory analysis, including failure attribution, trajectory-based debugging… See the full description on the dataset page: https://huggingface.co/datasets/RobinChen2001/A-Survey-for-LLM-Agent-Trajectory-Analysis.documentn<1K2 likes222 downloads3mo agoHugging Face12tecnologiactc /automl_llm_agent_m1 AutoML-LLM Agent Module 1 Benchmark This dataset contains the Module 1 benchmark for evaluating an AutoML assistant that interprets user requests, selects tabular modeling settings, produces an auditable AutoGluon Tabular plan, and synthesizes the compact Module 1 recipe consumed by Module 2 through the mandatory final LLM writer used by all A-E variants. The repository is scoped to Module 1 only. Tables cases: one row per Module 1 evaluation case. queries: one… See the full description on the dataset page: https://huggingface.co/datasets/tecnologiactc/automl_llm_agent_m1.texttabular-classificationn<1K2 likes180 downloads2mo agoHugging Face13open-llm-leaderboard-old /details_llm-agents__tora-70b-v1.0 Dataset Card for Evaluation run of llm-agents/tora-70b-v1.0 Dataset automatically created during the evaluation run of model llm-agents/tora-70b-v1.0 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-70b-v1.0.1 likes148 downloads3y agoHugging Face14synonym /aiwolf-nlp-agent-llm AIWolfDial 2026 Power Play Evaluation Public data release: 2026-09-14. This dataset is available at synonym/aiwolf-nlp-agent-llm, with the snapshot tag release-20260914. The matching code distribution is 1.0.0-rc.3, commit d427dc299bacf4eb4cb41c114c8af476b71ac7ed. The code distribution uses a single root commit; this dataset is separate and is not included in that repository. Paper publication identifiers are still pending. The dataset is distributed under the MIT license in… See the full description on the dataset page: https://huggingface.co/datasets/synonym/aiwolf-nlp-agent-llm.texttext-generation1K<n<10K1 likes133 downloads27d agoHugging Face15agentlans /llm-prompt-collection LLM Prompt Collection Prompts from the first 100 000 rows of each dataset were collected then deduplicated and shuffled. The prompts_k* configs are semantically clustered subsets of the all config for diversity and coverage. The screened_prompts config is a subset of safe, high-quality prompts for the all config as classified using agentlans/bge-small-en-v1.5-prompt-screener Source Rows agentlans/chatgpt all 100 000 agentlans/magpie all 99 982… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/llm-prompt-collection.text1M<n<10M1 likes86 downloads6mo agoHugging Face16open-llm-leaderboard-old /details_llm-agents__tora-7b-v1.0 Dataset Card for Evaluation run of llm-agents/tora-7b-v1.0 Dataset Summary Dataset automatically created during the evaluation run of model llm-agents/tora-7b-v1.0 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-7b-v1.0.1 likes84 downloads3y agoHugging Face17DiscoPosse /agent-llm-traces Multi-Benchmark LLM Agent Traces A comprehensive dataset of OpenTelemetry traces capturing LLM inference behavior across multiple agent frameworks, benchmarks, and model providers. This dataset enables research into LLM performance analysis, agent behavior patterns, and inference optimization. Collected by Exgentic - A platform for LLM observability and performance optimization. Dataset Overview This dataset contains 1,781 execution traces capturing detailed agent… See the full description on the dataset page: https://huggingface.co/datasets/DiscoPosse/agent-llm-traces.tabulartext-generation1K<n<10K1 likes84 downloads4mo agoHugging Face18GXMZU /llm-rag-agent-papers llm-rag-agent-papers Research papers on LLM, RAG, and AI Agents - Knowledge base for RAG pipeline Dataset Structure This dataset contains three subsets: llm: Large Language Model related content rag: Retrieval-Augmented Generation related content agent: AI Agent related content Usage from datasets import load_dataset # Load all subsets dataset = load_dataset("GXMZU/llm-rag-agent-papers") # Load specific subset llm_data =… See the full description on the dataset page: https://huggingface.co/datasets/GXMZU/llm-rag-agent-papers.tabulartext-generation1K<n<10K3 likes77 downloads9mo agoHugging Face19fineset-io /llm-agent-papers LLM Agent & Tool-Use Papers — FineSet A research-paper dataset on LLM Agent & Tool-Use Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-12. It is not auto-updated. Research on LLM Agent & Tool-Use Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/llm-agent-papers.tabulartext-classification1K<n<10K1 likes73 downloads4mo agoHugging Face20LLM-OS-Models /ECHO-Terminal-Agent-Prepared-Data ECHO-style Terminal Agent Prepared Data for LFM RLVR Prepared on 2026-06-09 for local no-Docker LFM terminal RLVR experiments. This dataset converts public terminal-agent task archives into two formats: echo_terminal_tasks_*.parquet: ECHO/SkyRL-style rows with prompt, path, and task_binary. lfm_live_tasks_mixed.*: local LFM no-Docker trainer rows with prompt, task_id, source, task_binary_b64, and metadata. Current manifest: total rows: 1500 Endless Terminals: 772 rows… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/ECHO-Terminal-Agent-Prepared-Data.text1K<n<10K0 likes69 downloads4mo agoHugging Face21open-llm-leaderboard /EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-details.tabular10K<n<100K4 likes62 downloads2y agoHugging Face22open-llm-leaderboard /Josephgflowers__Tinyllama-STEM-Cinder-Agent-v1-detailsgated Dataset Card for Evaluation run of Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1 Dataset automatically created during the evaluation run of model Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__Tinyllama-STEM-Cinder-Agent-v1-details.tabular10K<n<100K0 likes59 downloads2y agoHugging Face23open-llm-leaderboard /agentlans__Llama-3.2-1B-Instruct-CrashCourse12K-detailsgated Dataset Card for Evaluation run of agentlans/Llama-3.2-1B-Instruct-CrashCourse12K Dataset automatically created during the evaluation run of model agentlans/Llama-3.2-1B-Instruct-CrashCourse12K The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/agentlans__Llama-3.2-1B-Instruct-CrashCourse12K-details.tabular10K<n<100K0 likes54 downloads2y agoHugging Face24open-llm-leaderboard-old /details_llm-agents__tora-code-13b-v1.0 Dataset Card for Evaluation run of llm-agents/tora-code-13b-v1.0 Dataset automatically created during the evaluation run of model llm-agents/tora-code-13b-v1.0 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_llm-agents__tora-code-13b-v1.0.1 likes52 downloads3y agoHugging Face25open-llm-leaderboard /agentlans__Llama3.1-Daredevilish-Instruct-detailsgated Dataset Card for Evaluation run of agentlans/Llama3.1-Daredevilish-Instruct Dataset automatically created during the evaluation run of model agentlans/Llama3.1-Daredevilish-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/agentlans__Llama3.1-Daredevilish-Instruct-details.tabular10K<n<100K0 likes52 downloads2y agoHugging Face26Gene829 /gene-llm-agents-corpus llm-agents-corpus v92 Auto-built (demand): 1 open request(s) and 0 recent download(s) for 'llm-agents' with no dataset newer than 14 days Kind: scraped Domain: llm-agents Records: 702 Created: 2026-07-08T17:36:14+00:00 SHA-256: 74c00af747e66d1d4abf248168263fb9d5f1e424182d7adcdef160533c32dae2 Pipeline: v2.0.0 Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": null, "min_judge": null} Sources huggingface: 301 papers: 197 arxiv: 145 github:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-llm-agents-corpus.text-generationn<1K1 likes50 downloads3mo agoHugging Face27Dharun72 /llm-agentic-swiss-legal-checkpoints LLM Agentic Legal Information Retrieval — Checkpoints Public artifacts from competing in the Kaggle competition. Best public LB Submission LB v6 LightGBM baseline 0.0709 v9 DeepSeek paragraph injection 0.13167 v11 court sibling expansion 0.13665 v12 Qwen2.5-14B LoRA 0.13204 Structure submissions/ — final submission CSVs per version picks/ — per-query LLM output caches (V4-Pro picks, LoRA picks, profiles) training/ — LEXam fine-tuning data… See the full description on the dataset page: https://huggingface.co/datasets/Dharun72/llm-agentic-swiss-legal-checkpoints.textn<1K0 likes40 downloads5mo agoHugging Face28agentlans /Estwld-empathetic_dialogues_llmReformatted version of Estwld/empathetic_dialogues_llm. Changes: Added a random system prompt for the AI to be empathetic Truncated conversations that don't end with the AI's turn Removed extra fields not needed in the conversation Limitations: The dialogues aren't very long No background info for the user and AI English only texttext-generation10K<n<100K0 likes39 downloads2y agoHugging Face29open-llm-leaderboard /EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-detailsgated Dataset Card for Evaluation run of EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Meta-Llama-3.1-8B-Instruct-Agent-0.005-128K-code-COT-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face30GXMZU /llm-rag-agent-blogs llm-rag-agent-blogs Technical blogs on LLM, RAG, and AI Agents - Knowledge base for RAG pipeline Dataset Structure This dataset contains three subsets: llm: Large Language Model related content rag: Retrieval-Augmented Generation related content agent: AI Agent related content Usage from datasets import load_dataset # Load all subsets dataset = load_dataset("GXMZU/llm-rag-agent-blogs") # Load specific subset llm_data =… See the full description on the dataset page: https://huggingface.co/datasets/GXMZU/llm-rag-agent-blogs.texttext-generation1K<n<10K1 likes37 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.