Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kth8 /python-toolcallsLogs from run_python_code tool used for benchmarking. tabular10K<n<100K0 likes11k downloads5mo agoHugging Face02nvidia /Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 Dataset Description: We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.tabular10K<n<100K42 likes1.5k downloads11d agoHugging Face03clem /hf-coding-tools-traces_april26 HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 32 sessions, one per (tool, model, effort, thinking) configuration 9,130 query → response turns total (≈18,260 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/clem/hf-coding-tools-traces_april26.tabularn<1K0 likes569 downloads5mo agoHugging Face04davidkling /hf-coding-tools-traces-run-april12 HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 8,875 query → response turns total (≈17,750 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-run-april12.tabularn<1K0 likes532 downloads5mo agoHugging Face05davidkling /hf-coding-tools-traces HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 32 sessions, one per (tool, model, effort, thinking) configuration 9,130 query → response turns total (≈18,260 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces.tabularn<1K3 likes486 downloads5mo agoHugging Face06davidkling /hf-coding-tools-traces-all HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 9,603 query → response turns total (≈19,206 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-all.tabularn<1K0 likes442 downloads5mo agoHugging Face07cfahlgren1 /hf-coding-tools-traces HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 32 sessions, one per (tool, model, effort, thinking) configuration 9,130 query → response turns total (≈18,260 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/hf-coding-tools-traces.tabularn<1K0 likes436 downloads5mo agoHugging Face08clem /hf-coding-tools-traces HF Coding Tools — Agent Traces This dataset rehydrates the benchmark run in davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 8,881 query → response turns total (≈17,762 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6 (Copilot), gpt-4.1… See the full description on the dataset page: https://huggingface.co/datasets/clem/hf-coding-tools-traces.tabularn<1K10 likes414 downloads6mo agoHugging Face09davidkling /hf-coding-tools-traces-discovery HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 9,022 query → response turns total (≈18,044 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-discovery.tabularn<1K1 likes287 downloads5mo agoHugging Face10while-ai /tool-call-efficiency tool-call-efficiency Made with the whileai SDK · Collections: Efficiency, Start here: foundational post-training datasets Teach an agent to make every tool call count. An agent that calls a tool twice with the same arguments, looks up what the user just told it, or keeps calling after the task is done is slow, expensive, and harder to trust. Ask a base Qwen3-4B to work through 1,133 tool-using tasks across six agents and it does this a lot: only 52% of its 6,681 rollouts finish… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tool-call-efficiency.tabulartext-generation1K<n<10K0 likes237 downloads18d agoHugging Face11ToolGym /long-horizon-eval long-horizon-eval Evaluation results for long-horizon agent performance Dataset Description This dataset contains evaluation results for agent trajectories, including quality assessments and performance metrics. Dataset Structure The dataset is organized by model name, with each model having separate JSONL files for different experimental passes. long-horizon-eval/ ├── model-1/ │ ├── pass@1.jsonl │ ├── pass@2.jsonl │ └── pass@3.jsonl ├── model-2/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/ToolGym/long-horizon-eval.tabulartext-generation1K<n<10K0 likes219 downloads9mo agoHugging Face12davidkling /hf-coding-tools-traces-run-april22-v2 HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 5 sessions, one per (tool, model, effort, thinking) configuration 728 query → response turns total (≈1,456 events) Tools covered: claude_code, codex, cursor Models: claude-opus-4-6, claude-sonnet-4-6, composer-2, gpt-4.1, gpt-4.1-mini Each… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-run-april22-v2.tabularn<1K0 likes204 downloads5mo agoHugging Face13MasterVito /swe-agent-tool-rubrics-860 SWE Agent 逐 turn 工具调用评判数据集(860 个决策点) 本数据集来自 2026-08-06 的一次实验:**从真实 SWE agent 轨迹中归纳"怎么判断一次工具调用的好坏"**。 包含两个文件: 文件 行数 大小 内容 cases.jsonl 860 5.0 MB 决策点原始数据(题目、历史、两个候选命令、执行结果、现役判官打分) map_io.jsonl 860 9.6 MB 每个决策点喂给 GPT-5.6 的完整 prompt 原文与完整回复 两个文件通过 case_id 一一对应。 背景:为什么是"按动作分类"而不是"按工具分类" 轨迹来自 slime 的 minimal harness,该 harness 只暴露一个工具 bash (slime/agent/harness/minimal.py 里的 BASH_TOOL),全部 328,270 次调用的工具名都是 bash。 所以"不同工具用不同 rubric"无法按工具名实现,只能按命令在干什么分类。… See the full description on the dataset page: https://huggingface.co/datasets/MasterVito/swe-agent-tool-rubrics-860.tabular1K<n<10K1 likes181 downloads2mo agoHugging Face14mcp-tool-shop /jam-rollout-arc-evals Rollout arc — raw generations Every model generation behind the write-ups in mcp-tool-shop-org/ai-jam-sessions under experiments/rollout-arc/p4/. Two things you can do with this. Check our arithmetic. The repo has the readout scripts, the preregistrations and the intervals — but the generations they were computed from are ~51 MB and were never committed, so a clone got the conclusions and no way to recompute them. These are those files, unfiltered. Or run the loop yourself. The… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-rollout-arc-evals.tabulartext-generation1K<n<10K0 likes152 downloads26d agoHugging Face15build-small-hackathon /agenda-parser-tool-traces Agenda Parser — tool-calling reasoning traces ReAct tool-calling traces for the Agenda Parser agents: each row is one agent step — a {system, user, assistant} chat example where the assistant emits a single JSON action {"thought", "tool", "args"}. Two agents are covered (tagged by meta.domain): agenda — the uploaded-packet research agent, over real public-meeting agenda packets (tools: list/read items, semantic + exact search, summarize, report). Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.documenttext-generation1K<n<10K0 likes142 downloads4mo agoHugging Face16evalstate /model-toolcall-research Model Toolcall Research tabularn<1K2 likes140 downloads5mo agoHugging Face17rdubwiley /agenda-parser-tool-traces Agenda Parser — tool-calling reasoning traces ReAct tool-calling traces for the Agenda Parser agents: each row is one agent step — a {system, user, assistant} chat example where the assistant emits a single JSON action {"thought", "tool", "args"}. Two agents are covered (tagged by meta.domain): agenda — the uploaded-packet research agent, over real public-meeting agenda packets (tools: list/read items, semantic + exact search, summarize, report). Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/rdubwiley/agenda-parser-tool-traces.documenttext-generation1K<n<10K0 likes127 downloads4mo agoHugging Face18cfahlgren1 /model-toolcall-research Model Toolcall Research This dataset stores newline-delimited agent traces from bounded research runs on model repository tool-schema support. The Dataset Viewer is configured to index only .jsonl files: toolcall_traces loads trace files under traces/**/*.jsonl. research_session loads top-level provenance/session traces from *.jsonl. The archive/ directory preserves the earlier .trace.json uploads for reference, but those files are newline-delimited JSON streams rather than… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/model-toolcall-research.tabularn<1K0 likes118 downloads5mo agoHugging Face19jinjinyien /CLASHBench-ToolMind CLASHBench ToolMind training data The frozen processed ToolMind input used by CLASHBench GPU workloads. File: toolmind50k_direct_plain.json Records: 50000 Format: ShareGPT messages with role and content Cases using toolmind_fullfilter50k_direct_plain_train.json can use a local filename alias. All 50,000 training conversations are preserved; local source-path metadata is excluded. dataset_info.json supplies both LlamaFactory dataset names. Verify downloaded bytes with sha256sum… See the full description on the dataset page: https://huggingface.co/datasets/jinjinyien/CLASHBench-ToolMind.tabular10K<n<100K0 likes114 downloads22d agoHugging Face20bilalabic /turkish-tool-calling Türkçe Tool-Calling Veri Seti 56.247 kayıt. xLAM/APIGen 60k ve NVIDIA When2Call'dan türetilmiş, üç davranış sınıfı içeren Türkçe function-calling veri seti. from datasets import load_dataset ds = load_dataset("bilalabic/turkish-tool-calling") # mesaj listesi ds = load_dataset("bilalabic/turkish-tool-calling", "table") # düz tablo ds = load_dataset("bilalabic/turkish-tool-calling", "sharegpt") # ShareGPT İçerik Kayıt 56.247… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/turkish-tool-calling.tabulartext-generation100K<n<1M0 likes108 downloads2mo agoHugging Face21spade-rl /SPADE-Environment-Pool-GPT5.5-ToolUse SPARE GPT-5.5 Multi-Turn Tool-Use Games v1 A public static pool of 11,039 validated multi-turn tool-use environments generated by GPT-5.5 for SPARE actor training. Training alignment Source recipe: Qwen3-30B-A3B 0624 tool-use GAMES configuration 400 rollouts x 24 games/rollout = 9,600 no-reuse games required 11,039 validated games provide 1,439 games of headroom Six balanced skills: API orchestration, data retrieval, state modification, error recovery, tool… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-ToolUse.tabularreinforcement-learning10K<n<100K1 likes107 downloads2mo agoHugging Face22dougalldeepmind /2026-07-31-toolcalling-tulu-sft-run Run record — Qwen3.6-27B tool-calling 20/80 SFT Everything the training run produced except the weights: the TRL log history, the resolved config, the environment, the loss/accuracy figure and its greppable markdown mirror. The adapter is at LASR-Callum/2026-07-31-wrongly-trained-qwen36-toolcalling-tulu-lora-20-80; the training data is at LASR-Callum/2026-07-31-toolcalling-tulu-20-80-mixture. Required metadata field value experiment One bf16 LoRA SFT… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-toolcalling-tulu-sft-run.tabularn<1K0 likes104 downloads1mo agoHugging Face23davidkling /hf-coding-tools-traces-builder HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 5 sessions, one per (tool, model, effort, thinking) configuration 581 query → response turns total (≈1,162 events) Tools covered: claude_code, codex, cursor Models: claude-opus-4-6, claude-sonnet-4-6, composer-2, gpt-4.1, gpt-4.1-mini Each… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-builder.tabularn<1K0 likes97 downloads5mo agoHugging Face24RUC-AIBOX /STILL-3-TOOL-32B-Datatabularn<1K3 likes88 downloads2y agoHugging Face25lokahq /drug-tool-sft Drug Tool Calling SFT Dataset Summary Drug Tool Calling SFT is a supervised fine-tuning dataset for biomedical agent tool use. Each example contains a user prompt, the available drug-discovery tool schemas, and a serialized assistant/tool transcript with at least one structured tool call. The dataset was generated from the Loka drug-discovery copilot prompt bank and executed through the OpenAI-compatible Strands model provider against OpenRouter. It is intended… See the full description on the dataset page: https://huggingface.co/datasets/lokahq/drug-tool-sft.tabulartext-generation1K<n<10K0 likes84 downloads3mo agoHugging Face26asingh15 /qwen35-2b-tool-use-qwen36-27b-curation-candidates Full candidate collections: 2B tool use + 27B data curation This public Dataset contains two complete, unredacted, exact-40 candidate collections: Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and 233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL, Spider, and TravelPlanner. Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021 targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.tabulartext-generation100K<n<1M0 likes81 downloads2mo agoHugging Face27gradients-io-tournaments /pvp-tool-calling-sft PvP tool-calling SFT cold-start data Claude-vs-Claude games played through the G.O.D PvP tool-calling harness. Each row is one model turn (or post-game reflection): the system+user prompt the harness built, the assistant response (content + tool_calls), and the tools schemas — i.e. the OpenAI messages+tools format consumed by tokenizer.apply_chat_template(messages, tools=tools). On a move turn the assistant co-emits any memory-tool edits and a game_action committing a legal… See the full description on the dataset page: https://huggingface.co/datasets/gradients-io-tournaments/pvp-tool-calling-sft.tabularn<1K0 likes79 downloads4mo agoHugging Face28celerity-labs /celeritybench-tool-choice Which small model should run a Mac launcher Celeritas is a Spotlight-style launcher that turns what somebody types into tool calls on their own machine. Picking the model to put behind it meant measuring them, and the numbers were going on a public page, so the runs behind them are here. The question is narrow on purpose: for an agent with about thirty tools on a desktop, which model picks the right one? Not reasoning, not code, not knowledge. Tool choice, on short everyday… See the full description on the dataset page: https://huggingface.co/datasets/celerity-labs/celeritybench-tool-choice.tabulartext-generation1K<n<10K0 likes74 downloads21d agoHugging Face29ComparEdge /ai-tools-pricing-2026 AI Tools Pricing & Features Dataset 2026 A structured dataset of 104 AI tools across 9 categories — pricing plans, user ratings, and feature lists. Built for market analysis, recommendation systems, and pricing research. Dataset Description This dataset covers the AI software landscape in 2026, including LLMs, coding assistants, image generators, and more. Each entry contains real pricing data, user ratings, and feature sets. Source Curated from the live… See the full description on the dataset page: https://huggingface.co/datasets/ComparEdge/ai-tools-pricing-2026.tabulartext-classificationn<1K1 likes61 downloads6mo agoHugging Face30Sumukh66 /toolcall-datatabular100K<n<1M0 likes53 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.