datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apigen-function-calling
Dataset card for argilla/apigen-function-calling
This dataset is a merge of argilla/Synth-APIGen-v0.1
and Salesforce/xlam-function-calling-60k, making
over 100K function calling examples following the APIGen recipe.
Prepare for training
This version is not ready to do fine tuning, but you can run a script like prepare_for_sft.py
to prepare it, and run the same recipe that can be found in
argilla/Llama-3.2-1B-Instruct-APIGen-FC-v0.1#training-procedure.
Modify the prompt… See the full description on the dataset page: https://huggingface.co/datasets/argilla/apigen-function-calling.black-box-api-challenges
Dataset Card
Paper: On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research
Abstract: Perception of toxicity evolves over time and often differs between geographies and cultural backgrounds. Similarly, black-box commercially available APIs for detecting toxicity, such as the Perspective API, are not static, but frequently retrained to address any unattended weaknesses and biases. We evaluate the implications of these changes on the reproducibility of findings… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/black-box-api-challenges.Synth-APIGen-v0.1
Dataset card for Synth-APIGen-v0.1
This dataset has been created with distilabel.
Pipeline script: pipeline_apigen_train.py.
Dataset creation
It has been created with distilabel==1.4.0 version.
This dataset is an implementation of APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets in distilabel,
generated from synthetic functions. The process can be summarized as follows:
Generate (or in this case modify) python… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Synth-APIGen-v0.1.APIGen-MT-5k
Summary
APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications
This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Code: https://github.com/apigen-mt/apigen-mt.github.io
The repo contains 5000 multi-turn trajectories collected by APIGen-MT
This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.deepseek-v2-codder-minecraft-apisynth-apigen-qwen
Dataset Card for argilla-warehouse/synth-apigen-qwen
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
synth_apigen.py.
Dataset creation
This dataset is a replica in distilabel of the framework
defined in: APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets.
Using the seed dataset of synthetic python functions in argilla-warehouse/python-seed-tools,
the… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/synth-apigen-qwen.api_gateway_jwt_oauth_revocation_engine_teaser
🚀 Cloud Architecture - API Gateway JWT & OAuth Token Revocation Engine (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Purchase Full Production Master Dataset on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 Domain Focus & Capabilities
Distributed Redis token blacklisting, JTI invalidation race condition mitigation, and… See the full description on the dataset page: https://huggingface.co/datasets/emgena/api_gateway_jwt_oauth_revocation_engine_teaser.api-sequencing-training-pool
API sequencing training pool
Requests paired with the sequence of API calls that answers them, from three public datasets read
at the pinned revisions named below and from a fourth that the builder of this pool generated,
laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 136292 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
request… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/api-sequencing-training-pool.apigen-smollm-trl-FC
Dataset card for argilla-warehouse/apigen-smollm-trl-FC
This dataset is a merge of argilla/Synth-APIGen-v0.1
and Salesforce/xlam-function-calling-60k, and was prepared for training using the script
prepare_for_sft.py that can be found in the repository files.
References
@article{liu2024apigen,
title={APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets},
author={Liu, Zuxin and Hoang, Thai and Zhang, Jianguo and Zhu, Ming and… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/apigen-smollm-trl-FC.synth-apigen-llama
Dataset Card for argilla-warehouse/synth-apigen-llama
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
synth_apigen.py.
Dataset creation
This dataset is a replica in distilabel of the framework
defined in: APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets.
Using the seed dataset of synthetic python functions in argilla-warehouse/python-seed-tools,
the… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/synth-apigen-llama.Open-Router-API-Pricing-Analysis
OpenRouter API Pricing Analysis Dataset
Overview
This dataset provides a point-in-time capture of pricing and parameters for LLMs available through the OpenRouter API for inference.
Contents
Raw Data (raw/)
Contains the original data extracted from the OpenRouter API, including:
Model pricing (input/output token costs)
Model parameters and specifications
Computed fields such as output/input token price ratios
Enhanced Data (hf-enhanced/)… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Open-Router-API-Pricing-Analysis.WM-ENT-API-DUMP_FR_2026.07
Jeux De Données : Dump WikiMedia Français Juillet 2026 : Extraction et Nettoyage Complet
Description
Ce JDD contient des articles extraits du dump complet de Wikimedia d'août 2026, nettoyés et structurés pour l'entraînement de modèles d'apprentissage automatique.
Source originale : Wikimedia Enterprise
Licence : CC BY 4.0
Volume initial : 46 Go en tar.gz : ~120-150 Go décompressés
Volume final nettoyé : 14.6 Go : 144 fichiers .jsonl
Fichiers traités : 144… See the full description on the dataset page: https://huggingface.co/datasets/MisterAI/WM-ENT-API-DUMP_FR_2026.07.api-calling-training-pool
API calling training pool
Public API-calling data from five datasets, read at the pinned revisions named below and laid out
twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 199186 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
query
the user's request, as its source publishes it
functions
the declarations offered with the request, as a list… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/api-calling-training-pool.apigen-mt-5k-parsed
[PARSED] APIGen-MT-5k
The data in this dataset is a full of the original Salesforce/APIGen-MT-5k
Subset name
multi-turn
parallel
multiple definition
Last turn type
number of dataset
apigen-mt-5k
yes
no
yes
complex
5k
This is a re-parsing formatting dataset for the APIGen-MT-5k official dataset.
Load the dataset
from datasets import load_dataset
ds = load_dataset("minpeter/apigen-mt-5k-parsed")
print(ds)
# DatasetDict({
# train: Dataset({
#… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/apigen-mt-5k-parsed.apigen-synth-trl
Dataset card
This dataset is a version of argilla/Synth-APIGen-v0.1 prepared for
fine-tuning using trl. To generate it, the following script was run:
from datasets import load_dataset
from jinja2 import Template
SYSTEM_PROMPT = """
You are an expert in composing functions. You are given a question and a set of possible functions.
Based on the question, you will need to make one or more function/tool calls to achieve the purpose.
If none of the functions can be used, point it out… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/apigen-synth-trl.backend-api-instruction-dataset
Backend & API Development Dataset
Instruction dataset focused on RESTful API design, WebSocket real-time communication, microservices patterns, and API best practices.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Backend Api topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/backend-api-instruction-dataset.api-eval-20
APIEval-20: A Benchmark for Black-Box API Test Suite Generation
Motivation
Testing APIs thoroughly is one of the most critical, yet consistently underserved, activities in software engineering. Despite a rich ecosystem of API testing tools — Postman, RestAssured, Schemathesis, Dredd, and others — we found ourselves asking a deceptively simple question:
Given only the schema and an example payload of an API request — no source code, no documentation, no prior knowledge —… See the full description on the dataset page: https://huggingface.co/datasets/kusho-ai/api-eval-20.agentic-tool-use-multi-api-orchestration-2026
⚡ Agentic Tool-Use, Multi-API Calling & Autonomous Function Orchestration (2026)
Official 100-sample production preview of the Agentic Tool-Use & Multi-API Orchestration Suite (2026) by BeatsProm AI Research Lab. Engineered for parallel tool calling (<tool_call>), strict JSON-schema enforcement, stateful cursor pagination, and self-healing API error recovery.
🏛️ THE 20 AGENTIC OPERATIONAL CORES:
Parallel Portfolio Rebalancing: Multi-leg execution with… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/agentic-tool-use-multi-api-orchestration-2026.eplan-2027-api-qa
EPLAN 2027 Platform API Q&A for C# Automation
EPLAN 2027 API question-answer dataset for C#/.NET automation, EPLAN
scripting, electrical engineering assistants, code generation, RAG, and LLM
instruction tuning.
This dataset contains grounded English question-answer pairs generated from
excerpts of the EPLAN Platform API 2027 documentation. It is designed for
building assistants that help engineers understand and write EPLAN API scripts,
actions, add-ins, and automation code.… See the full description on the dataset page: https://huggingface.co/datasets/covaga/eplan-2027-api-qa.drupal7-dataset-api
Drupal 7 Dataset API
Drupal 7 reached end of life in early 2025, but plenty of sites still run on it, and the people maintaining them still need good answers. This dataset is meant to help with that: 7,952 examples about the Drupal 7 API, taken only from api.drupal.org. You can use it to fine-tune a model, to test how well a model knows Drupal 7, or as reference material for a retrieval (RAG) setup.
Everything is in English. It was put together by Daniel Ricardo Ramirez Marin as… See the full description on the dataset page: https://huggingface.co/datasets/danemarparceros/drupal7-dataset-api.unity_api_2022_3
Unity3d 2022.3 LTS API & Manual
In this dataset, you'll find a series of Q&A for the Unity3d API and Manual.
Dataset Creation
Download the unity offline documentation.
Process documentation, extract title, and description. Clean documentation.
Process each title and description item in llama3-8B-Instruct in order to generate several questions that capture the meaning of the API.
Re-process in llama3-8B-Instruct with question and API to generate the answer.… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/unity_api_2022_3.exercise-api
Exercise API — Dataset
Dataset de 104 ejercicios de gimnasio (bilingüe ES/EN) derivado de la
Exercise API. Cada ejercicio incluye grupo muscular,
equipamiento, músculos principal/secundario, instrucciones paso a paso e ilustración
masculina y femenina (208 imágenes en total).
Configuraciones
images — 1 fila por imagen (208). Etiquetas (grupo, equipamiento,
músculos, género) + caption_es/caption_en. Para clasificación de imagen y multimodal
(image-to-text / VQA).… See the full description on the dataset page: https://huggingface.co/datasets/natzx94/exercise-api.APIGen-MT-5k
Summary
APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications
This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Code: https://github.com/apigen-mt/apigen-mt.github.io
The repo contains 5000 multi-turn trajectories collected by APIGen-MT
This dataset is a subset of the data used to train the xLAM-2… See the full description on the dataset page: https://huggingface.co/datasets/WalterWangtao/APIGen-MT-5k.quantum-api-drift
Quantum API Drift
Quantum API Drift is an evaluation benchmark for measuring whether
LLM-generated quantum code targets the requested Qiskit SDK version. It
accompanies the paper
Benchmarking API Drift in LLM-Generated Quantum Code Across Successive SDK Versions.
The benchmark evaluates version fidelity, cross-version compatibility, failure
modes, and documentation-guided repair across Qiskit 0.43, 1.3, and 2.0.
Dataset Configurations
benchmark
The… See the full description on the dataset page: https://huggingface.co/datasets/arasyi/quantum-api-drift.API-Upgrade
API-Upgrade
API-Upgrade evaluates repository-level Rust synthesis against newer dependency APIs.
This dataset was used for evaluation in the paper
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code.
You can find the corresponding evaluation code in
the project GitHub repository.
Example Usage
from datasets import load_dataset
import json
dataset = load_dataset("eth-sri/API-Upgrade")
for instance in dataset["test"]:… See the full description on the dataset page: https://huggingface.co/datasets/eth-sri/API-Upgrade.api-vulnerability-dataset-10k
API Vulnerability Dataset (10K)
A dataset of 10,000 API-specific vulnerability samples used to fine-tune harsharajkumar273/api-security-qlora — a QLoRA adapter on CodeLlama-7b for automated API security analysis.
Dataset Summary
Each sample contains a vulnerable or clean API endpoint code snippet paired with a structured security analysis covering vulnerability type, severity, CWE ID, and a remediated version.
Language & Framework Distribution
Language… See the full description on the dataset page: https://huggingface.co/datasets/harsharajkumar273/api-vulnerability-dataset-10k.apigen-inferred
apigen-inferred
A verified, GPT-5.5-distilled subset of the
argilla/apigen-function-calling
dataset (109k rows in the upstream), with every golden tool-call argument
labelled as literal or dependency-derived to enable a clean
function-calling benchmark.
Pipeline
Filter the upstream to rows where every called API actually works
(replay each tool call against the real implementation — distilabel
Python functions or live RapidAPI / cached responses) → 45,984 rows.
Distill… See the full description on the dataset page: https://huggingface.co/datasets/gdgc-metacong/apigen-inferred.Research-Enterprise-Synth-API
EnterpriseSynth Public API Specs and Generated Artifacts
EnterpriseSynth converts OpenAPI/Swagger specifications into synthetic tool-use
training and evaluation artifacts without executing live API calls.
This dataset repository contains the public, redistributable dataset artifacts
from anote-ai/Research-Enterprise-Synth-API:
data/specs/: public OpenAPI/Swagger specs used by the experiments.
data/specs/phase3/: additional public held-out API specs.
data/generated/: generated… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/Research-Enterprise-Synth-API.triage-french
Triage médical français (FRENCH — SFMU)
Dataset synthétique pour entraîner un agent conversationnel de triage médical aux
urgences, fondé sur la grille FRENCH (SFMU, 5 niveaux d'urgence).
⚠️ Ce dataset est synthétique : il a été généré par un LLM à partir de règles
déterministes, sans validation clinique. Il est destiné à la recherche uniquement,
et non à un usage médical en production.
Construction
La grille FRENCH (196 règles, 16 catégories) sert de source de… See the full description on the dataset page: https://huggingface.co/datasets/apinsun/triage-french.SO-Python_QA-API_Usage-tanh_score
Stack Overflow Python Q&A Dataset
Description
Filtered Python Q&A with API_Usage subcategory without:
Images
Links
Blocks of code
Scores in Q1-Q3 scaled with MaxAbsScaler. Tanh function applyed to joint Scores.
