Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01argilla /apigen-function-calling Dataset card for argilla/apigen-function-calling This dataset is a merge of argilla/Synth-APIGen-v0.1 and Salesforce/xlam-function-calling-60k, making over 100K function calling examples following the APIGen recipe. Prepare for training This version is not ready to do fine tuning, but you can run a script like prepare_for_sft.py to prepare it, and run the same recipe that can be found in argilla/Llama-3.2-1B-Instruct-APIGen-FC-v0.1#training-procedure. Modify the prompt… See the full description on the dataset page: https://huggingface.co/datasets/argilla/apigen-function-calling.texttext-generation100K<n<1M20 likes22k downloads2y agoHugging Face02CohereLabs /black-box-api-challenges Dataset Card Paper: On the Challenges of Using Black-Box APIs for Toxicity Evaluation in Research Abstract: Perception of toxicity evolves over time and often differs between geographies and cultural backgrounds. Similarly, black-box commercially available APIs for detecting toxicity, such as the Perspective API, are not static, but frequently retrained to address any unattended weaknesses and biases. We evaluate the implications of these changes on the reproducibility of findings… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/black-box-api-challenges.texttext-classification10 likes8.6k downloads3y agoHugging Face03argilla /Synth-APIGen-v0.1 Dataset card for Synth-APIGen-v0.1 This dataset has been created with distilabel. Pipeline script: pipeline_apigen_train.py. Dataset creation It has been created with distilabel==1.4.0 version. This dataset is an implementation of APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets in distilabel, generated from synthetic functions. The process can be summarized as follows: Generate (or in this case modify) python… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Synth-APIGen-v0.1.texttext-generation10K<n<100K65 likes5.1k downloads2y agoHugging Face04Salesforce /APIGen-MT-5k Summary APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay Code: https://github.com/apigen-mt/apigen-mt.github.io The repo contains 5000 multi-turn trajectories collected by APIGen-MT This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.textquestion-answering1K<n<10K115 likes4.9k downloads1y agoHugging Face05SkillFi /deepseek-v2-codder-minecraft-apitexttext-generationn<1K0 likes1.4k downloads1y agoHugging Face06argilla-warehouse /synth-apigen-qwen Dataset Card for argilla-warehouse/synth-apigen-qwen This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: synth_apigen.py. Dataset creation This dataset is a replica in distilabel of the framework defined in: APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets. Using the seed dataset of synthetic python functions in argilla-warehouse/python-seed-tools, the… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/synth-apigen-qwen.texttext-generation10K<n<100K7 likes1.1k downloads2y agoHugging Face07emgena /api_gateway_jwt_oauth_revocation_engine_teaser 🚀 Cloud Architecture - API Gateway JWT & OAuth Token Revocation Engine (Evaluation Teaser) ⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Purchase Full Production Master Dataset on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout! 🌟 Domain Focus & Capabilities Distributed Redis token blacklisting, JTI invalidation race condition mitigation, and… See the full description on the dataset page: https://huggingface.co/datasets/emgena/api_gateway_jwt_oauth_revocation_engine_teaser.texttext-generationn<1K0 likes893 downloads9d agoHugging Face08Emulated-Inc /api-sequencing-training-pool API sequencing training pool Requests paired with the sequence of API calls that answers them, from three public datasets read at the pinned revisions named below and from a fourth that the builder of this pool generated, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 136292 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file request… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/api-sequencing-training-pool.texttext-generation100K<n<1M0 likes583 downloads28d agoHugging Face09argilla-warehouse /apigen-smollm-trl-FC Dataset card for argilla-warehouse/apigen-smollm-trl-FC This dataset is a merge of argilla/Synth-APIGen-v0.1 and Salesforce/xlam-function-calling-60k, and was prepared for training using the script prepare_for_sft.py that can be found in the repository files. References @article{liu2024apigen, title={APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets}, author={Liu, Zuxin and Hoang, Thai and Zhang, Jianguo and Zhu, Ming and… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/apigen-smollm-trl-FC.texttext-generation100K<n<1M2 likes380 downloads2y agoHugging Face10argilla-warehouse /synth-apigen-llama Dataset Card for argilla-warehouse/synth-apigen-llama This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: synth_apigen.py. Dataset creation This dataset is a replica in distilabel of the framework defined in: APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets. Using the seed dataset of synthetic python functions in argilla-warehouse/python-seed-tools, the… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/synth-apigen-llama.texttext-generation10K<n<100K3 likes243 downloads2y agoHugging Face11danielrosehill /Open-Router-API-Pricing-Analysis OpenRouter API Pricing Analysis Dataset Overview This dataset provides a point-in-time capture of pricing and parameters for LLMs available through the OpenRouter API for inference. Contents Raw Data (raw/) Contains the original data extracted from the OpenRouter API, including: Model pricing (input/output token costs) Model parameters and specifications Computed fields such as output/input token price ratios Enhanced Data (hf-enhanced/)… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Open-Router-API-Pricing-Analysis.texttext-generation1K<n<10K0 likes209 downloads11mo agoHugging Face12MisterAI /WM-ENT-API-DUMP_FR_2026.07 Jeux De Données : Dump WikiMedia Français Juillet 2026 : Extraction et Nettoyage Complet Description Ce JDD contient des articles extraits du dump complet de Wikimedia d'août 2026, nettoyés et structurés pour l'entraînement de modèles d'apprentissage automatique. Source originale : Wikimedia Enterprise Licence : CC BY 4.0 Volume initial : 46 Go en tar.gz : ~120-150 Go décompressés Volume final nettoyé : 14.6 Go : 144 fichiers .jsonl Fichiers traités : 144… See the full description on the dataset page: https://huggingface.co/datasets/MisterAI/WM-ENT-API-DUMP_FR_2026.07.texttext-generation100K<n<1M0 likes200 downloads17d agoHugging Face13Emulated-Inc /api-calling-training-pool API calling training pool Public API-calling data from five datasets, read at the pinned revisions named below and laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 199186 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file query the user's request, as its source publishes it functions the declarations offered with the request, as a list… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/api-calling-training-pool.texttext-generation10K<n<100K0 likes175 downloads28d agoHugging Face14minpeter /apigen-mt-5k-parsed [PARSED] APIGen-MT-5k The data in this dataset is a full of the original Salesforce/APIGen-MT-5k Subset name multi-turn parallel multiple definition Last turn type number of dataset apigen-mt-5k yes no yes complex 5k This is a re-parsing formatting dataset for the APIGen-MT-5k official dataset. Load the dataset from datasets import load_dataset ds = load_dataset("minpeter/apigen-mt-5k-parsed") print(ds) # DatasetDict({ # train: Dataset({ #… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/apigen-mt-5k-parsed.textquestion-answering1K<n<10K0 likes135 downloads1y agoHugging Face15argilla-warehouse /apigen-synth-trl Dataset card This dataset is a version of argilla/Synth-APIGen-v0.1 prepared for fine-tuning using trl. To generate it, the following script was run: from datasets import load_dataset from jinja2 import Template SYSTEM_PROMPT = """ You are an expert in composing functions. You are given a question and a set of possible functions. Based on the question, you will need to make one or more function/tool calls to achieve the purpose. If none of the functions can be used, point it out… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/apigen-synth-trl.texttext-generation10K<n<100K11 likes124 downloads2y agoHugging Face16bernabepuente /backend-api-instruction-dataset Backend & API Development Dataset Instruction dataset focused on RESTful API design, WebSocket real-time communication, microservices patterns, and API best practices. Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Backend Api topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/backend-api-instruction-dataset.texttext-generationn<1K0 likes101 downloads5mo agoHugging Face17kusho-ai /api-eval-20 APIEval-20: A Benchmark for Black-Box API Test Suite Generation Motivation Testing APIs thoroughly is one of the most critical, yet consistently underserved, activities in software engineering. Despite a rich ecosystem of API testing tools — Postman, RestAssured, Schemathesis, Dredd, and others — we found ourselves asking a deceptively simple question: Given only the schema and an example payload of an API request — no source code, no documentation, no prior knowledge —… See the full description on the dataset page: https://huggingface.co/datasets/kusho-ai/api-eval-20.text-generation6 likes91 downloads5mo agoHugging Face18beatsprom /agentic-tool-use-multi-api-orchestration-2026 ⚡ Agentic Tool-Use, Multi-API Calling & Autonomous Function Orchestration (2026) Official 100-sample production preview of the Agentic Tool-Use & Multi-API Orchestration Suite (2026) by BeatsProm AI Research Lab. Engineered for parallel tool calling (<tool_call>), strict JSON-schema enforcement, stateful cursor pagination, and self-healing API error recovery. 🏛️ THE 20 AGENTIC OPERATIONAL CORES: Parallel Portfolio Rebalancing: Multi-leg execution with… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/agentic-tool-use-multi-api-orchestration-2026.texttext-generationn<1K0 likes89 downloads1mo agoHugging Face19covaga /eplan-2027-api-qa EPLAN 2027 Platform API Q&A for C# Automation EPLAN 2027 API question-answer dataset for C#/.NET automation, EPLAN scripting, electrical engineering assistants, code generation, RAG, and LLM instruction tuning. This dataset contains grounded English question-answer pairs generated from excerpts of the EPLAN Platform API 2027 documentation. It is designed for building assistants that help engineers understand and write EPLAN API scripts, actions, add-ins, and automation code.… See the full description on the dataset page: https://huggingface.co/datasets/covaga/eplan-2027-api-qa.textquestion-answering10K<n<100K0 likes59 downloads18d agoHugging Face20danemarparceros /drupal7-dataset-api Drupal 7 Dataset API Drupal 7 reached end of life in early 2025, but plenty of sites still run on it, and the people maintaining them still need good answers. This dataset is meant to help with that: 7,952 examples about the Drupal 7 API, taken only from api.drupal.org. You can use it to fine-tune a model, to test how well a model knows Drupal 7, or as reference material for a retrieval (RAG) setup. Everything is in English. It was put together by Daniel Ricardo Ramirez Marin as… See the full description on the dataset page: https://huggingface.co/datasets/danemarparceros/drupal7-dataset-api.texttext-generation1K<n<10K0 likes58 downloads9d agoHugging Face21Hypersniper /unity_api_2022_3 Unity3d 2022.3 LTS API & Manual In this dataset, you'll find a series of Q&A for the Unity3d API and Manual. Dataset Creation Download the unity offline documentation. Process documentation, extract title, and description. Clean documentation. Process each title and description item in llama3-8B-Instruct in order to generate several questions that capture the meaning of the API. Re-process in llama3-8B-Instruct with question and API to generate the answer.… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/unity_api_2022_3.texttext-generation100K<n<1M12 likes56 downloads2y agoHugging Face22natzx94 /exercise-api Exercise API — Dataset Dataset de 104 ejercicios de gimnasio (bilingüe ES/EN) derivado de la Exercise API. Cada ejercicio incluye grupo muscular, equipamiento, músculos principal/secundario, instrucciones paso a paso e ilustración masculina y femenina (208 imágenes en total). Configuraciones images — 1 fila por imagen (208). Etiquetas (grupo, equipamiento, músculos, género) + caption_es/caption_en. Para clasificación de imagen y multimodal (image-to-text / VQA).… See the full description on the dataset page: https://huggingface.co/datasets/natzx94/exercise-api.imageimage-classification1K<n<10K0 likes51 downloads3mo agoHugging Face23WalterWangtao /APIGen-MT-5k Summary APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay Code: https://github.com/apigen-mt/apigen-mt.github.io The repo contains 5000 multi-turn trajectories collected by APIGen-MT This dataset is a subset of the data used to train the xLAM-2… See the full description on the dataset page: https://huggingface.co/datasets/WalterWangtao/APIGen-MT-5k.textquestion-answering1K<n<10K0 likes43 downloads1mo agoHugging Face24arasyi /quantum-api-drift Quantum API Drift Quantum API Drift is an evaluation benchmark for measuring whether LLM-generated quantum code targets the requested Qiskit SDK version. It accompanies the paper Benchmarking API Drift in LLM-Generated Quantum Code Across Successive SDK Versions. The benchmark evaluates version fidelity, cross-version compatibility, failure modes, and documentation-guided repair across Qiskit 0.43, 1.3, and 2.0. Dataset Configurations benchmark The… See the full description on the dataset page: https://huggingface.co/datasets/arasyi/quantum-api-drift.texttext-generationn<1K0 likes40 downloads3mo agoHugging Face25eth-sri /API-Upgrade API-Upgrade API-Upgrade evaluates repository-level Rust synthesis against newer dependency APIs. This dataset was used for evaluation in the paper Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code. You can find the corresponding evaluation code in the project GitHub repository. Example Usage from datasets import load_dataset import json dataset = load_dataset("eth-sri/API-Upgrade") for instance in dataset["test"]:… See the full description on the dataset page: https://huggingface.co/datasets/eth-sri/API-Upgrade.texttext-generationn<1K0 likes38 downloads3mo agoHugging Face26harsharajkumar273 /api-vulnerability-dataset-10k API Vulnerability Dataset (10K) A dataset of 10,000 API-specific vulnerability samples used to fine-tune harsharajkumar273/api-security-qlora — a QLoRA adapter on CodeLlama-7b for automated API security analysis. Dataset Summary Each sample contains a vulnerable or clean API endpoint code snippet paired with a structured security analysis covering vulnerability type, severity, CWE ID, and a remediated version. Language & Framework Distribution Language… See the full description on the dataset page: https://huggingface.co/datasets/harsharajkumar273/api-vulnerability-dataset-10k.texttext-classification10K<n<100K0 likes29 downloads6mo agoHugging Face27gdgc-metacong /apigen-inferred apigen-inferred A verified, GPT-5.5-distilled subset of the argilla/apigen-function-calling dataset (109k rows in the upstream), with every golden tool-call argument labelled as literal or dependency-derived to enable a clean function-calling benchmark. Pipeline Filter the upstream to rows where every called API actually works (replay each tool call against the real implementation — distilabel Python functions or live RapidAPI / cached responses) → 45,984 rows. Distill… See the full description on the dataset page: https://huggingface.co/datasets/gdgc-metacong/apigen-inferred.tabulartext-generation10K<n<100K1 likes26 downloads5mo agoHugging Face28anote-ai /Research-Enterprise-Synth-API EnterpriseSynth Public API Specs and Generated Artifacts EnterpriseSynth converts OpenAPI/Swagger specifications into synthetic tool-use training and evaluation artifacts without executing live API calls. This dataset repository contains the public, redistributable dataset artifacts from anote-ai/Research-Enterprise-Synth-API: data/specs/: public OpenAPI/Swagger specs used by the experiments. data/specs/phase3/: additional public held-out API specs. data/generated/: generated… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/Research-Enterprise-Synth-API.tabulartext-generationn<1K0 likes24 downloads2mo agoHugging Face29apinsun /triage-french Triage médical français (FRENCH — SFMU) Dataset synthétique pour entraîner un agent conversationnel de triage médical aux urgences, fondé sur la grille FRENCH (SFMU, 5 niveaux d'urgence). ⚠️ Ce dataset est synthétique : il a été généré par un LLM à partir de règles déterministes, sans validation clinique. Il est destiné à la recherche uniquement, et non à un usage médical en production. Construction La grille FRENCH (196 règles, 16 catégories) sert de source de… See the full description on the dataset page: https://huggingface.co/datasets/apinsun/triage-french.text-generation1K<n<10K0 likes24 downloads2d agoHugging Face30Myashka /SO-Python_QA-API_Usage-tanh_score Stack Overflow Python Q&A Dataset Description Filtered Python Q&A with API_Usage subcategory without: Images Links Blocks of code Scores in Q1-Q3 scaled with MaxAbsScaler. Tanh function applyed to joint Scores. tabulartext-generation1K<n<10K0 likes19 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.