Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Salesforce /APIGen-MT-5k Summary APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay Code: https://github.com/apigen-mt/apigen-mt.github.io The repo contains 5000 multi-turn trajectories collected by APIGen-MT This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.textquestion-answering1K<n<10K115 likes4.9k downloads1y agoHugging Face02mmathys /openai-moderation-api-evaluation Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection" The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper. Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label. Category Label Definition sexual S Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.tabulartext-classification1K<n<10K39 likes2.8k downloads3y agoHugging Face03rmems /api-contract-migration-trajectories Api Contract Migration Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/api-contract-migration-trajectories.text1K<n<10K0 likes467 downloads18d agoHugging Face04ibm-research /SQL-API-Bench Dataset Card for Dataset Name This dataset contains QA that requires DB and API access at the same time. It is composed of two new benchmarks consisting of questions whose answers require a combination of database and API calls, both of which are augmentations of the popular Spider dataset and benchmark. Benchmark I replaces a fraction of the real Spider database tables with equivalents that are executed via APIs. This allows us to directly test the mechanism by which database and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SQL-API-Bench.textquestion-answering1K<n<10K5 likes342 downloads1y agoHugging Face05wicai24 /api_audit_dataThis repository contains code for auditing Large Language Models (LLMs) to verify service integrity, as described in the paper Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs. Github repository: https://github.com/willsdca/llm_api_audit text100K<n<1M1 likes252 downloads1y agoHugging Face06AYI-NEDJIMI /oauth-api-security-en OAuth & API Security Dataset (EN) Comprehensive English dataset covering OAuth 2.0 vulnerabilities, API attacks (OWASP API Top 10 2023), security controls, and Q&A pairs for training cybersecurity-specialized language models. Dataset Contents Category Entries Description OAuth 2.0 Vulnerabilities 20 Authorization Code Interception, CSRF, PKCE bypass, JWT attacks, token leakage API Attacks 25 BOLA, BFLA, BOPLA, SSRF, GraphQL DoS, gRPC injection, CORS… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/oauth-api-security-en.textquestion-answeringn<1K0 likes246 downloads8mo agoHugging Face07AYI-NEDJIMI /oauth-api-security-fr Dataset OAuth & Securite API (FR) Dataset francophone complet sur les vulnerabilites OAuth 2.0, les attaques API (OWASP API Top 10 2023), les controles de securite, et les questions-reponses pour l'entrainement de modeles de langage specialises en cybersecurite. Contenu du Dataset Categorie Nombre d'entrees Description Vulnerabilites OAuth 2.0 20 Authorization Code Interception, CSRF, PKCE bypass, JWT attacks, token leakage Attaques API 25 BOLA, BFLA, BOPLA… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/oauth-api-security-fr.textquestion-answeringn<1K0 likes214 downloads8mo agoHugging Face08Emulated-Inc /api-calling-training-pool API calling training pool Public API-calling data from five datasets, read at the pinned revisions named below and laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 199186 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file query the user's request, as its source publishes it functions the declarations offered with the request, as a list… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/api-calling-training-pool.texttext-generation10K<n<100K0 likes175 downloads29d agoHugging Face09alvaroluceroo /llm-api-pricing LLM API pricing dataset Prices of the current large language model APIs, as published by their makers, with the day each price was last verified and the maker's page it was read from, plus a record of every change to those prices. This is the data behind the LLM API pricing table of AI Signal, published here as files so it can be versioned, diffed and cited. The same files are served at aisignalhq.com/data/ and versioned on GitHub at alvaroluceroo/llm-api-pricing-dataset. This… See the full description on the dataset page: https://huggingface.co/datasets/alvaroluceroo/llm-api-pricing.tabularothern<1K1 likes142 downloads4d agoHugging Face10bernabepuente /backend-api-instruction-dataset Backend & API Development Dataset Instruction dataset focused on RESTful API design, WebSocket real-time communication, microservices patterns, and API best practices. Dataset Details Dataset Description This is a high-quality instruction-tuning dataset focused on Backend Api topics. Each entry includes: A clear instruction/question Optional input context A detailed response/solution Chain-of-thought reasoning process Curated by: CloudKernel.IO… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/backend-api-instruction-dataset.texttext-generationn<1K0 likes101 downloads5mo agoHugging Face11MapEval /MapEval-API MapEval-API MapEval-API is created using MapQaTor. Usage from datasets import load_dataset # Load dataset ds = load_dataset("MapEval/MapEval-API", name="benchmark") # Generate better prompts for item in ds["test"]: # Start with a clear task description prompt = ( "You are a highly intelligent assistant. " "Answer the multiple-choice question by selecting the correct option.\n\n" "Question:\n" + item["question"] + "\n\n"… See the full description on the dataset page: https://huggingface.co/datasets/MapEval/MapEval-API.tabularquestion-answeringn<1K2 likes91 downloads2y agoHugging Face12nixiesearch /embed-api-latencyTODO tabular100K<n<1M2 likes86 downloads2y agoHugging Face13dusersad12 /api-assistant-tracetextn<1K0 likes61 downloads22d agoHugging Face14danemarparceros /drupal7-dataset-api Drupal 7 Dataset API Drupal 7 reached end of life in early 2025, but plenty of sites still run on it, and the people maintaining them still need good answers. This dataset is meant to help with that: 7,952 examples about the Drupal 7 API, taken only from api.drupal.org. You can use it to fine-tune a model, to test how well a model knows Drupal 7, or as reference material for a retrieval (RAG) setup. Everything is in English. It was put together by Daniel Ricardo Ramirez Marin as… See the full description on the dataset page: https://huggingface.co/datasets/danemarparceros/drupal7-dataset-api.texttext-generation1K<n<10K0 likes58 downloads9d agoHugging Face15NerdOptimize /nerd-knowledge-api NerdOptimize Dataset (v1.0.0) English dataset for SEO (Data‑Driven) and AI Search / AEO by NerdOptimize (Bangkok, TH).Built for GitHub, Hugging Face, and on‑site deployment, so LLMs can learn/cite the brand. Structure data/*.json → core machine‑readable data (ICPs, services, case studies, frameworks, articles, labels, metadata, processing steps) server.js / openapi.json → tiny Express API to serve the dataset schema-dataset.jsonld → Dataset JSON‑LD for Google Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NerdOptimize/nerd-knowledge-api.textzero-shot-classificationn<1K0 likes57 downloads11mo agoHugging Face16Hypersniper /unity_api_2022_3 Unity3d 2022.3 LTS API & Manual In this dataset, you'll find a series of Q&A for the Unity3d API and Manual. Dataset Creation Download the unity offline documentation. Process documentation, extract title, and description. Clean documentation. Process each title and description item in llama3-8B-Instruct in order to generate several questions that capture the meaning of the API. Re-process in llama3-8B-Instruct with question and API to generate the answer.… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/unity_api_2022_3.texttext-generation100K<n<1M12 likes56 downloads2y agoHugging Face17vGassen /Dutch-Tweede-Kamer-APItext10K<n<100K0 likes55 downloads1y agoHugging Face18nativeport /web-access-api-benchmarks NativePort Web-Access API Benchmarks Measured quality, latency, cost and error-rate figures for 22 commercial web-access APIs — search, SERP, scraping, crawling, extraction, sourced answers, screenshots, document parsing, browser actions and change watching — scored per capability on a fixed task corpus. This is the 2026-08-05 run: 67 provider × capability scorecards across 13 capabilities, flattened into 297 metric rows. It exists for one practical decision: when an AI agent… See the full description on the dataset page: https://huggingface.co/datasets/nativeport/web-access-api-benchmarks.tabularn<1K0 likes52 downloads2mo agoHugging Face19natzx94 /exercise-api Exercise API — Dataset Dataset de 104 ejercicios de gimnasio (bilingüe ES/EN) derivado de la Exercise API. Cada ejercicio incluye grupo muscular, equipamiento, músculos principal/secundario, instrucciones paso a paso e ilustración masculina y femenina (208 imágenes en total). Configuraciones images — 1 fila por imagen (208). Etiquetas (grupo, equipamiento, músculos, género) + caption_es/caption_en. Para clasificación de imagen y multimodal (image-to-text / VQA).… See the full description on the dataset page: https://huggingface.co/datasets/natzx94/exercise-api.imageimage-classification1K<n<10K0 likes51 downloads3mo agoHugging Face20data-sci-project /ev-count-google-apitabular1K<n<10K0 likes51 downloads18d agoHugging Face21Hoglet-33 /APIGen-50k50,000 samples from argilla/apigen-function-calling text10K<n<100K0 likes46 downloads6mo agoHugging Face22WalterWangtao /APIGen-MT-5k Summary APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay Code: https://github.com/apigen-mt/apigen-mt.github.io The repo contains 5000 multi-turn trajectories collected by APIGen-MT This dataset is a subset of the data used to train the xLAM-2… See the full description on the dataset page: https://huggingface.co/datasets/WalterWangtao/APIGen-MT-5k.textquestion-answering1K<n<10K0 likes43 downloads1mo agoHugging Face23fullstack /cik_sec_api_jsonltext1K<n<10K0 likes42 downloads2y agoHugging Face24arasyi /quantum-api-drift Quantum API Drift Quantum API Drift is an evaluation benchmark for measuring whether LLM-generated quantum code targets the requested Qiskit SDK version. It accompanies the paper Benchmarking API Drift in LLM-Generated Quantum Code Across Successive SDK Versions. The benchmark evaluates version fidelity, cross-version compatibility, failure modes, and documentation-guided repair across Qiskit 0.43, 1.3, and 2.0. Dataset Configurations benchmark The… See the full description on the dataset page: https://huggingface.co/datasets/arasyi/quantum-api-drift.texttext-generationn<1K0 likes40 downloads3mo agoHugging Face25elenagroundwork /llm-api-pricing-latency-2026 LLM Inference Unit Economics & Architecture Engine Empirical benchmark dataset by Groundwork Research (https://gworky.com). Full interactive decision engine available at: https://gworky.com/tools/llm-token-cost-calculator. Description Full-stack inference cost and latency estimator comparing frontier proprietary models (Claude 3.7, GPT-4.5) against open-weight hosted providers (Groq, DeepSeek R1, Together AI). Primary source authority: https://gworky.com/tech tabulartabular-classificationn<1K0 likes36 downloads1mo agoHugging Face26LisandraMoura /API_gen_mt5k_fittext1K<n<10K0 likes29 downloads1y agoHugging Face27apirrone /spaceship-game-leaderboard Spaceship Game - Leaderboard This dataset contains leaderboard entries for the Spaceship Game on Reachy Mini. Stats Entries: 1 Top Score: 350 by Antoijne Last Updated: 2026-03-12 Published by: apirrone Format The leaderboard.json file contains an array of entries: Field Type Description score int Final game score name string Player name date string ISO 8601 timestamp waves_completed int? Number of waves completed Top 10… See the full description on the dataset page: https://huggingface.co/datasets/apirrone/spaceship-game-leaderboard.tabularothern<1K0 likes29 downloads7mo agoHugging Face28harsharajkumar273 /api-vulnerability-dataset-10k API Vulnerability Dataset (10K) A dataset of 10,000 API-specific vulnerability samples used to fine-tune harsharajkumar273/api-security-qlora — a QLoRA adapter on CodeLlama-7b for automated API security analysis. Dataset Summary Each sample contains a vulnerable or clean API endpoint code snippet paired with a structured security analysis covering vulnerability type, severity, CWE ID, and a remediated version. Language & Framework Distribution Language… See the full description on the dataset page: https://huggingface.co/datasets/harsharajkumar273/api-vulnerability-dataset-10k.texttext-classification10K<n<100K0 likes29 downloads6mo agoHugging Face29yinita /ps4mas-api-rollouts-0922text10K<n<100K0 likes29 downloads18d agoHugging Face30SYSUSELab /APIKG4Syn-HarmonyOS-Dataset Framework-Aware Code Generation with API Knowledge Graph–Constructed Data: A Study on HarmonyOS 🗂️The Dataset OHBen.json: For the integrated version of the two aforementioned files, which constitutes the final dataset used for fine-tuning the LLM. text1K<n<10K3 likes27 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.