datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apigen-inferred
apigen-inferred
A verified, GPT-5.5-distilled subset of the
argilla/apigen-function-calling
dataset (109k rows in the upstream), with every golden tool-call argument
labelled as literal or dependency-derived to enable a clean
function-calling benchmark.
Pipeline
Filter the upstream to rows where every called API actually works
(replay each tool call against the real implementation — distilabel
Python functions or live RapidAPI / cached responses) → 45,984 rows.
Distill… See the full description on the dataset page: https://huggingface.co/datasets/gdgc-metacong/apigen-inferred.Research-Enterprise-Synth-API
EnterpriseSynth Public API Specs and Generated Artifacts
EnterpriseSynth converts OpenAPI/Swagger specifications into synthetic tool-use
training and evaluation artifacts without executing live API calls.
This dataset repository contains the public, redistributable dataset artifacts
from anote-ai/Research-Enterprise-Synth-API:
data/specs/: public OpenAPI/Swagger specs used by the experiments.
data/specs/phase3/: additional public held-out API specs.
data/generated/: generated… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/Research-Enterprise-Synth-API.SO-Python_QA-API_Usage-tanh_score
Stack Overflow Python Q&A Dataset
Description
Filtered Python Q&A with API_Usage subcategory without:
Images
Links
Blocks of code
Scores in Q1-Q3 scaled with MaxAbsScaler. Tanh function applyed to joint Scores.
self-reflective-apis
Self-Reflective APIs Benchmark
Dataset accompanying the paper "Self-Reflective APIs: Enhancing AI Agent Efficiency Through Structured Semantic Feedback" (Canedo, Grama — Siemens DI SW).
Overview
This dataset contains the benchmark tasks, experiment results, and tidy analysis table used to produce every table and figure in the paper. It covers two experimental domains (recipe conversion and billing/refund policy) and three LLM conditions across adversarial… See the full description on the dataset page: https://huggingface.co/datasets/arquicanedo/self-reflective-apis.
