datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apigen-function-calling
Dataset card for argilla/apigen-function-calling
This dataset is a merge of argilla/Synth-APIGen-v0.1
and Salesforce/xlam-function-calling-60k, making
over 100K function calling examples following the APIGen recipe.
Prepare for training
This version is not ready to do fine tuning, but you can run a script like prepare_for_sft.py
to prepare it, and run the same recipe that can be found in
argilla/Llama-3.2-1B-Instruct-APIGen-FC-v0.1#training-procedure.
Modify the prompt… See the full description on the dataset page: https://huggingface.co/datasets/argilla/apigen-function-calling.jev-ai-api-guide-assetsjlw-apiZeus-API-forecasts
Bittensor Subnet Zeus Archive Dataset
Interactive Tutorial: Want to dive right in? We have provided a fully standalone Jupyter Notebook tutorials. Go to the Files and versions tab, click on trustless_verification_tutorial.ipynb, and click "Open in Colab" to learn how to download and verify this data.
Dataset Summary
This dataset contains historical weather forecasts generated by the Zeus API. It enables clients to perform trustless verification of the API's… See the full description on the dataset page: https://huggingface.co/datasets/orpheus-zeus/Zeus-API-forecasts.API-BankAPIGen-MT-5k
Summary
APIGen-MT is an automated agentic data generation pipeline designed to synthesize verifiable, high-quality, realistic datasets for agentic applications
This dataset was released as part of APIGen-MT: Agentic PIpeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Code: https://github.com/apigen-mt/apigen-mt.github.io
The repo contains 5000 multi-turn trajectories collected by APIGen-MT
This dataset is a subset of the data used to train the xLAM-2 model… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/APIGen-MT-5k.Synth-APIGen-v0.1
Dataset card for Synth-APIGen-v0.1
This dataset has been created with distilabel.
Pipeline script: pipeline_apigen_train.py.
Dataset creation
It has been created with distilabel==1.4.0 version.
This dataset is an implementation of APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets in distilabel,
generated from synthetic functions. The process can be summarized as follows:
Generate (or in this case modify) python… See the full description on the dataset page: https://huggingface.co/datasets/argilla/Synth-APIGen-v0.1.synth-apigen-qwen-exec
Dataset Card for synth-apigen-qwen-exec
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
apigen_exec.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/plaguss/synth-apigen-qwen-exec/raw/main/apigen_exec.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/synth-apigen-qwen-exec.kaggle-api-test
kaggle-api-test — dev artifact mirror + CLIFFX visual gallery
Working log for the anchor-decay reconstruction campaign (Kaggle notebook
v9k7 series -> flush here; pulled + hash-audited + mirrored each round).
Status 2026-09-25 (post-CLIFFX): pre-registered stop triggered. Ship
artifact = champion weights 33c735603c9f (weights/model_conv_g112.pt)
deployed stack policy (ROUTER_T 0.4622 / k=2 / TAU 0.70-0.80).
All search arms closed under the re-anchored deployed gate:
0/284… See the full description on the dataset page: https://huggingface.co/datasets/favvnna/kaggle-api-test.jev-api-article-assetsopenai-moderation-api-evaluation
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.Hitek_ImCR_APIAPIGen-MT-5k-with-cot-v1-deepseek_deepseekai-api-pricing
AI API Pricing Dataset
Source-linked AI API pricing data covering token, cache, batch, tiered, multimodal, and non-token pricing across multiple providers, including OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, and Cohere. These are examples, not an exhaustive provider list.
Live dataset and documentation
Source repository
Fixed v1.1.0 release (snapshot 2026-09-30)
Version DOI
Methodology
This Hugging Face dataset is the machine-readable distribution of the public AI API… See the full description on the dataset page: https://huggingface.co/datasets/aicostbudget-ai/ai-api-pricing.deepseek-v2-codder-minecraft-apisynth-apigen-llama-exec
Dataset Card for synth-apigen-llama-exec
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
apigen_exec.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/plaguss/synth-apigen-llama-exec/raw/main/apigen_exec.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/synth-apigen-llama-exec.django-rest-api-2024APIBench
Gorilla: Large Language Model Connected with Massive APIs
By Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez (Project Website)
Gorilla enables LLMs to use tools by invoking APIs. Given a natural language query, Gorilla can write a semantically- and syntactically- correct API to invoke. With Gorilla, we are the first to demonstrate how to use LLMs to invoke 1,600+ (and growing) API calls accurately while reducing hallucination. We also release APIBench, the… See the full description on the dataset page: https://huggingface.co/datasets/gorilla-llm/APIBench.llm-api-prices
LLM API Price Index — daily snapshot
Live, verified LLM API prices from apipriceindex.com.
763 tracked endpoints, index as of 2026-10-06T08:15:04Z. Updated daily by an
automated sync (source).
Every price is read from the provider's official pricing page, cross-checked,
and carries its verification date and source URL — no hand-waved numbers.
Files
File
Contents
models.csv
One row per tracked endpoint: provider, model, USD per 1M… See the full description on the dataset page: https://huggingface.co/datasets/apipriceindex/llm-api-prices.synth-apigen-qwen
Dataset Card for argilla-warehouse/synth-apigen-qwen
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
synth_apigen.py.
Dataset creation
This dataset is a replica in distilabel of the framework
defined in: APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets.
Using the seed dataset of synthetic python functions in argilla-warehouse/python-seed-tools,
the… See the full description on the dataset page: https://huggingface.co/datasets/argilla-warehouse/synth-apigen-qwen.phonkquantum-video
Dataset Card for Dataset Name
quantum suite video
quantum suite video dataset, used to train the quantum suite video model.
Dataset Details
will upload, collecting data
Dataset Description
This dataset is aimed to be curated for the quantum suite video dataset with video from wikimedia (Cc allowing commercial use), this dataset allows commercial use, and will become useful for you to use in your video models,
the size is aimed to be 2.5tb, after we… See the full description on the dataset page: https://huggingface.co/datasets/ai-api-key-free-finder/quantum-video.FINDER_API_KEY_AI_SEARCH_2023
FINDER_API_KEY_AI_SEARCH_2023
tags: data collection, machine learning, API performance
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.api_gateway_jwt_oauth_revocation_engine_teaser
🚀 Cloud Architecture - API Gateway JWT & OAuth Token Revocation Engine (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Purchase Full Production Master Dataset on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 Domain Focus & Capabilities
Distributed Redis token blacklisting, JTI invalidation race condition mitigation, and… See the full description on the dataset page: https://huggingface.co/datasets/emgena/api_gateway_jwt_oauth_revocation_engine_teaser.NexusRaven_API_evaluation
NexusRaven API Evaluation dataset
Please see blog post or NexusRaven Github repo for more information.
License
The evaluation data in this repository consists primarily of our own curated evaluation data that only uses open source commercializable models. However, we include general domain data from the ToolLLM and ToolAlpaca papers. Since the data in the ToolLLM and ToolAlpaca works use OpenAI's GPT models for the generated content, the data is not commercially… See the full description on the dataset page: https://huggingface.co/datasets/Nexusflow/NexusRaven_API_evaluation.taco-api-fixtures
TACO API Fixtures
A deterministic collection of 50 small TACO datasets whose payloads are all
Rumi files. The fixtures exercise TACO contracts, metadata hierarchies,
FOLDER and ZIP containers, ZIP partitioning, TACOCAT consolidation, generated
locations, and local or remote range reads.
They are API fixtures, not training data or a scientific benchmark. Spatial
and temporal metadata generated here is intentionally synthetic.
Matrix
The repository combines ten… See the full description on the dataset page: https://huggingface.co/datasets/asterisk-labs/taco-api-fixtures.api-contract-migration-trajectories
Api Contract Migration Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/api-contract-migration-trajectories.starcoderdata-apishermes_salesforce_apigen_tool_usestarcoder-apis-0
