datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openai-moderation-api-evaluation
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.llm-api-pricing
LLM API pricing dataset
Prices of the current large language model APIs, as published by their makers, with the day each price was last verified and the maker's page it was read from, plus a record of every change to those prices. This is the data behind the LLM API pricing table of AI Signal, published here as files so it can be versioned, diffed and cited.
The same files are served at aisignalhq.com/data/ and versioned on GitHub at alvaroluceroo/llm-api-pricing-dataset. This… See the full description on the dataset page: https://huggingface.co/datasets/alvaroluceroo/llm-api-pricing.MapEval-API
MapEval-API
MapEval-API is created using MapQaTor.
Usage
from datasets import load_dataset
# Load dataset
ds = load_dataset("MapEval/MapEval-API", name="benchmark")
# Generate better prompts
for item in ds["test"]:
# Start with a clear task description
prompt = (
"You are a highly intelligent assistant. "
"Answer the multiple-choice question by selecting the correct option.\n\n"
"Question:\n" + item["question"] + "\n\n"… See the full description on the dataset page: https://huggingface.co/datasets/MapEval/MapEval-API.embed-api-latencyTODO
web-access-api-benchmarks
NativePort Web-Access API Benchmarks
Measured quality, latency, cost and error-rate figures for 22 commercial web-access
APIs — search, SERP, scraping, crawling, extraction, sourced answers, screenshots,
document parsing, browser actions and change watching — scored per capability on a
fixed task corpus. This is the 2026-08-05 run: 67 provider × capability
scorecards across 13 capabilities, flattened into 297 metric rows.
It exists for one practical decision: when an AI agent… See the full description on the dataset page: https://huggingface.co/datasets/nativeport/web-access-api-benchmarks.ev-count-google-apillm-api-pricing-latency-2026
LLM Inference Unit Economics & Architecture Engine
Empirical benchmark dataset by Groundwork Research (https://gworky.com).
Full interactive decision engine available at: https://gworky.com/tools/llm-token-cost-calculator.
Description
Full-stack inference cost and latency estimator comparing frontier proprietary models (Claude 3.7, GPT-4.5) against open-weight hosted providers (Groq, DeepSeek R1, Together AI).
Primary source authority: https://gworky.com/tech
spaceship-game-leaderboard
Spaceship Game - Leaderboard
This dataset contains leaderboard entries for the Spaceship Game on Reachy Mini.
Stats
Entries: 1
Top Score: 350 by Antoijne
Last Updated: 2026-03-12
Published by: apirrone
Format
The leaderboard.json file contains an array of entries:
Field
Type
Description
score
int
Final game score
name
string
Player name
date
string
ISO 8601 timestamp
waves_completed
int?
Number of waves completed
Top 10… See the full description on the dataset page: https://huggingface.co/datasets/apirrone/spaceship-game-leaderboard.ai-api-catalog-2026
ai-api-catalog-2026
AI API catalog. live probed
Records: 15 | Real data | Updated: 2026-10-06
API: GET https://api.legion-api.com/api-catalog
Bundle: gemmo.gumroad.com/l/mdevxu
Gated — auto-approved.
Research-Enterprise-Synth-API
EnterpriseSynth Public API Specs and Generated Artifacts
EnterpriseSynth converts OpenAPI/Swagger specifications into synthetic tool-use
training and evaluation artifacts without executing live API calls.
This dataset repository contains the public, redistributable dataset artifacts
from anote-ai/Research-Enterprise-Synth-API:
data/specs/: public OpenAPI/Swagger specs used by the experiments.
data/specs/phase3/: additional public held-out API specs.
data/generated/: generated… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/Research-Enterprise-Synth-API.SO-Python_QA-Networking_and_APIs_classSO-Python_QA-API_USAGE_classapi_graph_reflectionvarroa_mmdet_runs_fcos_dgfe_apiguardrails-api-test-resultsprism-alignment-apicleanedapi_tree_of_thought
