datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openai-moderation-api-evaluation
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.ai-api-pricing
AI API Pricing Dataset
Source-linked AI API pricing data covering token, cache, batch, tiered, multimodal, and non-token pricing across multiple providers, including OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, and Cohere. These are examples, not an exhaustive provider list.
Live dataset and documentation
Source repository
Fixed v1.1.0 release (snapshot 2026-09-30)
Version DOI
Methodology
This Hugging Face dataset is the machine-readable distribution of the public AI API… See the full description on the dataset page: https://huggingface.co/datasets/aicostbudget-ai/ai-api-pricing.FINDER_API_KEY_AI_SEARCH_2023
FINDER_API_KEY_AI_SEARCH_2023
tags: data collection, machine learning, API performance
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.incidentes-ia-espanol
Incidentes de IA en espanol
Corpus editorial de incidentes de IA traducidos y clasificados con taxonomia cerrada por LaAutopsIA (ApisDom Intelligence Group). Fuente principal: AI Incident Database (CC-BY-SA-4.0).
Cifras del snapshot actual
11 incidentes publicados en este snapshot.
Mes archivado: 2026-09.
Ultima edicion: 2026-10-01T07:58:18.068Z.
Frecuencia: sincronizacion mensual.
Para que sirve este dataset
Corpus editorial mensual con… See the full description on the dataset page: https://huggingface.co/datasets/apisdom/incidentes-ia-espanol.indice-fallos-ia-espanol
Indice de Fallos IA en espanol
Snapshots mensuales del Indice de Fallos IA producido por el observatorio La AutopsIA (ApisDom Intelligence Group). Mide la fiabilidad de modelos LLM con benchmarks oficiales independientes, en formato citable y trazable.
Cifras del snapshot actual
780 mediciones en este snapshot.
Mes archivado: 2026-10.
Recomputado: 2026-09-29T19:17:29.575Z.
Frecuencia: sincronizacion mensual.
Para que sirve este dataset
Datos… See the full description on the dataset page: https://huggingface.co/datasets/apisdom/indice-fallos-ia-espanol.vulnerabilidades-ia-espanol
Vulnerabilidades CVE en sistemas de IA (espanol)
Corpus de advisories CVE/GHSA que afectan a paquetes y SDKs de IA (langchain, openai, anthropic, llamaindex, etc.) traducido al espanol por LaAutopsIA (ApisDom Intelligence Group). Fuente principal: GitHub Advisory Database (CC-BY-4.0).
Cifras del snapshot actual
18 vulnerabilidades publicadas en este snapshot.
Mes archivado: 2026-09.
Ultima edicion: 2026-10-01T07:58:25.270Z.
Frecuencia: sincronizacion mensual.… See the full description on the dataset page: https://huggingface.co/datasets/apisdom/vulnerabilidades-ia-espanol.mozart-api-demo-pages
Dataset Card for Dataset Name
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DoctorSlimm/mozart-api-demo-pages.llm-api-pricing
LLM API pricing dataset
Prices of the current large language model APIs, as published by their makers, with the day each price was last verified and the maker's page it was read from, plus a record of every change to those prices. This is the data behind the LLM API pricing table of AI Signal, published here as files so it can be versioned, diffed and cited.
The same files are served at aisignalhq.com/data/ and versioned on GitHub at alvaroluceroo/llm-api-pricing-dataset. This… See the full description on the dataset page: https://huggingface.co/datasets/alvaroluceroo/llm-api-pricing.ai-api-prices-brazil
AI API Prices in Brazil 🇧🇷
Source-linked observations of public LLM API prices from Brazil-focused gateways, other gateways, and official model providers. Prices retain their published currency (BRL or USD in this snapshot); no foreign-exchange conversion or cross-currency price ranking is applied.
Current snapshot: sources fetched on 2026-10-01 UTC, corresponding to 2026-10-02 in Singapore (UTC+8). The current prices.csv contains 1,625 price observations from 12 platforms.… See the full description on the dataset page: https://huggingface.co/datasets/fanpuch/ai-api-prices-brazil.MapEval-API
MapEval-API
MapEval-API is created using MapQaTor.
Usage
from datasets import load_dataset
# Load dataset
ds = load_dataset("MapEval/MapEval-API", name="benchmark")
# Generate better prompts
for item in ds["test"]:
# Start with a clear task description
prompt = (
"You are a highly intelligent assistant. "
"Answer the multiple-choice question by selecting the correct option.\n\n"
"Question:\n" + item["question"] + "\n\n"… See the full description on the dataset page: https://huggingface.co/datasets/MapEval/MapEval-API.embed-api-latencyTODO
ai-api-pricing-snapshot
AI API Model Pricing Snapshot — Qubax AI
Per-token public pricing for 211 ready models served by the Qubax AI API (OpenAI-compatible), exported from the public /v1/models endpoint.
Columns
Column
Description
model_id
API model identifier
model_name
Display name
owned_by
Publisher namespace
context_length
Max context window (tokens)
input_usd_per_1m_tokens
Input price, USD per 1M tokens
output_usd_per_1m_tokens
Output price, USD per 1M tokens… See the full description on the dataset page: https://huggingface.co/datasets/QubaxAI/ai-api-pricing-snapshot.agenttune-apigen-SFT-qwen3-0.6b-tracesevalap-comparing-albert-api-models-v11-12-2025-107
Comparing Albert-API models v11-12-2025 (ID: 107)
Comparing albert models on MFS-AIA datasets
Overview
This dataset contains 20 experiments
from the EvalAP evaluation platform.
Datasets: Assistant IA - QA, MFS_questions_v01
Models evaluated: albert-large, albert-small, openweight-large, openweight-medium, openweight-small
Metrics: generation_time, judge_notator, judge_precision, nb_tokens_completion, nb_tokens_prompt, output_length
Scores
Assistant IA… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/evalap-comparing-albert-api-models-v11-12-2025-107.APIGen-MT-46kprototype-jenga-pickup-longThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 612,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/apinguen/prototype-jenga-pickup-long.web-access-api-benchmarks
NativePort Web-Access API Benchmarks
Measured quality, latency, cost and error-rate figures for 22 commercial web-access
APIs — search, SERP, scraping, crawling, extraction, sourced answers, screenshots,
document parsing, browser actions and change watching — scored per capability on a
fixed task corpus. This is the 2026-08-05 run: 67 provider × capability
scorecards across 13 capabilities, flattened into 297 metric rows.
It exists for one practical decision: when an AI agent… See the full description on the dataset page: https://huggingface.co/datasets/nativeport/web-access-api-benchmarks.ev-count-google-apiv3-eval-rubric-v2-apiso100_api_11This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 2,
"total_frames": 490,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/so100_api_11.so100_api_20This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 298,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/so100_api_20.so100_api_10This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 2,
"total_frames": 576,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/so100_api_10.eval_so100_api_17This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 2,
"total_frames": 580,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pepijn223/eval_so100_api_17.llm-api-pricing-latency-2026
LLM Inference Unit Economics & Architecture Engine
Empirical benchmark dataset by Groundwork Research (https://gworky.com).
Full interactive decision engine available at: https://gworky.com/tools/llm-token-cost-calculator.
Description
Full-stack inference cost and latency estimator comparing frontier proprietary models (Claude 3.7, GPT-4.5) against open-weight hosted providers (Groq, DeepSeek R1, Together AI).
Primary source authority: https://gworky.com/tech
polarison-ai-api-pricing
AI API pricing and capability dataset
Official list prices, context limits and independent capability scores for 24 AI models from OpenAI, Anthropic, Google, xAI, DeepSeek, maintained at polarison.com.
Every price is copied by hand from the provider's official pricing page and carries the date it was last verified. An automated job re-checks those pages daily, so the dataset tracks real changes instead of estimates.
Rows: 24 models
Last verified: 2026-10-03
Live copies:… See the full description on the dataset page: https://huggingface.co/datasets/Goveia/polarison-ai-api-pricing.jenga_push3
jenga_push3
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
record-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 2083,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/apinguen/record-test.app-store-data-api-sample-data
Apple App Store API — Apps, Reviews, Ratings & ASO
Unofficial Apple App Store API in one Apify actor. 10 endpoints: app details, search, reviews, top charts, similar apps, developer profiles, autocomplete, rating histograms, privacy labels, version history. Pure HTTP, sub-3s cold start, batch & parallel. For iOS devs, ASO and AI tools.
What the actor scrapes
🍎 Apple App Store API — Scrape iOS Apps, Reviews, Ratings & ASO Data Unofficial Apple App Store API in a… See the full description on the dataset page: https://huggingface.co/datasets/logiover/app-store-data-api-sample-data.spaceship-game-leaderboard
Spaceship Game - Leaderboard
This dataset contains leaderboard entries for the Spaceship Game on Reachy Mini.
Stats
Entries: 1
Top Score: 350 by Antoijne
Last Updated: 2026-03-12
Published by: apirrone
Format
The leaderboard.json file contains an array of entries:
Field
Type
Description
score
int
Final game score
name
string
Player name
date
string
ISO 8601 timestamp
waves_completed
int?
Number of waves completed
Top 10… See the full description on the dataset page: https://huggingface.co/datasets/apirrone/spaceship-game-leaderboard.africa-api-security-dataset
API Security Posture (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-api-security-dataset.
