datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Launchergov_reportGovReport long document summarization dataset.
There are three configs:
- plain_text: plain text document-to-summary pairs
- plain_text_with_recommendations: plain text doucment-summary pairs, with "What GAO recommends" included in the summary
- structure: data with section structurespacex-launches
SpaceX Launch History
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Complete record of every SpaceX launch from spacex.com, including mission descriptions, pre/post-launch timelines, and photo galleries. Covers Falcon 1, Falcon 9, Falcon Heavy, and Starship missions.
The data is sourced from the official SpaceX content API and organized into three tables that can be joined on the slug field: launches (one row per mission… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/spacex-launches.pumpfun-launches
pump.fun launches with outcomes
(a) Rows in launches_legacy (days up to 2026-09-26) predate the 2026-09-27 decoder fix: about 48% of buy rows and 32% of the SOL entering curves were not recorded, so buy-derived features and the inputs of the rug rule are undercounted there (decoder_complete = 0), while the launches subset (days from 2026-09-27) is complete, apart from its first 620 rows (launched before the fix on 2026-09-27), which carry decoder_complete = 0. (b) Early… See the full description on the dataset page: https://huggingface.co/datasets/loopholetape/pumpfun-launches.e621_sample_clonexetAll images of all ratings from e621.net from the date it was generated, at sample resolution where possible.
This includes the following additional metadata:
post ID
created at
updated at
tags (stored as IDs you can cross-reference from an e621 tags dump)
rating (0 = safe, 1 = questionable, 2 = explicit)
favorite count
comment count
up score
down score
Note that this dataset excludes images that are, at the time of scraping:
pending
tagged with tags indicating that it is illegal to possess… See the full description on the dataset page: https://huggingface.co/datasets/launch-calcium/e621_sample_clonexet.space-launch-log
Space Launch Log
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Complete global launch history from Jonathan McDowell's General Catalog of Artificial Space Objects (GCAT) at the Harvard-Smithsonian Center for Astrophysics. Every orbital and suborbital launch attempt is cataloged with its vehicle type, launch site, mission objective, operating agency, and outcome code.
McDowell, an astrophysicist at the Harvard-Smithsonian… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/space-launch-log.ExpertLongBench
🎓 ExpertLongBench: Expert-Level Benchmark for Long-Form Generation with Structured Checklists
📊 The leaderboard for ExpertLongBench is hosted here: 🔗 https://huggingface.co/spaces/launch/ExpertLongBench
This is the public portion of the ExpertLongBench dataset, introduced in the paper:
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured ChecklistsJie Ruan, Inderjeet Jayakumar Nair, Shuyang Cao, Amy Liu, Sheza Munir, Micah… See the full description on the dataset page: https://huggingface.co/datasets/launch/ExpertLongBench.MCLASH
MCLASH: Multilingual CLASH
Paper | Code
MCLASH is the multilingual extension of CLASH (Character perspective-based LLM Assessments in Situations with High-stakes), a benchmark of long-form, human-written high-stakes dilemmas evaluated from multiple character perspectives. MCLASH carries the same dilemmas and character-perspective methodology into 5 additional languages — Spanish, Hindi, Korean, Malay, and Chinese.
See CLASH dataset card for detailed explanation of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/launch/MCLASH.rocket-lab-launches
Rocket Lab Launch Log
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Complete Rocket Lab launch manifest — past and upcoming missions flown or planned by Peter Beck's small-launch company, sourced from The Space Devs Launch Library 2 API.
Covers Electron (small-lift two-stage rocket using Rutherford 3D-printed engines, operational since 2017 from Mahia Peninsula in New Zealand and from LC-2 at Wallops Island, Virginia) and… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/rocket-lab-launches.ManyICLBenchThis dataset contains 21 many-shot ICL tasks that are designed to evaluate the long-context capability of LLMs, as introduced in the paper On Many-Shot In-Context Learning for Long-Context Evaluation. We categorize the tasks into similar-sample learning (SSL) and all-sample learning (ASL) groups.
SSL Tasks: banking77, dialogRE, TREC50, CLINC150, and BBH_geometric_shapes
ASL Tasks: GSM8K, MATH-algebra, MATH-counting_and_probability, MATH-geometry, MATH-number_theory, XLSUM, GPQA_cot… See the full description on the dataset page: https://huggingface.co/datasets/launch/ManyICLBench.ula-launches
ULA Launch Log
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Complete United Launch Alliance (ULA) launch manifest — past and upcoming flights of Atlas V, Delta II, Delta IV, Delta IV Heavy, and Vulcan Centaur — sourced from The Space Devs Launch Library 2 API.
ULA is the Boeing-Lockheed Martin joint venture formed in 2006 to consolidate US national security space launches under a single EELV (Evolved Expendable Launch… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/ula-launches.blue-origin-launches
Blue Origin Launch Log
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Complete Blue Origin launch manifest — past and upcoming missions flown or planned by Jeff Bezos's space company, sourced from The Space Devs Launch Library 2 API.
Covers New Shepard (suborbital reusable vehicle used for research and crewed space tourism) and New Glenn (heavy-lift reusable orbital rocket that first flew in 2025). Each row captures mission… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/blue-origin-launches.Chud-Launcherchahuadev-dev-launcher
Chahuadev Dev Launcher
Portable desktop command runner for development projects.
Purpose
Drop dev-launcher.exe inside a project folder.
Auto-detect project type from package.json and scripts.
Run grouped commands from one UI with full terminal output.
Optionally switch project root from the top bar (Select Project / Use Launcher Project).
Manage Cloudflare R2 upload flow for build artifacts (dev mode only).
Repository… See the full description on the dataset page: https://huggingface.co/datasets/chahuadev/chahuadev-dev-launcher.logs_launcher_backtestengineapi
rough
gov_report_qsGovReport-QS hierarchical question-summary generation dataset.
There are two configs:
- paragraph: paragraph-level annotated data
- document: aggregated paragraph-level annotated data for the same documentwz_launcherlaunch-vehicles
Launch Vehicles Database
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Orbital and suborbital launch vehicles from around the world, sourced from Wikidata. Covers every space launch vehicle recorded in Wikidata's structured knowledge base, including historical rockets, active workhorses, and vehicles under development.
From the German V-2 through the Saturn V to the Falcon 9 and Starship, launch vehicles have defined… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/launch-vehicles.ampere
Dataset Card for AMPERE
Dataset Description
This dataset is released together with our NAACL 2019 Paper "Argument Mining for Understanding Peer Reviews". If you find our work useful, please cite:
@inproceedings{hua-etal-2019-argument,
title = "Argument Mining for Understanding Peer Reviews",
author = "Hua, Xinyu and
Nikolov, Mitko and
Badugu, Nikhil and
Wang, Lu",
booktitle = "Proceedings of the 2019 Conference of the North {A}merican… See the full description on the dataset page: https://huggingface.co/datasets/launch/ampere.FactBench
FactBench Leaderboard
VERIFY: A Pipeline for Factuality Evaluation
Language models (LMs) are widely used by an increasing number of users, underscoring the challenge of maintaining factual accuracy across a broad range of topics. We present VERIFY (Verification and Evidence Retrieval for Factuality evaluation), a pipeline to evaluate LMs' factual accuracy in real-world user interactions.
Content Categorization
VERIFY considers the verifiability of LM-generated… See the full description on the dataset page: https://huggingface.co/datasets/launch/FactBench.agent-launch-pad-trajectories
agent-launch-pad trajectories
Multi-bench trajectory dataset collected by agent-launch-pad.
Each row is one (agent × model × task) cell with the full sharegpt-format conversation
and a grade_pass signal from the bench's own verifier (pytest, reward.txt, etc).
Coverage
Total trajectories: 1380
grade_pass=True: 145 (10.5%)
Per benchmark
terminal-bench-2: 1204 cells, 138 grade_pass (11.5%)
scienceagentbench: 176 cells, 7 grade_pass (4.0%)
Per model… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/agent-launch-pad-trajectories.ai-product-launches-2026
Ai Product Launches 2026
Part of the LEGION Intelligence dataset collection.
Provider: LEGION Systems
Access: Requires approval — submit request below
Usage
from datasets import load_dataset
dataset = load_dataset("gemmozero/ai-product-launches-2026")
API Access
Real-time access via LEGION API:
curl https://api.legion-api.com/incidents
API Docs · Pro Access €29/mo
License
CC BY-NC 4.0 — Research and non-commercial use only.… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-product-launches-2026.LudoBench
LudoBench
LLMs as Rules Oracles: Exploring Real-World Multimodal Reasoning in Tabletop Strategy Game Environments
ICLR 2026
A multimodal board-game reasoning benchmark evaluating LLM/VLM reasoning
across 5 strategy games and 3 difficulty tiers.
Fields
Field
Description
ID
Unique question identifier
Game
Board game name
tier
Difficulty tier (1, 2, or 3)
Question
The question text
Answer
Expected answer
game_state_url
Path(s) to game state… See the full description on the dataset page: https://huggingface.co/datasets/launch/LudoBench.reddit_qgReddit question generation dataset.spacex-launches
SpaceX Launch History
Credit: NASA
Part of a dataset collection on Hugging Face.
Dataset description
Complete record of every SpaceX launch from spacex.com, including mission descriptions, pre/post-launch timelines, and photo galleries. Covers Falcon 1, Falcon 9, Falcon Heavy, and Starship missions.
The data is sourced from the official SpaceX content API and organized into three tables that can be joined on the slug field: launches (one row per mission… See the full description on the dataset page: https://huggingface.co/datasets/windsurfer1511/spacex-launches.open_question_typeOpen-ended question type annotated dataset.e621-2024
e621-2024
e621-2024 is a large-scale furry image dataset retrieved from e621, a mature furry imageboard. This dataset is heavily inspired by the danbooru2023 dataset.
Similar to danbooru2023, the images in this dataset are bucketed into 1000 subdirectories (0000-0999), which is the E621 ID modulo 1000 (so all images in 0999/ have an ID ending in '999').
Currently there is no loading script for this dataset, so loading it with HuggingFace datasets is not supported.
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/launch-calcium/e621-2024.Launchspace-launches
Launches and the satellites they carried, 1957 to 2026
This is a copy of the dataset at mcval.org/launches/data, updated there every Monday and here straight after. Cite it by its DOI (see Citing).
Version 2026-10-07. 7,634 launches from 1957-10-04 to 2026-10-07, and 27,969 payloads, every one linked to the launch that carried it.
Made by McVal for Every launch since Sputnik, a 3D globe of the same data. Licensed CC BY 4.0: use it for anything, with credit (see Citing).… See the full description on the dataset page: https://huggingface.co/datasets/mcval/space-launches.CLASH
CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives
Paper: CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple PerspectivesContact: leeay@umich.edu
Overview
CLASH (Character perspective-based LLM Assessments in Situations with High-stakes) is a benchmark consisting of 345 long-form, human-written dilemmas spanning high-impact domains.
Each dilemma includes a pair of value-related rationales… See the full description on the dataset page: https://huggingface.co/datasets/launch/CLASH.
