datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
college-roi-data
College ROI Data — what U.S. colleges and majors actually pay back
Clean, citable tables on the lifetime financial return of U.S. colleges and majors —
30-year net present value by school and state, ROI by major category, the out-of-state
premium, and how exposed each major's career paths are to today's AI. Maintained by
LE TEEN, a college-ROI data project. Every number traces to a
public source; nothing is scraped, modeled behind closed doors, or vibes.
The headline the… See the full description on the dataset page: https://huggingface.co/datasets/le-teen/college-roi-data.record-test-blueboxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 202,
"total_frames": 64086,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:202"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mark-Teeratorn/record-test-bluebox.teenytinystories
TeenyTinyStories (TTS)
A second-order TinyStories-style corpus: short, simple English stories generated by a small open model rather than by a frontier model.
Provenance
Generator: roneneldan/TinyStories-Instruct-33M (GPT-Neo architecture, 4 layers x 768 wide), run from random-free inference with the GPT-Neo tokenizer (EleutherAI/gpt-neo-125M).
Seeding: each story is conditioned on a structured header with three words (verb, noun, adjective, in the order used by… See the full description on the dataset page: https://huggingface.co/datasets/ableian/teenytinystories.record-test-3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 40,
"total_frames": 8800,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mark-Teeratorn/record-test-3.screwdriver_95_TEEEESTThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "mcx",
"total_episodes": 1,
"total_frames": 1751,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/antwoor/screwdriver_95_TEEEEST.record-test-snack2_20260701_221026This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Mark-Teeratorn/record-test-snack2_20260701_221026.record-test-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 3443,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mark-Teeratorn/record-test-2.record-test-snack_20260701_220626This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Mark-Teeratorn/record-test-snack_20260701_220626.tee_20260725_002010This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/mr-mph/tee_20260725_002010.eval_put-red-cube-in-blue-trayThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mark-Teeratorn/eval_put-red-cube-in-blue-tray.record-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 10,
"total_frames": 4039,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Mark-Teeratorn/record-test.Code_Opt_Triton
Overview
This dataset, TEEN-D/Code_Opt_Triton, is an extended version of the publicly available GPUMODE/Inductor_Created_Data_Permissive dataset. It contains pairs of original (PyTorch or Triton) programs and their equivalent Triton code (generated by torch inductor), intended for training models in PyTorch-to-Triton code translation and optimization.
The primary modification in this extended version is that each optimized Triton code snippet is paired with both its original source… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/Code_Opt_Triton.put-red-cube-in-blue-box_20260521_165414This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Mark-Teeratorn/put-red-cube-in-blue-box_20260521_165414.europe-ilo-ees-tees-sex-ind-jbc-nb-employees-by-ilo-sector-and-sex-and-type-of-job-co
Employees by ILO sector and sex and type of job contract (thousands) | Europe (ILOSTAT)
🇪🇺 44,599 observations · 15 Europe countries · 2001–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 44,599 observations of Employees data across 15 Europe countries, spanning 2001–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-ilo-ees-tees-sex-ind-jbc-nb-employees-by-ilo-sector-and-sex-and-type-of-job-co.abuja-housing-prices-v1
Abuja Housing Prices — Groundwork Data v1
Dataset Description
This is the first public, structured dataset of Abuja residential property prices — covering both rental and sale listings across 16 areas of Nigeria's Federal Capital Territory.
Collected by Groundwork Data, a Nigerian civic data project building the structured, domain-specific datasets that Nigerian ML engineers, researchers, and startups actually need.
Nigeria has a severe structured data gap. If you… See the full description on the dataset page: https://huggingface.co/datasets/BABA-TEE/abuja-housing-prices-v1.details_TeetouchQQ__model_mergev2
Dataset Card for Evaluation run of TeetouchQQ/model_mergev2
Dataset automatically created during the evaluation run of model TeetouchQQ/model_mergev2.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_TeetouchQQ__model_mergev2.TeeZee__DoubleBagel-57B-v1.0-details
Dataset Card for Evaluation run of TeeZee/DoubleBagel-57B-v1.0
Dataset automatically created during the evaluation run of model TeeZee/DoubleBagel-57B-v1.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TeeZee__DoubleBagel-57B-v1.0-details.Code_Opt_Triton_Shuffled
TEEN-D/Code_Opt_Triton_Shuffled
Overview
This dataset, TEEN-D/Code_Opt_Triton_Shuffled, is a shuffled version of the extended TEEN-D/Code_Opt_Triton dataset (which itself is an extension of GPUMODE/Inductor_Created_Data_Permissive). It provides a collection of pairs of original (PyTorch or Triton) programs and their corresponding optimized Triton code, designed for training machine learning models for code translation and optimization tasks targeting GPUs.
The key… See the full description on the dataset page: https://huggingface.co/datasets/Teen-Different/Code_Opt_Triton_Shuffled.emilia-yodas-english-neucodec-2000asia-ilo-ees-tees-eco-ocu-nb-employees-by-economic-activity-and-occupation-thou
Employees by economic activity and occupation (thousands) | Asia (ILOSTAT)
🌏 176,729 observations · 37 Asia countries · 1996–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 176,729 observations of Employees data across 37 Asia countries, spanning 1996–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ees-tees-eco-ocu-nb-employees-by-economic-activity-and-occupation-thou.africa-ilo-ees-tees-sex-eco-mts-nb-employees-by-sex-economic-activity-and-marital-sta
Employees by sex, economic activity and marital status (thousands) | Africa (ILOSTAT) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ilo-ees-tees-sex-eco-mts-nb-employees-by-sex-economic-activity-and-marital-sta.asia-ilo-ees-tees-sex-ocu-est-nb-employees-by-sex-occupation-and-establishment-size
Employees by sex, occupation and establishment size (thousands) | Asia (ILOSTAT)
🌏 44,856 observations · 25 Asia countries · 2000–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 44,856 observations of Employees data across 25 Asia countries, spanning 2000–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ees-tees-sex-ocu-est-nb-employees-by-sex-occupation-and-establishment-size.asia-ilo-ees-tees-sex-mts-dsb-nb-employees-by-sex-marital-status-and-disability-sta
Employees by sex, marital status and disability status (thousands) | Asia (ILOSTAT)
🌏 8,482 observations · 20 Asia countries · 1996–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 8,482 observations of Employees data across 20 Asia countries, spanning 1996–2024, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ees-tees-sex-mts-dsb-nb-employees-by-sex-marital-status-and-disability-sta.asia-ilo-ees-tees-sex-est-nb-employees-by-sex-and-establishment-size-thousands
Employees by sex and establishment size (thousands) | Asia (ILOSTAT)
🌏 4,444 observations · 26 Asia countries · 2000–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 4,444 observations of Employees data across 26 Asia countries, spanning 2000–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across employment… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ees-tees-sex-est-nb-employees-by-sex-and-establishment-size-thousands.europe-ilo-ees-tees-sex-ocu-est-nb-employees-by-sex-occupation-and-establishment-size
Employees by sex, occupation and establishment size (thousands) | Europe (ILOSTAT)
🇪🇺 61,558 observations · 13 Europe countries · 1993–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 61,558 observations of Employees data across 13 Europe countries, spanning 1993–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-ilo-ees-tees-sex-ocu-est-nb-employees-by-sex-occupation-and-establishment-size.qwen3-30b-traces-16x1024-cot
Qwen3-30B-A3B Thinking OpenR1-Math Traces
Generated with qwen3-30b-a3b-thinking from problems sampled from open-r1/OpenR1-Math-220k.
source_dataset: open-r1/OpenR1-Math-220k
source_config: default
source_split: train
sample_seed: 42
num_problems: 4096
traces_per_problem: 4
total_samples: 16384
samples_per_shard: 1024
num_shards: 16
temperature: 0.6
top_p: 0.95
top_k: 20
max_tokens: 32768
generated_at_utc: 2026-05-03T01:27:14.461250+00:00
Each row is one generated reasoning trace.… See the full description on the dataset page: https://huggingface.co/datasets/TeenSpirit/qwen3-30b-traces-16x1024-cot.asia-ilo-ees-tees-age-ec2-nb-employees-by-age-and-economic-activity-isic-level
Employees by age and economic activity - ISIC level 2 (thousands) | Asia (ILOSTAT)
🌏 71,386 observations · 35 Asia countries · 1996–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 71,386 observations of Employees data across 35 Asia countries, spanning 1996–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ees-tees-age-ec2-nb-employees-by-age-and-economic-activity-isic-level.asia-ilo-ees-tees-sex-oc2-nb-employees-by-sex-and-occupation-isco-level-2-thous
Employees by sex and occupation - ISCO level 2 (thousands) | Asia (ILOSTAT)
🌏 36,717 observations · 35 Asia countries · 1996–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 36,717 observations of Employees data across 35 Asia countries, spanning 1996–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ees-tees-sex-oc2-nb-employees-by-sex-and-occupation-isco-level-2-thous.asia-ilo-ees-tees-sex-est-dsb-nb-employees-by-sex-establishment-size-and-disability
Employees by sex, establishment size and disability status (thousands) | Asia (ILOSTAT)
🌏 1,898 observations · 13 Asia countries · 2006–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 1,898 observations of Employees data across 13 Asia countries, spanning 2006–2024, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-ees-tees-sex-est-dsb-nb-employees-by-sex-establishment-size-and-disability.europe-ilo-ees-tees-sex-ifl-dsb-nb-employees-by-sex-informal-formal-job-and-disabilit
Employees by sex, informal/formal job and disability status (thousands) | Europe (ILOSTAT)
🇪🇺 223 observations · 2 Europe countries · 2007–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 223 observations of Informal economy data across 2 Europe countries, spanning 2007–2025, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-ilo-ees-tees-sex-ifl-dsb-nb-employees-by-sex-informal-formal-job-and-disabilit.
