datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset_with_scriptThis is a test dataset.aloha_sim_insertion_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_scripted.aloha_sim_transfer_cube_scriptedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_scripted.repro-scriptsaloha_sim_insertion_scripted_imageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_scripted_image.aloha_sim_transfer_cube_scripted_imageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_scripted_image.ocr
OCR UV Scripts
Part of uv-scripts: self-contained UV scripts you run on Hugging Face Jobs in one command.
One script per OCR model. Each script runs the model on a GPU with Hugging Face Jobs and writes the text as markdown: as a new column in a Hub dataset, as .md files in a Bucket, or as resumable parquet parts (the -saturate recipes). A few scripts return JSON from a schema, detect layout regions, or compare the output of two models.
Quick Start
First… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/ocr.ScriptsForVoxBlinkResourcesimage-text_medieval-scripts_xiv-xv-xvi
Dataset Card for image-text_medieval-scripts_xiv-xv-xvi
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 548322 samples across 1 split(s).
Geographical scope: BelgiumPeriod: 1350-1550Languages: FlemishType of document: ProtocolProvenance: State Archives in Leuven
Projects Included
Itinera Nova
Parts of Charters from Königsfelden
SAL7304_full
SAL7305_full
SAL7306_full
SAL7307
SAL7307_full… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_medieval-scripts_xiv-xv-xvi.Notebook_ScriptsDayZ_Scriptscommon-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.text-dataset-tiny-code-script-py-format
USED of tahamajs/medicine_ds_persian for .parquet file
USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file
USED of Abirate/english_quotes for .jsonl file
NEW FILES (05/12/2025)
NEW FILES (12/26/2025)
NEW FILES (02/15/2026)
jupyter-scripts-smollm3
The Stack v2 Jupyter Notebooks as Scripts
This dataset contains script representations of the Jupyter notebooks in
The Stack v2. It was
created from the materialized Jupyter_Notebook split in
jordangong/the-stack-v2-smollm3.
The output schema follows the Jupyter-script schema used by
bigcode/starcoderdata,
but this release is not deduplicated, PII-filtered, or otherwise equivalent
to StarCoderData's filtered split.
Relationship to the SmolLM3 training mix
This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.simpsons_script_linesscriptsjobs-scriptsgated_dataset_with_scriptThis is a test dataset.92k-real-world-call-center-scripts-englishArXiv Paper Publication Here: "Real-World En Call Center Transcripts Dataset with PII Redaction"
This dataset includes 91,706 high-quality transcriptions corresponding to approximately 10,500 hours of real-world call center conversations in English, collected across various industries and global regions. The dataset features both inbound and outbound calls and spans multiple accents, including Indian, American, and Filipino English. All transcripts have been carefully redacted for PII and… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/92k-real-world-call-center-scripts-english.simpsons_script_lines_parsedaloha_sim_transfer_cube_scripted_rawwolf-community-scriptsWOLF - AI-Powered Grasshopper Assistant 🐺
Transform your parametric design workflow with AI-powered script generation and sharing.
WOLF revolutionizes Grasshopper by turning your scripts into an intelligent, searchable library and enabling instant script generation through natural language descriptions.
📥 Installation
Step 1: Unblock the Plugin
⚠️ IMPORTANT: Do this FIRST before copying files
Right-click on the downloaded Wolf.gha file
Select "Properties"
Check "Unblock" at the bottom if… See the full description on the dataset page: https://huggingface.co/datasets/WolfParametric/wolf-community-scripts.kronos-scripts
Kronos — Open-source coding-model training pipeline
Live status: https://kronos.menustudioai.com · R0 adapter: jaivial/kronos-round0-qwen15coder-lora · Logs: this dataset
Kronos is an experiment in autonomously training an Opus-4.6-comparable coding LLM on €100 of seed capital, driven 24/7 by a self-evolving agent ("brain") running on systemd. Every cycle the brain orients on its own SQLite memory, picks ONE action (research / train / benchmark / report / improve / spend / idle)… See the full description on the dataset page: https://huggingface.co/datasets/jaivial/kronos-scripts.dynamic_robot_bench_dr_scripted_14k
dynamic_robot_bench_dr_scripted_14k
14,400 scripted-expert demonstrations across the 72 evaluated task
families of dynamic-robot-bench — a
conveyor-belt dynamic-manipulation benchmark (Franka Panda + wrist camera, ManiSkill 3 / SAPIEN
GPU sim). One LeRobot v2.1 dataset: 200 episodes per family,
success-filtered, every domain-randomization knob on, and belt speed uniform over
0.10–0.40 m/s.
1,118,617 frames · 209 distinct language instructions · 20 fps
The belt speed… See the full description on the dataset page: https://huggingface.co/datasets/Damin3927/dynamic_robot_bench_dr_scripted_14k.BridgeData-V2-Scripted-Images
BridgeData V2 Image Triplets Dataset
This dataset contains image triplets from BridgeData V2 trajectories in ImageFolder format.
Derived From
This dataset is a derivative of the 30 GB scripted subset of BridgeData V2 from RAIL-Berkeley. All rights and original licensing apply.
Dataset Structure
initial_images/: Contains first frame images (initial state)
intermediate_images/: Contains intermediate frame images (frame 38)
final_images/: Contains final frame… See the full description on the dataset page: https://huggingface.co/datasets/VyoJ/BridgeData-V2-Scripted-Images.ScriptsForVoxBlink2
The VoxBlink2 Dataset
The VoxBlink2 dataset is a Large Scale speaker recognition dataset with 100K+ speakers obtained from YouTube platform. This repository provides guidelines to build the corpus and relative resources to reproduce the results in our article . For more introduction, please see cite. If you find this repository helpful to your research, don't forget to give us star🌟.
Resource
Let's start with obtaining the resource files and decompressing… See the full description on the dataset page: https://huggingface.co/datasets/adollahamini1998/ScriptsForVoxBlink2.ancient-scripts-datasets
Ancient Scripts Decipherment Datasets
Collated datasets for the paper:
Deciphering Undersegmented Ancient Scripts Using Phonetic Prior
Jiaming Luo, Frederik Hartmann, Enrico Santus, Regina Barzilay, Yuan Cao
Transactions of the Association for Computational Linguistics, 2021
arXiv:2010.11054
This repository gathers the training datasets used in the paper — both those hosted in the authors' GitHub repos and the external cited sources.
Repository Structure
data/
├──… See the full description on the dataset page: https://huggingface.co/datasets/Nacryos/ancient-scripts-datasets.common-voice-scripted-speech-kab-26-huge
Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned)
Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips.
Source
Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12)
Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective)
License: CC0-1.0
Generated: 2026-07-12
Cleaning Pipeline
Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/common-voice-scripted-speech-kab-26-huge.Atlatic_train_lang_script_idanswer_scripts
Answer Scripts Dataset
This dataset contains handwritten answer scripts along with extracte code text.
Structure
images/: Contains scanned answer sheets.
annotations.parquet: Contains corresponding text for each image.
Usage
from datasets import load_dataset
dataset = load_dataset("gopika13/answer_scripts")
print(dataset)
