Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /dataset_with_scriptThis is a test dataset.textn<1K0 likes116k downloads2y agoHugging Face02lerobot /aloha_sim_insertion_scriptedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 50, "total_frames": 20000, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_scripted.tabularrobotics10K<n<100K5 likes35k downloads4mo agoHugging Face03lerobot /aloha_sim_transfer_cube_scriptedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aloha", "total_episodes": 50, "total_frames": 20000, "total_tasks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4", "features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_scripted.tabularrobotics10K<n<100K7 likes21k downloads4mo agoHugging Face04abhishekkataria16 /repro-scripts0 likes8.6k downloads2mo agoHugging Face05lerobot /aloha_sim_insertion_scripted_imageThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 50, "total_frames": 20000, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path": null, "features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_scripted_image.imagerobotics10K<n<100K2 likes7.7k downloads4mo agoHugging Face06lerobot /aloha_sim_transfer_cube_scripted_imageThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "aloha", "total_episodes": 50, "total_frames": 20000, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path": null, "features": {… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_scripted_image.imagerobotics10K<n<100K0 likes6.3k downloads4mo agoHugging Face07uv-scripts /ocr OCR UV Scripts Part of uv-scripts: self-contained UV scripts you run on Hugging Face Jobs in one command. One script per OCR model. Each script runs the model on a GPU with Hugging Face Jobs and writes the text as markdown: as a new column in a Hub dataset, as .md files in a Bucket, or as resumable parquet parts (the -saturate recipes). A few scripts return JSON from a schema, detect layout regions, or compare the output of two models. Quick Start First… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/ocr.163 likes5.9k downloads6d agoHugging Face08adollahamini1998 /ScriptsForVoxBlinkResourcestext1M<n<10M0 likes3.9k downloads26d agoHugging Face09dh-unibe /image-text_medieval-scripts_xiv-xv-xvi Dataset Card for image-text_medieval-scripts_xiv-xv-xvi This dataset was created using pagexml-hf converter from Transkribus PageXML data. Dataset Summary This dataset contains 548322 samples across 1 split(s). Geographical scope: BelgiumPeriod: 1350-1550Languages: FlemishType of document: ProtocolProvenance: State Archives in Leuven Projects Included Itinera Nova Parts of Charters from Königsfelden SAL7304_full SAL7305_full SAL7306_full SAL7307 SAL7307_full… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_medieval-scripts_xiv-xv-xvi.image100K<n<1M1 likes3k downloads5mo agoHugging Face10Mightys /Notebook_Scriptstext100K<n<1M0 likes2.7k downloads2d agoHugging Face11FelonieZ /DayZ_Scriptstextquestion-answering1K<n<10K0 likes2k downloads2y agoHugging Face12Peacockery /common-voice-scripted-speech-26 Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tabularautomatic-speech-recognition100K<n<1M0 likes1.7k downloads3mo agoHugging Face13ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.6k downloads4mo agoHugging Face14jordangong /jupyter-scripts-smollm3 The Stack v2 Jupyter Notebooks as Scripts This dataset contains script representations of the Jupyter notebooks in The Stack v2. It was created from the materialized Jupyter_Notebook split in jordangong/the-stack-v2-smollm3. The output schema follows the Jupyter-script schema used by bigcode/starcoderdata, but this release is not deduplicated, PII-filtered, or otherwise equivalent to StarCoderData's filtered split. Relationship to the SmolLM3 training mix This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.tabulartext-generation1M<n<10M0 likes1.4k downloads27d agoHugging Face15k19862217 /simpsons_script_linestext10K<n<100K0 likes1.1k downloads3y agoHugging Face16datablations /scriptstabular1M<n<10M0 likes904 downloads3y agoHugging Face17davanstrien /jobs-scripts0 likes888 downloads2mo agoHugging Face18hf-internal-testing /gated_dataset_with_scriptgatedThis is a test dataset.0 likes794 downloads2y agoHugging Face19AIxBlock /92k-real-world-call-center-scripts-englishArXiv Paper Publication Here: "Real-World En Call Center Transcripts Dataset with PII Redaction" This dataset includes 91,706 high-quality transcriptions corresponding to approximately 10,500 hours of real-world call center conversations in English, collected across various industries and global regions. The dataset features both inbound and outbound calls and spans multiple accents, including Indian, American, and Filipino English. All transcripts have been carefully redacted for PII and… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/92k-real-world-call-center-scripts-english.10K<n<100K35 likes710 downloads1y agoHugging Face20Aurel-test /simpsons_script_lines_parsedtext10K<n<100K0 likes645 downloads5mo agoHugging Face21cadene /aloha_sim_transfer_cube_scripted_rawimagen<1K0 likes626 downloads2y agoHugging Face22WolfParametric /wolf-community-scriptsWOLF - AI-Powered Grasshopper Assistant 🐺 Transform your parametric design workflow with AI-powered script generation and sharing. WOLF revolutionizes Grasshopper by turning your scripts into an intelligent, searchable library and enabling instant script generation through natural language descriptions. 📥 Installation Step 1: Unblock the Plugin ⚠️ IMPORTANT: Do this FIRST before copying files Right-click on the downloaded Wolf.gha file Select "Properties" Check "Unblock" at the bottom if… See the full description on the dataset page: https://huggingface.co/datasets/WolfParametric/wolf-community-scripts.text-classification1M<n<10M3 likes620 downloads1mo agoHugging Face23jaivial /kronos-scripts Kronos — Open-source coding-model training pipeline Live status: https://kronos.menustudioai.com · R0 adapter: jaivial/kronos-round0-qwen15coder-lora · Logs: this dataset Kronos is an experiment in autonomously training an Opus-4.6-comparable coding LLM on €100 of seed capital, driven 24/7 by a self-evolving agent ("brain") running on systemd. Every cycle the brain orients on its own SQLite memory, picks ONE action (research / train / benchmark / report / improve / spend / idle)… See the full description on the dataset page: https://huggingface.co/datasets/jaivial/kronos-scripts.n<1K0 likes617 downloads5mo agoHugging Face24Damin3927 /dynamic_robot_bench_dr_scripted_14k dynamic_robot_bench_dr_scripted_14k 14,400 scripted-expert demonstrations across the 72 evaluated task families of dynamic-robot-bench — a conveyor-belt dynamic-manipulation benchmark (Franka Panda + wrist camera, ManiSkill 3 / SAPIEN GPU sim). One LeRobot v2.1 dataset: 200 episodes per family, success-filtered, every domain-randomization knob on, and belt speed uniform over 0.10–0.40 m/s. 1,118,617 frames · 209 distinct language instructions · 20 fps The belt speed… See the full description on the dataset page: https://huggingface.co/datasets/Damin3927/dynamic_robot_bench_dr_scripted_14k.imagerobotics1M<n<10M0 likes562 downloads25d agoHugging Face25VyoJ /BridgeData-V2-Scripted-Images BridgeData V2 Image Triplets Dataset This dataset contains image triplets from BridgeData V2 trajectories in ImageFolder format. Derived From This dataset is a derivative of the 30 GB scripted subset of BridgeData V2 from RAIL-Berkeley. All rights and original licensing apply. Dataset Structure initial_images/: Contains first frame images (initial state) intermediate_images/: Contains intermediate frame images (frame 38) final_images/: Contains final frame… See the full description on the dataset page: https://huggingface.co/datasets/VyoJ/BridgeData-V2-Scripted-Images.imageimage-to-image1K<n<10K0 likes521 downloads1y agoHugging Face26adollahamini1998 /ScriptsForVoxBlink2 The VoxBlink2 Dataset The VoxBlink2 dataset is a Large Scale speaker recognition dataset with 100K+ speakers obtained from YouTube platform. This repository provides guidelines to build the corpus and relative resources to reproduce the results in our article . For more introduction, please see cite. If you find this repository helpful to your research, don't forget to give us star🌟. Resource Let's start with obtaining the resource files and decompressing… See the full description on the dataset page: https://huggingface.co/datasets/adollahamini1998/ScriptsForVoxBlink2.2 likes515 downloads2mo agoHugging Face27Nacryos /ancient-scripts-datasets Ancient Scripts Decipherment Datasets Collated datasets for the paper: Deciphering Undersegmented Ancient Scripts Using Phonetic Prior Jiaming Luo, Frederik Hartmann, Enrico Santus, Regina Barzilay, Yuan Cao Transactions of the Association for Computational Linguistics, 2021 arXiv:2010.11054 This repository gathers the training datasets used in the paper — both those hosted in the authors' GitHub repos and the external cited sources. Repository Structure data/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Nacryos/ancient-scripts-datasets.tabulartext-classification10M<n<100M1 likes494 downloads7mo agoHugging Face28taqbaylit /common-voice-scripted-speech-kab-26-huge Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned) Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips. Source Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12) Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective) License: CC0-1.0 Generated: 2026-07-12 Cleaning Pipeline Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/common-voice-scripted-speech-kab-26-huge.audio100K<n<1M0 likes484 downloads3mo agoHugging Face29yiyic /Atlatic_train_lang_script_idtext1M<n<10M0 likes467 downloads2y agoHugging Face30gopika13 /answer_scripts Answer Scripts Dataset This dataset contains handwritten answer scripts along with extracte code text. Structure images/: Contains scanned answer sheets. annotations.parquet: Contains corresponding text for each image. Usage from datasets import load_dataset dataset = load_dataset("gopika13/answer_scripts") print(dataset) 0 likes455 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.