Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /dataset_with_scriptThis is a test dataset.textn<1K0 likes112k downloads2y agoHugging Face02adollahamini1998 /ScriptsForVoxBlinkResourcestext1M<n<10M0 likes4k downloads29d agoHugging Face03dh-unibe /image-text_medieval-scripts_xiv-xv-xvi Dataset Card for image-text_medieval-scripts_xiv-xv-xvi This dataset was created using pagexml-hf converter from Transkribus PageXML data. Dataset Summary This dataset contains 548322 samples across 1 split(s). Geographical scope: BelgiumPeriod: 1350-1550Languages: FlemishType of document: ProtocolProvenance: State Archives in Leuven Projects Included Itinera Nova Parts of Charters from Königsfelden SAL7304_full SAL7305_full SAL7306_full SAL7307 SAL7307_full… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_medieval-scripts_xiv-xv-xvi.image100K<n<1M1 likes3k downloads6mo agoHugging Face04Mightys /Notebook_Scriptstext100K<n<1M0 likes2.7k downloads4d agoHugging Face05Peacockery /common-voice-scripted-speech-26 Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tabularautomatic-speech-recognition100K<n<1M0 likes1.8k downloads3mo agoHugging Face06ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.6k downloads4mo agoHugging Face07uisp /pali-commentary-thai-script-siamrath-version Multi-File CSV Dataset คำอธิบาย อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์ 01/010001.csv: เล่ม 1 หน้า 1 01/010002.csv: เล่ม 1 หน้า 2 ... 02/020001.csv: เล่ม 2 หน้า 1 คำอธิบายของแต่ละเล่ม เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑) เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒) เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓) เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑) เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒) เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.tabular100K<n<1M3 likes1.2k downloads2y agoHugging Face08k19862217 /simpsons_script_linestext10K<n<100K0 likes1.1k downloads3y agoHugging Face09FelonieZ /DayZ_Scriptstextquestion-answering1K<n<10K0 likes991 downloads2y agoHugging Face10jordangong /jupyter-scripts-smollm3 The Stack v2 Jupyter Notebooks as Scripts This dataset contains script representations of the Jupyter notebooks in The Stack v2. It was created from the materialized Jupyter_Notebook split in jordangong/the-stack-v2-smollm3. The output schema follows the Jupyter-script schema used by bigcode/starcoderdata, but this release is not deduplicated, PII-filtered, or otherwise equivalent to StarCoderData's filtered split. Relationship to the SmolLM3 training mix This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.tabulartext-generation1M<n<10M0 likes977 downloads29d agoHugging Face11datablations /scriptstabular1M<n<10M0 likes941 downloads3y agoHugging Face12Aurel-test /simpsons_script_lines_parsedtext10K<n<100K0 likes644 downloads5mo agoHugging Face13shimokawatoko /parc-track3-scripted-demos PARC Track3 Scripted Demos PARC2026 Track3(言語・構成的把持運搬)学習用の決定論的スクリプト方策デモ集。 Franka Emika Panda 7-DoF × LIBERO-plus キッチンシーン(MuJoCo、EGL、128px記録、 最大300ステップ/試行、制御20Hz)。 成功 = タスクdone かつ非注目物体の変位1mm以下。 生成コードはPARC2026プロジェクトの work/infra/remote/scripted_demo.py(+テスト)。 プランナーは決定論的:同一BDDL+同一.pruned_init+同一コード=同一デモ。 乱数源なし(random/np.random不使用)。物体配置はBDDL領域サンプリングによる。 バッチ一覧(timestamp dir) dir 本数 方策 proxy smooth* 備考 20260828(3 dirs) 264 phases既定 k1 0.327 baseline… See the full description on the dataset page: https://huggingface.co/datasets/shimokawatoko/parc-track3-scripted-demos.imageroboticsn<1K0 likes601 downloads9h agoHugging Face14Nacryos /ancient-scripts-datasets Ancient Scripts Decipherment Datasets Collated datasets for the paper: Deciphering Undersegmented Ancient Scripts Using Phonetic Prior Jiaming Luo, Frederik Hartmann, Enrico Santus, Regina Barzilay, Yuan Cao Transactions of the Association for Computational Linguistics, 2021 arXiv:2010.11054 This repository gathers the training datasets used in the paper — both those hosted in the authors' GitHub repos and the external cited sources. Repository Structure data/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Nacryos/ancient-scripts-datasets.tabulartext-classification10M<n<100M1 likes521 downloads7mo agoHugging Face15VyoJ /BridgeData-V2-Scripted-Images BridgeData V2 Image Triplets Dataset This dataset contains image triplets from BridgeData V2 trajectories in ImageFolder format. Derived From This dataset is a derivative of the 30 GB scripted subset of BridgeData V2 from RAIL-Berkeley. All rights and original licensing apply. Dataset Structure initial_images/: Contains first frame images (initial state) intermediate_images/: Contains intermediate frame images (frame 38) final_images/: Contains final frame… See the full description on the dataset page: https://huggingface.co/datasets/VyoJ/BridgeData-V2-Scripted-Images.imageimage-to-image1K<n<10K0 likes511 downloads1y agoHugging Face16taqbaylit /common-voice-scripted-speech-kab-26-huge Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned) Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips. Source Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12) Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective) License: CC0-1.0 Generated: 2026-07-12 Cleaning Pipeline Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/common-voice-scripted-speech-kab-26-huge.audio100K<n<1M0 likes510 downloads3mo agoHugging Face17mgprogm /pali-tripitaka-thai-script-siamrath-version 📚 พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฐ (๔๕ เล่ม) ต้นฉบับสามารถเข้าถึงได้ที่: 84000 พระธรรมขันธ์ 84000.org พระไตรปิฎก 🧾 รายการพระไตรปิฎก 📘 เล่ม ๑–๘: วินัยปิฎก เล่ม ๑: มหาวิภงฺโค (๑) เล่ม ๒: มหาวิภงฺโค (๒) เล่ม ๓: ภิกฺขุนีวิภงฺโค เล่ม ๔: มหาวคฺโค (๑) เล่ม ๕: มหาวคฺโค (๒) เล่ม ๖: จุลฺลวคฺโค (๑) เล่ม ๗: จุลฺลวคฺโค (๒) เล่ม ๘: ปริวาโร 📗 เล่ม ๙–๒๕: สุตตันตปิฎก เล่ม ๙–๑๑: ทีฆนิกาย เล่ม ๑๒–๑๔: มัชฌิมนิกาย เล่ม ๑๕–๑๙: สังยุตตนิกาย… See the full description on the dataset page: https://huggingface.co/datasets/mgprogm/pali-tripitaka-thai-script-siamrath-version.tabular100K<n<1M1 likes426 downloads1y agoHugging Face18aneeshas /imsdb-genre-movie-scripts Dataset Card for "imsdb-genre-movie-scripts" More Information needed textn<1K3 likes417 downloads3y agoHugging Face19yiyic /Atlatic_train_lang_script_idtext1M<n<10M0 likes398 downloads2y agoHugging Face20uv-scripts /ocr-demo OCR demo: Food for Space Flight Seven scanned pages from NASA's Food for Space Flight booklet, with headings, columns, photographs, food lists and tables. This small dataset is an input for trying OCR recipes and inspecting their results. The images are PDF pages 3-9 (printed pages 2-8) of the original booklet. The selection omits the reproduction disclaimer and dark cover. The complete original PDF and a matching seven-page extract are in the OCR demo Bucket. Use as… See the full description on the dataset page: https://huggingface.co/datasets/uv-scripts/ocr-demo.imageimage-to-textn<1K2 likes350 downloads24d agoHugging Face21windfromthenorth /scripted_atomic_step_pose_0.6This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "ur5_wsg50_lego_atomic_step", "total_episodes": 955, "total_frames": 159935, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:955" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_pose_0.6.tabularrobotics100K<n<1M0 likes347 downloads3mo agoHugging Face22ScriptSmith /sponsorblock-youtube-metadata-2024 SponsorBlock YouTube Metadata Dataset A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos. Contains the top videos from the SponsorBlock database that had data added in the year 2024. Quick Stats Metric Value Total videos 154,536 Videos with subtitles 62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.imagetext-classification10M<n<100M0 likes341 downloads2mo agoHugging Face23anurag-chand /vivekananda-scriptures-lancedbtext100K<n<1M0 likes317 downloads4mo agoHugging Face24badrex /anv_data_ke_kikuyu_scriptedaudio100K<n<1M0 likes314 downloads1y agoHugging Face25uisp /pali-tripitaka-thai-script-siamrath-version Multi-File CSV Dataset คำอธิบาย พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์ 01/010001.csv: เล่ม 1 หน้า 1 01/010002.csv: เล่ม 1 หน้า 2 ... 02/020001.csv: เล่ม 2 หน้า 1 ... คำอธิบายของแต่ละเล่ม เล่ม ๑: วินย. มหาวิภงฺโค (๑) เล่ม ๒: วินย. มหาวิภงฺโค (๒) เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค เล่ม ๔: วินย. มหาวคฺโค (๑) เล่ม ๕: วินย. มหาวคฺโค (๒) เล่ม ๖: วินย. จุลฺลวคฺโค (๑) เล่ม ๗: วินย. จุลฺลวคฺโค (๒) เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.tabular100K<n<1M3 likes312 downloads2y agoHugging Face26loubnabnl /kaggle_scripts_new_format_subset Dataset Card for "kaggle_scripts_new_format_subset" More Information needed tabular1M<n<10M0 likes305 downloads3y agoHugging Face27windfromthenorth /scripted_atomic_step_train_frac0.2_largeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "ur5_wsg50_lego_atomic_step", "total_episodes": 512, "total_frames": 107566, "total_tasks": 1, "total_videos": 1024, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:512" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_train_frac0.2_large.tabularrobotics100K<n<1M0 likes303 downloads3mo agoHugging Face28anhnv125 /movie-scriptstext1K<n<10K0 likes291 downloads1y agoHugging Face29windfromthenorth /scripted_atomic_step_train_frac0.3_large_blindThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "ur5_wsg50_lego_atomic_step", "total_episodes": 664, "total_frames": 116214, "total_tasks": 1, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 20, "splits": { "train": "0:664" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_train_frac0.3_large_blind.tabularrobotics100K<n<1M0 likes272 downloads3mo agoHugging Face30osunlp /QUEST-SFT-Data-Objective-Script QUEST SFT Data Objective Script Project Page | Paper | GitHub Supervised fine-tuning split for QUEST / DeepResearch objective tasks. Each row includes the user prompt, a rule-style reward_model, extra_info, and the objective task category. The corresponding objective evaluation scripts are provided separately under eval_scripts/. This dataset follows the same broad schema style as osunlp/QUEST-RL-Data: each row includes prompt, reward_model, extra_info, and rl_task_category. The… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/QUEST-SFT-Data-Objective-Script.texttext-generation1K<n<10K1 likes233 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.