Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01k19862217 /simpsons_script_linestext10K<n<100K0 likes1.1k downloads3y agoHugging Face02sgogoi /movie-scriptstext100K<n<1M0 likes223 downloads11mo agoHugging Face03andrewkroening /Star-wars-scripts-dialogue-IV-VI Dataset Contents This dataset contains the concatenated scripts from the original (and best) Star Wars trilogy. The scripts are reduced to dialogue only, and are tagged with a line number and speaker. Dataset Disclaimer I don't own this data; or Star Wars. But it would be cool if I did. Star Wars is owned by Lucasfilms. I do not own any of the rights to this information. The scripts are derived from a couple sources: This GitHub Repo with raw files A Kaggle Dataset put… See the full description on the dataset page: https://huggingface.co/datasets/andrewkroening/Star-wars-scripts-dialogue-IV-VI.text1K<n<10K7 likes102 downloads4y agoHugging Face04NickIBrody /1c-enterprise-script-clean-corpus 1C Enterprise Script Clean Corpus Кратко Это датасет для обучения моделей работе с кодом на 1C:Enterprise Script. В репозитории есть два независимых поднабора: pretrain: чистый корпус реального кода для continued pretraining / domain adaptation sft_strict: instruction/chat датасет для SFT, построенный поверх очищенного корпуса Состав pretrain/train.jsonl pretrain/validation.jsonl sft_strict/train.jsonl sft_strict/validation.jsonl manifests/repos.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/1c-enterprise-script-clean-corpus.texttext-generation10K<n<100K1 likes83 downloads5mo agoHugging Face05rajivmehtapy /shell-script-specialist-dataset Full-Spectrum Shell Script Specialist Dataset This dataset contains 1,000 curated, unique ChatML conversation records engineered to fine-tune a specialist language model for Production-Grade Shell Scripting (Bash 5+, POSIX /bin/sh, jq, awk, sed). It was used to train the rajivmehtapy/gemma-4-e4b-shell-specialist model using Unsloth. Dataset Splits Split File Records Description train train.jsonl 900 Core training set across all 4 production modules… See the full description on the dataset page: https://huggingface.co/datasets/rajivmehtapy/shell-script-specialist-dataset.texttext-generation1K<n<10K0 likes54 downloads16d agoHugging Face06Gabrui /source_scripts_data_ayatextn<1K0 likes50 downloads2y agoHugging Face07cmu-lti /agents_vs_scripttext10K<n<100K3 likes39 downloads2y agoHugging Face08script-langid /sinhala-script-lid Sinhala-script language identification: Sinhala, Pali, Sanskrit Sentence-level instances of three languages written in Sinhala script, with a leakage-free document-blocked split. Labels: sin_Sinh, pli_Sinh, san_Sinh. Rows split sin_Sinh pli_Sinh san_Sinh total train 83105 58309 10565 151979 validation 10345 7317 1337 18999 test 10567 7105 1310 18982 Construction Built by scripts/maintainer/resplit_target.py (config: target_split… See the full description on the dataset page: https://huggingface.co/datasets/script-langid/sinhala-script-lid.texttext-classification100K<n<1M0 likes36 downloads4d agoHugging Face09hellosindh /indus-script-synthetic Synthetic Indus Script Dataset This dataset contains 5,000 synthetic Indus Script sequences produced by a two-stage training and generation pipeline built on 3,310 real archaeological inscriptions. Stage 1 — Train on real inscriptions: Four models were trained independently on the 3,310 real sequences. TinyBERT was trained as both a masked language model (predicting missing signs) and a sequence classifier (valid vs corrupted). An N-gram RTL model was trained to learn right-to-left… See the full description on the dataset page: https://huggingface.co/datasets/hellosindh/indus-script-synthetic.texttext-generationn<1K0 likes30 downloads6mo agoHugging Face10trentmkelly /disturban-scripts Disturban scripts This dataset contains scripts from the YouTuber Disturban. The records were parsed with Gemini 3 Flash to generate both voiceover text and video descriptions, so this dataset is not just audio transcripts. It is structured for workflows that need paired narration and visual-scene descriptions. Files disturban.jsonl: JSONL records containing parsed Disturban script content and generated descriptions. Possible uses Narration-to-video or… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/disturban-scripts.texttext-generationn<1K0 likes23 downloads5mo agoHugging Face11databounty-io /python-data-transformation-scripts-cmskdijz Python Data Transformation Scripts Each item is a python data transformation scripts example providing Input sample, What the transform should do, Transformation script, Expected output. Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones. About This dataset was produced by the DataBounty community and published here as part of an open, karma-only program. Contributor items exported: 1000 Language: Python Framework:… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/python-data-transformation-scripts-cmskdijz.text1K<n<10K0 likes23 downloads1mo agoHugging Face12tmobley96 /black_mirror_scripts_S1-5Black Mirror Scripts Dataset (Seasons 1-5) This dataset, titled 'black_mirror_scripts_S1-5.csv', contains the meticulously compiled transcripts of the critically acclaimed anthology series Black Mirror, covering Seasons 1 through 5. Each entry in this dataset is categorized by unique identifiers including Script ID, Title, Scene, Dialogue, and Timestamp, making it an ideal resource for natural language processing tasks, script analysis, sentiment analysis, and more. Dataset Composition Our… See the full description on the dataset page: https://huggingface.co/datasets/tmobley96/black_mirror_scripts_S1-5.text10K<n<100K2 likes21 downloads3y agoHugging Face13sergiopaniego /ttt-scripted-smoke OpenEnv rollouts Collected with OpenEnv (openenv collect). Episodes: 20 Run metadata: key value env openspiel:tic_tac_toe env_base_url https://sergiopaniego-openspiel-ttt-env.hf.space provider scripted model None num_episodes_requested 20 temperature 0.2 keep_losses False Schema Each line of results.jsonl is one episode: episode_id (string) messages (chat transcript; TRL SFTTrainer-compatible) reward (float) done (bool) tool_trace (list of… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/ttt-scripted-smoke.textn<1K0 likes19 downloads6mo agoHugging Face14sidddd625 /adaption-business-compliance-qa-in-devanagri-script This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-Business compliance QA in devanagri script This dataset contains question-and-answer pairs addressing Indian business compliance, tax regulations, and government funding schemes for startups and SMEs. Each sample features a specific entrepreneur query followed by a detailed response that corrects misconceptions, cites relevant legal statutes, and outlines actionable compliance steps.… See the full description on the dataset page: https://huggingface.co/datasets/sidddd625/adaption-business-compliance-qa-in-devanagri-script.textn<1K0 likes14 downloads3mo agoHugging Face15aditya-chandras-4 /ETL_SCRIPTStextn<1K1 likes13 downloads3y agoHugging Face16TitleOS /scripture_1500_pairs_gemini_flash_lite Torah & Quran Theological Reasoning Dataset This dataset contains 1,500 highly complex, synthetic question-and-answer pairs designed to train language models in theological, physical, and metaphysical reasoning. It serves as the foundational training data for the Elohim-3.8B reasoning model. Dataset Sources The dataset is built upon the combined text of two foundational religious scriptures. To ensure clean text extraction, the sources were curated before being processed:… See the full description on the dataset page: https://huggingface.co/datasets/TitleOS/scripture_1500_pairs_gemini_flash_lite.text1K<n<10K0 likes13 downloads6mo agoHugging Face17freococo /pali-words-myanmar-script Pali Words in Myanmar Script (Master Index) This dataset is a master index of 220,252 unique Pali words written in Myanmar (Burmese) Unicode script, intended for reuse across linguistic, religious, and computational workflows. Data Fields Each record contains: word_id: A stable, sequential integer identifier. pali_word: A Pali lexical item rendered in Myanmar Unicode script. Data Processing Methodology The dataset was constructed using the following steps:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/pali-words-myanmar-script.text100K<n<1M0 likes12 downloads8mo agoHugging Face18Chaser-cz /roleplay_scriptstextn<1K0 likes9 downloads2y agoHugging Face19build-small-hackathon /midnight-static-scriptstextn<1K0 likes7 downloads4mo agoHugging Face20mateowilliam /Star-wars-scripts-dialogue-IV-VI Dataset Contents This dataset contains the concatenated scripts from the original (and best) Star Wars trilogy. The scripts are reduced to dialogue only, and are tagged with a line number and speaker. Dataset Disclaimer I don't own this data; or Star Wars. But it would be cool if I did. Star Wars is owned by Lucasfilms. I do not own any of the rights to this information. The scripts are derived from a couple sources: This GitHub Repo with raw files A Kaggle Dataset put… See the full description on the dataset page: https://huggingface.co/datasets/mateowilliam/Star-wars-scripts-dialogue-IV-VI.text1K<n<10K0 likes6 downloads7mo agoHugging Face21mfdoom987 /youtube_script_generationtextn<1K0 likes5 downloads1y agoHugging Face22arpitkr0 /script_datatextn<1K0 likes5 downloads1y agoHugging Face23fc9999 /develop_script_codetextn<1K0 likes4 downloads2y agoHugging Face24cs23s036 /native_script_codemixedgatedtext1M<n<10M0 likes4 downloads1y agoHugging Face25Mathieu-Thomas-JOSSET /scripttextn<1K0 likes4 downloads1y agoHugging Face26Penbox /pen-script-instructtextn<1K0 likes3 downloads3y agoHugging Face27pedrocastellanos /custom-scripts-for-youtubetextn<1K0 likes3 downloads2y agoHugging Face28sanattt /ScriptAutomationtextn<1K0 likes3 downloads2y agoHugging Face29MECoban /fatih-script-dataset Fatih Çoban YouTube Script Dataset This dataset contains 270 training pairs for fine-tuning a language model to generate YouTube scripts in Fatih Çoban's style. Each example consists of: Input: Analysis of successful YouTube videos, including strategic analysis and intro techniques Output: A complete script for a YouTube video Usage from datasets import load_dataset # Load from Hugging Face dataset = load_dataset("your-username/fatih-script-dataset") # Replace with… See the full description on the dataset page: https://huggingface.co/datasets/MECoban/fatih-script-dataset.textn<1K0 likes3 downloads1y agoHugging Face30LEEyoungyo /Script_2textn<1K0 likes3 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.