datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
simpsons_script_linesmovie-scriptsStar-wars-scripts-dialogue-IV-VI
Dataset Contents
This dataset contains the concatenated scripts from the original (and best) Star Wars trilogy. The scripts are reduced to dialogue only, and are tagged with a line number and speaker.
Dataset Disclaimer
I don't own this data; or Star Wars. But it would be cool if I did.
Star Wars is owned by Lucasfilms. I do not own any of the rights to this information.
The scripts are derived from a couple sources:
This GitHub Repo with raw files
A Kaggle Dataset put… See the full description on the dataset page: https://huggingface.co/datasets/andrewkroening/Star-wars-scripts-dialogue-IV-VI.1c-enterprise-script-clean-corpus
1C Enterprise Script Clean Corpus
Кратко
Это датасет для обучения моделей работе с кодом на 1C:Enterprise Script.
В репозитории есть два независимых поднабора:
pretrain: чистый корпус реального кода для continued pretraining / domain adaptation
sft_strict: instruction/chat датасет для SFT, построенный поверх очищенного корпуса
Состав
pretrain/train.jsonl
pretrain/validation.jsonl
sft_strict/train.jsonl
sft_strict/validation.jsonl
manifests/repos.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/1c-enterprise-script-clean-corpus.shell-script-specialist-dataset
Full-Spectrum Shell Script Specialist Dataset
This dataset contains 1,000 curated, unique ChatML conversation records engineered to fine-tune a specialist language model for Production-Grade Shell Scripting (Bash 5+, POSIX /bin/sh, jq, awk, sed).
It was used to train the rajivmehtapy/gemma-4-e4b-shell-specialist model using Unsloth.
Dataset Splits
Split
File
Records
Description
train
train.jsonl
900
Core training set across all 4 production modules… See the full description on the dataset page: https://huggingface.co/datasets/rajivmehtapy/shell-script-specialist-dataset.source_scripts_data_ayapython-data-transformation-scripts-cmskdijz
Python Data Transformation Scripts
Each item is a python data transformation scripts example providing Input sample, What the transform should do, Transformation script, Expected output. Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Contributor items exported: 1000
Language: Python
Framework:… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/python-data-transformation-scripts-cmskdijz.agents_vs_scriptindus-script-synthetic
Synthetic Indus Script Dataset
This dataset contains 5,000 synthetic Indus Script sequences produced by a two-stage training and generation pipeline built on 3,310 real archaeological inscriptions.
Stage 1 — Train on real inscriptions:
Four models were trained independently on the 3,310 real sequences. TinyBERT was trained as both a masked language model (predicting missing signs) and a sequence classifier (valid vs corrupted). An N-gram RTL model was trained to learn right-to-left… See the full description on the dataset page: https://huggingface.co/datasets/hellosindh/indus-script-synthetic.black_mirror_scripts_S1-5Black Mirror Scripts Dataset (Seasons 1-5)
This dataset, titled 'black_mirror_scripts_S1-5.csv', contains the meticulously compiled transcripts of the critically acclaimed anthology series Black Mirror, covering Seasons 1 through 5. Each entry in this dataset is categorized by unique identifiers including Script ID, Title, Scene, Dialogue, and Timestamp, making it an ideal resource for natural language processing tasks, script analysis, sentiment analysis, and more.
Dataset Composition
Our… See the full description on the dataset page: https://huggingface.co/datasets/tmobley96/black_mirror_scripts_S1-5.disturban-scripts
Disturban scripts
This dataset contains scripts from the YouTuber Disturban.
The records were parsed with Gemini 3 Flash to generate both voiceover text and video descriptions, so this dataset is not just audio transcripts. It is structured for workflows that need paired narration and visual-scene descriptions.
Files
disturban.jsonl: JSONL records containing parsed Disturban script content and generated descriptions.
Possible uses
Narration-to-video or… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/disturban-scripts.ttt-scripted-smoke
OpenEnv rollouts
Collected with OpenEnv (openenv collect).
Episodes: 20
Run metadata:
key
value
env
openspiel:tic_tac_toe
env_base_url
https://sergiopaniego-openspiel-ttt-env.hf.space
provider
scripted
model
None
num_episodes_requested
20
temperature
0.2
keep_losses
False
Schema
Each line of results.jsonl is one episode:
episode_id (string)
messages (chat transcript; TRL SFTTrainer-compatible)
reward (float)
done (bool)
tool_trace (list of… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/ttt-scripted-smoke.test_upload_scriptETL_SCRIPTSARC-data-scriptadaption-business-compliance-qa-in-devanagri-script
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-Business compliance QA in devanagri script
This dataset contains question-and-answer pairs addressing Indian business compliance, tax regulations, and government funding schemes for startups and SMEs. Each sample features a specific entrepreneur query followed by a detailed response that corrects misconceptions, cites relevant legal statutes, and outlines actionable compliance steps.… See the full description on the dataset page: https://huggingface.co/datasets/sidddd625/adaption-business-compliance-qa-in-devanagri-script.pali-words-myanmar-script
Pali Words in Myanmar Script (Master Index)
This dataset is a master index of 220,252 unique Pali words written in Myanmar (Burmese) Unicode script, intended for reuse across linguistic, religious, and computational workflows.
Data Fields
Each record contains:
word_id: A stable, sequential integer identifier.
pali_word: A Pali lexical item rendered in Myanmar Unicode script.
Data Processing Methodology
The dataset was constructed using the following steps:… See the full description on the dataset page: https://huggingface.co/datasets/freococo/pali-words-myanmar-script.roleplay_scriptsscripture_1500_pairs_gemini_flash_lite
Torah & Quran Theological Reasoning Dataset
This dataset contains 1,500 highly complex, synthetic question-and-answer pairs designed to train language models in theological, physical, and metaphysical reasoning. It serves as the foundational training data for the Elohim-3.8B reasoning model.
Dataset Sources
The dataset is built upon the combined text of two foundational religious scriptures. To ensure clean text extraction, the sources were curated before being processed:… See the full description on the dataset page: https://huggingface.co/datasets/TitleOS/scripture_1500_pairs_gemini_flash_lite.midnight-static-scriptsStar-wars-scripts-dialogue-IV-VI
Dataset Contents
This dataset contains the concatenated scripts from the original (and best) Star Wars trilogy. The scripts are reduced to dialogue only, and are tagged with a line number and speaker.
Dataset Disclaimer
I don't own this data; or Star Wars. But it would be cool if I did.
Star Wars is owned by Lucasfilms. I do not own any of the rights to this information.
The scripts are derived from a couple sources:
This GitHub Repo with raw files
A Kaggle Dataset put… See the full description on the dataset page: https://huggingface.co/datasets/mateowilliam/Star-wars-scripts-dialogue-IV-VI.youtube_script_generationscript_datadevelop_script_codenative_script_codemixedscriptfatih-script-dataset
Fatih Çoban YouTube Script Dataset
This dataset contains 270 training pairs for fine-tuning a language model to generate YouTube scripts in Fatih Çoban's style.
Each example consists of:
Input: Analysis of successful YouTube videos, including strategic analysis and intro techniques
Output: A complete script for a YouTube video
Usage
from datasets import load_dataset
# Load from Hugging Face
dataset = load_dataset("your-username/fatih-script-dataset") # Replace with… See the full description on the dataset page: https://huggingface.co/datasets/MECoban/fatih-script-dataset.pen-script-instructcustom-scripts-for-youtubescripts-dataset
