datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-dataset-tiny-code-script-py-format
USED of tahamajs/medicine_ds_persian for .parquet file
USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file
USED of Abirate/english_quotes for .jsonl file
NEW FILES (05/12/2025)
NEW FILES (12/26/2025)
NEW FILES (02/15/2026)
jupyter-scripts-smollm3
The Stack v2 Jupyter Notebooks as Scripts
This dataset contains script representations of the Jupyter notebooks in
The Stack v2. It was
created from the materialized Jupyter_Notebook split in
jordangong/the-stack-v2-smollm3.
The output schema follows the Jupyter-script schema used by
bigcode/starcoderdata,
but this release is not deduplicated, PII-filtered, or otherwise equivalent
to StarCoderData's filtered split.
Relationship to the SmolLM3 training mix
This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.sponsorblock-youtube-metadata-2024
SponsorBlock YouTube Metadata Dataset
A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos.
Contains the top videos from the SponsorBlock database that had data added in the year 2024.
Quick Stats
Metric
Value
Total videos
154,536
Videos with subtitles
62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.QUEST-SFT-Data-Objective-Script
QUEST SFT Data Objective Script
Project Page | Paper | GitHub
Supervised fine-tuning split for QUEST / DeepResearch objective tasks. Each row includes the user prompt, a rule-style reward_model, extra_info, and the objective task category. The corresponding objective evaluation scripts are provided separately under eval_scripts/.
This dataset follows the same broad schema style as osunlp/QUEST-RL-Data: each row includes prompt, reward_model, extra_info, and rl_task_category. The… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/QUEST-SFT-Data-Objective-Script.1c-enterprise-script-clean-corpus
1C Enterprise Script Clean Corpus
Кратко
Это датасет для обучения моделей работе с кодом на 1C:Enterprise Script.
В репозитории есть два независимых поднабора:
pretrain: чистый корпус реального кода для continued pretraining / domain adaptation
sft_strict: instruction/chat датасет для SFT, построенный поверх очищенного корпуса
Состав
pretrain/train.jsonl
pretrain/validation.jsonl
sft_strict/train.jsonl
sft_strict/validation.jsonl
manifests/repos.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/1c-enterprise-script-clean-corpus.DayZ_Scriptsadaption-konkani-script-12k
Konkani Script Conversion
Konkani script tasks: Devanagari to Kannada script, Kannada script back to Devanagari, ISO 15919 romanisation, and script identification.
Rows
12,000
Domain
Konkani language
Format
data.parquet, one row per example
Licence
cc-by-4.0
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-konkani-script-12k.shell-script-specialist-dataset
Full-Spectrum Shell Script Specialist Dataset
This dataset contains 1,000 curated, unique ChatML conversation records engineered to fine-tune a specialist language model for Production-Grade Shell Scripting (Bash 5+, POSIX /bin/sh, jq, awk, sed).
It was used to train the rajivmehtapy/gemma-4-e4b-shell-specialist model using Unsloth.
Dataset Splits
Split
File
Records
Description
train
train.jsonl
900
Core training set across all 4 production modules… See the full description on the dataset page: https://huggingface.co/datasets/rajivmehtapy/shell-script-specialist-dataset.astra-skills-scripts
🛠️ ASTRA Skills Scripts: Script-Backed Skills from the Agent Skill Tool-use Repository Atlas
Script-backed skills from the Agent Skill Tool-use Repository Atlas
ASTRA Skills Scripts contains only skill directories from zhangdw/astra-skills that include a scripts/ subdirectory, making it easier to study agent skills that pair written instructions with runnable helper code.
Quick Start ·
At a Glance ·
Subset Definition ·… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/astra-skills-scripts.Structured-Scripture-for-AI
::PROJECT{Structured_Scripture_for_AI}
::PURPOSE{ENABLE(AI) → UNDERSTAND(Christian_theology) ∧ EXPLAIN(→ ∀ @HUMAN, ∀ culture, ∀ language, ∀ education_level) ∧ ZERO(friction)}
::TYPE{¬digitized_Bible ⇒ structured_encoding(three_layers)}
::ARCHITECTURE
::LAYER{text}
WHAT(happened) — narrative ∧ events ∧ cause_effect ∧ speech
::LAYER{theology}
WHAT(it_means) — within(Christian_doctrine) | logic ∧ paradox ∧ moral_principles ∧ emotion… See the full description on the dataset page: https://huggingface.co/datasets/monetise/Structured-Scripture-for-AI.rick-and-morty-scripts-vicuna-1
Rick and Morty scripts in Vicuna 1 format
license: other
License as in https://www.kaggle.com/datasets/andradaolteanu/rickmorty-scripts
Original dataset by Andrada, adjusted to Llama 2 format by Jędrzej Paweł Maczan for C-137 project - Llama 2 7B on Apple M2 fine-tuned to revive Rick
indus-script-synthetic
Synthetic Indus Script Dataset
This dataset contains 5,000 synthetic Indus Script sequences produced by a two-stage training and generation pipeline built on 3,310 real archaeological inscriptions.
Stage 1 — Train on real inscriptions:
Four models were trained independently on the 3,310 real sequences. TinyBERT was trained as both a masked language model (predicting missing signs) and a sequence classifier (valid vs corrupted). An N-gram RTL model was trained to learn right-to-left… See the full description on the dataset page: https://huggingface.co/datasets/hellosindh/indus-script-synthetic.disturban-scripts
Disturban scripts
This dataset contains scripts from the YouTuber Disturban.
The records were parsed with Gemini 3 Flash to generate both voiceover text and video descriptions, so this dataset is not just audio transcripts. It is structured for workflows that need paired narration and visual-scene descriptions.
Files
disturban.jsonl: JSONL records containing parsed Disturban script content and generated descriptions.
Possible uses
Narration-to-video or… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/disturban-scripts.rick-and-morty-script
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/jmaczan/rick-and-morty-script.albedo-mining-scripts
albedo-mining-scripts
Training scripts and shared-fail SFT/DPO mix for Albedo SN97 (Qwen3.6-35B-A3B),
rebuilt for king 119. See RUNBOOK.md for the GPU run.
Start weights:
comguys/albedo-qwen3.6-35b-rzzi5i6@0c5dd11b7be00aae874e63d53d0e454f81c3e4a1
datasets/mix/sft.jsonl — 1679 upsampled SFT rows (reference trajectories)
datasets/mix/dpo.jsonl — 332 unique DPO pairs (chat vs chat; rejected is king 119 when the eval has those turns)
evals/ — source eval artifacts, including 119's… See the full description on the dataset page: https://huggingface.co/datasets/athena2634/albedo-mining-scripts.
