Team Ai
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.7k downloads4mo agoHugging Face02jordangong /jupyter-scripts-smollm3 The Stack v2 Jupyter Notebooks as Scripts This dataset contains script representations of the Jupyter notebooks in The Stack v2. It was created from the materialized Jupyter_Notebook split in jordangong/the-stack-v2-smollm3. The output schema follows the Jupyter-script schema used by bigcode/starcoderdata, but this release is not deduplicated, PII-filtered, or otherwise equivalent to StarCoderData's filtered split. Relationship to the SmolLM3 training mix This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.tabulartext-generation1M<n<10M0 likes378 downloads1mo agoHugging Face03ScriptSmith /sponsorblock-youtube-metadata-2024 SponsorBlock YouTube Metadata Dataset A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos. Contains the top videos from the SponsorBlock database that had data added in the year 2024. Quick Stats Metric Value Total videos 154,536 Videos with subtitles 62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.imagetext-classification10M<n<100M0 likes315 downloads3mo agoHugging Face04osunlp /QUEST-SFT-Data-Objective-Script QUEST SFT Data Objective Script Project Page | Paper | GitHub Supervised fine-tuning split for QUEST / DeepResearch objective tasks. Each row includes the user prompt, a rule-style reward_model, extra_info, and the objective task category. The corresponding objective evaluation scripts are provided separately under eval_scripts/. This dataset follows the same broad schema style as osunlp/QUEST-RL-Data: each row includes prompt, reward_model, extra_info, and rl_task_category. The… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/QUEST-SFT-Data-Objective-Script.texttext-generation1K<n<10K1 likes278 downloads3mo agoHugging Face05NickIBrody /1c-enterprise-script-clean-corpus 1C Enterprise Script Clean Corpus Кратко Это датасет для обучения моделей работе с кодом на 1C:Enterprise Script. В репозитории есть два независимых поднабора: pretrain: чистый корпус реального кода для continued pretraining / domain adaptation sft_strict: instruction/chat датасет для SFT, построенный поверх очищенного корпуса Состав pretrain/train.jsonl pretrain/validation.jsonl sft_strict/train.jsonl sft_strict/validation.jsonl manifests/repos.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/1c-enterprise-script-clean-corpus.texttext-generation10K<n<100K1 likes83 downloads5mo agoHugging Face06FelonieZ /DayZ_Scriptstextquestion-answering1K<n<10K0 likes75 downloads2y agoHugging Face07rodriguescarson /adaption-konkani-script-12k Konkani Script Conversion Konkani script tasks: Devanagari to Kannada script, Kannada script back to Devanagari, ISO 15919 romanisation, and script identification. Rows 12,000 Domain Konkani language Format data.parquet, one row per example Licence cc-by-4.0 Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-konkani-script-12k.tabulartext-generation10K<n<100K0 likes68 downloads14d agoHugging Face08rajivmehtapy /shell-script-specialist-dataset Full-Spectrum Shell Script Specialist Dataset This dataset contains 1,000 curated, unique ChatML conversation records engineered to fine-tune a specialist language model for Production-Grade Shell Scripting (Bash 5+, POSIX /bin/sh, jq, awk, sed). It was used to train the rajivmehtapy/gemma-4-e4b-shell-specialist model using Unsloth. Dataset Splits Split File Records Description train train.jsonl 900 Core training set across all 4 production modules… See the full description on the dataset page: https://huggingface.co/datasets/rajivmehtapy/shell-script-specialist-dataset.texttext-generation1K<n<10K0 likes54 downloads15d agoHugging Face09zhangdw /astra-skills-scripts 🛠️ ASTRA Skills Scripts: Script-Backed Skills from the Agent Skill Tool-use Repository Atlas Script-backed skills from the Agent Skill Tool-use Repository Atlas &nbsp;&nbsp;&nbsp;&nbsp; ASTRA Skills Scripts contains only skill directories from zhangdw/astra-skills that include a scripts/ subdirectory, making it easier to study agent skills that pair written instructions with runnable helper code. Quick Start · At a Glance · Subset Definition ·… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/astra-skills-scripts.text-generation10K<n<100K1 likes46 downloads3mo agoHugging Face10monetise /Structured-Scripture-for-AI ::PROJECT{Structured_Scripture_for_AI} ::PURPOSE{ENABLE(AI) → UNDERSTAND(Christian_theology) ∧ EXPLAIN(→ ∀ @HUMAN, ∀ culture, ∀ language, ∀ education_level) ∧ ZERO(friction)} ::TYPE{¬digitized_Bible ⇒ structured_encoding(three_layers)} ::ARCHITECTURE ::LAYER{text} WHAT(happened) — narrative ∧ events ∧ cause_effect ∧ speech ::LAYER{theology} WHAT(it_means) — within(Christian_doctrine) | logic ∧ paradox ∧ moral_principles ∧ emotion… See the full description on the dataset page: https://huggingface.co/datasets/monetise/Structured-Scripture-for-AI.text-generation0 likes43 downloads6mo agoHugging Face11jmaczan /rick-and-morty-scripts-vicuna-1 Rick and Morty scripts in Vicuna 1 format license: other License as in https://www.kaggle.com/datasets/andradaolteanu/rickmorty-scripts Original dataset by Andrada, adjusted to Llama 2 format by Jędrzej Paweł Maczan for C-137 project - Llama 2 7B on Apple M2 fine-tuned to revive Rick text-generation1K<n<10K0 likes40 downloads3y agoHugging Face12hellosindh /indus-script-synthetic Synthetic Indus Script Dataset This dataset contains 5,000 synthetic Indus Script sequences produced by a two-stage training and generation pipeline built on 3,310 real archaeological inscriptions. Stage 1 — Train on real inscriptions: Four models were trained independently on the 3,310 real sequences. TinyBERT was trained as both a masked language model (predicting missing signs) and a sequence classifier (valid vs corrupted). An N-gram RTL model was trained to learn right-to-left… See the full description on the dataset page: https://huggingface.co/datasets/hellosindh/indus-script-synthetic.texttext-generationn<1K0 likes30 downloads6mo agoHugging Face13trentmkelly /disturban-scripts Disturban scripts This dataset contains scripts from the YouTuber Disturban. The records were parsed with Gemini 3 Flash to generate both voiceover text and video descriptions, so this dataset is not just audio transcripts. It is structured for workflows that need paired narration and visual-scene descriptions. Files disturban.jsonl: JSONL records containing parsed Disturban script content and generated descriptions. Possible uses Narration-to-video or… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/disturban-scripts.texttext-generationn<1K0 likes23 downloads5mo agoHugging Face14jmaczan /rick-and-morty-script Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/jmaczan/rick-and-morty-script.texttext-generation1K<n<10K2 likes20 downloads2y agoHugging Face15athena2634 /albedo-mining-scripts albedo-mining-scripts Training scripts and shared-fail SFT/DPO mix for Albedo SN97 (Qwen3.6-35B-A3B), rebuilt for king 119. See RUNBOOK.md for the GPU run. Start weights: comguys/albedo-qwen3.6-35b-rzzi5i6@0c5dd11b7be00aae874e63d53d0e454f81c3e4a1 datasets/mix/sft.jsonl — 1679 upsampled SFT rows (reference trajectories) datasets/mix/dpo.jsonl — 332 unique DPO pairs (chat vs chat; rejected is king 119 when the eval has those turns) evals/ — source eval artifacts, including 119's… See the full description on the dataset page: https://huggingface.co/datasets/athena2634/albedo-mining-scripts.text-generation0 likes12 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.