Team Ai
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.7k downloads4mo agoHugging Face02jordangong /jupyter-scripts-smollm3 The Stack v2 Jupyter Notebooks as Scripts This dataset contains script representations of the Jupyter notebooks in The Stack v2. It was created from the materialized Jupyter_Notebook split in jordangong/the-stack-v2-smollm3. The output schema follows the Jupyter-script schema used by bigcode/starcoderdata, but this release is not deduplicated, PII-filtered, or otherwise equivalent to StarCoderData's filtered split. Relationship to the SmolLM3 training mix This… See the full description on the dataset page: https://huggingface.co/datasets/jordangong/jupyter-scripts-smollm3.tabulartext-generation1M<n<10M0 likes378 downloads1mo agoHugging Face03ScriptSmith /sponsorblock-youtube-metadata-2024 SponsorBlock YouTube Metadata Dataset A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos. Contains the top videos from the SponsorBlock database that had data added in the year 2024. Quick Stats Metric Value Total videos 154,536 Videos with subtitles 62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.imagetext-classification10M<n<100M0 likes315 downloads3mo agoHugging Face04osunlp /QUEST-SFT-Data-Objective-Script QUEST SFT Data Objective Script Project Page | Paper | GitHub Supervised fine-tuning split for QUEST / DeepResearch objective tasks. Each row includes the user prompt, a rule-style reward_model, extra_info, and the objective task category. The corresponding objective evaluation scripts are provided separately under eval_scripts/. This dataset follows the same broad schema style as osunlp/QUEST-RL-Data: each row includes prompt, reward_model, extra_info, and rl_task_category. The… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/QUEST-SFT-Data-Objective-Script.texttext-generation1K<n<10K1 likes278 downloads3mo agoHugging Face05NickIBrody /1c-enterprise-script-clean-corpus 1C Enterprise Script Clean Corpus Кратко Это датасет для обучения моделей работе с кодом на 1C:Enterprise Script. В репозитории есть два независимых поднабора: pretrain: чистый корпус реального кода для continued pretraining / domain adaptation sft_strict: instruction/chat датасет для SFT, построенный поверх очищенного корпуса Состав pretrain/train.jsonl pretrain/validation.jsonl sft_strict/train.jsonl sft_strict/validation.jsonl manifests/repos.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/1c-enterprise-script-clean-corpus.texttext-generation10K<n<100K1 likes83 downloads5mo agoHugging Face06FelonieZ /DayZ_Scriptstextquestion-answering1K<n<10K0 likes75 downloads2y agoHugging Face07rodriguescarson /adaption-konkani-script-12k Konkani Script Conversion Konkani script tasks: Devanagari to Kannada script, Kannada script back to Devanagari, ISO 15919 romanisation, and script identification. Rows 12,000 Domain Konkani language Format data.parquet, one row per example Licence cc-by-4.0 Built for supervised fine-tuning (SFT) experiments on Adaption AutoScientist Columns Column Description original_prompt The prompt (user turn) as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-konkani-script-12k.tabulartext-generation10K<n<100K0 likes68 downloads14d agoHugging Face08rajivmehtapy /shell-script-specialist-dataset Full-Spectrum Shell Script Specialist Dataset This dataset contains 1,000 curated, unique ChatML conversation records engineered to fine-tune a specialist language model for Production-Grade Shell Scripting (Bash 5+, POSIX /bin/sh, jq, awk, sed). It was used to train the rajivmehtapy/gemma-4-e4b-shell-specialist model using Unsloth. Dataset Splits Split File Records Description train train.jsonl 900 Core training set across all 4 production modules… See the full description on the dataset page: https://huggingface.co/datasets/rajivmehtapy/shell-script-specialist-dataset.texttext-generation1K<n<10K0 likes54 downloads16d agoHugging Face09hellosindh /indus-script-synthetic Synthetic Indus Script Dataset This dataset contains 5,000 synthetic Indus Script sequences produced by a two-stage training and generation pipeline built on 3,310 real archaeological inscriptions. Stage 1 — Train on real inscriptions: Four models were trained independently on the 3,310 real sequences. TinyBERT was trained as both a masked language model (predicting missing signs) and a sequence classifier (valid vs corrupted). An N-gram RTL model was trained to learn right-to-left… See the full description on the dataset page: https://huggingface.co/datasets/hellosindh/indus-script-synthetic.texttext-generationn<1K0 likes30 downloads6mo agoHugging Face10trentmkelly /disturban-scripts Disturban scripts This dataset contains scripts from the YouTuber Disturban. The records were parsed with Gemini 3 Flash to generate both voiceover text and video descriptions, so this dataset is not just audio transcripts. It is structured for workflows that need paired narration and visual-scene descriptions. Files disturban.jsonl: JSONL records containing parsed Disturban script content and generated descriptions. Possible uses Narration-to-video or… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/disturban-scripts.texttext-generationn<1K0 likes23 downloads5mo agoHugging Face11jmaczan /rick-and-morty-script Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/jmaczan/rick-and-morty-script.texttext-generation1K<n<10K2 likes20 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.