Team Ai
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NickIBrody /1c-enterprise-script-clean-corpus 1C Enterprise Script Clean Corpus Кратко Это датасет для обучения моделей работе с кодом на 1C:Enterprise Script. В репозитории есть два независимых поднабора: pretrain: чистый корпус реального кода для continued pretraining / domain adaptation sft_strict: instruction/chat датасет для SFT, построенный поверх очищенного корпуса Состав pretrain/train.jsonl pretrain/validation.jsonl sft_strict/train.jsonl sft_strict/validation.jsonl manifests/repos.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/1c-enterprise-script-clean-corpus.texttext-generation10K<n<100K1 likes83 downloads5mo agoHugging Face02rajivmehtapy /shell-script-specialist-dataset Full-Spectrum Shell Script Specialist Dataset This dataset contains 1,000 curated, unique ChatML conversation records engineered to fine-tune a specialist language model for Production-Grade Shell Scripting (Bash 5+, POSIX /bin/sh, jq, awk, sed). It was used to train the rajivmehtapy/gemma-4-e4b-shell-specialist model using Unsloth. Dataset Splits Split File Records Description train train.jsonl 900 Core training set across all 4 production modules… See the full description on the dataset page: https://huggingface.co/datasets/rajivmehtapy/shell-script-specialist-dataset.texttext-generation1K<n<10K0 likes54 downloads16d agoHugging Face03hellosindh /indus-script-synthetic Synthetic Indus Script Dataset This dataset contains 5,000 synthetic Indus Script sequences produced by a two-stage training and generation pipeline built on 3,310 real archaeological inscriptions. Stage 1 — Train on real inscriptions: Four models were trained independently on the 3,310 real sequences. TinyBERT was trained as both a masked language model (predicting missing signs) and a sequence classifier (valid vs corrupted). An N-gram RTL model was trained to learn right-to-left… See the full description on the dataset page: https://huggingface.co/datasets/hellosindh/indus-script-synthetic.texttext-generationn<1K0 likes30 downloads6mo agoHugging Face04trentmkelly /disturban-scripts Disturban scripts This dataset contains scripts from the YouTuber Disturban. The records were parsed with Gemini 3 Flash to generate both voiceover text and video descriptions, so this dataset is not just audio transcripts. It is structured for workflows that need paired narration and visual-scene descriptions. Files disturban.jsonl: JSONL records containing parsed Disturban script content and generated descriptions. Possible uses Narration-to-video or… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/disturban-scripts.texttext-generationn<1K0 likes23 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.