datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
1c-enterprise-script-clean-corpus
1C Enterprise Script Clean Corpus
Кратко
Это датасет для обучения моделей работе с кодом на 1C:Enterprise Script.
В репозитории есть два независимых поднабора:
pretrain: чистый корпус реального кода для continued pretraining / domain adaptation
sft_strict: instruction/chat датасет для SFT, построенный поверх очищенного корпуса
Состав
pretrain/train.jsonl
pretrain/validation.jsonl
sft_strict/train.jsonl
sft_strict/validation.jsonl
manifests/repos.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/1c-enterprise-script-clean-corpus.shell-script-specialist-dataset
Full-Spectrum Shell Script Specialist Dataset
This dataset contains 1,000 curated, unique ChatML conversation records engineered to fine-tune a specialist language model for Production-Grade Shell Scripting (Bash 5+, POSIX /bin/sh, jq, awk, sed).
It was used to train the rajivmehtapy/gemma-4-e4b-shell-specialist model using Unsloth.
Dataset Splits
Split
File
Records
Description
train
train.jsonl
900
Core training set across all 4 production modules… See the full description on the dataset page: https://huggingface.co/datasets/rajivmehtapy/shell-script-specialist-dataset.indus-script-synthetic
Synthetic Indus Script Dataset
This dataset contains 5,000 synthetic Indus Script sequences produced by a two-stage training and generation pipeline built on 3,310 real archaeological inscriptions.
Stage 1 — Train on real inscriptions:
Four models were trained independently on the 3,310 real sequences. TinyBERT was trained as both a masked language model (predicting missing signs) and a sequence classifier (valid vs corrupted). An N-gram RTL model was trained to learn right-to-left… See the full description on the dataset page: https://huggingface.co/datasets/hellosindh/indus-script-synthetic.disturban-scripts
Disturban scripts
This dataset contains scripts from the YouTuber Disturban.
The records were parsed with Gemini 3 Flash to generate both voiceover text and video descriptions, so this dataset is not just audio transcripts. It is structured for workflows that need paired narration and visual-scene descriptions.
Files
disturban.jsonl: JSONL records containing parsed Disturban script content and generated descriptions.
Possible uses
Narration-to-video or… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/disturban-scripts.
