Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes550 downloads2y agoHugging Face02projenix /tinysynth-reasoning TinySynth Reasoning Primitives Synthetic training data for teaching small language models stable state representation and controlled reasoning operations — entity/attribute binding, state persistence, mutation, transfer, reference resolution, current-vs-cumulative distinctions, and claim validation — in a systems/computing vocabulary. Every example is generated from a hidden symbolic world and verified by a symbolic solver before any natural language is produced: semantic state… See the full description on the dataset page: https://huggingface.co/datasets/projenix/tinysynth-reasoning.tabulartext-generation1M<n<10M0 likes359 downloads26d agoHugging Face03malaiwah /k2-horizon-tiny-cpu-repro-v1 K2-Horizon MoVA tiny random CPU fixture Complete untrained K2HorizonForCausalLM, not Moonshot Kimi despite the K2 name. Architecture source: IFM/K2-Horizon-MoVA-36B-A4B at 05cab0a4d7150c1c460a000b37ff40cc1af2feaa. No pretrained weights, original tokenizer, training data, paid GPU or cloud compute used. Complete text-only K2HorizonForCausalLM, not Kimi: three-layer dense prefix followed by two real MoVA+MoE layers, grouped RMSNorm, sigmoid top-k routing with selection-only bias… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes314 downloads1mo agoHugging Face04malaiwah /deepseek-v4-tiny-cpu-repro-v1 DeepSeek-V4 tiny corrected-native-primitives CPU text fixture Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction. No upstream weights, paid GPU/cloud compute or useful-model claim. This is not unmodified native Transformers or the complete production release. Architecture and scope Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes278 downloads1mo agoHugging Face05malaiwah /glm5-next-tiny-cpu-repro-v1This repository is an evidence bundle, not one root-format dataset at repository root. first/ and repeat/ are separate complete sealed QFS root datasets; comparison/ holds the comparison receipt and tokenwise result. panel/ is the sealed input panel. Other files are provenance, logs and reproduction tools. Do not pass the bundle root as a QFS dataset. GLM5-Next tiny native CPU fixture This is a complete untrained random-initialized native Glm5NextForConditionalGeneration wrapper… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes253 downloads1mo agoHugging Face06malaiwah /qwen3-5-tiny-cpu-repro-v1 Qwen3.5 tiny native random CPU fixture Complete randomly initialized, untrained Qwen3_5ForConditionalGeneration checkpoint. This is a pipeline/reproducibility fixture, not a useful language model, distillation, quantization, quality benchmark, or claim about the performance of Qwen3.8-27B. No upstream model weights or training data were used. No paid GPU/cloud compute. Architecture and lineage Architecture lineage: Qwen/Qwen3.8-27B at… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-5-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes215 downloads1mo agoHugging Face07algerian-nlp /TinyStories-Algerian-Darija TinyStories Algerian Darija Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous). The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.tabulartext-generation10K<n<100K0 likes211 downloads23d agoHugging Face08malaiwah /minimax-m2-tiny-cpu-repro-v1 minimax-m2 complete native tiny random CPU fixture Complete untrained MiniMaxM2ForCausalLM checkpoint with an untied full LM head, a real 272-entry byte tokenizer and every native state tensor. Architecture lineage: MiniMaxAI/MiniMax-M2.7@d494266a4affc0d2995ba1fa35c8481cbd84294b. No upstream weights, training data, paid GPU or cloud compute were used. Complete native text causal LM: sigmoid/top-k MoE routing with correction bias, per-layer flattened Q/K RMSNorm and half-head… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m2-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes210 downloads1mo agoHugging Face09malaiwah /minimax-m3-tiny-cpu-repro-v1 minimax-m3 complete native tiny random CPU fixture Complete untrained MiniMaxM3SparseForConditionalGeneration checkpoint with an untied full LM head, a real 272-entry byte tokenizer and every native state tensor. Architecture lineage: MiniMaxAI/MiniMax-M3@f0e1c1e04d40177e4673a22097036854f536e9c0. No upstream weights, training data, paid GPU or cloud compute were used. Complete native image/text wrapper with real shrunk Conv3D vision, nonempty 3D RoPE, patch-merge projector and… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m3-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes201 downloads1mo agoHugging Face10touati-kamel /TinyStories-Algerian-Darija TinyStories Algerian Darija: Parallel Children Stories Corpus & Cultural Adaptation Pipeline A large-scale, high-fidelity parallel corpus of 11,326 synthetically generated children stories translated from Microsoft's roneneldan/TinyStories and culturally localized into authentic Algerian Arabic (الدارجة الجزائرية) in clean Arabic script. The dataset is engineered to train and evaluate Small Language Models (SLMs) and Low-Resource Dialectal LLMs on reasoning, narrative… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/TinyStories-Algerian-Darija.tabulartranslation10K<n<100K0 likes151 downloads12d agoHugging Face11nampdn-ai /tiny-code-textbooksgated Code Explanation Textbooks A collection of 207k synthetic code with explanation as a tiny textbook. Filtered from the-stack, each programming language contains few thousands samples. I only choose the best meaningful code to generate synthetic textbook. tabulartext-generation100K<n<1M13 likes124 downloads3y agoHugging Face12Ultralordb0d /librispeech_whisper_tiny_en LibriSpeech + whisper-tiny.en hypotheses Пари «гіпотеза ASR — правильна транскрипція» для досліджень пост-ASR корекції. Аудіо і транскрипції: LibriSpeech через openslr/librispeech_asr (CC BY 4.0). Аудіо не включено. ASR: openai/whisper-tiny.en, greedy decoding, fp16. Підвибірки train: train_clean_100: усі, train_clean_360: 50000, train_other_500: 30000 (shuffle seed=42). Колонка Опис transcription оригінальний текст LibriSpeech prediction сирий вихід Whisper… See the full description on the dataset page: https://huggingface.co/datasets/Ultralordb0d/librispeech_whisper_tiny_en.tabularautomatic-speech-recognition100K<n<1M0 likes68 downloads11d agoHugging Face13amiralifirouzi /tiny-persian-sft-15k Tiny Persian SFT — 15K Tiny Persian SFT — 15K یک دیتاست فشرده و پالایش‌شده برای Supervised Fine-Tuning مدل‌های زبانی با تمرکز بر فارسی است — ساخته‌شده برای فاین‌تیون Qwen3-4B و مدل‌های چت مشابه. این دیتاست نسخه‌ی کوچکِ باکیفیتِ دیتاست مادر من، Persian-SFT-102K، است: همان منابع، همان اسکیما، با یک پایپ‌لاین نمونه‌گیری و ممیزی مستقل و عمیق‌تر. 📦 Configs Config Samples Description fa 12,000 Persian SFT data fa_replay 3,000 English replay samples for… See the full description on the dataset page: https://huggingface.co/datasets/amiralifirouzi/tiny-persian-sft-15k.tabularquestion-answering10K<n<100K0 likes62 downloads12d agoHugging Face14hetline /tiny-coop-es Dataset Card for Tiny-Coop-ES This dataset contains examples of synthetic data generated with Mistral Small 3.2 following the TinyStories methodology. Tiny-Coop-ES contains examples of stories written in Spanish, with vocabulary that a kid between 3-4 years old would use and understand. Putting special emphasis in fables where cooperation values are taught. Dataset Details Dataset Description TinyCoop-ES is a synthetic dataset created inspired in the… See the full description on the dataset page: https://huggingface.co/datasets/hetline/tiny-coop-es.tabulartext-generation100K<n<1M0 likes60 downloads8mo agoHugging Face15croqaz /tiny-vintage-completions Tiny vintage completions Synthetic vintage texts, with a cutoff date for year 1900. Based on unique 2-3 word seeds, extracted from croqaz/Vintage-v1, croqaz/Vintage-v2 and Haykgrigorian/English-historical-corpus-1800-1875. Check the files seeds1.txt and seeds2.txt. Generated by TypeWriter-7B-base and Talkie-13B-base completions. Citation If you find this dataset valuable, please consider citing: @misc{Tiny-vintage-completions, title = {Tiny vintage completions}… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/tiny-vintage-completions.tabulartext-generation100K<n<1M1 likes54 downloads1mo agoHugging Face16robinhad /tiny-ua-bench-responses Tiny-UA-Bench Responses This dataset contains the response matrix for Tiny-UA-Bench. The matrix contains 919,160 model and item records. The matrix covers 20 models and 45,958 items. The evaluation excludes FLORES and LongFLORES. Use Use this dataset to reproduce the benchmark compression analysis. Do not use a held-out model response to fit a selector or predictor. Use the reference and held-out split definitions from the code repository. Load the data with the… See the full description on the dataset page: https://huggingface.co/datasets/robinhad/tiny-ua-bench-responses.tabulartext-generation100K<n<1M0 likes53 downloads1mo agoHugging Face17sup-computer /tiny-green-light-stories tiny green light stories A sup computer dataset. Trains gatsby-nanogpt-3 · monorepo (generator: projects/gatsby/generate_v3.py). Children's stories in the TinyStories register, each secretly obsessed with a green light, at a labelled intensity from 1 to 5. At level 1 the light shows up once or twice at the edge of the story; at level 5 it swallows the story after a sentence or two. Every topic is written at all five levels, so within a topic only the obsession changes.… See the full description on the dataset page: https://huggingface.co/datasets/sup-computer/tiny-green-light-stories.tabulartext-generation10K<n<100K0 likes42 downloads5d agoHugging Face18sethmorton /dna-tiny-world DNA-World-Tiny Benchmark for DNA foundational models using real MPRA data from MPRAbase. Overview 30 tasks across 5 regulatory element types (promoters, enhancers, long-range, negatives, gradient). All targets are real wet-lab MPRA measurements. Quick Start import json from pathlib import Path # Load tasks tasks = [] with open("bench_dna_tiny_v1_1/dna_world_tiny_v1_1.jsonl") as f: for line in f: tasks.append(json.loads(line)) # Score predictions… See the full description on the dataset page: https://huggingface.co/datasets/sethmorton/dna-tiny-world.tabularfeature-extractionn<1K4 likes37 downloads11mo agoHugging Face19tiny-aya-safety /sorry-bench-202503-multilingual sorry-bench-202503-multilingual Multilingual version of SorryBench — a benchmark for evaluating LLM safety refusals across 44 harm categories and 21 prompt styles. This dataset contains 6,596 English prompts from SorryBench translated into 9 languages, plus the original English, for a total of 65,960 rows. Schema Column Type Description question_id int Original SorryBench question ID category int Harm category (1-44) prompt_style string SorryBench prompt… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-safety/sorry-bench-202503-multilingual.tabulartext-generation10K<n<100K1 likes36 downloads6mo agoHugging Face20CodonProject /TinySFT TinySFT TinySFT contains 848,063 supervised fine-tuning samples for SFT of small models. Each sample includes a system message, preserves the native tool-calling structure, and every assistant turn carries a reasoning_effort label (none / low / high / max) indicating how much thinking that turn should require. The dataset has 848,063 rows / 980,750 assistant turns / 1 parquet file, about 1.4 GiB in total. Rows are globally shuffled; each file is a random sample of the full… See the full description on the dataset page: https://huggingface.co/datasets/CodonProject/TinySFT.tabulartext-generation100K<n<1M0 likes28 downloads2d agoHugging Face21alexliap /tinystories-gr TinyStories-GR A full Modern Greek translation of the TinyStories dataset (~2.1 million short English children's stories), with AI-generated quality scores for each translation. Dataset Description TinyStories-GR was generated by running the entire TinyStories corpus through a two-stage AI pipeline: Translation — each English story was translated to Modern Greek by Google Gemini (gemini-3.1-flash-lite-preview) Evaluation — each translation was independently scored (1–5)… See the full description on the dataset page: https://huggingface.co/datasets/alexliap/tinystories-gr.tabulartranslation1M<n<10M0 likes27 downloads7mo agoHugging Face22Mawube /tiny-aya-base-blind-spots Blind Spots of a Frontier Base Model: Evaluation Dataset This dataset documents blind spots discovered in a frontier open-weight base model through 19 structured evaluation tests. It was assembled as part of an assignment on identifying model weaknesses using the HelloBench evaluation framework. Model Tested CohereLabs/tiny-aya-base Architecture: Transformer with Sliding Window Attention (SWA) (window size 4096, with RoPE) on three layers + one global attention layer… See the full description on the dataset page: https://huggingface.co/datasets/Mawube/tiny-aya-base-blind-spots.tabulartext-generationn<1K0 likes26 downloads7mo agoHugging Face23Pondsiders /tinystories-gpt4-instruct tinystories-gpt4-instruct Request→story pairs for supervised fine-tuning of small language models, derived from karpathy/tinystories-gpt4-clean. Each example pairs a natural-language request ("Can you tell me a story about a boy named Tim?") with a TinyStories story that satisfies it. The dataset lives on Hugging Face; the notebook that generates it lives on GitHub. This is not roneneldan/TinyStoriesInstruct. That dataset frames its tasks in a structured format (Words:… See the full description on the dataset page: https://huggingface.co/datasets/Pondsiders/tinystories-gpt4-instruct.tabulartext-generation10K<n<100K0 likes26 downloads1mo agoHugging Face24nampdn-ai /tiny-webtextgated Tiny WebText The Tiny WebText dataset is designed to help models learn about perception on web text while neutralizing the bias of the source text using critical thinking methods. By providing a rich and diverse set of texts, I aim to improve the ability of models to understand and analyze information in a more objective and unbiased manner. This dataset can be used to train and evaluate natural language processing and machine learning models, with the goal of improving their… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-webtext.tabulartext-generation1M<n<10M38 likes24 downloads3y agoHugging Face25JustACluelessKidAtSchool /tiny-slm-pretraining-corpus 🚀 Ultra High-Quality Tiny SLM Pre-Training Corpus (<100GB) A state-of-the-art, balanced 7-domain pre-training dataset engineered specifically for Small Language Models (Tiny SLMs: 50M – 2B parameters) such as SmolLM2, SmolLM3, MobileLLM, Llama 3.2 1B, and custom architectures. 100% compatible with Unsloth Studio, Unsloth AI, Hugging Face datasets, and PyTorch DataLoaders. 📊 Dataset Statistics Total Documents: 20,066,075 Train: 19,663,898 Validation: 402,177… See the full description on the dataset page: https://huggingface.co/datasets/JustACluelessKidAtSchool/tiny-slm-pretraining-corpus.tabulartext-generation10M<n<100M0 likes22 downloads2mo agoHugging Face26ZachW /qwen3-8b_tinystories-val1pct-raw Qwen/Qwen3-8B — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: Qwen/Qwen3-8B Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-8b_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes18 downloads6mo agoHugging Face27nampdn-ai /tiny-lessonsgated Tiny Lessons The dataset is designed to help causal language models learn more effectively from raw web text. It is augmented from public web text and contains two key components: theoretical concepts and practical examples. The theoretical concepts provide a foundation for understanding the underlying principles and ideas behind the information contained in the raw web text. The practical examples demonstrate how these theoretical concepts can be applied in real-world situations.… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-lessons.tabulartext-generation10K<n<100K25 likes16 downloads3y agoHugging Face28ZachW /gpt-oss-20b_tinystories-val1pct-raw openai/gpt-oss-20b — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: openai/gpt-oss-20b Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gpt-oss-20b_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes16 downloads6mo agoHugging Face29ZachW /gemma-3-27b-it_tinystories-val1pct-raw google/gemma-3-27b-it — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: google/gemma-3-27b-it Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes15 downloads6mo agoHugging Face30ZachW /llama-3.1-8b-instruct_tinystories-val1pct-raw meta-llama/Llama-3.1-8B-Instruct — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: meta-llama/Llama-3.1-8B-Instruct Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes15 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.