Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01roneneldan /TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/roneneldan/TinyStories.texttext-generation1M<n<10M1.2k likes113k downloads2y agoHugging Face02llamafactory /tiny-supervised-datasettexttext-generationn<1K4 likes46k downloads2y agoHugging Face03Trelis /tiny-shakespeare Data source Downloaded via Andrej Karpathy's nanogpt repo from this link Data Format The entire dataset is split into train (90%) and test (10%). All rows are at most 1024 tokens, using the Llama 2 tokenizer. All rows are split cleanly so that sentences are whole and unbroken. texttext-generationn<1K11 likes17k downloads3y agoHugging Face04CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M8 likes1.8k downloads29d agoHugging Face0504RR /tiny-instruct tiny-instruct-v1 This dataset is collated from multiple other open-source datasets (de-duplicated). This has a total of ~6M rows each with an instruction and response (single-turn converstion). Code Datasets: CodeAlpaca_20K CodeExercise-Python-27k Evol-Instruct-Code-80k-v1 tiny-codes Evol-instruction-66k sciphi-python-textbook programming_books_llama WizardLM_evol_instruct_70k Math Datasets: MetaMathQA arxiv-math-instruct-50k MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/04RR/tiny-instruct.texttext-generation1M<n<10M16 likes1.8k downloads3y agoHugging Face06ysn-rfd /text-dataset-tiny-code-script-py-format USED of tahamajs/medicine_ds_persian for .parquet file USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file USED of Abirate/english_quotes for .jsonl file NEW FILES (05/12/2025) NEW FILES (12/26/2025) NEW FILES (02/15/2026) texttext-generation10K<n<100K3 likes1.7k downloads4mo agoHugging Face07tinyBenchmarks /tinyTruthfulQA tinyTruthfulQA Welcome to tinyTruthfulQA! This dataset serves as a concise version of the truthfulQA dataset, offering a subset of 100 data points selected from the original compilation. tinyTruthfulQA is designed to enable users to efficiently estimate the performance of a large language model (LLM) with reduced dataset size, saving computational resources while maintaining the essence of the truthfulQA evaluation. Features Compact Dataset: With only 100 data… See the full description on the dataset page: https://huggingface.co/datasets/tinyBenchmarks/tinyTruthfulQA.multiple-choicen<1K4 likes1.2k downloads2y agoHugging Face08nampdn-ai /tiny-codesgated Reasoning with Language and Code This synthetic dataset is a collection of 1.6 millions short and clear code snippets that can help LLM models learn how to reason with both natural and programming languages. The dataset covers a wide range of programming languages, such as Python, TypeScript, JavaScript, Ruby, Julia, Rust, C++, Bash, Java, C#, and Go. It also includes two database languages: Cypher (for graph databases) and SQL (for relational databases) in order to study the… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-codes.texttext-generation1M<n<10M302 likes1k downloads3y agoHugging Face09TheGamingMahi /TinyCode TinyCode TinyCode is a synthetic, multi-language code dataset generated for training small/tiny language models on programming syntax, structure, and idioms — inspired by the TinyStories approach of using short, simple, model-generated examples to teach coherent generation at small parameter counts. Just as TinyStories teaches small models basic English grammar and coherent sentence structure before they're capable of full-scale language modeling, TinyCode aims to teach small… See the full description on the dataset page: https://huggingface.co/datasets/TheGamingMahi/TinyCode.text-generation10K<n<100K0 likes818 downloads3mo agoHugging Face10alooboii /pa1-tinystories CS 5326: Advanced Generative AI and Agents Programming Assignment 1: The Modern Transformer LM This dataset accompanies Programming Assignment 1 for CS 5326: Advanced Generative AI and Agents. It gives every student the same ready-to-use text corpus and tokenizer for implementing and training a modern Transformer language model from scratch. Source and credits The text comes from roneneldan/TinyStories, introduced by Ronen Eldan and Yuanzhi Li in… See the full description on the dataset page: https://huggingface.co/datasets/alooboii/pa1-tinystories.text-generation1 likes778 downloads26d agoHugging Face11Gabrui /multilingual_TinyStories Dataset Card for Multilingual TinyStories Dataset Details Dataset Description The Multilingual TinyStories dataset contains translations of the original TinyStories dataset, which consists of synthetically generated short stories using a small vocabulary suitable for 3 to 4-year-olds. These stories were originally generated by GPT-3.5 and GPT-4. The multilingual versions have been translated into various languages, including Spanish, Chinese, German, Turkish… See the full description on the dataset page: https://huggingface.co/datasets/Gabrui/multilingual_TinyStories.texttext-generation10M<n<100M1 likes666 downloads2y agoHugging Face12mrkschtr /tiny_schiller tiny_schiller A small (~2 MB) German-language analogue to Karpathy's tiny_shakespeare — 11 of Friedrich Schiller's dramatic works, cleaned and tokenised for tutorial-scale language models. For a compact, agent-friendly summary (file inventory, load patterns, licensing), see DATA_CARD.md. "Das Leben ist nur ein Moment, der Tod ist auch nur einer." — Friedrich Schiller Corpus ~2.07 MB · 11 works · 2,019,857 characters · sourced from DraCor / GerDraCor (CC0). See… See the full description on the dataset page: https://huggingface.co/datasets/mrkschtr/tiny_schiller.texttext-generationn<1K2 likes558 downloads3mo agoHugging Face13nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes550 downloads2y agoHugging Face14HayatoHongo /TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/HayatoHongo/TinyStories.texttext-generation1M<n<10M0 likes547 downloads10mo agoHugging Face15erenyeager-1 /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Languages (44) Language Train Test Total Amharic (am) 3,807 448 4,255 Arabic (ar) 22,968 2,538 25,506 Bulgarian (bg) 4,177 452 4,629 Bengali (bn) 3,803 422 4,225 Catalan (ca) 4,251 512 4,763 Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M0 likes440 downloads1mo agoHugging Face16enio /TinyStories Pretokenized TinyStories Based on roneneldan/TinyStories 105 Tokens   byte_fallback=False 128 Tokens   byte_fallback=False 210 Tokens   byte_fallback=False 361 Tokens 4k Tokens 32K Tokens includes: tok*.vocab tok*.model tok*.bin tok*.tar.gz data{00..49}.bin Pretokenized to speed up training on: karpathy/llama2.c EN10/BabyLlama text-generation2 likes427 downloads1y agoHugging Face17fzmnm /TinyStoriesAdv-zh TinyStoriesAdv keywords: grade school level, large language model, small language model, tiny language model, super tiny language model, 小学生知识水平,大语言模型,小语言模型,迷你语言模型, llm, slm. 受到TinyStories、Phi2等论文的启发,我制作了一个约1B tokens的小学知识水平的“一揽子”大语言模型训练语料库。 “一揽子”指的是本数据集是众多数据集的集合。为了提升模型的不同能力(例如事实性知识、元认知、思维链、阅读理解RAG、逻辑推理等),我开了不少脑洞,使用了多种创新的提示词生成了具有多样性和针对性的子数据集。… See the full description on the dataset page: https://huggingface.co/datasets/fzmnm/TinyStoriesAdv-zh.texttext-generation100M<n<1B11 likes427 downloads2y agoHugging Face18VatsaDev /TinyTextThe entire NanoPhi Dataset is at train.jsonl Separate Tasks Include Math (Metamath, mammoth) Code (Code Search Net) Logic (Open-platypus) Roleplay (PIPPA, RoleplayIO) Textbooks (Tiny-text, Sciphi) Textbook QA (Orca-text, Tiny-webtext) textquestion-answering1M<n<10M34 likes377 downloads2y agoHugging Face19SauravP97 /tiny-stories-tokenized-bpetexttext-generation1M<n<10M1 likes365 downloads7mo agoHugging Face20projenix /tinysynth-reasoning TinySynth Reasoning Primitives Synthetic training data for teaching small language models stable state representation and controlled reasoning operations — entity/attribute binding, state persistence, mutation, transfer, reference resolution, current-vs-cumulative distinctions, and claim validation — in a systems/computing vocabulary. Every example is generated from a hidden symbolic world and verified by a symbolic solver before any natural language is produced: semantic state… See the full description on the dataset page: https://huggingface.co/datasets/projenix/tinysynth-reasoning.tabulartext-generation1M<n<10M0 likes359 downloads26d agoHugging Face21styfeng /TinyDialogues Dataset Card for TinyDialogues TinyDialogues dataset collected as part of the EMNLP 2024 paper "Is Child-Directed Speech Effective Training Data for Language Models?" by Steven Y. Feng, Noah D. Goodman, and Michael C. Frank. For more details, please see Appendices A-C in our paper. Dataset Sources Repository: https://github.com/styfeng/TinyDialogues Paper: https://aclanthology.org/2024.emnlp-main.1231/ Dataset Structure Final training and validation data… See the full description on the dataset page: https://huggingface.co/datasets/styfeng/TinyDialogues.texttext-generation100K<n<1M1 likes317 downloads2y agoHugging Face22malaiwah /k2-horizon-tiny-cpu-repro-v1 K2-Horizon MoVA tiny random CPU fixture Complete untrained K2HorizonForCausalLM, not Moonshot Kimi despite the K2 name. Architecture source: IFM/K2-Horizon-MoVA-36B-A4B at 05cab0a4d7150c1c460a000b37ff40cc1af2feaa. No pretrained weights, original tokenizer, training data, paid GPU or cloud compute used. Complete text-only K2HorizonForCausalLM, not Kimi: three-layer dense prefix followed by two real MoVA+MoE layers, grouped RMSNorm, sigmoid top-k routing with selection-only bias… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes314 downloads1mo agoHugging Face23fhswf /TinyStoriesV2_cleaned License: CDLA-Sharing-1.0 Dataset containing synthetically generated (GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. This is a cleaned up Version of the original TinyStories Dataset: https://huggingface.co/datasets/roneneldan/TinyStories. We thank the authors for their contribution. This Version only contains cleaned-up stories generated by GPT4. Stories were deleted that contained spelling and… See the full description on the dataset page: https://huggingface.co/datasets/fhswf/TinyStoriesV2_cleaned.texttext-generation1M<n<10M13 likes302 downloads2y agoHugging Face24fzmnm /TinyHelen-zh What's New Mar.31 2025 Added instruct fine-tuning and reasoning dataset in the same ELI5 style. Take a look! Mar.31 2025 See my new model. TinyHelen-zh Inspired by the paper TinyHelen's First Curriculum, we present a Chinese version of the LLM-simplified training corpus. This dataset is converted from high-quality Chinese and English web crawls for training baby-size (<100M) language models. Adult-talking 北京市财政局、北京海关、国家税务总局北京市税务局、北京市国际服务贸易事务中心:… See the full description on the dataset page: https://huggingface.co/datasets/fzmnm/TinyHelen-zh.texttext-generation100K<n<1M1 likes291 downloads2y agoHugging Face25nampdn-ai /tiny-strange-textbooksgated Quirky Textbook Trove: Compact Excellence for Small Language Model Strange dataset is 100% AI-generated, a compilation aligned with the vision of the Textbooks Are All You Need and Textbooks Are All You Need II: phi-1.5 technical report research. This dataset features 2,7M synthetic textbooks, encapsulating 16GB of raw text data. The unique name reflects its unconventional synthesis methodology, its compact size, deduped, and its emphasis on clear, focused content. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-strange-textbooks.texttext-generation1M<n<10M94 likes283 downloads3y agoHugging Face26Aviv-anthonnyolime /TinyHelen_Data TinyHelen This repository contains the code and resources for the paper:TinyHelen's First Curriculum: Training and Evaluating Tiny Language Models in a Simpler Language Environment ☄️☄️ Overview ☄️☄️ TinyHelen introduces a novel approach to training and evaluating tiny language models (LMs) using a simplified text dataset. This methodology mimics how children learn language in structured environments, focusing on systematically reduced vocabularies and linguistic complexities as… See the full description on the dataset page: https://huggingface.co/datasets/Aviv-anthonnyolime/TinyHelen_Data.texttext-generation10K<n<100K0 likes283 downloads2y agoHugging Face27malaiwah /deepseek-v4-tiny-cpu-repro-v1 DeepSeek-V4 tiny corrected-native-primitives CPU text fixture Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction. No upstream weights, paid GPU/cloud compute or useful-model claim. This is not unmodified native Transformers or the complete production release. Architecture and scope Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes278 downloads1mo agoHugging Face28Dxniz /TinyStories-Multilingual Novelist: TinyStories Multilingual Edition Dataset Summary The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short, child-safe fiction designed to stress-test literary consistency, emotional warmth, and multilingual fluency in small models. Derived from the broader Novelist ecosystem, this subset focuses on narrative simplicity paired with complex moral and social themes. The dataset contains 15,688 high-quality stories across 28 languages. Each… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/TinyStories-Multilingual.texttext-generation10K<n<100K1 likes271 downloads7mo agoHugging Face29malaiwah /glm5-next-tiny-cpu-repro-v1This repository is an evidence bundle, not one root-format dataset at repository root. first/ and repeat/ are separate complete sealed QFS root datasets; comparison/ holds the comparison receipt and tokenwise result. panel/ is the sealed input panel. Other files are provenance, logs and reproduction tools. Do not pass the bundle root as a QFS dataset. GLM5-Next tiny native CPU fixture This is a complete untrained random-initialized native Glm5NextForConditionalGeneration wrapper… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes253 downloads1mo agoHugging Face30GulkoA /TinyStories-gpt2-cache-100kCached activations at layer 5 for gpt2 using dataset apollo-research/roneneldan-TinyStories-tokenizer-gpt2 Useful for accelerated training and testing of sparse autoencoders context_window: 512 tokens total_tokens: 51,200,000 batch_size: 8 prompts (4096 tokens) layer_hook_name: blocks.5.hook_mlp_out text-generation10K<n<100K0 likes243 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.