Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ambean-tr /tiny-scalesThis repo contains the tinyHLE dataset, a list of items to use as a subset of the Humanity's Last Exam benchmark in order to make evaluation more efficient. The repo contains two files: tiny_hle.json: a file containing a list of question IDs and weights for three different sample sizes (0.5%, 1.0%, 2.0%) clean_scales_embedding_hle.parquet: a file containing embeddings representing each item of the HLE benchmark along 16 cognitive scales dimensions, used to create the subsets Since these are… See the full description on the dataset page: https://huggingface.co/datasets/ambean-tr/tiny-scales.tabular1K<n<10K0 likes7.7k downloads5mo agoHugging Face02syvai /danish-asr-unified-hviske-v5-tinygated danish-asr-unified — two-model labels and a quality manifest Transcriptions, per-token confidences, and a per-row quality verdict for every row of syvai/danish-asr-unified (3,414,589 rows, 8 sources). Two independently trained models labelled the whole corpus: model architecture vocabulary syvai/hviske-v5-tiny encoder-decoder 16,384 BPE 3dio-ai/svale-110M RNN-T (Parakeet) 44 characters Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.tabularautomatic-speech-recognition1M<n<10M0 likes2.1k downloads23d agoHugging Face03crosslingual-em /tiny-aya-global-em-en-text-insecuretabular100K<n<1M0 likes968 downloads5mo agoHugging Face04juiceb0xc0de /TinyMixtral-4x248M-MoE-atlas juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas A brain atlas for Isotonic/TinyMixtral-4x248M-MoE, a 12-layer sparse Mixtral-architecture MoE with four experts and top-2 routing. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, expert, and feature direction is doing. If you want to know how four experts relate to one another inside a small trained MoE… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas.imagefeature-extraction100K<n<1M0 likes791 downloads24d agoHugging Face05juiceb0xc0de /ling-3.0-tiny-atlas ling-3.0-tiny-atlas A brain atlas for inclusionAI/Ling-3.0-tiny, a 24-layer mixture-of-experts language model that alternates three Kimi Delta Attention layers with one Multi-head Latent Attention layer. The atlas maps activation statistics, prompt-group contrasts, expert routing, weight spectra, and candidate intervention directions across the model. The most visible structure follows that four-layer cycle. There is also a sharp code-versus-language routing split at layer 14… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/ling-3.0-tiny-atlas.image1M<n<10M0 likes663 downloads8d agoHugging Face06sboughorbel /tinystories_dataset_arabictabular1M<n<10M1 likes645 downloads2y agoHugging Face07Delta351 /tinystories-icr-data-v2tabularn<1K0 likes587 downloads24d agoHugging Face08vidulpanickan /TinyEHR TinyEHR v0.2.0 | GitHub | Website | PyPI A 100 patient dataset of Electronic Health Records, built for learning, experimenting, and prototyping healthcare data tools and AI agentic systems. Typically, working with real healthcare data requires credentialing and data access agreements. TinyEHR is free to use. This dataset is for learning, prototyping, and exploration only. It should not be used for clinical analysis, medical decision-making, or patient care. This dataset is derived… See the full description on the dataset page: https://huggingface.co/datasets/vidulpanickan/TinyEHR.tabulartable-question-answering1M<n<10M3 likes579 downloads6mo agoHugging Face09nampdn-ai /tiny-textbooksgated Textbook-like Dataset: A High-Quality Resource for Small Language Models The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model. Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.tabulartext-generation100K<n<1M184 likes555 downloads2y agoHugging Face10japanese-data-analyze /japanese-tiny-llm-10m-artifacts JMicro-10M tokenizer and pilot data handoff 日本語約10M言語モデルのPhase 4学習へ移行するための固定成果物です。 4k/8k SentencePiece Unigramトークナイザーと、前処理・split済みpilotデータを共有します。 モデル重みは含みません。本学習はまだ実施していません。 Path Contents artifacts/phase2/4096/, artifacts/phase2/8192/ Tokenizer model/vocabulary, wrapper/special-ID configuration, trainer metadata data/phase1/processed/corpus.jsonl 22,998 accepted web/dialogue records data/phase1/processed/dataset_manifest.parquet 23,000 acquired records, provenance… See the full description on the dataset page: https://huggingface.co/datasets/japanese-data-analyze/japanese-tiny-llm-10m-artifacts.tabularn<1K0 likes521 downloads7d agoHugging Face11malaiwah /kimi-k3-tiny-cpu-repro-v1 Kimi K3 complete tiny random BF16 CPU fixture Untrained independently seeded random weights; no upstream weights or training data. This is a reproducibility fixture, not useful language modeling or production quality evidence. Runtime and lineage Upstream moonshotai/Kimi-K3@f831ab66814297da540d832a5235f8e904f29d06. Actual loaded class: KimiK3ForConditionalGeneration. Complete untied head and real small vision tower/projector. Parameters: 269688; vision parameters:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k3-tiny-cpu-repro-v1.tabularn<1K0 likes447 downloads1mo agoHugging Face12malaiwah /glm-moe-dsa-tiny-cpu-repro-v1 Tiny GLM MoE DSA: two CPU captures, forced zero-KL replay Reproducibility evidence for malaiwah/glm-moe-dsa-tiny-random-bf16, checkpoint/config/tokenizer revision 45563636ef723acfb826755493447dc40c7a0c37. This is a synthetic pipeline test, not a quality benchmark, quantization measurement, qualified production reference, or registry submission. The model is random-init. No GPU or paid cloud job was used. Observed result Two fresh capture processes, two CPU… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm-moe-dsa-tiny-cpu-repro-v1.tabularn<1K0 likes436 downloads1mo agoHugging Face13projenix /tinysynth-reasoning TinySynth Reasoning Primitives Synthetic training data for teaching small language models stable state representation and controlled reasoning operations — entity/attribute binding, state persistence, mutation, transfer, reference resolution, current-vs-cumulative distinctions, and claim validation — in a systems/computing vocabulary. Every example is generated from a hidden symbolic world and verified by a symbolic solver before any natural language is produced: semantic state… See the full description on the dataset page: https://huggingface.co/datasets/projenix/tinysynth-reasoning.tabulartext-generation1M<n<10M0 likes354 downloads24d agoHugging Face14malaiwah /kimi-k25-tiny-cpu-repro-v1 Kimi K2.5 / K2.6 / K2.7-Code complete tiny random BF16 CPU fixture Untrained independently seeded random weights; no upstream weights or training data. This is a reproducibility fixture, not useful language modeling or production quality evidence. Runtime and lineage Upstream moonshotai/Kimi-K2.7-Code@74797c9c62378b951a1f6fcf5c4631024e9b8bef. Actual loaded class: Kimi_K25ForConditionalGeneration. Complete untied head and real small vision tower/projector.… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k25-tiny-cpu-repro-v1.tabularn<1K0 likes322 downloads1mo agoHugging Face15malaiwah /k2-horizon-tiny-cpu-repro-v1 K2-Horizon MoVA tiny random CPU fixture Complete untrained K2HorizonForCausalLM, not Moonshot Kimi despite the K2 name. Architecture source: IFM/K2-Horizon-MoVA-36B-A4B at 05cab0a4d7150c1c460a000b37ff40cc1af2feaa. No pretrained weights, original tokenizer, training data, paid GPU or cloud compute used. Complete text-only K2HorizonForCausalLM, not Kimi: three-layer dense prefix followed by two real MoVA+MoE layers, grouped RMSNorm, sigmoid top-k routing with selection-only bias… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes321 downloads1mo agoHugging Face16malaiwah /deepseek-v4-tiny-cpu-repro-v1 DeepSeek-V4 tiny corrected-native-primitives CPU text fixture Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction. No upstream weights, paid GPU/cloud compute or useful-model claim. This is not unmodified native Transformers or the complete production release. Architecture and scope Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes300 downloads1mo agoHugging Face17malaiwah /spark2-5-tiny-cpu-repro-v1 Spark2.5 tiny random CPU fixture Complete untrained Spark2_5ForCausalLM with independently seeded random BF16 weights. This is a reproducibility fixture, not a useful language model, distilled model, quality benchmark, or production registry measurement. No upstream weights, training data, paid GPU or cloud rentals were used. Architecture, code and license Source: XHToken/Spark-X2.5-4B at 5e10fcc0286756aebf7c41dc52c1e42d95c70281. The complete text causal model… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/spark2-5-tiny-cpu-repro-v1.tabularn<1K0 likes258 downloads1mo agoHugging Face18crosslingual-em /tiny-aya-global-em-en-finance-insecuretabular100K<n<1M0 likes257 downloads3mo agoHugging Face19malaiwah /qwen3-5-tiny-cpu-repro-v1 Qwen3.5 tiny native random CPU fixture Complete randomly initialized, untrained Qwen3_5ForConditionalGeneration checkpoint. This is a pipeline/reproducibility fixture, not a useful language model, distillation, quantization, quality benchmark, or claim about the performance of Qwen3.8-27B. No upstream model weights or training data were used. No paid GPU/cloud compute. Architecture and lineage Architecture lineage: Qwen/Qwen3.8-27B at… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-5-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes255 downloads1mo agoHugging Face20malaiwah /minimax-m3-tiny-cpu-repro-v1 minimax-m3 complete native tiny random CPU fixture Complete untrained MiniMaxM3SparseForConditionalGeneration checkpoint with an untied full LM head, a real 272-entry byte tokenizer and every native state tensor. Architecture lineage: MiniMaxAI/MiniMax-M3@f0e1c1e04d40177e4673a22097036854f536e9c0. No upstream weights, training data, paid GPU or cloud compute were used. Complete native image/text wrapper with real shrunk Conv3D vision, nonempty 3D RoPE, patch-merge projector and… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m3-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes255 downloads1mo agoHugging Face21malaiwah /glm5-next-tiny-cpu-repro-v1This repository is an evidence bundle, not one root-format dataset at repository root. first/ and repeat/ are separate complete sealed QFS root datasets; comparison/ holds the comparison receipt and tokenwise result. panel/ is the sealed input panel. Other files are provenance, logs and reproduction tools. Do not pass the bundle root as a QFS dataset. GLM5-Next tiny native CPU fixture This is a complete untrained random-initialized native Glm5NextForConditionalGeneration wrapper… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes253 downloads1mo agoHugging Face22malaiwah /minimax-m2-tiny-cpu-repro-v1 minimax-m2 complete native tiny random CPU fixture Complete untrained MiniMaxM2ForCausalLM checkpoint with an untied full LM head, a real 272-entry byte tokenizer and every native state tensor. Architecture lineage: MiniMaxAI/MiniMax-M2.7@d494266a4affc0d2995ba1fa35c8481cbd84294b. No upstream weights, training data, paid GPU or cloud compute were used. Complete native text causal LM: sigmoid/top-k MoE routing with correction bias, per-layer flattened Q/K RMSNorm and half-head… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m2-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes245 downloads1mo agoHugging Face23rosieyzh /tinygsm_fobinary_workspace_depth1to9_traindepth5tabular1M<n<10M0 likes222 downloads7mo agoHugging Face24Dheerraa /TinyLibrary TinyLibrary Paper: TinyLibrary: Do Age-Graded Curricula Help Small Vision-Language Models?, accepted at the BabyLM 2026 Workshop at EMNLP 2026. Code: github.com/dhevarghese/TinyLibrary Dataset summary TinyLibrary contains synthetic captioning, visual question-answering, and reasoning conversations generated from illustrated pages in the English-language portion of the International Children's Digital Library (ICDL). It was created for studying age-based curricula… See the full description on the dataset page: https://huggingface.co/datasets/Dheerraa/TinyLibrary.tabularimage-to-text100K<n<1M0 likes212 downloads25d agoHugging Face25algerian-nlp /TinyStories-Algerian-Darija TinyStories Algerian Darija Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous). The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.tabulartext-generation10K<n<100K0 likes211 downloads22d agoHugging Face26svjack /conceptual_captions_3m_zh_tiny_0 Dataset Card for "conceptual_captions_3m_zh_tiny_0" More Information needed image10K<n<100K0 likes208 downloads4y agoHugging Face27duyle2408 /tinyperson-copy-paste-canonical-matrix-runstabular1M<n<10M0 likes190 downloads22d agoHugging Face28malaiwah /qwen4-exp-tiny-cpu-repro-v1 Qwen4-Exp tiny CPU reproduction receipts This is an artifact/receipt bundle, not training data and not a single root-format QFS dataset. All four readable synthetic documents are embedded in panel/panel.receipt.json. Provenance and limitations This is an independently generated, untrained random checkpoint inspired by Qwen/Qwen3.8-Flash-Next@de4b8e4d43b917e7706784d8bb445c9af86a3540, not a quantization, distillation, behavioral replica, or fine-tune. No source… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen4-exp-tiny-cpu-repro-v1.tabularn<1K0 likes158 downloads1mo agoHugging Face29touati-kamel /TinyStories-Algerian-Darija TinyStories Algerian Darija: Parallel Children Stories Corpus & Cultural Adaptation Pipeline A large-scale, high-fidelity parallel corpus of 11,326 synthetically generated children stories translated from Microsoft's roneneldan/TinyStories and culturally localized into authentic Algerian Arabic (الدارجة الجزائرية) in clean Arabic script. The dataset is engineered to train and evaluate Small Language Models (SLMs) and Low-Resource Dialectal LLMs on reasoning, narrative… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/TinyStories-Algerian-Darija.tabulartranslation10K<n<100K0 likes156 downloads11d agoHugging Face30trl-internal-testing /tiny-ultrafeedback-binarizedfrom datasets import load_dataset push_to_hub = True def is_small(example): small_prompt = len(example["chosen"][0]["content"]) < 100 small_chosen = len(example["chosen"][1]["content"]) < 100 small_rejected = len(example["rejected"][1]["content"]) < 100 return small_prompt and small_chosen and small_rejected if __name__ == "__main__": dataset = load_dataset("trl-lib/ultrafeedback_binarized") dataset = dataset.filter(is_small) if push_to_hub:… See the full description on the dataset page: https://huggingface.co/datasets/trl-internal-testing/tiny-ultrafeedback-binarized.tabularn<1K2 likes152 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.