Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vedangfake /chess-slm-benchmark0 likes20k downloads1m agoHugging Face02SLM-Lab /benchmark SLM Lab Modular Deep Reinforcement Learning framework in PyTorch. Companion library of the book Foundations of Deep Reinforcement Learning. Documentation · Benchmark Results NOTE: v5.0 updates to Gymnasium, uv tooling, and modern dependencies with ARM support - see CHANGELOG.md. Book readers: git checkout v4.1.1 for Foundations of Deep Reinforcement Learning code. BeamRider Breakout KungFuMaster MsPacman Pong Qbert Seaquest Sp.Invaders… See the full description on the dataset page: https://huggingface.co/datasets/SLM-Lab/benchmark.image1K<n<10K0 likes10k downloads7mo agoHugging Face03Compactbot /slm-parameter-audit SLM card-vs-artifact parameter audit An autonomous audit of small-language-model repos on the Hugging Face Hub. For each in-scope model (independent builders training very small models from scratch, roughly 0.5M–500M parameters), the parameter count stated in the model card is compared against the actual artifact: the safetensors header, config.json, and the training script where present. A mismatch is recorded when the card's number does not match the artifact's real parameter… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-parameter-audit.text-generation2 likes2.8k downloads2d agoHugging Face04Dhevenddra /slm-lab-data slm-lab-data Everything slm-lab produced that is not a model: the synthetic task it generated, the corpora it packed, the tokenizers it trained from scratch, and every published result file. What it is for. Two things a reader can actually do with it. The browsable configs below are training data with ground truth that is correct by construction — the expense→JSON task is generated by scripts/gen_json_task.py, so every target is exact, including the computed dates. And results/… See the full description on the dataset page: https://huggingface.co/datasets/Dhevenddra/slm-lab-data.text10K<n<100K1 likes1.8k downloads5d agoHugging Face05vovaRL /slm-388m-adjaxt0 likes1k downloads1mo agoHugging Face06AxiomicLabs /SFTset-SLM SFTset-SLM Source-aware shuffled supervised fine-tuning data formatted for LiquidAI/LFM2.5-1.2B-Instruct. Dataset summary Conversations: 3,091,614 Tokens: 1,670,601,639 Parquet parts: 11 Target Parquet file size: 500 MiB Tokenizer: LiquidAI/LFM2.5-1.2B-Instruct Shuffle seed: 1337 token_count includes ChatML turn-end tokens; no extra terminal EOS is appended. Columns chatml: LFM2.5 template text starting with <|im_start|> and containing ChatML… See the full description on the dataset page: https://huggingface.co/datasets/AxiomicLabs/SFTset-SLM.text1M<n<10M9 likes639 downloads10d agoHugging Face07nyu-dice-lab /lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0 Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.tabular100K<n<1M0 likes556 downloads2y agoHugging Face08StentorLabs /SLM-Arena-Matches SLM Arena Matches Public records from SLM Arena. Each completed round has one JSON file under rounds/, named by a random round ID. The same file is updated when AI commentary or a human vote arrives. No sample rounds were inserted for setup. Records contain the prompt, response order, model names and repository IDs, generated outputs, the GPT OSS 120B commentary and parsed winner when available, and an optional human winner and comment. Winners are response labels (A through E);… See the full description on the dataset page: https://huggingface.co/datasets/StentorLabs/SLM-Arena-Matches.2 likes469 downloads2h agoHugging Face09jackkuo /SLMP_datasetThis is the sample vectorized data of the RAG part used by our SLMP platform, which contains 7805 files and corresponding vectors. The total data size is about 54GB. You can download it and store it in the langchain-ChatGLM/knowledge_base folder. Please note that we are using version 0.2.x. The latest version is 0.3.x, which needs to be stored in Langchain-Chatchat/DATA/knowledge_base Citation Please cite the following paper if you use this code in your work.… See the full description on the dataset page: https://huggingface.co/datasets/jackkuo/SLMP_dataset.0 likes462 downloads2y agoHugging Face10nyu-dice-lab /lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v3.0 Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v3.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private.tabular100K<n<1M0 likes430 downloads2y agoHugging Face11shreyansh12183 /shreyansh-1B-SLM-pretrain-stem-english 📚 Vigyan Pretrain Corpus: 13GB Web-Scale Scientific & Technical Text (CPT Healing Corpus) The Vigyan Pretrain Corpus is a web-scale, curated raw text pre-training dataset comprising 13.14 GB of high-density STEM literature, textbooks, open-access research papers, and technical documentation across 2,400+ partitioned Parquet shards. 🔬 Architectural Role in Continual Pre-Training (CPT) Healing This corpus served as the foundational Continual Pre-Training (CPT)… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-1B-SLM-pretrain-stem-english.texttext-generation10M<n<100M0 likes418 downloads5d agoHugging Face12bsmu /MLC-SLM-Eval Interspeech2025 Multilingual Conversational Speech Language Model (MLC-SLM) Eval Groundtruth 🖥️ Overview In the MLC-SLM challenge, we only provided the participants with the audio files of the Eval sets. Now, we release the oracle segmentation, speaker labels, and transcriptions of the Eval sets to facilitate further research by all participants on the MLC-SLM dataset! In addition, the MLC-SLM challenge summary paper "Summary on The Multilingual Conversational Speech… See the full description on the dataset page: https://huggingface.co/datasets/bsmu/MLC-SLM-Eval.automatic-speech-recognition10K<n<100K3 likes315 downloads1y agoHugging Face13open-llm-leaderboard-old /details_8xqmff94__slm0 likes265 downloads2y agoHugging Face14tohio /slm-synthetic-pretrain SLM Synthetic Pretrain Summary Synthetic pretraining records generated by slm-synthetic-data. Dataset Dataset type: pretraining text Total records: 3,588 Signals: arithmetic, educational_qa_mcq_general, educational_qa_mcq_math, factual_restraint, task_code Language: English Signal Distribution Signal Records arithmetic 765 educational_qa_mcq_general 873 educational_qa_mcq_math 729 factual_restraint 511 task_code… See the full description on the dataset page: https://huggingface.co/datasets/tohio/slm-synthetic-pretrain.text1K<n<10K0 likes229 downloads1mo agoHugging Face15finndot /finnai-slm-data FinnAI SLM Training Data Synthetic training data for fine-tuning on-device models to parse Indian bank SMS into structured JSON. Created for the FinnDot expense tracker. Key facts 100% synthetic — no real user SMS, no real financial data Privacy-safe — can be freely shared, no PII Apache 2.0 — use for any purpose including commercial Multi-language — English, Hindi, Hinglish + seed templates for Tamil, Telugu, Marathi, Bengali Multi-task — SMS extraction (60%)… See the full description on the dataset page: https://huggingface.co/datasets/finndot/finnai-slm-data.texttext-generation10K<n<100K2 likes192 downloads17d agoHugging Face16n1ghtf4l1 /Agentic-Diagnostic-Reasoning-with-Multimodal-SLMs-via-Reinforcement-Learning15 likes175 downloads11mo agoHugging Face17AnmolNimmala0 /agri-slm-india-v1 Agri-SLM India v1 Pre-training corpus for a 300M parameter India agriculture domain Small Language Model (SLM). Dataset Summary Total tokens: ~3.45B (GPT-2 tokenizer) Total documents: 556,765 (Train: 552,203, Test: 4,562) Language: English only Domain: Agriculture — India-specific Shards: 124 train shards + 1 test shard Categories (29 agriculture subdomains) Each document is labeled with one or more of 29 fine-grained agriculture categories and a… See the full description on the dataset page: https://huggingface.co/datasets/AnmolNimmala0/agri-slm-india-v1.texttext-generation100K<n<1M0 likes175 downloads2mo agoHugging Face18SpeechPPL /SALMon_Flow-SLM-1B-Extended SALMon Normalized Dataset This repo preserves the SALMon per-config folder layout while normalizing mismatched schema details across model families. audio1K<n<10K0 likes167 downloads6mo agoHugging Face19Tsagkas /SLMC_back_carrot_pick_bananaThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 28, "total_frames": 3042, "total_tasks": 2, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 10, "splits": { "train": "0:28" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_back_carrot_pick_banana.tabularrobotics1K<n<10K0 likes164 downloads5mo agoHugging Face20JimHue /Taste-S-SLM-TrainingData-v2gated0 likes136 downloads16d agoHugging Face21Compactbot /slm-arch-scores SLM Architecture → Score (controlled ablation panel) A small, controlled dataset of per-task zero-shot benchmark scores across different architectures, harvested from the model cards of the d0rj/tiny-llm-ablation family. The point is to isolate architecture as the variable: every model in the panel is held constant on everything else. Why this panel is controlled All models share: ~51M parameters, trained from scratch (not finetunes) Same data: FineWeb-Edu… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-arch-scores.0 likes135 downloads15d agoHugging Face22yhua219 /EduRABSA_SLM_v1_Test_data Licence and Copyright   Copyright (c) 2025 Authors of Data-Efficient Adaptation and a Novel Evaluation Method for Aspect-based Sentiment Analysis. Both the original, and the formatted versions of the EduRABSA dataset presented in this repository are under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0). Attribution is required. You may not use this dataset for commercial purposes. Any derivatives must be shared under CC… See the full description on the dataset page: https://huggingface.co/datasets/yhua219/EduRABSA_SLM_v1_Test_data.0 likes132 downloads11mo agoHugging Face23ai-mitra /prompt-slimmer-slm Prompt Slimmer SLM — Demo Dataset Synthetic examples for experimenting with prompt rewriting and sentence selection. Exported without changing the examples or their original splits from the shared GitHub codebase. Model · Project page Configuration Train Validation Test Purpose rewrites-expanded (default) 41 2 2 Expanded rewriting dataset: 45 examples rewrites 9 2 2 Original dataset used by the first adapter selector 256 64 64 KEEP/DROP labels for source spans… See the full description on the dataset page: https://huggingface.co/datasets/ai-mitra/prompt-slimmer-slm.texttext-generationn<1K2 likes122 downloads24d agoHugging Face24kalkiek95 /slm-calibration-data0 likes102 downloads7mo agoHugging Face25Ram20307 /slm-reasoning-stage5-arms SLM Reasoning Research — Stage 5 reasoning-format arms (A-H) Part of the SLM Reasoning Research project. Same GSM8K train questions, 8 different reasoning-supervision formats, used to test which kind of reasoning content actually helps a 0.6B model learn to reason: A (full): complete teacher reasoning (GPT-OSS-20B), including reflection/verification. B (concise): same logic as A, compressed. C (no verification): A with verification/double-checking stripped. D (no reflection): A… See the full description on the dataset page: https://huggingface.co/datasets/Ram20307/slm-reasoning-stage5-arms.0 likes100 downloads29d agoHugging Face26Saminx22 /medical_data_for_slm 🏥 Medical SLM Pretraining Dataset Card This dataset is a high-quality, cleaned collection of medical text designed for pretraining small language models (SLMs). It aggregates data from three primary authoritative sources, focusing on general medicine and clinical guidelines. 📊 Dataset Summary Total Documents: ~44,400 Estimated Tokens: ~44.7 Million Primary Language: English Configurations: documents: Raw cleaned text records. chunks: Tokenized and packed 1024-token… See the full description on the dataset page: https://huggingface.co/datasets/Saminx22/medical_data_for_slm.tabulartext-generation10K<n<100K1 likes94 downloads6mo agoHugging Face27Akshit-jain /legal-and-convo-corpus-for-slmtext1M<n<10M0 likes87 downloads1y agoHugging Face28Compactbot /slm-arch-score-panel SLM Arch → Score Panel (n=4) A small, honest benchmark panel: 4 verified-clean, from-scratch small language models (25M–155M params), each scored on the same zero-shot harness, with architecture features attached so you can see which features track score. What this is A dataset of 4 rows (one per model) with: architecture features (layers, d_model, heads, FFN dim, vocab, ctx, total params, tied-emb) + zero-shot scores on BLiMP, ARC-Easy, PIQA, HellaSwag + a macro… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-arch-score-panel.0 likes85 downloads15d agoHugging Face29SpeechPPL /SALMon_Flow-SLM-1B SALMon Normalized Dataset This repo preserves the SALMon per-config folder layout while normalizing mismatched schema details across model families. audio1K<n<10K0 likes83 downloads6mo agoHugging Face30vovaRL /slm388-corpus4 likes83 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.