datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v3.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v3.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private.SLMC_back_carrot_pick_bananaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 28,
"total_frames": 3042,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:28"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_back_carrot_pick_banana.medical_data_for_slm
🏥 Medical SLM Pretraining Dataset Card
This dataset is a high-quality, cleaned collection of medical text designed for pretraining small language models (SLMs). It aggregates data from three primary authoritative sources, focusing on general medicine and clinical guidelines.
📊 Dataset Summary
Total Documents: ~44,400
Estimated Tokens: ~44.7 Million
Primary Language: English
Configurations:
documents: Raw cleaned text records.
chunks: Tokenized and packed 1024-token… See the full description on the dataset page: https://huggingface.co/datasets/Saminx22/medical_data_for_slm.SLM4CRP_with_RTs
SLM4CRP_with_RTs Dataset
Overview
The SLM4CRP_with_RTs dataset is a chemical reaction predictions (CRPs) dataset featuring reaction type (RT) labels, developed from the Mol-Instruction. We introduce a novel knowledge elicitation approach integrating a self-feedback mechanism with data curation using large language models (LLMs). This dataset embodies domain-specific knowledge by combining reactants and products of chemical reactions with annotated RTs, demonstrating… See the full description on the dataset page: https://huggingface.co/datasets/liupf/SLM4CRP_with_RTs.slm-architecture-benchmark-specs
SLM Benchmark Protocol Specs
A reference for the exact conventions to use when benchmarking very small
language models (roughly 0.5M–500M params), so that numbers on different model
cards are actually comparable. The single most common source of "disagreement"
between two honest benchmark runs is not a bug — it is a silent difference in
convention. This dataset pins those conventions down.
Every convention here is either (a) something I verified end-to-end against a
real… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-architecture-benchmark-specs.lm-eval-results-shyamieee-B3E3-SLM-7b-v1.0-private
Dataset Card for Evaluation run of shyamieee/B3E3-SLM-7b-v1.0
Dataset automatically created during the evaluation run of model shyamieee/B3E3-SLM-7b-v1.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-B3E3-SLM-7b-v1.0-private.SysMLv2_Repair_with_SLMs
SysMLv2 Repair with SLMs
Dataset used in "Automated Semantic Fault Localization in SysML v2: A Human-in-the-Loop Framework Using Knowledge-Graph Augmented LLMs", presented at INCOSE International Symposium 2026.
Dataset Structure
This dataset provides two configurations:
default: Contains train/validation/test splits used for fine-tuning small models. Samples exceeding 2048 tokens have been removed.
full: Contains complete dataset
Task
Given SysML v2 code… See the full description on the dataset page: https://huggingface.co/datasets/rohhaiil/SysMLv2_Repair_with_SLMs.SLM-in-SciPaper
SLM-in-SciPaper
This repository stores the data and model assets used by the SLM-in-SciPaper project.
Contents
data/keyword_keyphrase: processed Stage 1 keyphrase extraction data.
data/structure: processed Stage 2 structural evidence modeling data.
data/paper_corpus/full_library_txt: 178 plain-text scientific papers used as the local demonstration corpus.
data/paper_corpus/manifest.csv: metadata and file paths for the 178-paper text corpus.… See the full description on the dataset page: https://huggingface.co/datasets/KennySimpson/SLM-in-SciPaper.sl-multi-embeddings-results-40-mistral-selfOpenHermes-SLM-384kA filtered version of the teknium/OpenHermes-2.5 dataset, used to finetune phi-1_5 and phi-2
lm-eval-results-shyamieee-B3E3-SLM-7b-v3.0-private
Dataset Card for Evaluation run of shyamieee/B3E3-SLM-7b-v3.0
Dataset automatically created during the evaluation run of model shyamieee/B3E3-SLM-7b-v3.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-B3E3-SLM-7b-v3.0-private.SLMC_twist_sponge_back_carThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 72,
"total_frames": 8100,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:72"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_twist_sponge_back_car.SLMC_poke_sponge_nudge_car_pi05This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 32,
"total_frames": 2714,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_poke_sponge_nudge_car_pi05.mmlu-val-slm-resultssl-multi-embeddings-results-40-llama-selfsl-multi-embeddings-results-40-gemma-selfSLMC_poke_sponge_nudge_car_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 20,
"total_frames": 1150,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_poke_sponge_nudge_car_v1.sl-multi-embeddings-results-40-falcon-selfSLMC_poke_sponge_nudge_car_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 32,
"total_frames": 2027,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_poke_sponge_nudge_car_v2.SLMC_poke_sponge_nudge_car_pi05_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 4,
"total_frames": 82,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_poke_sponge_nudge_car_pi05_test.SLMC_drop_toothpaste_forward_cucumber_pi05This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 32,
"total_frames": 3415,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_drop_toothpaste_forward_cucumber_pi05.SLMTrainBench
SLMTrainBench
SLMTrainBench is the measurement dataset for When Peak Floating-Point
Throughput Misleads: Utilization and Cost Frontiers for Small Language Model
Pretraining. It maps batch-saturated, single-GPU training performance for
nine dense decoder-only models from 150 million to 8 billion parameters across
ten NVIDIA GPUs and context lengths from 512 to 32,768 tokens.
The dataset contains 2,963 tested batch configurations, including successful
measurements and… See the full description on the dataset page: https://huggingface.co/datasets/FAIRC/SLMTrainBench.lm-eval-results-shyamieee-B3E3-SLM-7b-v2.0-private
Dataset Card for Evaluation run of shyamieee/B3E3-SLM-7b-v2.0
Dataset automatically created during the evaluation run of model shyamieee/B3E3-SLM-7b-v2.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-B3E3-SLM-7b-v2.0-private.SLMC_drop_toothpaste_forward_cucumberThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 32,
"total_frames": 2864,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_drop_toothpaste_forward_cucumber.SLMC_plate_corn_bowl_eggplant_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 32,
"total_frames": 3305,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_plate_corn_bowl_eggplant_v2.SLMC_left_potato_right_tomatoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 32,
"total_frames": 2849,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_left_potato_right_tomato.SLMC_plate_corn_bowl_eggplantThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 32,
"total_frames": 3239,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_plate_corn_bowl_eggplant.SLMC_plate_corn_bowl_eggplant_pi05This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 32,
"total_frames": 3869,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_plate_corn_bowl_eggplant_pi05.SLMC_left_lemon_right_coke_pi05This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 32,
"total_frames": 3578,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:32"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_left_lemon_right_coke_pi05.
