datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v3.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v3.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private.SLM4CRP_with_RTs
SLM4CRP_with_RTs Dataset
Overview
The SLM4CRP_with_RTs dataset is a chemical reaction predictions (CRPs) dataset featuring reaction type (RT) labels, developed from the Mol-Instruction. We introduce a novel knowledge elicitation approach integrating a self-feedback mechanism with data curation using large language models (LLMs). This dataset embodies domain-specific knowledge by combining reactants and products of chemical reactions with annotated RTs, demonstrating… See the full description on the dataset page: https://huggingface.co/datasets/liupf/SLM4CRP_with_RTs.slm-architecture-benchmark-specs
SLM Benchmark Protocol Specs
A reference for the exact conventions to use when benchmarking very small
language models (roughly 0.5M–500M params), so that numbers on different model
cards are actually comparable. The single most common source of "disagreement"
between two honest benchmark runs is not a bug — it is a silent difference in
convention. This dataset pins those conventions down.
Every convention here is either (a) something I verified end-to-end against a
real… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-architecture-benchmark-specs.lm-eval-results-shyamieee-B3E3-SLM-7b-v1.0-private
Dataset Card for Evaluation run of shyamieee/B3E3-SLM-7b-v1.0
Dataset automatically created during the evaluation run of model shyamieee/B3E3-SLM-7b-v1.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-B3E3-SLM-7b-v1.0-private.SysMLv2_Repair_with_SLMs
SysMLv2 Repair with SLMs
Dataset used in "Automated Semantic Fault Localization in SysML v2: A Human-in-the-Loop Framework Using Knowledge-Graph Augmented LLMs", presented at INCOSE International Symposium 2026.
Dataset Structure
This dataset provides two configurations:
default: Contains train/validation/test splits used for fine-tuning small models. Samples exceeding 2048 tokens have been removed.
full: Contains complete dataset
Task
Given SysML v2 code… See the full description on the dataset page: https://huggingface.co/datasets/rohhaiil/SysMLv2_Repair_with_SLMs.lm-eval-results-shyamieee-B3E3-SLM-7b-v3.0-private
Dataset Card for Evaluation run of shyamieee/B3E3-SLM-7b-v3.0
Dataset automatically created during the evaluation run of model shyamieee/B3E3-SLM-7b-v3.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-B3E3-SLM-7b-v3.0-private.lm-eval-results-shyamieee-B3E3-SLM-7b-v2.0-private
Dataset Card for Evaluation run of shyamieee/B3E3-SLM-7b-v2.0
Dataset automatically created during the evaluation run of model shyamieee/B3E3-SLM-7b-v2.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-B3E3-SLM-7b-v2.0-private.SLM-Math-Bench-1-mediumslm-rl-colab-dataslm-rl-colab-dataslm-rl-colab-dataslm-rl-colab-dataslm-rl-colab-datablazing-audio-slm-v7-4-dataset-dev
Blazing Audio SLM V7.4 development dataset
This is the exact joint-replay optimizer dataset used for the V7.4 development checkpoint. It is
not a claim that the resulting model passed release: V7.4 failed the minimum-family,
general-adversarial, and strict measurement-deferral V1 gates. Manifest V2 remains sealed and is
not included.
Splits and composition
Split
Rows
Calculate
Explain
Abstain
Defer measurement
Train
3,923
2,621
582
360
360
Validation… See the full description on the dataset page: https://huggingface.co/datasets/audiuphile/blazing-audio-slm-v7-4-dataset-dev.
