datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SLMTrainBench
SLMTrainBench
SLMTrainBench is the measurement dataset for When Peak Floating-Point
Throughput Misleads: Utilization and Cost Frontiers for Small Language Model
Pretraining. It maps batch-saturated, single-GPU training performance for
nine dense decoder-only models from 150 million to 8 billion parameters across
ten NVIDIA GPUs and context lengths from 512 to 32,768 tokens.
The dataset contains 2,963 tested batch configurations, including successful
measurements and… See the full description on the dataset page: https://huggingface.co/datasets/FAIRC/SLMTrainBench.slm-125m-qa-dataset
slm-125m QA dataset (SFT)
24,713 reviewed grounded-QA pairs (LLM-judged, kept >=4) over US case law, SEC filings,
and educational web text. Columns: source, source_file, type, question, answer, context,
llm_judge_score, llm_judge_verdict, llm_judge_reason. Used to fine-tune Sudhanshu1985/slm-125m-sft.
Types: lookup / reasoning / unanswerable (refusals).
slmDatasetSLMDatasetTest
servicenow_incidents_25k_slm4servicenow_incidents_slm6SLMArtDatasetSLMArtDatasetTest
servicenow_incidents_25k_slm4test_slmservicenow_incidents_slm1SLMFTKMREPOSAP_SLMSAP_SLM_DataSLMFTKMREPO2servicenow_incidents_slm5
