sLLM
Datasets
All datasets matching “sLLM”RocketEval-sLLMs
🚀 RocketEval 🚀
🚀 [ICLR '25] RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
Github | OpenReview | Colab
This dataset contains the queries, generated checklist data, and responses data from 4 public benchmark datasets:
Dataset
No. of Queries
Comments
MT-Bench
160
Each 2-turn dialogue is split into 2 queries.
AlpacaEval
805
Arena-Hard
500
WildBench
1,000
To fit the context window of lightweight LLMs, we use a subset of WildBench including 1000… See the full description on the dataset page: https://huggingface.co/datasets/wjkim9653/RocketEval-sLLMs.sllm-amazonia-saude-sft
sLLM Amazônia Saúde SFT
Conjunto de dados sintético de 1.000 diálogos clínicos em português brasileiro para fine-tuning (SFT/QLoRA) de modelos de linguagem de pequeno porte (sLLMs) que atuam como assistentes clínicos offline em comunidades isoladas da Amazônia.
Contexto
Profissionais de saúde e agentes comunitários que atuam em áreas ribeirinhas, quilombolas e indígenas da Amazônia frequentemente operam sem conectividade. Este dataset adapta modelos de linguagem… See the full description on the dataset page: https://huggingface.co/datasets/admin-lima/sllm-amazonia-saude-sft.sllm
Dataset Card for GEM/viggo
Link to Main Data Card
You can find the main data card on the GEM Website.
Dataset Summary
ViGGO is an English data-to-text generation dataset in the video game domain, with target responses being more conversational than information-seeking, yet constrained to the information presented in a meaning representation. The dataset is relatively small with about 5,000 datasets but very clean, and can thus serve for evaluating transfer… See the full description on the dataset page: https://huggingface.co/datasets/AlexFromSynlabs/sllm.step4_sllm_v2
Step 4 PII 후보 판정 데이터셋 — 최종 10,000개
RAG 답변에서 상위 NER 단계가 추출한 후보가 문맥상 특정 자연인의 개인정보인지 PII 또는 NOT_PII로 판정하도록 Qwen을 SFT하기 위한 합성 데이터셋이다.
바로 사용하는 파일
step4_final_10000_qwen_train.jsonl: 학습 8,000개
step4_final_10000_qwen_valid.jsonl: 검증 1,000개
step4_final_10000_qwen_test.jsonl: 최종 평가 1,000개
step4_final_10000_qwen_all.jsonl: 전체 확인용 10,000개
각 행의 최상위 필드는 messages 하나뿐이며 system, user, assistant 순서다. 학습 시 Qwen tokenizer의 chat template를 적용하고 assistant 응답 부분에만 loss를 계산한다.… See the full description on the dataset page: https://huggingface.co/datasets/bbanany/step4_sllm_v2.qa_BBQ_trans_gender
Dataset Card for "qa_BBQ_trans_gender"
More Information needed
AnimalQA
Dataset Card for SAKURA-AnimalQA
This dataset contains the audio and the single/multi-hop questions/answers of the animal track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information".
The fields of the dataset are:
file: The filename of the audio files.
audio: The audio recordings.
attribute_label: The attribute labels (i.e., the kinds of animal making the sounds) of the audio… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/AnimalQA.
