Team Ai
20 results

sLLM

wjkim9653 /RocketEval-sLLMs 🚀 RocketEval 🚀 🚀 [ICLR '25] RocketEval: Efficient Automated LLM Evaluation via Grading Checklist Github | OpenReview | Colab This dataset contains the queries, generated checklist data, and responses data from 4 public benchmark datasets: Dataset No. of Queries Comments MT-Bench 160 Each 2-turn dialogue is split into 2 queries. AlpacaEval 805 Arena-Hard 500 WildBench 1,000 To fit the context window of lightweight LLMs, we use a subset of WildBench including 1000… See the full description on the dataset page: https://huggingface.co/datasets/wjkim9653/RocketEval-sLLMs.texttext-generation1K<n<10K0 likes239 downloads2y agoHugging Faceadmin-lima /sllm-amazonia-saude-sft sLLM Amazônia Saúde SFT Conjunto de dados sintético de 1.000 diálogos clínicos em português brasileiro para fine-tuning (SFT/QLoRA) de modelos de linguagem de pequeno porte (sLLMs) que atuam como assistentes clínicos offline em comunidades isoladas da Amazônia. Contexto Profissionais de saúde e agentes comunitários que atuam em áreas ribeirinhas, quilombolas e indígenas da Amazônia frequentemente operam sem conectividade. Este dataset adapta modelos de linguagem… See the full description on the dataset page: https://huggingface.co/datasets/admin-lima/sllm-amazonia-saude-sft.texttext-generation1K<n<10K0 likes75 downloads4mo agoHugging FaceAlexFromSynlabs /sllm Dataset Card for GEM/viggo Link to Main Data Card You can find the main data card on the GEM Website. Dataset Summary ViGGO is an English data-to-text generation dataset in the video game domain, with target responses being more conversational than information-seeking, yet constrained to the information presented in a meaning representation. The dataset is relatively small with about 5,000 datasets but very clean, and can thus serve for evaluating transfer… See the full description on the dataset page: https://huggingface.co/datasets/AlexFromSynlabs/sllm.table-to-text0 likes61 downloads3y agoHugging Facebbanany /step4_sllm_v2 Step 4 PII 후보 판정 데이터셋 — 최종 10,000개 RAG 답변에서 상위 NER 단계가 추출한 후보가 문맥상 특정 자연인의 개인정보인지 PII 또는 NOT_PII로 판정하도록 Qwen을 SFT하기 위한 합성 데이터셋이다. 바로 사용하는 파일 step4_final_10000_qwen_train.jsonl: 학습 8,000개 step4_final_10000_qwen_valid.jsonl: 검증 1,000개 step4_final_10000_qwen_test.jsonl: 최종 평가 1,000개 step4_final_10000_qwen_all.jsonl: 전체 확인용 10,000개 각 행의 최상위 필드는 messages 하나뿐이며 system, user, assistant 순서다. 학습 시 Qwen tokenizer의 chat template를 적용하고 assistant 응답 부분에만 loss를 계산한다.… See the full description on the dataset page: https://huggingface.co/datasets/bbanany/step4_sllm_v2.texttext-classification10K<n<100K0 likes52 downloads3mo agoHugging FaceSLLMBias /qa_BBQ_trans_gender Dataset Card for "qa_BBQ_trans_gender" More Information needed audio1K<n<10K0 likes49 downloads2y agoHugging FaceSLLM-multi-hop /AnimalQA Dataset Card for SAKURA-AnimalQA This dataset contains the audio and the single/multi-hop questions/answers of the animal track of the SAKURA benchmark from Interspeech 2025 paper, "SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information". The fields of the dataset are: file: The filename of the audio files. audio: The audio recordings. attribute_label: The attribute labels (i.e., the kinds of animal making the sounds) of the audio… See the full description on the dataset page: https://huggingface.co/datasets/SLLM-multi-hop/AnimalQA.audion<1K0 likes41 downloads1y agoHugging Face