datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
Korean-YouTube-Comment-Sentiment-Dataset
Korean YouTube Comment Sentiment Dataset
Data Overview
Summary
본 데이터셋은 유튜브에서 수집된 한국어 댓글 5,482개와 이에 대응하는 감정 레이블(긍정, 부정, 중립, 불명확)로 구성된 감정 분류용 데이터셋입니다.
주요 레이블: 긍정, 부정, 중립, 불명확
Features
수집 대상: 요리, 뷰티, 게임, 여행, 쇼핑 등 분야의 10만 명 이상 구독자를 보유한 유튜브 채널
형식: JSON (id, text, label)
검수: 한국인 검수자에 의한 수작업 라벨링 및 교차 검토
본 데이터셋은 구어체, 이모지, 줄임말 등 실제 사용자 표현이 반영되어 있습니다.
Dataset Structure
Dataset Fields
Field
Type
Description
id
string
각 댓글의… See the full description on the dataset page: https://huggingface.co/datasets/LLM-SocialMedia/Korean-YouTube-Comment-Sentiment-Dataset.valkompass-2026-llms
Valkompass 2026 × LLMs
How do 50 popular large language models answer the 35
questions in SVT's Swedish election compass (Valkompass 2026, Riksdag) — and which
of the 8 Riksdag parties does each model end up closest to?
Unlike comparisons that query chat products (ChatGPT, Gemini, Claude, Grok web UIs),
which have web-search / tools and act as agents, this dataset probes the raw model
weights only, via the OpenRouter API, with no system prompt, no tools, no web access.… See the full description on the dataset page: https://huggingface.co/datasets/nordan-ai/valkompass-2026-llms.repro-evaluating-llms-comparative-signals-traces
Agent traces
Agent sessions published from a Trackio Logbook.
llm-security-leaderboard-requeststts_llmsr_data
tts_llmsr_data
repro-codetaste-can-llms-generate-human-level-code-refactorings-traces
Agent traces
Agent sessions published from a Trackio Logbook.
sud_resh_evaluated_llms_answers
📊 Результаты Оценки Больших Языковых Моделей на Бенчмарке Судебных Решений
В данном документе представлен анализ производительности 15 больших языковых моделей (LLM), протестированных на специализированном бенчмарке, который включает 105 000 записей из судебных решений России. Оценка проводилась по 10 различным категориям права (например, трудовое, уголовное, гражданское) и 7 типам инструкций (например, изложение исковых требований, анализ доказательств, итоговое решение).
Ответы… See the full description on the dataset page: https://huggingface.co/datasets/lawful-good-project/sud_resh_evaluated_llms_answers.1-800-LLMs__Qwen-2.5-14B-Hindi-details
Dataset Card for Evaluation run of 1-800-LLMs/Qwen-2.5-14B-Hindi
Dataset automatically created during the evaluation run of model 1-800-LLMs/Qwen-2.5-14B-Hindi
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/1-800-LLMs__Qwen-2.5-14B-Hindi-details.1-800-LLMs__Qwen-2.5-14B-Hindi-Custom-Instruct-details
Dataset Card for Evaluation run of 1-800-LLMs/Qwen-2.5-14B-Hindi-Custom-Instruct
Dataset automatically created during the evaluation run of model 1-800-LLMs/Qwen-2.5-14B-Hindi-Custom-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/1-800-LLMs__Qwen-2.5-14B-Hindi-Custom-Instruct-details.LLMSRRestaurant_Review_LLMs_AnnotationC_C_plus_plus_Patch_Dataset_for_Secure_Code_LLMs
C/C++ Security Vulnerability & Automated Patch Dataset for llms fine tuning
A large-scale corpus of 23,826 function-level vulnerability repair pairs ($Vulnerable \rightarrow Fixed$) mined from 64 major C/C++ open-source projects spanning 12,130 unique commits identified by the security-focused mining pipeline.
This dataset is specifically structured for Automated Program Repair (APR), Vulnerability Assessment, and Instruction Fine-Tuning of Code LLMs (e.g., Qwen2.5-Coder… See the full description on the dataset page: https://huggingface.co/datasets/prabhatl0dhi/C_C_plus_plus_Patch_Dataset_for_Secure_Code_LLMs.
