datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IfEvalCode-testsetMMM-datasets-TestsetMultilingual Mutual Reinforcement Effect Mix Datasets
This is a Training set of OIELLM.
This Train set already formatted by OIELLM's format. The test set is in the another page in huggingface.
The MMM support 3 languages (English, Chinese and Japanese). And you must use task instruct words to define kind of task.
Mutual Reinforcement Effect.
OIELLM's input and output
MMM Dataset
The following is input and output format:
{
"input": "In 1953, filming of "On the Waterfront" starring… See the full description on the dataset page: https://huggingface.co/datasets/ganchengguang/MMM-datasets-Testset.video-SALMONN_2_testset
video-SALMONN 2 Benchmark
Generate the caption corresponding to the video and the audio with video_salmonn2_test.json
Organize your results in the format like the following example:
[
{
"id": ["0.mp4"],
"pred": "Generated Caption"
}
]
Replace res_file in eval.py with your result file.
Run python3 eval.pySpatialGen-Testset
SpatialGen Testset
This repository contains the test set for SPATIALGEN: Layout-guided 3D Indoor Scene Generation, a novel multi-view multi-modal diffusion model for generating realistic and semantically consistent 3D indoor scenes.
Project page | Paper | Code
We provide a test set of 48 preprocessed point clouds and their corresponding GT layouts, multi-view images are cropped from the high-resolution panoramic images.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialGen-Testset.NVVSpeech-Challenge-Track1-Test-Set
NVVSpeech Challenge Track 1 Test Set
Track 1 test set for the NVVSpeech Challenge at ISCSLP 2026.
Task
Given a speech recording, produce a transcript that contains the spoken content and the non-verbal vocalization (NVV) tags at their corresponding positions.
Dataset Summary
Language
Samples
Chinese
985
English
961
Total
1,946
Files
.
├── README.md
├── SUBMISSION_GUIDE.txt
├── test.jsonl
├── ground_truth.jsonl
├──… See the full description on the dataset page: https://huggingface.co/datasets/NVVSpeech-Challenge/NVVSpeech-Challenge-Track1-Test-Set.NVVSpeech-Challenge-Track2-Test-Set
NVVSpeech Challenge Track 2 Test Set
Track 2 test set for the NVVSpeech Challenge at ISCSLP 2026.
Task
Given a transcript containing one or more non-verbal vocalization (NVV) tags, synthesize speech that naturally realizes the requested NVVs while preserving intelligibility, naturalness, and audio quality.
Dataset Summary
Language
Samples
Chinese
800
English
800
Total
1,600
Files
.
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/NVVSpeech-Challenge/NVVSpeech-Challenge-Track2-Test-Set.ddro-testsetsrede_saude_publica_test_set
Rede Saude Publica Test Set
This dataset is the public-health transfer benchmark for the released Text-to-SQL agent artifact. It is a synthetic Brazilian public-health schema and test set used to measure cross-database generalization: the fine-tuned model was not trained on trajectories from this schema.
Code and reproducibility repository:
https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br
Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/rede_saude_publica_test_set.CISA-testsetCISA-testset from İbrahim Temo's Memoir: Cross-Individual Sentiment Analysis Test Dataset for Historical Turkish
This test dataset is specifically designed for evaluating Cross-Individual Sentiment Analysis (CISA) performance on historical Turkish texts from İbrahim Temo's memoirs.
📚 Dataset Description
This test dataset contains 200 sentences extracted from the first 66 pages of İbrahim Temo's original memoirs "İttihad ve Terakki Cemiyetinin Teşekkülü ve Hidematı Vataniye ve İnkılâbı Milliye… See the full description on the dataset page: https://huggingface.co/datasets/dbbiyte/CISA-testset.SpatialGen-Testset
SpatialGen Testset
This repository contains the test set for SPATIALGEN: Layout-guided 3D Indoor Scene Generation, a novel multi-view multi-modal diffusion model for generating realistic and semantically consistent 3D indoor scenes.
Project page | Paper | Code
We provide a test set of 48 preprocessed point clouds and their corresponding GT layouts, multi-view images are cropped from the high-resolution panoramic images.
Folder Structure
Outlines of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/BenjaminChai579/SpatialGen-Testset.HatEval_2019_Test_Set_Task5Test Set from HatEval (Basile et al, 2019), SemEval-2019 Task 5
llama-3.1-medprm-reward-raw-test-setKoala-test-setThis dataset is taken from https://github.com/arnav-gudibande/koala-test-set
testsetenvironmental_registry_test_set
Environmental Registry Test Set
This dataset is the anonymized primary benchmark used for evaluating agentic Portuguese Text-to-SQL over a real PostgreSQL/PostGIS environmental-registry database. The underlying production database is not released, but the benchmark metadata and gold labels are provided for transparency and comparison.
Code and reproducibility repository:
https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br
Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/environmental_registry_test_set.LL144-Test-SetHate_Political_Opponent_2021_Test_SetTest set from "Hate Towards the Political Opponent"(Grimminger et al., 2021)
Vidyaapati-Hindi-Konkani-Testsettestset5stereoset_test_setOLID_2019_Test_SetTest set from OLID dataset (Zampieri et al., 2019), SemEval 2019
TestSetGenkeval-testset
keval_test
The keval-testset is a dataset designed for training and validating the keval model.
The keval model follows the LLM-as-a-judge approach, which evaluates LLMs by assessing their responses to prompts from the ko-bench dataset. In other words, the keval model assigns scores to LLM-generated responses based on predefined evaluation criteria.
The keval-testset serves as a crucial resource for training and validating the keval model, enabling precise benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/davidkim205/keval-testset.testset_ragasembedding-testsettestset4postvalid-v2-test-set
LongShOTBench (Test Split)
Benchmark accompanying the NeurIPS 2026 submission
"A Benchmark for Omni-Modal Reasoning in Long Videos."
This dataset is shared anonymously to support double-blind review.
Purpose
LongShOTBench evaluates multimodal LLMs on long-form video understanding
across vision, speech, and non-speech audio, using intent-driven questions
and weighted criterion-level rubrics. Intended for evaluation only, not
training.
License
CC BY-NC-SA 4.0.… See the full description on the dataset page: https://huggingface.co/datasets/anonymsubs/postvalid-v2-test-set.cv10-uk-testset-clean-zipaGenerated by https://github.com/lingjzhu/zipa (zipa_large_crctc_500000_avg10.pth) and https://github.com/dmort27/epitran
10_1injection_test_setkgrammar-testset
kgrammar-testset
The kgrammar-testset is a dataset designed for training and validating the kgrammar model, which identifies grammatical errors in Korean text and outputs the number of detected errors.
The kgrammar-testset was generated using GPT-4o. To create error-containing documents, a predefined prompt was used to introduce grammatical mistakes into responses when given a question. The dataset is structured to ensure a balanced distribution, consisting of 50% general questions… See the full description on the dataset page: https://huggingface.co/datasets/davidkim205/kgrammar-testset.
