datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-bench-dummy-test-datasetdatasets-tests-compressiontests-raw-jsonlraw_jsonlcompressed_filesner-jsonlcircleci-test-resultsIfEvalCode-testsetps4mas-final-test-rollouts-0813
PS4MAS Final Test Rollouts (0813)
Source split: ps4mas-0521-splits final_test_scenarios.jsonl
Each traces/<model>/<model>.jsonl contains the agent-tool-loop output for 200 final_test scenarios × 4 topologies. Most baseline/oracle files are raw traces. GiGPO 0805-r2 step20/40/60/80 evals include OSS-120B scores and summary.json.
Files
Model
Rows
Path
best_rl_gigpo_debate_step40
800
traces/best_rl_gigpo_debate_step40/best_rl_gigpo_debate_step40.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-final-test-rollouts-0813.Lora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.fav_db_test_0trl-test-instructionTested-143k-Python-AlpacaContributors: Nicolas Mejia Petit
Vezora's CodeTester Dataset
Introduction
Today, on March 6, 2024, we are excited to release our internal Python dataset with 143,327 examples of code. These examples have been meticulously tested and verified as working. Our dataset was created using a script we developed.
Dataset Creation
Our script operates by extracting Python code from the output section of Alpaca-formatted datasets. It tests each extracted piece of code… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/Tested-143k-Python-Alpaca.ICCV_workshop_testRealEstate10K_testRSCC-RSEdit-Test-Split
RSCC-RSEdit-Test-Split
This directory contains the test split for RSCC-RSEdit dataset.
Directory Structure
RSCC-RSEdit-Test-Split/
├── images/ # Original images (676 PNG files)
├── masks/ # Original grayscale masks (338 PNG files)
│ └── [mask files with pixel values 0,1,2,3,4]
├── masks_colorful/ # Colorful RGBA visualization masks (338 PNG files)
│ └── [same filenames as masks/, but in RGBA format with colors]
├──… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/RSCC-RSEdit-Test-Split.Tested-22k-Python-AlpacaContributors: Nicolas Mejia Petit
Vezora's CodeTester Dataset
Introduction
Today, on November 2, 2023, we are excited to release our internal Python dataset with 22,600 examples of code. These examples have been meticulously tested and verified as working. Our dataset was created using a script we developed.
Dataset Creation
Our script operates by extracting Python code from the output section of Alpaca-formatted datasets. It tests each extracted piece of… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/Tested-22k-Python-Alpaca.rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.Video_Reality_Test
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization.
Benchmark Structure
This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:
real_hard: 100 samples.… See the full description on the dataset page: https://huggingface.co/datasets/kolerk/Video_Reality_Test.pile_val_test
The Pile: Validation and Test Splits
This repo contains the validation and test splits of The Pile, an 825 GiB English text dataset designed for training large language models.
Files
File
Split
Size
val.jsonl
Validation
1.4 GB
test.jsonl
Test
1.3 GB
Format
Each line is a JSON object with two fields:
{"text": "The document text...", "meta": {"pile_set_name": "Pile-CC"}}
The meta.pile_set_name field indicates which of the 22 constituent… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/pile_val_test.faiss-integration-testtest-traces
Test Traces
Codex-style rollout JSONL traces exported from fast-agent for validation against the Hugging Face Agent Trace Viewer.
WideSeek-R1-test-data
Testing Dataset
We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required.
test
Medical Question Classification Dataset
Dataset Summary
This dataset is designed for medical language models evaluation. It merges several of the most important medical QA datasets into a common format and classifies them into 35 distinct medical categories. This structure enables users to identify any specific categories where the model's performance may be lacking and address these areas accordingly.
Dataset Structure
Data Fields
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/martagm17/test.teich-test-v1
hy3-preview coding agent traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by tencent/hy3-preview:free.
Training-ready tools
Use this tools payload when rendering converted examples through your training chat template.
The same structure is emitted on each converted example as the tools field.
[
{
"type": "function",
"function": {
"name": "bash",
"description": "Execute bash… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/teich-test-v1.OmniVideo-Test
OmniVideo-Test
Official repository for OmniVideo-Test, the human-verified test set introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains".
This repository includes:
videos/: Raw video files.
test_505.jsonl: The test set containing 505 multiple-choice QA pairs, complete with task taxonomies, ground-truth answers, and options.
OmniVideo-Test serves as the evaluation companion to the OmniVideo-100K… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-Test.fav_db_test_16Long-video-test-dataAgentChat-Test
Test Set Description
This directory contains the test set used for tool-use evaluation. The JSON files under Test-JSON/ are organized by task type:
SingleTaskProcessing/tool-select_test.json: single-tool selection tasks.
ParallelProcessing/parallel-call_test.json: parallel tool-call tasks.
ProactiveSeeking/searchTools_test_predictions_kept.json: proactive tool-search tasks.
TaskDecomposition/muti-tool-select_test.json: multi-tool task decomposition tasks.… See the full description on the dataset page: https://huggingface.co/datasets/leungtianle/AgentChat-Test.FineFineWeb-test
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-test.
