datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ps4mas-final-test-rollouts-0813
PS4MAS Final Test Rollouts (0813)
Source split: ps4mas-0521-splits final_test_scenarios.jsonl
Each traces/<model>/<model>.jsonl contains the agent-tool-loop output for 200 final_test scenarios × 4 topologies. Most baseline/oracle files are raw traces. GiGPO 0805-r2 step20/40/60/80 evals include OSS-120B scores and summary.json.
Files
Model
Rows
Path
best_rl_gigpo_debate_step40
800
traces/best_rl_gigpo_debate_step40/best_rl_gigpo_debate_step40.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-final-test-rollouts-0813.fav_db_test_0rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.faiss-integration-testtest-traces
Test Traces
Codex-style rollout JSONL traces exported from fast-agent for validation against the Hugging Face Agent Trace Viewer.
test
Medical Question Classification Dataset
Dataset Summary
This dataset is designed for medical language models evaluation. It merges several of the most important medical QA datasets into a common format and classifies them into 35 distinct medical categories. This structure enables users to identify any specific categories where the model's performance may be lacking and address these areas accordingly.
Dataset Structure
Data Fields
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/martagm17/test.teich-test-v1
hy3-preview coding agent traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by tencent/hy3-preview:free.
Training-ready tools
Use this tools payload when rendering converted examples through your training chat template.
The same structure is emitted on each converted example as the tools field.
[
{
"type": "function",
"function": {
"name": "bash",
"description": "Execute bash… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/teich-test-v1.OmniVideo-Test
OmniVideo-Test
Official repository for OmniVideo-Test, the human-verified test set introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains".
This repository includes:
videos/: Raw video files.
test_505.jsonl: The test set containing 505 multiple-choice QA pairs, complete with task taxonomies, ground-truth answers, and options.
OmniVideo-Test serves as the evaluation companion to the OmniVideo-100K… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-Test.fav_db_test_16FineFineWeb-test
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-test.fav_db_test_11fav_db_test_1fav_db_test_9fav_db_test_19fav_db_test_15fav_db_test_7fav_db_test_4ch4-test-eval
Ch4 HER2Match TEST eval results (public)
JSON shards only. No PNGs. Operating checkpoints: ugustander/ch4-operating-ckpts.
Data tiles: ugustander/her2match-full (unchanged).
test
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/boxin-wbx/test.rag_instruct_benchmark_tester
Dataset Card for RAG-Instruct-Benchmark-Tester
Dataset Summary
This is an updated benchmarking test dataset for "retrieval augmented generation" (RAG) use cases in the enterprise, especially for financial services, and legal. This test dataset includes 200 questions with context passages pulled from common 'retrieval scenarios', e.g., financial news, earnings releases,
contracts, invoices, technical articles, general news and short texts.
The questions are segmented… See the full description on the dataset page: https://huggingface.co/datasets/llmware/rag_instruct_benchmark_tester.fav_db_test_12ur5fail_test_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.fav_db_test_17chi-bench
Clinical Healthcare In-Situ Environment
Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark
What is in this dataset
CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/test-alexpouliquen/chi-bench.processed_test_splitslm-eval-results-alnrg2arg-blockchainlabs_test3_seminar-private
Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_test3_seminar
Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_test3_seminar
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-alnrg2arg-blockchainlabs_test3_seminar-private.Trajectory-Stitching-Test-7M
Dataset Creation & Methodology
Building this dataset required a highly optimized pipeline running on a dual-H100 NVL GPU cluster. The stitching process operates autonomously without relying on external LLM calls, using a specialized two-pass algorithm.
1. High-Information Keyword Extraction
Instead of relying on simple word counts, the pipeline dynamically builds a dataset-specific stopword list by analyzing Document Frequency (DF) to banish words appearing in more than… See the full description on the dataset page: https://huggingface.co/datasets/pythonformer/Trajectory-Stitching-Test-7M.test-braindecode-integration
EEG Dataset
This dataset was created using braindecode, a library for deep learning with EEG/MEG/ECoG signals.
Dataset Information
Number of recordings: 1
Number of channels: 26
Sampling frequency: 250.0 Hz
Data type: Windowed (from Epochs object)
Number of windows: 48
Total size: 0.04 MB
Storage format: zarr
Usage
To load this dataset:
from braindecode.datasets import BaseConcatDataset
# Load dataset from Hugging Face Hub
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Kkuntal990/test-braindecode-integration.vlm_tsr_test_1
vlm_tsr_test_1
vlm_tsr_test의 1/4 파트. scene 그룹 50000~50004 포함.
전체 테스트셋은 4개 레포로 나뉘어 있습니다:
vlm_tsr_test_1
vlm_tsr_test_2
vlm_tsr_test_3
vlm_tsr_test_4
코드 및 전체 파이프라인: Lim-Sung-Jun/vlm_training_template
구조
각 샘플은 3개 파일 세트로 구성됩니다:
test/source/T01_C01/{id}.jpg # 테이블 이미지
test/source/T01_C01/{id}.json # 정답 HTML + 메타데이터
test/label/T01_C01/{id}.html # 렌더링용 GT HTML
평가 메트릭
메트릭
설명
TEDS
Tree-Edit Distance 기반 구조 유사도 (0~1)
TEDS-Structure
텍스트 제외 구조만… See the full description on the dataset page: https://huggingface.co/datasets/sungjun12/vlm_tsr_test_1.bnci-windows-test
EEG Dataset
This dataset was created using braindecode, a library for deep learning with EEG/MEG/ECoG signals.
Dataset Information
Number of recordings: 1
Number of channels: 26
Sampling frequency: 250.0 Hz
Data type: Windowed (from Epochs object)
Number of windows: 48
Total size: 0.04 MB
Storage format: zarr
Usage
To load this dataset:
from braindecode.datasets import BaseConcatDataset
# Load dataset from Hugging Face Hub
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Kkuntal990/bnci-windows-test.
