Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yinita /ps4mas-final-test-rollouts-0813 PS4MAS Final Test Rollouts (0813) Source split: ps4mas-0521-splits final_test_scenarios.jsonl Each traces/<model>/<model>.jsonl contains the agent-tool-loop output for 200 final_test scenarios × 4 topologies. Most baseline/oracle files are raw traces. GiGPO 0805-r2 step20/40/60/80 evals include OSS-120B scores and summary.json. Files Model Rows Path best_rl_gigpo_debate_step40 800 traces/best_rl_gigpo_debate_step40/best_rl_gigpo_debate_step40.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-final-test-rollouts-0813.tabular1K<n<10K0 likes4.3k downloads24d agoHugging Face02kyler0941 /fav_db_test_0tabular100K<n<1M0 likes2.3k downloads2y agoHugging Face03paulpacaud /rlbenchfail_test_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes892 downloads8mo agoHugging Face04gunnybd01 /faiss-integration-testtabularn<1K0 likes737 downloads5mo agoHugging Face05evalstate /test-traces Test Traces Codex-style rollout JSONL traces exported from fast-agent for validation against the Hugging Face Agent Trace Viewer. tabularn<1K2 likes677 downloads5mo agoHugging Face06martagm17 /test Medical Question Classification Dataset Dataset Summary This dataset is designed for medical language models evaluation. It merges several of the most important medical QA datasets into a common format and classifies them into 35 distinct medical categories. This structure enables users to identify any specific categories where the model's performance may be lacking and address these areas accordingly. Dataset Structure Data Fields id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/martagm17/test.tabularquestion-answering100K<n<1M1 likes615 downloads2y agoHugging Face07armand0e /teich-test-v1 hy3-preview coding agent traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by tencent/hy3-preview:free. Training-ready tools Use this tools payload when rendering converted examples through your training chat template. The same structure is emitted on each converted example as the tools field. [ { "type": "function", "function": { "name": "bash", "description": "Execute bash… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/teich-test-v1.tabularn<1K0 likes594 downloads5mo agoHugging Face08MiG-NJU /OmniVideo-Test OmniVideo-Test Official repository for OmniVideo-Test, the human-verified test set introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains". This repository includes: videos/: Raw video files. test_505.jsonl: The test set containing 505 multiple-choice QA pairs, complete with task taxonomies, ground-truth answers, and options. OmniVideo-Test serves as the evaluation companion to the OmniVideo-100K… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-Test.tabularvideo-text-to-textn<1K4 likes579 downloads4mo agoHugging Face09kyler0941 /fav_db_test_16tabular100K<n<1M0 likes535 downloads2y agoHugging Face10m-a-p /FineFineWeb-test FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-test.tabulartext-classification1M<n<10M5 likes501 downloads2y agoHugging Face11kyler0941 /fav_db_test_11tabular100K<n<1M0 likes492 downloads2y agoHugging Face12kyler0941 /fav_db_test_1tabular100K<n<1M0 likes466 downloads2y agoHugging Face13kyler0941 /fav_db_test_9tabular100K<n<1M0 likes389 downloads2y agoHugging Face14kyler0941 /fav_db_test_19tabular100K<n<1M0 likes373 downloads2y agoHugging Face15kyler0941 /fav_db_test_15tabular100K<n<1M0 likes344 downloads2y agoHugging Face16kyler0941 /fav_db_test_7tabular100K<n<1M0 likes324 downloads2y agoHugging Face17kyler0941 /fav_db_test_4tabular10K<n<100K0 likes323 downloads2y agoHugging Face18augustander /ch4-test-eval Ch4 HER2Match TEST eval results (public) JSON shards only. No PNGs. Operating checkpoints: ugustander/ch4-operating-ckpts. Data tiles: ugustander/her2match-full (unchanged). tabularn<1K0 likes295 downloads11d agoHugging Face19boxin-wbx /test Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/boxin-wbx/test.tabulartext-classification10K<n<100K0 likes277 downloads3y agoHugging Face20llmware /rag_instruct_benchmark_tester Dataset Card for RAG-Instruct-Benchmark-Tester Dataset Summary This is an updated benchmarking test dataset for "retrieval augmented generation" (RAG) use cases in the enterprise, especially for financial services, and legal. This test dataset includes 200 questions with context passages pulled from common 'retrieval scenarios', e.g., financial news, earnings releases, contracts, invoices, technical articles, general news and short texts. The questions are segmented… See the full description on the dataset page: https://huggingface.co/datasets/llmware/rag_instruct_benchmark_tester.tabularn<1K59 likes263 downloads3y agoHugging Face21kyler0941 /fav_db_test_12tabular100K<n<1M0 likes219 downloads2y agoHugging Face22paulpacaud /ur5fail_test_dataset Guardian Failure Detection Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.tabularvisual-question-answeringn<1K1 likes203 downloads8mo agoHugging Face23kyler0941 /fav_db_test_17tabular100K<n<1M0 likes194 downloads2y agoHugging Face24test-alexpouliquen /chi-bench Clinical Healthcare In-Situ Environment Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark What is in this dataset CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/test-alexpouliquen/chi-bench.documenttext-generationn<1K0 likes186 downloads5mo agoHugging Face25starsofchance /processed_test_splitstabular10K<n<100K0 likes185 downloads1y agoHugging Face26nyu-dice-lab /lm-eval-results-alnrg2arg-blockchainlabs_test3_seminar-private Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_test3_seminar Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_test3_seminar The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-alnrg2arg-blockchainlabs_test3_seminar-private.tabular100K<n<1M0 likes181 downloads2y agoHugging Face27pythonformer /Trajectory-Stitching-Test-7M Dataset Creation & Methodology Building this dataset required a highly optimized pipeline running on a dual-H100 NVL GPU cluster. The stitching process operates autonomously without relying on external LLM calls, using a specialized two-pass algorithm. 1. High-Information Keyword Extraction Instead of relying on simple word counts, the pipeline dynamically builds a dataset-specific stopword list by analyzing Document Frequency (DF) to banish words appearing in more than… See the full description on the dataset page: https://huggingface.co/datasets/pythonformer/Trajectory-Stitching-Test-7M.tabular1M<n<10M0 likes180 downloads5mo agoHugging Face28Kkuntal990 /test-braindecode-integration EEG Dataset This dataset was created using braindecode, a library for deep learning with EEG/MEG/ECoG signals. Dataset Information Number of recordings: 1 Number of channels: 26 Sampling frequency: 250.0 Hz Data type: Windowed (from Epochs object) Number of windows: 48 Total size: 0.04 MB Storage format: zarr Usage To load this dataset: from braindecode.datasets import BaseConcatDataset # Load dataset from Hugging Face Hub dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Kkuntal990/test-braindecode-integration.tabularn<1K0 likes167 downloads11mo agoHugging Face29sungjun12 /vlm_tsr_test_1 vlm_tsr_test_1 vlm_tsr_test의 1/4 파트. scene 그룹 50000~50004 포함. 전체 테스트셋은 4개 레포로 나뉘어 있습니다: vlm_tsr_test_1 vlm_tsr_test_2 vlm_tsr_test_3 vlm_tsr_test_4 코드 및 전체 파이프라인: Lim-Sung-Jun/vlm_training_template 구조 각 샘플은 3개 파일 세트로 구성됩니다: test/source/T01_C01/{id}.jpg # 테이블 이미지 test/source/T01_C01/{id}.json # 정답 HTML + 메타데이터 test/label/T01_C01/{id}.html # 렌더링용 GT HTML 평가 메트릭 메트릭 설명 TEDS Tree-Edit Distance 기반 구조 유사도 (0~1) TEDS-Structure 텍스트 제외 구조만… See the full description on the dataset page: https://huggingface.co/datasets/sungjun12/vlm_tsr_test_1.imageimage-to-text1K<n<10K0 likes162 downloads5mo agoHugging Face30Kkuntal990 /bnci-windows-test EEG Dataset This dataset was created using braindecode, a library for deep learning with EEG/MEG/ECoG signals. Dataset Information Number of recordings: 1 Number of channels: 26 Sampling frequency: 250.0 Hz Data type: Windowed (from Epochs object) Number of windows: 48 Total size: 0.04 MB Storage format: zarr Usage To load this dataset: from braindecode.datasets import BaseConcatDataset # Load dataset from Hugging Face Hub dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Kkuntal990/bnci-windows-test.tabularn<1K0 likes147 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.