Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01klieret /swe-bench-dummy-test-datasettextn<1K0 likes50k downloads1y agoHugging Face02albertvillanova /datasets-tests-compressiontextn<1K0 likes43k downloads5y agoHugging Face03albertvillanova /tests-raw-jsonltext10K<n<100K1 likes31k downloads5y agoHugging Face04hf-internal-testing /raw_jsonltext10K<n<100K0 likes28k downloads5y agoHugging Face05hf-internal-testing /compressed_filestextn<1K0 likes11k downloads5y agoHugging Face06hf-internal-testing /ner-jsonltext10K<n<100K0 likes7.3k downloads1y agoHugging Face07transformers-community /circleci-test-resultstextn<1K4 likes5.3k downloads4mo agoHugging Face08Multilingual-Multimodal-NLP /IfEvalCode-testsettextn<1K2 likes4.3k downloads1y agoHugging Face09yinita /ps4mas-final-test-rollouts-0813 PS4MAS Final Test Rollouts (0813) Source split: ps4mas-0521-splits final_test_scenarios.jsonl Each traces/<model>/<model>.jsonl contains the agent-tool-loop output for 200 final_test scenarios × 4 topologies. Most baseline/oracle files are raw traces. GiGPO 0805-r2 step20/40/60/80 evals include OSS-120B scores and summary.json. Files Model Rows Path best_rl_gigpo_debate_step40 800 traces/best_rl_gigpo_debate_step40/best_rl_gigpo_debate_step40.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-final-test-rollouts-0813.tabular1K<n<10K0 likes4.3k downloads24d agoHugging Face10lvesucces /Lora_Cloud_Dataset_Test VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件 VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包 本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。 本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。 一、Mac 端文件布局自动识别(针对您的 iild 结构) 一、核心架构与流水线 评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器: 在本次评测中,整条上行与闭环流水线严格遵循您的设想: 上游双塔一致性(In-Domain Consistency): 输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.imagen<1K0 likes2.8k downloads24d agoHugging Face11kyler0941 /fav_db_test_0tabular100K<n<1M0 likes2.3k downloads2y agoHugging Face12philschmid /trl-test-instructiontextn<1K0 likes2.2k downloads3y agoHugging Face13Vezora /Tested-143k-Python-AlpacaContributors: Nicolas Mejia Petit Vezora's CodeTester Dataset Introduction Today, on March 6, 2024, we are excited to release our internal Python dataset with 143,327 examples of code. These examples have been meticulously tested and verified as working. Our dataset was created using a script we developed. Dataset Creation Our script operates by extracting Python code from the output section of Alpaca-formatted datasets. It tests each extracted piece of code… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/Tested-143k-Python-Alpaca.text100K<n<1M55 likes2k downloads3y agoHugging Face14Louischong /ICCV_workshop_testtextn<1K0 likes1.7k downloads1y agoHugging Face15Hualingchu /RealEstate10K_testtextn<1K0 likes1.1k downloads4mo agoHugging Face16BiliSakura /RSCC-RSEdit-Test-Split RSCC-RSEdit-Test-Split This directory contains the test split for RSCC-RSEdit dataset. Directory Structure RSCC-RSEdit-Test-Split/ ├── images/ # Original images (676 PNG files) ├── masks/ # Original grayscale masks (338 PNG files) │ └── [mask files with pixel values 0,1,2,3,4] ├── masks_colorful/ # Colorful RGBA visualization masks (338 PNG files) │ └── [same filenames as masks/, but in RGBA format with colors] ├──… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/RSCC-RSEdit-Test-Split.imagen<1K0 likes1k downloads5mo agoHugging Face17Vezora /Tested-22k-Python-AlpacaContributors: Nicolas Mejia Petit Vezora's CodeTester Dataset Introduction Today, on November 2, 2023, we are excited to release our internal Python dataset with 22,600 examples of code. These examples have been meticulously tested and verified as working. Our dataset was created using a script we developed. Dataset Creation Our script operates by extracting Python code from the output section of Alpaca-formatted datasets. It tests each extracted piece of… See the full description on the dataset page: https://huggingface.co/datasets/Vezora/Tested-22k-Python-Alpaca.text10K<n<100K68 likes1k downloads3y agoHugging Face18paulpacaud /rlbenchfail_test_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes892 downloads8mo agoHugging Face19kolerk /Video_Reality_Test VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization. Benchmark Structure This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation: real_hard: 100 samples.… See the full description on the dataset page: https://huggingface.co/datasets/kolerk/Video_Reality_Test.texttext-to-videon<1K11 likes834 downloads5mo agoHugging Face20EleutherAI /pile_val_test The Pile: Validation and Test Splits This repo contains the validation and test splits of The Pile, an 825 GiB English text dataset designed for training large language models. Files File Split Size val.jsonl Validation 1.4 GB test.jsonl Test 1.3 GB Format Each line is a JSON object with two fields: {"text": "The document text...", "meta": {"pile_set_name": "Pile-CC"}} The meta.pile_set_name field indicates which of the 22 constituent… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/pile_val_test.texttext-generation100K<n<1M0 likes803 downloads8mo agoHugging Face21gunnybd01 /faiss-integration-testtabularn<1K0 likes737 downloads5mo agoHugging Face22evalstate /test-traces Test Traces Codex-style rollout JSONL traces exported from fast-agent for validation against the Hugging Face Agent Trace Viewer. tabularn<1K2 likes677 downloads5mo agoHugging Face23WideSeek-R1 /WideSeek-R1-test-data Testing Dataset We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required. texttext-generationn<1K0 likes640 downloads5mo agoHugging Face24martagm17 /test Medical Question Classification Dataset Dataset Summary This dataset is designed for medical language models evaluation. It merges several of the most important medical QA datasets into a common format and classifies them into 35 distinct medical categories. This structure enables users to identify any specific categories where the model's performance may be lacking and address these areas accordingly. Dataset Structure Data Fields id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/martagm17/test.tabularquestion-answering100K<n<1M1 likes615 downloads2y agoHugging Face25armand0e /teich-test-v1 hy3-preview coding agent traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by tencent/hy3-preview:free. Training-ready tools Use this tools payload when rendering converted examples through your training chat template. The same structure is emitted on each converted example as the tools field. [ { "type": "function", "function": { "name": "bash", "description": "Execute bash… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/teich-test-v1.tabularn<1K0 likes594 downloads5mo agoHugging Face26MiG-NJU /OmniVideo-Test OmniVideo-Test Official repository for OmniVideo-Test, the human-verified test set introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains". This repository includes: videos/: Raw video files. test_505.jsonl: The test set containing 505 multiple-choice QA pairs, complete with task taxonomies, ground-truth answers, and options. OmniVideo-Test serves as the evaluation companion to the OmniVideo-100K… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-Test.tabularvideo-text-to-textn<1K4 likes579 downloads4mo agoHugging Face27kyler0941 /fav_db_test_16tabular100K<n<1M0 likes535 downloads2y agoHugging Face28sfsdfsafsddsfsdafsa /Long-video-test-datatextn<1K2 likes523 downloads3y agoHugging Face29leungtianle /AgentChat-Test Test Set Description This directory contains the test set used for tool-use evaluation. The JSON files under Test-JSON/ are organized by task type: SingleTaskProcessing/tool-select_test.json: single-tool selection tasks. ParallelProcessing/parallel-call_test.json: parallel tool-call tasks. ProactiveSeeking/searchTools_test_predictions_kept.json: proactive tool-search tasks. TaskDecomposition/muti-tool-select_test.json: multi-tool task decomposition tasks.… See the full description on the dataset page: https://huggingface.co/datasets/leungtianle/AgentChat-Test.audion<1K0 likes504 downloads3mo agoHugging Face30m-a-p /FineFineWeb-test FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-test.tabulartext-classification1M<n<10M5 likes501 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.