Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01programbench /ProgramBench-Tests ProgramBench Generated Tests This dataset contains the AI-generated behavioral test suites used to evaluate model solutions in ProgramBench. ProgramBench is a benchmark that evaluates whether language models can rebuild programs from scratch. Given only a compiled binary and its documentation, AI agents must architect and implement a complete codebase that reproduces the original program's behavior. These test suites are used to assess whether a candidate solution is behaviorally… See the full description on the dataset page: https://huggingface.co/datasets/programbench/ProgramBench-Tests.text-generationn<1K11 likes95k downloads5mo agoHugging Face02commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes6.4k downloads25d agoHugging Face03diyer22 /Panda-CVL-test Panda-CVL Test Split This dataset is based on the paper onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction. Code is available at GitHub. Overview Panda-CVL is a token-level correction dataset and benchmark annotated with the onPanda tool. Given a question-response pair, the model first judges whether the response is acceptable. If correction is needed, it must locate the first inappropriate token and replace it… See the full description on the dataset page: https://huggingface.co/datasets/diyer22/Panda-CVL-test.text-generationn<1K0 likes984 downloads17d agoHugging Face04Alga2025 /UltraData-Math-TEST UltraData-Math 🤗 Dataset | 💻 Source Code | 🇨🇳 中文 README UltraData-Math is a large-scale, high-quality mathematical pre-training dataset totaling 290B+ tokens across three progressive tiers—L1 (170.5B tokens web corpus), L2 (33.7B tokens quality-selected), and L3 (88B tokens multi-format refined)—designed to systematically enhance mathematical reasoning in LLMs. It has been applied to the mathematical pre-training of the MiniCPM Series models. 🆕 What's New… See the full description on the dataset page: https://huggingface.co/datasets/Alga2025/UltraData-Math-TEST.texttext-generation100M<n<1B0 likes916 downloads8mo agoHugging Face05EleutherAI /pile_val_test The Pile: Validation and Test Splits This repo contains the validation and test splits of The Pile, an 825 GiB English text dataset designed for training large language models. Files File Split Size val.jsonl Validation 1.4 GB test.jsonl Test 1.3 GB Format Each line is a JSON object with two fields: {"text": "The document text...", "meta": {"pile_set_name": "Pile-CC"}} The meta.pile_set_name field indicates which of the 22 constituent… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/pile_val_test.texttext-generation100K<n<1M0 likes803 downloads8mo agoHugging Face06neil-code /dialogsum-test Dataset Card for DIALOGSum Corpus Dataset Description Links Homepage: https://aclanthology.org/2021.findings-acl.449 Repository: https://github.com/cylnlp/dialogsum Paper: https://aclanthology.org/2021.findings-acl.449 Point of Contact: https://huggingface.co/knkarthick Dataset Summary DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 (Plus 100 holdout data for topic generation) dialogues with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/neil-code/dialogsum-test.textsummarization1K<n<10K15 likes770 downloads3y agoHugging Face07sigcp /hardtests_tests Dataset Card for HARDTESTS Tests HARDTESTS Tests is the test suite of HARDTESTS, a competitive programming dataset. The dataset contains multiple .parquet files. Each instance contains structured test case objects for one problem. This dataset is presented in the paper HardTests: Synthesizing High-Quality Test Cases for LLM Coding. Project Page Data Summary The test suite is generated using the HARDTESTSGEN pipeline. The dataset contains the generated test suites of… See the full description on the dataset page: https://huggingface.co/datasets/sigcp/hardtests_tests.text-generation2 likes671 downloads10mo agoHugging Face08WideSeek-R1 /WideSeek-R1-test-data Testing Dataset We provide test.jsonl, a testing split for evaluating WideSeek-R1 on the standard WideSearch dataset. All examples are sourced from WideSearch; we only convert them into a format that is directly compatible with the WideSeek-R1 evaluation scripts. This makes the dataset plug-and-play—no additional configuration required. texttext-generationn<1K0 likes640 downloads5mo agoHugging Face09InfiX-ai /android_control_test Processed Android Control Test Set for InfiGUI-R1 Evaluation Dataset Description This repository contains the processed test set derived from the Android Control dataset by Google Research. It has been specifically prepared for evaluating the performance of our model, InfiGUI-R1. The InfiGUI-R1 model is detailed in our paper: InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners This dataset facilitates standardized testing and… See the full description on the dataset page: https://huggingface.co/datasets/InfiX-ai/android_control_test.text-generation1K<n<10K2 likes612 downloads1y agoHugging Face10davanstrien /encyclopaedia-britannica-lance-test Encyclopaedia Britannica (1771-1860) - Lance Format This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading. Dataset Details Total Pages: 155,388 Total Volumes: 195 Format: Lance (columnar format with blob storage for images) Source: National Library of Scotland (NLS) License: Public Domain (CC0) Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia-britannica-lance-test.imageimage-to-text100K<n<1M0 likes556 downloads8mo agoHugging Face11m-a-p /FineFineWeb-test FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-test.tabulartext-classification1M<n<10M5 likes501 downloads2y agoHugging Face12lenadan /otel-test-snippet-jsonl ⚠️ TEST DATASET - DO NOT USE FOR PRODUCTION This is a small test snippet for internal validation purposes only. This dataset contains a subset of OpenTelemetry traces from various LLM inference benchmarks. It is intended for testing dataset infrastructure and should NOT be used for research, benchmarking, or production purposes. Dataset Structure The dataset contains OpenTelemetry traces organized by: Benchmark: appworld, tau2_telecom Agent Framework: openai_solo… See the full description on the dataset page: https://huggingface.co/datasets/lenadan/otel-test-snippet-jsonl.text-generationn<1K0 likes458 downloads5mo agoHugging Face13jitx /Methods2Test_java_unit_test_code Dataset Description Microsoft created this large dataset of Java Junit test cases with its corresponding focal methods. It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K Java open source project hosted on GitHub. The mapping between test case and focal methods are based heuristics rules and Java developer's best practice. More information could be found here: methods2test Github repo Methods2Test: A dataset of focal methods… See the full description on the dataset page: https://huggingface.co/datasets/jitx/Methods2Test_java_unit_test_code.texttext-generation100K<n<1M18 likes383 downloads3y agoHugging Face14breadlicker45 /Bread-chatbot-dataset-test Dataset Card for "Bread-chatbot-dataset-test" More Information needed texttext-generation1M<n<10M0 likes374 downloads3y agoHugging Face15test-time-compute /aime_2025 AIME 2025 - Unified Test-Time Scaling Format This is the AIME (American Invitational Mathematics Examination) 2025 dataset in a unified format for test-time scaling experiments. Dataset Description Source: MathArena/aime_2025 Size: 30 competition-level mathematics problems Format: Unified TTS format (question, answer, metadata) Dataset Structure Fields question (string): The mathematical problem statement answer (string): The numerical answer… See the full description on the dataset page: https://huggingface.co/datasets/test-time-compute/aime_2025.textquestion-answeringn<1K0 likes366 downloads1y agoHugging Face16JohnBeanerson /pi-mono-test Coding agent session traces DUPLICATED FROM badlogicgames/pi-mono FOR TESTING PURPOSES. ORIGINAL README BELOW This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-mono.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are… See the full description on the dataset page: https://huggingface.co/datasets/JohnBeanerson/pi-mono-test.text-generation0 likes347 downloads6mo agoHugging Face17JeremyAlain /123_testThe Fewshot Table dataset consists of tables that naturally occur on the web, that are formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. The dataset consists of approximately 413K tables that are extracted from the WDC Web Table Corpora 2015, which is released under the Apache-2.0 license. The WDC Web Table Corpora "contains vast amounts of HTML tables. [...] The Web Data Commons project extracts relational Web tables from the Common Crawl, the largest and most up-to-date Web corpus that is currently available to the public."textmultiple-choicen<1K2 likes336 downloads4y agoHugging Face18AIM-Intelligence /COMPASS-Policy-Alignment-Testbed-Dataset COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs This dataset evaluates how well Large Language Models (LLMs) follow organization-specific policies in realistic enterprise-style settings. What is COMPASS? COMPASS is a framework for evaluating policy alignment: given only an organization’s policy (e.g., allow/deny rules), it enables you to benchmark whether an LLM’s responses comply with that policy in structured, enterprise-like… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-Alignment-Testbed-Dataset.texttext-generation1K<n<10K12 likes319 downloads1mo agoHugging Face19google /code_x_glue_cc_cloze_testing_all Dataset Card for "code_x_glue_cc_cloze_testing_all" Dataset Summary CodeXGLUE ClozeTesting-all dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-all Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem. Here we… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_all.texttext-generation100K<n<1M6 likes290 downloads3y agoHugging Face20suul999922 /x_dataset_test Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/suul999922/x_dataset_test.texttext-classification10M<n<100M0 likes281 downloads2y agoHugging Face21boxin-wbx /test Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/boxin-wbx/test.tabulartext-classification10K<n<100K0 likes277 downloads3y agoHugging Face22Emulated-Inc /python-unit-test-training-pool Python unit test training pool A pool of public data for training a model to write tests for Python code. It is a straight collection of open datasets, not a new corpus: every row comes from one of the sources below, at the revision named, and the only rows removed are the ones an overlap filter flagged against held-out material this pool is kept separate from. Every row of the normalised layer pairs a program with tests for it. That is the point of the pool, and it is why the… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/python-unit-test-training-pool.texttext-generation1M<n<10M0 likes259 downloads28d agoHugging Face23AnonyCAD /test0611imagetext-generationn<1K0 likes249 downloads2y agoHugging Face24davanstrien /encyclopaedia-britannica-lance-test2 Encyclopaedia Britannica (1771-1860) - Lance Format This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading. Dataset Details Total Pages: 155,388 Total Volumes: 195 Format: Lance (columnar format with blob storage for images) Source: National Library of Scotland (NLS) License: Public Domain (CC0) Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia-britannica-lance-test2.imageimage-to-text100K<n<1M0 likes242 downloads8mo agoHugging Face25google /code_x_glue_cc_cloze_testing_maxmin Dataset Card for "code_x_glue_cc_cloze_testing_maxmin" Dataset Summary CodeXGLUE ClozeTesting-maxmin dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-maxmin Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem.… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_maxmin.texttext-generation1K<n<10K3 likes238 downloads3y agoHugging Face26FreedomIntelligence /huatuo26M-testdatasets Dataset Card for huatuo26M-testdatasets Dataset Summary We are pleased to announce the release of our evaluation dataset, a subset of the Huatuo-26M. This dataset contains 6,000 entries that we used for Natural Language Generation (NLG) experimentation in our associated research paper. We encourage researchers and developers to use this evaluation dataset to gauge the performance of their own models. This is not only a chance to assess the accuracy and relevancy of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo26M-testdatasets.texttext-generation1K<n<10K22 likes229 downloads3y agoHugging Face27rabbit-hmi /MM-Mind2Web-tilde_test_snapshot_20dist MultiModal-Mind2Web~ (MM-Mind2Web~) rabbit inc. [Leaderboard & Blogpost to be released] Configuration: test split, snapshot with seed 42, 20 distractors Multimodal-Mind2Web is a dataset proposed by Boyuan et al.. It's designed for the development and evaluation of generalist web agents and includes various action trajectories of humans on real websites. We've simplified the raw dump from both Multimodal-Mind2Web and Mind2Web into sequences of observation-action pairs. We've… See the full description on the dataset page: https://huggingface.co/datasets/rabbit-hmi/MM-Mind2Web-tilde_test_snapshot_20dist.texttext-generation1K<n<10K2 likes187 downloads2y agoHugging Face28test-alexpouliquen /chi-bench Clinical Healthcare In-Situ Environment Task fixtures for a long-horizon, policy-rich healthcare-workflow agent benchmark What is in this dataset CHI-Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1… See the full description on the dataset page: https://huggingface.co/datasets/test-alexpouliquen/chi-bench.documenttext-generationn<1K0 likes186 downloads5mo agoHugging Face29summykai /chem-rlvr-TEST ChemBench-RLVR: Comprehensive Chemistry Dataset for Reinforcement Learning from Verifiable Rewards Dataset Description ChemBench-RLVR is a high-quality, balanced dataset containing 7,001 question-answer pairs across 14 chemistry task types. This dataset is specifically designed for training language models using Reinforcement Learning from Verifiable Rewards (RLVR), where all answers are computationally verifiable using established cheminformatics tools. Key… See the full description on the dataset page: https://huggingface.co/datasets/summykai/chem-rlvr-TEST.tabularquestion-answering1K<n<10K0 likes182 downloads1y agoHugging Face30Infektyd /syntra-testing-evals-v4 SyntraTesting Evals v4 Complete benchmark suite for evaluating AI models on advanced reasoning tasks. Contents Split File Description prompts data/splits/prompts.tar.gz (~60KB) CMT prompts, coherence structures, drift resilience, logic, ethics benchmarks data/splits/benchmarks.tar.gz (~36KB) ARC, CMT, GSM8K benchmark data and utilities runs data/splits/runs.tar.gz (~4.4MB) Sample evaluation runs with graded results resources data/splits/resources.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/syntra-testing-evals-v4.text-generation0 likes181 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.