Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kelseye /test_data RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension 🌐 Project Page • 💻 GitHub • 📖 Paper RPC-Bench is a fine-grained benchmark for research paper comprehension. It is built from review-rebuttal exchanges of high-quality academic papers and supports both text-only and visual evaluation through complementary paper representations. Data Structure RPC-Bench is organized into train, dev, and test subsets. Split assignments… See the full description on the dataset page: https://huggingface.co/datasets/kelseye/test_data.question-answering0 likes1k downloads4mo agoHugging Face02suul999922 /x_dataset_test Bittensor Subnet 13 X (Twitter) Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/suul999922/x_dataset_test.texttext-classification10M<n<100M0 likes281 downloads2y agoHugging Face03ShareGPTVideo /test_raw_video_data ShareGPTVideo Raw Videos for Testing data All dataset and models can be found at ShareGPTVideo. Contents: In case of need, this contains raw videos corresponding to test frames in Test video frames textquestion-answering1K<n<10K2 likes115 downloads2y agoHugging Face04AIcell /war-test-dataset War Forecast Bench Dataset for the paper "When AI Navigates the Fog of War" (arXiv:2603.16642). Website: war-forecast-arena.com Overview A temporally grounded benchmark for evaluating LLM reasoning during an ongoing geopolitical conflict. The dataset covers the early stages of the 2026 Middle East conflict, which unfolded after the training cutoff of current frontier models, substantially mitigating training-data leakage concerns. Temporal Nodes… See the full description on the dataset page: https://huggingface.co/datasets/AIcell/war-test-dataset.textquestion-answering1K<n<10K2 likes104 downloads7mo agoHugging Face05GSKCM24 /reddit_dataset_128_test Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/GSKCM24/reddit_dataset_128_test.texttext-classification10K<n<100K0 likes88 downloads2y agoHugging Face06toolevalxm /MultiTaskNLP-TestDataset MultiTaskNLP-Dataset 1. Introduction The MultiTaskNLP-Dataset has undergone significant quality improvements through iterative data curation. In the latest version, we have substantially enhanced the data completeness and label accuracy by implementing rigorous annotation protocols and multi-stage quality assurance mechanisms. The dataset demonstrates outstanding quality metrics across various dimensions, including completeness, accuracy… See the full description on the dataset page: https://huggingface.co/datasets/toolevalxm/MultiTaskNLP-TestDataset.imagetext-classificationn<1K0 likes40 downloads7mo agoHugging Face07closerG /Dataset_Large_test PPU-Bench imagequestion-answeringn<1K0 likes37 downloads5mo agoHugging Face08holyknight101 /knowledgebase-electric_engineering_test_dataThis dataset are based on question answering iterations of this dataset: "STEM-AI-mtl/Electrical-engineering" Question answering using Deepseek R1 from TogetherAI API checkpoint Usage: Reasoning trace data to injecteed as CoT chain in SCIENCE related task. textquestion-answeringn<1K1 likes31 downloads1y agoHugging Face09MTNQLN /thales-dataset-testtextquestion-answering1K<n<10K0 likes23 downloads2y agoHugging Face10zzzzhhh /test_data Dataset Card for "super_glue" Dataset Summary SuperGLUE (https://super.gluebenchmark.com/) is a new benchmark styled after GLUE with a new set of more difficult language understanding tasks, improved resources, and a new public leaderboard. BoolQ (Boolean Questions, Clark et al., 2019a) is a QA task where each example consists of a short passage and a yes/no question about the passage. The questions are provided anonymously and unsolicited by users of the Google search… See the full description on the dataset page: https://huggingface.co/datasets/zzzzhhh/test_data.texttext-classificationn<1K0 likes20 downloads3y agoHugging Face11Kamitor /TestDataQuora Question Answer Dataset (Quora-QuAD) contains 56,402 question-answer pairs scraped from Quora. Usage: For instructions on fine-tuning a model (Flan-T5) with this dataset, please check out the article: https://www.toughdata.net/blog/post/finetune-flan-t5-question-answer-quora-dataset textquestion-answering10K<n<100K0 likes19 downloads2y agoHugging Face12Dockyin /testdata testdata 你好,我是testdata数据集 question-answering0 likes16 downloads3y agoHugging Face13modongsong /mds_test_data Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/modongsong/mds_test_data.text-classification1M<n<10M0 likes15 downloads3y agoHugging Face14Mannesmok /test_dataset_augmentation_reasoningtextquestion-answeringn<1K0 likes15 downloads2y agoHugging Face15JulianVelandia /unal-repository-dataset-test-instructTítulo: Grade Works UNAL Dataset Instruct Test (split 75/25) Descripción: Split 25% del dataset original. Este dataset contiene un formato estructurado de Pregunta: Respuesta generado a partir del contenido de los trabajos de grado del repositorio de la Universidad Nacional de Colombia. Cada registro incluye un fragmento del contenido del trabajo, una pregunta generada a partir de este y su respuesta correspondiente. Este dataset es ideal para tareas de fine-tuning en modelos de lenguaje para… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset-test-instruct.texttable-question-answering1K<n<10K0 likes14 downloads2y agoHugging Face16brcarry /mcp-pymilvus-code-generate-helper-test-dataset Overview This dataset is designed to generate Python code snippets for various functionalities related to Milvus. The test_dataset.json currently contains 139 test cases. Each test case represents a query along with the corresponding documents that have been identified as providing sufficient information to answer that query. Dataset Generation Steps Query Generation For each document, a set of queries is generated based on its length. The current rule is to… See the full description on the dataset page: https://huggingface.co/datasets/brcarry/mcp-pymilvus-code-generate-helper-test-dataset.question-answering1M<n<10M0 likes14 downloads2y agoHugging Face17Hasaranga85 /test-datatest dataset in Alpaca format textquestion-answeringn<1K0 likes12 downloads2y agoHugging Face18Egrigor /Coffee-Making-Test-Dataset Dataset Card for Coffee-Making-Test-Dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/Egrigor/Coffee-Making-Test-Dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Egrigor/Coffee-Making-Test-Dataset.texttext-generationn<1K0 likes12 downloads2y agoHugging Face19syafie /test_datasetdocumenttext-classificationn<1K0 likes10 downloads2y agoHugging Face20revflask /test-datatextquestion-answeringn<1K0 likes10 downloads2y agoHugging Face21alicezzjiang /test_data Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/alicezzjiang/test_data.textquestion-answering10K<n<100K0 likes9 downloads2mo agoHugging Face22MJannik /test-synthetic-dataset Dataset Card for test-synthetic-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/MJannik/test-synthetic-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/MJannik/test-synthetic-dataset.texttext-generationn<1K0 likes6 downloads2y agoHugging Face23Aleksmorshen /Testdataquestion-answering0 likes5 downloads2y agoHugging Face24L0CHINBEK /my-test-dataset Dataset Card for my-test-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/L0CHINBEK/my-test-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/L0CHINBEK/my-test-dataset.texttext-generationn<1K0 likes5 downloads2y agoHugging Face25SerhiiLebediuk /test_modern_datasettextquestion-answeringn<1K0 likes4 downloads2y agoHugging Face26dsguala /Test_dataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/dsguala/Test_dataset.textquestion-answeringn<1K0 likes3 downloads3y agoHugging Face27Yddvcodes /test_datasettextquestion-answeringn<1K0 likes3 downloads3y agoHugging Face28MatthewWhaley /test_dataset1tabularquestion-answeringn<1K0 likes2 downloads4y agoHugging Face29kevinlan888 /test_datatextquestion-answeringn<1K0 likes2 downloads3y agoHugging Face30Murphy2025 /testData1textquestion-answeringn<1K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.