Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google-research-datasets /mbpp Dataset Card for Mostly Basic Python Problems (mbpp) Dataset Summary The benchmark consists of around 1,000 crowd-sourced Python programming problems, designed to be solvable by entry level programmers, covering programming fundamentals, standard library functionality, and so on. Each problem consists of a task description, code solution and 3 automated test cases. As described in the paper, a subset of the data has been hand-verified by us. Released here as part of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/mbpp.text1K<n<10K260 likes268k downloads3y agoHugging Face02applied-ai-018 /peacock-data-public-datasets-idc0 likes250k downloads2y agoHugging Face03GokuScraper /seedance-2-prompts-datasets 🎞️ Seedance-2-prompts-datasets 🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators. This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.imagetext-to-video1K<n<10K48 likes199k downloads4d agoHugging Face04google-research-datasets /paws Dataset Card for PAWS: Paraphrase Adversaries from Word Scrambling Dataset Summary PAWS: Paraphrase Adversaries from Word Scrambling This dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of paraphrase identification. The dataset has two subsets, one based on Wikipedia and the other one based on the Quora Question Pairs (QQP) dataset. For further… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/paws.texttext-classification100K<n<1M41 likes134k downloads3y agoHugging Face05Goku-OpenLab /gpt-image-2-prompts-datasets 🖼️ GPT Image 2 Prompt Dataset 🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset. Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.imagetext-to-image10K<n<100K8 likes111k downloads4d agoHugging Face06autogluon /fev_datasets Forecast evaluation datasets This repository contains time series datasets that can be used for evaluation of univariate & multivariate forecasting models. The main focus of this repository is on datasets that reflect real-world forecasting scenarios, such as those involving covariates, missing values, and other practical complexities. The datasets follow a format that is compatible with the fev package. Data format and usage Each dataset satisfies the following… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/fev_datasets.tabulartime-series-forecasting100K<n<1M13 likes86k downloads9mo agoHugging Face07legacy-datasets /wikipediaWikipedia dataset containing cleaned articles of all languages. The datasets are built from the Wikipedia dump (https://dumps.wikimedia.org/) with one split per language. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.).text-generationn<1K675 likes56k downloads3y agoHugging Face08yifengzhu-hf /LIBERO-datasets LIBERO Datasets This is a repo that stores the LIBERO datasets. The structure of the dataset can be found below: libero_object/ libero_spatial/ libero_goal/ libero_90/ libero_10/ Demonstrations of each task is stored in a hdf5 file. Please refer to download script from the official LIBERO repo for more details. 67 likes43k downloads1y agoHugging Face09albertvillanova /datasets-tests-compressiontextn<1K0 likes43k downloads5y agoHugging Face10serialexperimentsleon /fish_datasets_real_electrodyn_expertsys_twodim_fourier0 likes42k downloads1y agoHugging Face11legacy-datasets /banking77 Dataset Card for BANKING77 Dataset Summary Deprecated: Dataset "banking77" is deprecated and will be deleted. Use "PolyAI/banking77" instead. Dataset composed of online banking queries annotated with their corresponding intents. BANKING77 dataset provides a very fine-grained set of intents in a banking domain. It comprises 13,083 customer service queries labeled with 77 intents. It focuses on fine-grained single-domain intent detection. Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/legacy-datasets/banking77.texttext-classification10K<n<100K52 likes30k downloads3y agoHugging Face12autogluon /chronos_datasets Chronos datasets Time series datasets used for training and evaluation of the Chronos forecasting models. Note that some Chronos datasets (ETTh, ETTm, brazilian_cities_temperature and spanish_energy_and_weather) that rely on a custom builder script are available in the companion repo autogluon/chronos_datasets_extra. See the paper for more information. Data format and usage The recommended way to use these datasets is via https://github.com/autogluon/fev. All datasets… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/chronos_datasets.tabulartime-series-forecasting10M<n<100M79 likes30k downloads2y agoHugging Face13detection-datasets /cocoimageobject-detection100K<n<1M92 likes26k downloads4y agoHugging Face14Goku-OpenLab /nano-banana-pro-prompts-datasets 🖼️ Nano Banana Pro Prompt Dataset 🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.imagetext-to-image10K<n<100K3 likes26k downloads2mo agoHugging Face15harborframework /harbor-datasets0 likes26k downloads6mo agoHugging Face16google-research-datasets /nq_open Dataset Card for nq_open Dataset Summary The NQ-Open task, introduced by Lee et.al. 2019, is an open domain question answering benchmark that is derived from Natural Questions. The goal is to predict an English answer string for an input English question. All questions can be answered using the contents of English Wikipedia. Supported Tasks and Leaderboards Open Domain Question-Answering, EfficientQA Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/nq_open.textquestion-answering10K<n<100K36 likes25k downloads3y agoHugging Face17ethz-vlg /mv3dpt-datasets Multi-View 3D Point Tracking Datasets This repository hosts the training and evaluation datasets associated with the paper Multi-View 3D Point Tracking. Project Page: https://ethz-vlg.github.io/mvtracker/ Code/Github Repository: https://github.com/ethz-vlg/mvtracker Abstract We introduce the first data-driven multi-view 3D point tracker, designed to track arbitrary points in dynamic scenes using multiple camera views. Unlike existing monocular trackers, which struggle… See the full description on the dataset page: https://huggingface.co/datasets/ethz-vlg/mv3dpt-datasets.keypoint-detection3 likes23k downloads7mo agoHugging Face18zalando-datasets /fashion_mnist Dataset Card for FashionMNIST Dataset Summary Fashion-MNIST is a dataset of Zalando's article images—consisting of a training set of 60,000 examples and a test set of 10,000 examples. Each example is a 28x28 grayscale image, associated with a label from 10 classes. We intend Fashion-MNIST to serve as a direct drop-in replacement for the original MNIST dataset for benchmarking machine learning algorithms. It shares the same image size and structure of training and testing… See the full description on the dataset page: https://huggingface.co/datasets/zalando-datasets/fashion_mnist.imageimage-classification10K<n<100K67 likes21k downloads2y agoHugging Face19RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M95 likes19k downloads1y agoHugging Face20datasets-maintainers /dataset-with-standalone-yamlThis is a test dataset used in the datasets library CI textn<1K0 likes19k downloads3y agoHugging Face21google-research-datasets /go_emotions Dataset Card for GoEmotions Dataset Summary The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral. The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test splits. Supported Tasks and Leaderboards This dataset is intended for multi-class, multi-label emotion classification. Languages The data is in English. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.tabulartext-classification100K<n<1M271 likes19k downloads3y agoHugging Face22legacy-datasets /c4A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset by AllenAI.text-generation100M<n<1B242 likes17k downloads3y agoHugging Face23rishitdagli /nerf-gs-datasetsI keep a collection compiled of existing datasets from various sources for training NeRFs or Splats. This dataset is most of that collection. All of the individual scenes also have a trained Gaussian Splat. https://rishit-dagli.github.io/2025/03/28/nerf-gs-datasets.html image1K<n<10K2 likes14k downloads2y agoHugging Face24toloka /crowdkit-datasets0 likes13k downloads3y agoHugging Face25google-research-datasets /conceptual_captions Dataset Card for Conceptual Captions Dataset Summary Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/conceptual_captions.imageimage-to-text1M<n<10M111 likes13k downloads2y agoHugging Face26google-research-datasets /natural_questions Dataset Card for Natural Questions Dataset Summary The NQ corpus contains questions from real users, and it requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read an entire page to find the answer, cause NQ to be a more realistic and challenging task than prior QA datasets. Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/natural_questions.textquestion-answering10K<n<100K128 likes13k downloads3y agoHugging Face27community-datasets /quarel Dataset Card for "quarel" Dataset Summary QuaRel is a crowdsourced dataset of 2771 multiple-choice story questions, including their logical forms. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 0.63 MB Size of the generated dataset: 1.53 MB Total amount of disk used: 2.17 MB An example of 'train'… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/quarel.text1K<n<10K2 likes12k downloads2y agoHugging Face28applied-ai-018 /peacock-data-public-datasets-ywang300 likes12k downloads2y agoHugging Face29defunct-datasets /eli5Explain Like I'm 5 long form QA dataset100K<n<1M52 likes12k downloads3y agoHugging Face30robometer /processed_datasets2 likes11k downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.