Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01uoft-cs /cifar10 Dataset Card for CIFAR-10 Dataset Summary The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images. The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar10.imageimage-classification10K<n<100K126 likes177k downloads3y agoHugging Face02uoft-cs /cifar100 Dataset Card for CIFAR-100 Dataset Summary The CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses. There are two labels per image - fine label (actual class) and coarse label (superclass). Supported Tasks and Leaderboards image-classification: The… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar100.imageimage-classification10K<n<100K70 likes34k downloads3y agoHugging Face03cimec /lambada Dataset Card for LAMBADA Dataset Summary The LAMBADA evaluates the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative passages sharing the characteristic that human subjects are able to guess their last word if they are exposed to the whole passage, but not if they only see the last sentence preceding the target word. To succeed on LAMBADA, computational models cannot simply rely on local… See the full description on the dataset page: https://huggingface.co/datasets/cimec/lambada.text10K<n<100K67 likes22k downloads3y agoHugging Face04google /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.tabulartext-classification1M<n<10M40 likes9.8k downloads3y agoHugging Face05Ciroc0 /dmi-aarhus-predictions DMI Aarhus Predictions Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0. Primary files File Purpose Produced by predictions_latest.parquet Current future + verified prediction store dmi-collector frontend_snapshot.json Primary integration contract for the Vercel frontend dmi-collector Compatibility files File Status Notes predictions.parquet Legacy Still read by compatibility… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-predictions.tabular1K<n<10K6 likes8.8k downloads33m agoHugging Face06CIawevy /TextPecker-1.5M TextPecker-1.5M: A Dataset for Training and evaluating TextPecker This repository contains the TextPecker-1.5M dataset, a new benchmark proposed in the paper "TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering". Code and Project Page The official implementation and project details for the TextPecker and TextPecker-1.5M dataset can be found on the GitHub repository: https://github.com/CIawevy/TextPecker Sample Usage You… See the full description on the dataset page: https://huggingface.co/datasets/CIawevy/TextPecker-1.5M.imageimage-to-text1M<n<10M0 likes3.5k downloads7mo agoHugging Face07fahadhafeezofficial /cissp-llmbench CISSP-LLMBench tabulartext-generation10K<n<100K0 likes3.1k downloads3mo agoHugging Face08tanganke /cifar100image10K<n<100K1 likes2.8k downloads2y agoHugging Face09Ciroc0 /dmi-aarhus-weather-data DMI Aarhus Weather Data Training data and model artifact dataset for the Aarhus weather pipeline. Maintained by Ciroc0. Primary files File Purpose Produced by training_matrix.parquet Current source of truth for training rows and causal observation context dmi-collector model_registry.json Active bucket registry per target dmi-ml-trainer model_meta.json Training timestamp, sample count and training window dmi-ml-trainer temperature_models.pkl… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-weather-data.tabular10K<n<100K1 likes2.6k downloads22h agoHugging Face10xinlinzz /cifar-10-cimage100K<n<1M0 likes2.5k downloads1y agoHugging Face11damo-da /ciaa-annual-reports CIAA Annual Reports — Nepali transcripts, ruled tables and chart data Machine-readable transcripts of the annual reports of Nepal's Commission for the Investigation of Abuse of Authority (अख्तियार दुरुपयोग अनुसन्धान आयोग, CIAA) — all 35 it has published to date. The 1st to 35th reports, fiscal years BS 2047/48 – 2081/82 (AD 1990–2025). The CIAA publishes these as PDFs whose text layer is, for several years, legacy pre-Unicode Devanagari that ordinary extractors turn into… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/ciaa-annual-reports.imagetext-retrieval100K<n<1M0 likes2.3k downloads2mo agoHugging Face12evaluate /glue-ci Dataset Card for GLUE Dataset Summary GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. Supported Tasks and Leaderboards The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: ax A manually-curated evaluation dataset for fine-grained analysis of system… See the full description on the dataset page: https://huggingface.co/datasets/evaluate/glue-ci.tabulartext-classification1M<n<10M1 likes2.1k downloads1y agoHugging Face13ibm-research /cif-dataset Cracks in the Foundation A civil-infrastructure visual inspection dataset for instance segmentation with 6 defect/condition categories: Algae · Crack · Net-Crack · Crack with Precipitation · Rust · Spalling Each sample is either a full-resolution inspection image or a 1024×1024 tile derived from one. Tiled samples carry extra fields (tile_row, tile_col, file_name_original, …) that are None for full-resolution samples. Splits Each split is its own parquet shard and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/cif-dataset.imageobject-detection100K<n<1M8 likes2k downloads5mo agoHugging Face14tanganke /cifar10image10K<n<100K1 likes1.8k downloads2y agoHugging Face15flwrlabs /cinic10 Dataset Card for CINIC-10 CINIC-10 has a total of 270,000 images equally split amongst three subsets: train, validate, and test. This means that CINIC-10 has 4.5 times as many samples than CIFAR-10. Dataset Details In each subset (90,000 images), there are ten classes (identical to CIFAR-10 classes). There are 9000 images per class per subset. Using the suggested data split (an equal three-way split), CINIC-10 has 1.8 times as many training samples as in CIFAR-10.… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/cinic10.imageimage-classification100K<n<1M2 likes1.8k downloads2y agoHugging Face16CIIRC-NLP /mmlu-cs Czech MMLU This is a Czech translation of the original MMLU dataset, created using the WMT 21 En-X model. The 'auxiliary_train' subset is not included. The translation was completed for use within the Czech-Bench evaluation framework. The script used for translation can be reviewed here. Citation Original dataset: @article{hendryckstest2021, title={Measuring Massive Multitask Language Understanding}, author={Dan Hendrycks and Collin Burns and Steven Basart and… See the full description on the dataset page: https://huggingface.co/datasets/CIIRC-NLP/mmlu-cs.textmultiple-choice10K<n<100K0 likes1.6k downloads2y agoHugging Face17Chris1 /cityscapesimage1K<n<10K5 likes1.5k downloads4y agoHugging Face18bvsam /cic-ids-2017 CIC-IDS-2017 Dataset This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use. Dataset Structure Configurations machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs). traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC. Raw Data The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/bvsam/cic-ids-2017.tabulartabular-classification1M<n<10M3 likes1.5k downloads10mo agoHugging Face19allenai /asta-summary-citation-counts Dataset Summary This dataset tracks which scientific papers are most often cited by Asta, an agentic research platform that uses retrieval-augmented generation (RAG) to answer scientific questions. Each record is a paper cited by Asta's Summarize Literature tool, ranked by the number of times the system cited that paper. Across more than 113,000 user queries, we track 4M citations to over 2M distinct papers. By making this data public, we aim to create a transparent, trackable… See the full description on the dataset page: https://huggingface.co/datasets/allenai/asta-summary-citation-counts.tabular100M<n<1B11 likes1.4k downloads23h agoHugging Face20inference-optimization /speculators-ci-datasets speculator-tutorial Raw vs. on-policy regenerated conversation data for training speculative-decoding drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside so you can see exactly what regeneration changes and why it matters. Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B. Why regenerate at all? A speculative-decoding drafter is trained to predict what the verifier would say next. If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.tabulartext-generation1K<n<10K0 likes1.2k downloads2mo agoHugging Face21noahshinn /cifar100_2_to_100_constant_size_dataset Dataset Card for "cifar100_2_to_100_constant_size_dataset" More Information needed image10K<n<100K0 likes1.1k downloads3y agoHugging Face22spaicom-lab /semasia-cifar100 Latents for cifar100 (timm) &nbsp;&nbsp;&nbsp; This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on cifar100, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability. Each config corresponds to a single model; only that model's Parquet files are read on load_dataset. Usage Load with datasets and… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-cifar100.tabularfeature-extraction100M<n<1B0 likes1.1k downloads3mo agoHugging Face23zhaochenyang20 /mmsu-ci-2000audio1K<n<10K0 likes1k downloads6mo agoHugging Face24google-research-datasets /circa Dataset Card for CIRCA Dataset Summary The Circa (meaning ‘approximately’) dataset aims to help machine learning systems to solve the problem of interpreting indirect answers to polar questions. The dataset contains pairs of yes/no questions and indirect answers, together with annotations for the interpretation of the answer. The data is collected in 10 different social conversational situations (eg. food preferences of a friend). The following are the situational… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/circa.texttext-classification10K<n<100K7 likes983 downloads3y agoHugging Face25zhaochenyang20 /mmmu-ci-50imagen<1K0 likes962 downloads6mo agoHugging Face26WNJXYK /TTA-Cityscapes-Cimage1K<n<10K1 likes942 downloads4mo agoHugging Face27samansmink /duckdb_ci_teststextn<1K0 likes929 downloads2y agoHugging Face28hyin-ustc /CIRCLE-40K Highlights 40,000 high-quality video-based spatial reasoning training samples Built primarily from five large-scale real-world indoor 3D scene datasets (ScanNet, ScanNet++, S3DIS, ARKitScenes, Aria Digital Twin), plus a small ProcTHOR simulated subset Covers diverse spatial skills: geometric perception, spatial relations, counting, and temporal / appearance-order reasoning over video Filtered with rejection sampling using Qwen3-VL-8B-Instruct to reduce ambiguous or low-quality… See the full description on the dataset page: https://huggingface.co/datasets/hyin-ustc/CIRCLE-40K.textvisual-question-answering10K<n<100K0 likes894 downloads3mo agoHugging Face29dragonintelligence /CIFAKE-image-datasetimage100K<n<1M0 likes828 downloads2y agoHugging Face30CircleRadon /EOC-Bench EOC-Bench : Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? 🔍 Overview we introduce EOC-Bench, an innovative benchmark designed to systematically evaluate object-centric embodied cognition in dynamic egocentric scenarios. Specially, EOC-Bench features 3,277 meticulously annotated QA pairs categorized into three temporal categories: Past, Present, and Future, covering 11 fine-grained evaluation dimensions and 3 visual object referencing… See the full description on the dataset page: https://huggingface.co/datasets/CircleRadon/EOC-Bench.tabular1K<n<10K6 likes802 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.