datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers_circleci_workflow_runscifar10
Dataset Card for CIFAR-10
Dataset Summary
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar10.Glot500
Glot500 Corpus
A dataset of natural language data collected by putting together more than 150
existing mono-lingual and multilingual datasets together and crawling known multilingual websites.
The focus of this dataset is on 500 extremely low-resource languages.
(More Languages still to be uploaded here)
This dataset is used to train the Glot500 model.
Homepage: homepage
Repository: github
Paper: acl, arxiv
This dataset has the identical data format as the Taxi1500 Raw Data… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/Glot500.cifar100
Dataset Card for CIFAR-100
Dataset Summary
The CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).
Supported Tasks and Leaderboards
image-classification: The… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar100.ciSciOL-CI
Scientific Openly-Licensed Publications
This repository contains companion material for the following publication:
Tim Tarsi, Heike Adel, Jan Hendrik Metzen, Dan Zhang, Matteo Finco, Annemarie Friedrich. SciOL and MuLMS-Img: Introducing A Large-Scale Multimodal Scientific Dataset and Models for Image-Text Tasks in the Scientific Domain. WACV 2024.
Please cite this paper if using the dataset, and direct any questions regarding the dataset
to Tim Tarsi
Summary
Scientific… See the full description on the dataset page: https://huggingface.co/datasets/Timbrt/SciOL-CI.lambada
Dataset Card for LAMBADA
Dataset Summary
The LAMBADA evaluates the capabilities of computational models
for text understanding by means of a word prediction task.
LAMBADA is a collection of narrative passages sharing the characteristic
that human subjects are able to guess their last word if
they are exposed to the whole passage, but not if they
only see the last sentence preceding the target word.
To succeed on LAMBADA, computational models cannot
simply rely on local… See the full description on the dataset page: https://huggingface.co/datasets/cimec/lambada.ciel-knowledge-basehfh_ci_scan_dataset_btransformers_flash_attn_cicivil_comments
Dataset Card for "civil_comments"
Dataset Summary
The comments in this dataset come from an archive of the Civil Comments
platform, a commenting plugin for independent news sites. These public comments
were created from 2015 - 2017 and appeared on approximately 50 English-language
news sites across the world. When Civil Comments shut down in 2017, they chose
to make the public comments available in a lasting open archive to enable future
research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.CIC-IoT-2023CircuitSense
CircuitSense
This dataset is a comprehensive multimodal circuit question-answering benchmark designed to evaluate visual reasoning and problem-solving capabilities across three main domains: Perception, Analysis, and Design. The dataset contains structured question-answer pairs with accompanying visual content, targeting different engineering cognitive levels and reasoning tasks.
Dataset Structure
The dataset is organized into three primary folders, each containing… See the full description on the dataset page: https://huggingface.co/datasets/armanakbari4/CircuitSense.dmi-aarhus-predictions
DMI Aarhus Predictions
Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0.
Primary files
File
Purpose
Produced by
predictions_latest.parquet
Current future + verified prediction store
dmi-collector
frontend_snapshot.json
Primary integration contract for the Vercel frontend
dmi-collector
Compatibility files
File
Status
Notes
predictions.parquet
Legacy
Still read by compatibility… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-predictions.CiQi-VQA
CiQi-Agent
Github | Model | Dataset | Paper
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
Accepted to ECCV 2026
🎯 Overview
CiQi-Agent has been accepted to ECCV 2026.
We present CiQi-Agent, a domain-specific multimodal agent for antique Chinese porcelain connoisseurship. The project is designed to combine fine-grained visual perception, tool-augmented reasoning, and cultural-heritage knowledge… See the full description on the dataset page: https://huggingface.co/datasets/SII-Monument-Valley/CiQi-VQA.transformers_pr_ciCIDM-v1cityscapesversion https://git-lfs.github.com/spec/v1
oid sha256:4bcf87ecfbbb8e07a01b21415a970c8b53a5283bf6872b657040d3f45c9241f7
size 31
transformers_daily_ciwm_imagined
Imagined Data
This repository hosts imagined interaction data generated by world models across different environments, tasks, and data sources. Data are organized into separate subdatasets, with additional types of imagined data to be added over time.
The repository currently contains only the RoboTwin2.0 subdataset. Storage formats, field definitions, and loading instructions are documented in the corresponding section for each subdataset.
Dataset Index… See the full description on the dataset page: https://huggingface.co/datasets/Cirquar-Tech/wm_imagined.CinC_challenge_2021fin_whale_1
Paper Informations
Self-Supervised vs Supervised Representation Learning for Fin Whale Vocalization Detection
by Adam Chareyre, Haodong Zhang, Shuwen Ge, Randall Balestriero, Sébastien Paris, Hervé Glotin
OpenReview link: https://openreview.net/forum?id=fCe4l0T320
Requirements
Update path files
Depending of the method you want to use (SSL or Supervised), you will need to fill some path.
.env file, SSL_PRETRAINED/DP_run.sh file, SSL_PRETRAINED/fine_tune.sh file… See the full description on the dataset page: https://huggingface.co/datasets/CIANLabxBROWNUniv/fin_whale_1.cifar-10-pythoncircleci-test-resultstransformers_daily_ciGeoBenchMeta
GeoBench: A Benchmark for Geometric Image Editing
This repository contains the GeoBench benchmark dataset, introduced in the paper Training-Free Diffusion for Geometric Image Editing.
Project Page & Code: https://github.com/CIawevy/FreeFine
GeoBench is designed to evaluate the capability of diffusion models in geometric image editing tasks. It supports various scenarios including object repositioning, reorientation, reshaping, fine-grained partial editing, structure completion… See the full description on the dataset page: https://huggingface.co/datasets/CIawevy/GeoBenchMeta.TextPecker-1.5M
TextPecker-1.5M: A Dataset for Training and evaluating TextPecker
This repository contains the TextPecker-1.5M dataset, a new benchmark proposed in the paper "TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering".
Code and Project Page
The official implementation and project details for the TextPecker and TextPecker-1.5M dataset can be found on the GitHub repository:
https://github.com/CIawevy/TextPecker
Sample Usage
You… See the full description on the dataset page: https://huggingface.co/datasets/CIawevy/TextPecker-1.5M.semasia-cifar10
Latents for cifar10 (timm)
This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on cifar10, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability.
Each config corresponds to a single model;
only that model's Parquet files are read on load_dataset.
Usage
Load with datasets and… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-cifar10.CI-VID
📄 CI-VID: A Coherent Interleaved Text-Video Dataset
CI-VID is a large-scale dataset designed to advance coherent multi-clip video generation. Unlike traditional text-to-video (T2V) datasets with isolated clip-caption pairs, CI-VID supports text-and-video-to-video (TV2V) generation by providing over 340,000 interleaved sequences of video clips and rich captions. It enables models to learn both intra-clip content and inter-clip transitions, fostering story-driven generation with… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CI-VID.grad_cilp0.28_100
