datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CIC-IoT-2023dmi-aarhus-predictions
DMI Aarhus Predictions
Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0.
Primary files
File
Purpose
Produced by
predictions_latest.parquet
Current future + verified prediction store
dmi-collector
frontend_snapshot.json
Primary integration contract for the Vercel frontend
dmi-collector
Compatibility files
File
Status
Notes
predictions.parquet
Legacy
Still read by compatibility… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-predictions.civil_comments
Dataset Card for "civil_comments"
Dataset Summary
The comments in this dataset come from an archive of the Civil Comments
platform, a commenting plugin for independent news sites. These public comments
were created from 2015 - 2017 and appeared on approximately 50 English-language
news sites across the world. When Civil Comments shut down in 2017, they chose
to make the public comments available in a lasting open archive to enable future
research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.CICIoT2023Small
CICIoT2023
This dataset provides a processed derivative of the CICIoT2023 traffic collection. The repository organizes truncated PCAP files and flow-based CSV extractions aligned to the original CICResearch folder hierarchy.
Processing Workflow
The processing pipeline follows four stages:
Source acquisition from the CICResearch CICIoT2023 portal.
Flow extraction from full PCAP files using TriFlowMeter.
PCAP size reduction by truncating packet payloads to 128 bytes with… See the full description on the dataset page: https://huggingface.co/datasets/somnath0100/CICIoT2023Small.glue-ci
Dataset Card for GLUE
Dataset Summary
GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems.
Supported Tasks and Leaderboards
The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks:
ax
A manually-curated evaluation dataset for fine-grained analysis of system… See the full description on the dataset page: https://huggingface.co/datasets/evaluate/glue-ci.dmi-aarhus-weather-data
DMI Aarhus Weather Data
Training data and model artifact dataset for the Aarhus weather pipeline. Maintained by Ciroc0.
Primary files
File
Purpose
Produced by
training_matrix.parquet
Current source of truth for training rows and causal observation context
dmi-collector
model_registry.json
Active bucket registry per target
dmi-ml-trainer
model_meta.json
Training timestamp, sample count and training window
dmi-ml-trainer
temperature_models.pkl… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-weather-data.ciaa-annual-reports
CIAA Annual Reports — Nepali transcripts, ruled tables and chart data
Machine-readable transcripts of the annual reports of Nepal's Commission for the
Investigation of Abuse of Authority (अख्तियार दुरुपयोग अनुसन्धान आयोग, CIAA) —
all 35 it has published to date. The 1st to 35th reports, fiscal years
BS 2047/48 – 2081/82 (AD 1990–2025).
The CIAA publishes these as PDFs whose text layer is, for several years, legacy
pre-Unicode Devanagari that ordinary extractors turn into… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/ciaa-annual-reports.GlotCC-V1
Dataset Summary
GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com/cisnlp/GlotCC.
List of Languages: See https://datasets-server.huggingface.co/splits?dataset=cis-lmu/GlotCC-V1 to get the list of splits available.
Usage (Huggingface Hub -- Recommended)… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotCC-V1.cissp-llmbench
CISSP-LLMBench
cic-ids-2017
CIC-IDS-2017 Dataset
This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use.
Dataset Structure
Configurations
machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs).
traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC.
Raw Data
The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/bvsam/cic-ids-2017.CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles:
Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.cardiac_cine_myops2020
MyoPS 2020 (Multi-Sequence CMR)
Processed NIfTI multi-sequence CMR data (C0/DE/T2) for myocardial pathology segmentation.
Dataset Summary
Modality: CMR (C0, DE, T2)
Task: Segmentation of myocardium/pathology
Splits: train / test
Data Structure (per example)
c0, de, t2, label
Metadata columns listed below
Columns
Imaging
pid
c0, de, t2, label
Metadata (all columns)
orig_spacing_x, orig_spacing_y, orig_spacing_z
n_slices
crop_lower_x… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_myops2020.Cifer-Fraud-Detection-Dataset-AF
📊 Cifer Fraud Detection Dataset
🧠 Overview
The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection.
This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/CiferAI/Cifer-Fraud-Detection-Dataset-AF.asta-summary-citation-counts
Dataset Summary
This dataset tracks which scientific papers are most often cited by Asta, an agentic research platform that uses retrieval-augmented generation (RAG) to answer scientific questions. Each record is a paper cited by Asta's Summarize Literature tool, ranked by the number of times the system cited that paper. Across more than 113,000 user queries, we track 4M citations to over 2M distinct papers. By making this data public, we aim to create a transparent, trackable… See the full description on the dataset page: https://huggingface.co/datasets/allenai/asta-summary-citation-counts.semasia-cifar100
Latents for cifar100 (timm)
This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on cifar100, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability.
Each config corresponds to a single model;
only that model's Parquet files are read on load_dataset.
Usage
Load with datasets and… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-cifar100.speculators-ci-datasets
speculator-tutorial
Raw vs. on-policy regenerated conversation data for training speculative-decoding
drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside
so you can see exactly what regeneration changes and why it matters.
Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B.
Why regenerate at all?
A speculative-decoding drafter is trained to predict what the verifier would say next.
If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.neuro2-neuroscience-datasets
Neuro2 Neuroscience Dataset Atlas (unofficial mirror)
Neuro2 is a catalog and 3D knowledge graph of open neuroscience datasets, created and maintained by Nataliya Kosmyna and Eugene Hauptmann. It indexes dataset records from OpenNeuro, Zenodo, DataCite, DANDI, OSF and dozens of other repositories and links them to authors, papers, tasks, institutions and funders.
This repository is a dated snapshot of that public catalog, reshaped into six Parquet tables you can load with… See the full description on the dataset page: https://huggingface.co/datasets/ciaochris/neuro2-neuroscience-datasets.cardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.msc-cifar100
MSC — Minimum Sufficient Compute
Artifacts for Is Compute Difficulty Architecture-Agnostic? Measuring and
Distilling Per-Sample Minimum Sufficient Computation.
Generated 2026-08-06T03:56:37Z by msc_lib v1.0.0.
Repositories
Shanmuk4622/msc-cifar100 — everything, one folder per run
What MSC is
The smallest cost-normalised configuration at which a network's decision has
stably settled to its full-compute decision, defined uniformly over depth… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/msc-cifar100.lca-ci-builds-repair
🏟️ Long Code Arena (CI builds repair)
This is the benchmark for CI builds repair task as part of the
🏟️ Long Code Arena benchmark.
🛠️ Task. Given the logs of a failed GitHub Actions workflow and the corresponding repository snapshot,
repair the repository contents in order to make the workflow pass.
All the data is collected from repositories published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause, and BSD-2-Clause). The datapoints can be removed upon request.
To… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-ci-builds-repair.cis5300-word-embeddings
Word Embeddings and Semantic Similarity (CIS 5300)
Dataset Description
This dataset supports learning about word embeddings — dense vector representations that capture word meaning. It includes a standard similarity benchmark, a word sense disambiguation task, and a Shakespeare corpus for training custom embeddings.
Configs
SimLex-999: Word Similarity Benchmark
SimLex-999 (Hill et al., 2015) is a gold-standard benchmark for evaluating word… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-word-embeddings.cifar100-224-adv-selfgencivbench-v1
CivBench
A benchmark of LLM agents playing full games of Civilization VI through an
MCP (Model Context Protocol) server. Each run captures the full per-turn
game state, every tool call the agent issued, and the agent's own
structured reflections.
Configs
tables — curated parquet tables, one row per logical record. Use this
for analysis: datasets.load_dataset("civbench/civbench-v1", "tables", split="games").
raw — byte-identical mirror of the on-disk telemetry JSONL… See the full description on the dataset page: https://huggingface.co/datasets/civbench/civbench-v1.cardiac_cine_mnms2
M&Ms2 (Cardiac Cine-MRI, RV Focus)
Processed NIfTI cine-MRI data derived from the M&Ms2 challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation (RV focus)
Views: SAX + LAX 4C (LAX 2C if present)
Splits: train / val / test
Data Structure (per example)
SAX: sax_ed, sax_ed_gt, sax_es, sax_es_gt
LAX 4C: lax_4c_ed, lax_4c_ed_gt, lax_4c_es, lax_4c_es_gt
LAX 2C (if present): lax_2c_ed, lax_2c_ed_gt, lax_2c_es, lax_2c_es_gt
Metadata columns… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms2.CICIOT2023-PARQUET
CICIoT2023 — ipfixprobe flow records (Parquet)
1,479,074,715 bidirectional network flows re-exported from the raw PCAPs of
CICIoT2023 (Canadian
Institute for Cybersecurity, University of New Brunswick) with
ipfixprobe 5.7.0, stored as 308 Parquet
files (~19.4 GB) covering 33 attack classes + benign traffic from the
105-device IoT testbed.
The original dataset ships ~587 GB of PCAPs and CSV features computed with a
closed pipeline. This conversion provides an alternative… See the full description on the dataset page: https://huggingface.co/datasets/Lystea/CICIOT2023-PARQUET.behavior-1k-2026-partial20-submission
BEHAVIOR-1K 2026 Pi0.5 Partial-20 Self-Evaluation
This repository contains an unedited partial self-evaluation for the 2026
BEHAVIOR Challenge.
Scope
BEHAVIOR-1K version: v3.9.1
Policy: Pi0.5 task-embedding policy derived from
IliaLarchenko/behavior-1k-solution
Robot: bundled R1Pro configuration
Evaluation wrapper: omnigibson.eval.wrappers.DefaultWrapper
Public instance indices: 0-9 (instance IDs 301-310)
Rollouts per instance: 1
Evaluated tasks: 20 / 100… See the full description on the dataset page: https://huggingface.co/datasets/circle-great/behavior-1k-2026-partial20-submission.cardiac_cine_mnms
M&Ms (Cardiac Cine-MRI)
Processed NIfTI cine-MRI data derived from the M&Ms (Multi-Centre, Multi-Vendor & Multi-Disease) challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation
Views: SAX (ED/ES)
Splits: train / val / test
Data Structure (per example)
sax_ed, sax_ed_gt
sax_es, sax_es_gt
Optional: sax_t (if present)
Metadata columns listed below
Columns
Imaging
pid
sax_ed, sax_ed_gt, sax_es, sax_es_gt
sax_t (if present)… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms.cinepile_10kcivitai-stable-diffusion-337k
How to Use
from datasets import load_dataset
dataset = load_dataset("thefcraft/civitai-stable-diffusion-337k")
print(dataset['train'][0])
download images
download zip files from images dir
https://huggingface.co/datasets/thefcraft/civitai-stable-diffusion-337k/tree/main/images
it contains some images with id
from zipfile import ZipFile
with ZipFile("filename.zip", 'r') as zObject: zObject.extractall()
Dataset Summary
GitHub URL:-… See the full description on the dataset page: https://huggingface.co/datasets/thefcraft/civitai-stable-diffusion-337k.DREAM-pyramid-circlesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/maedmatt/DREAM-pyramid-circles.
