datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CIC-IoT-2023imdb-ciCIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles:
Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.cardiac_cine_myops2020
MyoPS 2020 (Multi-Sequence CMR)
Processed NIfTI multi-sequence CMR data (C0/DE/T2) for myocardial pathology segmentation.
Dataset Summary
Modality: CMR (C0, DE, T2)
Task: Segmentation of myocardium/pathology
Splits: train / test
Data Structure (per example)
c0, de, t2, label
Metadata columns listed below
Columns
Imaging
pid
c0, de, t2, label
Metadata (all columns)
orig_spacing_x, orig_spacing_y, orig_spacing_z
n_slices
crop_lower_x… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_myops2020.Cifer-Fraud-Detection-Dataset-AF
📊 Cifer Fraud Detection Dataset
🧠 Overview
The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection.
This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/CiferAI/Cifer-Fraud-Detection-Dataset-AF.cardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.cis5300-word-embeddings
Word Embeddings and Semantic Similarity (CIS 5300)
Dataset Description
This dataset supports learning about word embeddings — dense vector representations that capture word meaning. It includes a standard similarity benchmark, a word sense disambiguation task, and a Shakespeare corpus for training custom embeddings.
Configs
SimLex-999: Word Similarity Benchmark
SimLex-999 (Hill et al., 2015) is a gold-standard benchmark for evaluating word… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-word-embeddings.FinQAcardiac_cine_mnms2
M&Ms2 (Cardiac Cine-MRI, RV Focus)
Processed NIfTI cine-MRI data derived from the M&Ms2 challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation (RV focus)
Views: SAX + LAX 4C (LAX 2C if present)
Splits: train / val / test
Data Structure (per example)
SAX: sax_ed, sax_ed_gt, sax_es, sax_es_gt
LAX 4C: lax_4c_ed, lax_4c_ed_gt, lax_4c_es, lax_4c_es_gt
LAX 2C (if present): lax_2c_ed, lax_2c_ed_gt, lax_2c_es, lax_2c_es_gt
Metadata columns… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms2.cardiac_cine_mnms
M&Ms (Cardiac Cine-MRI)
Processed NIfTI cine-MRI data derived from the M&Ms (Multi-Centre, Multi-Vendor & Multi-Disease) challenge.
Dataset Summary
Modality: CMR cine MRI
Task: LV/RV/MYO segmentation
Views: SAX (ED/ES)
Splits: train / val / test
Data Structure (per example)
sax_ed, sax_ed_gt
sax_es, sax_es_gt
Optional: sax_t (if present)
Metadata columns listed below
Columns
Imaging
pid
sax_ed, sax_ed_gt, sax_es, sax_es_gt
sax_t (if present)… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms.CICIDS-2017Raw network data was collected over a period of 5 days, Monday through Friday, and stored in PCAP files.
Monday was used to create most of the Benign data, while the Attack-Network implemented various types of attacks over the next 4 days,
such as Brute Force connections (FTP and SSH), several types of DoS attacks, as well as a Botnet attack, Infiltration attacks and subsequent Port-Scanning activity.
The PCAP data was processed using a tool developed by one of the authors of [1], called… See the full description on the dataset page: https://huggingface.co/datasets/bvk/CICIDS-2017.citiesThis datasets contains 47,605 cities from around the world. The latest version can be found and filtered differently on: https://www.workwithdata.com/datasets/cities
Similar datasets can be found on: https://www.workwithdata.com
circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations
CircuitLens & WeightLens: Transcoder Descriptions and Evaluations
This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods.
Methods
CircuitLens: https://github.com/egolimblevskaia/CircuitLens
WeightLens: https://github.com/egolimblevskaia/WeightLens
Dataset Structure
The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.udhr-lid
UDHR-LID
Why UDHR-LID?
You can access UDHR (Universal Declaration of Human Rights) here, but when a verse is missing, they have texts such as "missing" or "?". Also, about 1/3 of the sentences consist only of "articles 1-30" in different languages. We cleaned the entire dataset from XML files and selected only the paragraphs. We cleared any unrelated language texts from the data and also removed the cases that were incorrect.
Incorrect? Look at the ckb and kmr files in the UDHR.… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/udhr-lid.photonic-integrated-circuit-yield
🏭 Photonic Integrated Circuit Yield Dataset
📊 125,000 synthetic (yield query, yield reasoning response) pairs covering process variation, defect density, lithography, and metrology challenges in CMOS-compatible photonic integrated circuit (PIC) manufacturing.
⚠️ Disclaimer: All entries are synthetically generated. Yield figures are computed from textbook models over sampled inputs, and citations are placeholders styled after technical sources; none reference a real… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/photonic-integrated-circuit-yield.irish-census
Irish Census 1901 & 1926
Person-level records from the 1901 and 1926 censuses of Ireland, as published by
the National Archives of Ireland — every individual return, in flat CSV.
Year
Rows
Size
Coverage
1901
4,434,939
4.31 GB
All of Ireland (32 counties)
1926
2,973,480
0.56 GB
Saorstát Éireann (26 counties)
Total
7,408,419
4.87 GB
The 1926 census is the first taken by the Irish Free State and was released to
the public in 2026 under the 100-year rule. The… See the full description on the dataset page: https://huggingface.co/datasets/Cianmcnally/irish-census.CISLR
CISLR: Corpus for Indian Sign Language Recognition
This repository contains the Indian Sign Language Dataset proposed in the following paper
Paper: CISLR: Corpus for Indian Sign Language Recognition https://preview.aclanthology.org/emnlp-22-ingestion/2022.emnlp-main.707/
Authors: Abhinav Joshi, Ashwani Bhat, Pradeep S, Priya Gole, Shashwat Gupta, Shreyansh Agarwal, Ashutosh Modi
Abstract: Indian Sign Language, though used by a diverse community, still lacks well-annotated… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/CISLR.CIS435-CreditCardFraudDetectionCIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles:
Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/muzom/CIC-IDS2017.cil-regionalizationCISR24
CISR24
CISR24 is the controlled complex-baseband benchmark introduced in the paper
"CDTFNet: A Cross-Domain Token-Fusion Transformer for Waveform-Level
Electromagnetic Awareness in Low-Altitude Intelligent Networks." It defines one closed-set task over 24 waveform classes:
10 communication, 6 sensing, and 8 integrated sensing and communication (ISAC)
classes.
Paper protocol
Property
Setting
Sampling rate
10 MHz
Record
1024 complex samples (102.4 us)… See the full description on the dataset page: https://huggingface.co/datasets/okra123/CISR24.cardiac_cine_kaggle
Kaggle Cardiac Cine-MRI
Processed NIfTI cine-MRI sequences from the Kaggle cardiac dataset.
Dataset Summary
Modality: CMR cine MRI
Views: SAX, LAX 2C, LAX 4C (time series)
Splits: train / val (if present)
Data Structure (per example)
sax_t, lax_2c_t, lax_4c_t
Metadata columns listed below
Columns
Imaging
pid
sax_t, lax_2c_t, lax_4c_t
Metadata (all columns)
n_slices, n_frames
original_sax_spacing_x, original_sax_spacing_y… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_kaggle.adityaramachandran27_world-air-quality-index-by-city-and-coordinates
World Air Quality Index by City and Coordinates
A Comprehensive Dataset on Cities, Latitude, Longitude, and Pollution Levels
Dataset Info
Source: Kaggle
Original Size: 0.36 MB
Kaggle Downloads: 11,064
Files: 1
Files
AQI and Lat Long of Countries.csv
Mirrored from Kaggle
GlotStoryBook
Dataset Description
Story Books for 180 ISO-639-3 codes.
The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset.
This dataset consists of 2 subsets:
default, which consists of 4 publishers:
asp: African Storybook
pb: Pratham Books
lcb: Little Cree Books
lida: LIDA Stories
nalibali, which comes from Nal'ibali stories.
Usage (HF Loader)
default:
from datasets import load_dataset
dataset… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotStoryBook.k-beauty-ai-citation-dataset
K-Beauty AI Citation Dataset
Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research.
Canonical source: https://kbeautyanswers.com/dataset/
License: CC BY 4.0
Maintainer: K-Beauty Answers (site)
Initial release: 2026-05-23
What's in it
128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.CIMemories
CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs
Paper
Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introduces critical risks when sensitive information is revealed in inappropriate contexts. We present CIMemories, a benchmark for evaluating whether LLMs appropriately control information flow from memory based on task context.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/CIMemories.Harbor-Citation-Intake
Harbor Citation Intake Ledger
This card records citation packages submitted by regional archive partners.
Canonical register: published
Citation submissions
citation_key
collection
title
submitted_on
decision
pages
HC-ALPHA
Maps
Estuary Charts
2023-04-18
approved
42
HC-BRAVO
Oral History
Dockworkers Speak
2022-11-09
approved
118
HC-CHARLIE
Photographs
Beacon Repairs
2024-01-12
approved
36
HC-DELTA
Maps
Channel Soundings
2021-07-03
approved
75… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Harbor-Citation-Intake.twin-cities-public-records
Twin Cities public records, joined
25 datasets · 1,575,384 rows · free, CC BY 4.0 · mirrored from brickandmortar.dev
A city emits records constantly — parcels, recorded sales, assessments, permits, licences, inspections, 911 calls, cleanup sites, flood zones, federal loans, wages, census measures — and almost nobody joins them. These are the joined slices, published as files rather than as an API you have to ask for a key to. The join is the work; the data is free.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/brickandmortar/twin-cities-public-records.circor-heart-soundOriginal-circuit-discovery
