Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bencorn /CIC-IoT-2023tabular10M<n<100M0 likes13k downloads7mo agoHugging Face02evaluate /imdb-citextn<1K0 likes2.4k downloads4y agoHugging Face03c01dsnap /CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.tabular1M<n<10M5 likes1.4k downloads3y agoHugging Face04viennh2012 /cardiac_cine_myops2020 MyoPS 2020 (Multi-Sequence CMR) Processed NIfTI multi-sequence CMR data (C0/DE/T2) for myocardial pathology segmentation. Dataset Summary Modality: CMR (C0, DE, T2) Task: Segmentation of myocardium/pathology Splits: train / test Data Structure (per example) c0, de, t2, label Metadata columns listed below Columns Imaging pid c0, de, t2, label Metadata (all columns) orig_spacing_x, orig_spacing_y, orig_spacing_z n_slices crop_lower_x… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_myops2020.tabularimage-segmentationn<1K0 likes1.4k downloads8mo agoHugging Face05CiferAI /Cifer-Fraud-Detection-Dataset-AF 📊 Cifer Fraud Detection Dataset 🧠 Overview The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection. This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/CiferAI/Cifer-Fraud-Detection-Dataset-AF.tabulartabular-classification10M<n<100M14 likes1.3k downloads1y agoHugging Face06viennh2012 /cardiac_cine_acdc ACDC (Cardiac Cine-MRI) ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format. Dataset Summary Modality: Cardiac cine‑MRI (NIfTI) Task: Segmentation of LV, RV, and myocardium Frames: ED/ES + full SAX time series (sax_t) Labels: LV/RV cavities + myocardium Splits: train, test (as provided in processed output) Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.tabularimage-segmentationn<1K0 likes991 downloads8mo agoHugging Face07CCB /cis5300-word-embeddings Word Embeddings and Semantic Similarity (CIS 5300) Dataset Description This dataset supports learning about word embeddings — dense vector representations that capture word meaning. It includes a standard similarity benchmark, a word sense disambiguation task, and a Shakespeare corpus for training custom embeddings. Configs SimLex-999: Word Similarity Benchmark SimLex-999 (Hill et al., 2015) is a gold-standard benchmark for evaluating word… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-word-embeddings.tabularsentence-similarity1K<n<10K0 likes633 downloads5mo agoHugging Face08circircircle /FinQAtext100K<n<1M0 likes613 downloads3y agoHugging Face09viennh2012 /cardiac_cine_mnms2 M&Ms2 (Cardiac Cine-MRI, RV Focus) Processed NIfTI cine-MRI data derived from the M&Ms2 challenge. Dataset Summary Modality: CMR cine MRI Task: LV/RV/MYO segmentation (RV focus) Views: SAX + LAX 4C (LAX 2C if present) Splits: train / val / test Data Structure (per example) SAX: sax_ed, sax_ed_gt, sax_es, sax_es_gt LAX 4C: lax_4c_ed, lax_4c_ed_gt, lax_4c_es, lax_4c_es_gt LAX 2C (if present): lax_2c_ed, lax_2c_ed_gt, lax_2c_es, lax_2c_es_gt Metadata columns… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms2.tabularimage-segmentationn<1K0 likes496 downloads8mo agoHugging Face10viennh2012 /cardiac_cine_mnms M&Ms (Cardiac Cine-MRI) Processed NIfTI cine-MRI data derived from the M&Ms (Multi-Centre, Multi-Vendor & Multi-Disease) challenge. Dataset Summary Modality: CMR cine MRI Task: LV/RV/MYO segmentation Views: SAX (ED/ES) Splits: train / val / test Data Structure (per example) sax_ed, sax_ed_gt sax_es, sax_es_gt Optional: sax_t (if present) Metadata columns listed below Columns Imaging pid sax_ed, sax_ed_gt, sax_es, sax_es_gt sax_t (if present)… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms.tabularimage-segmentationn<1K0 likes460 downloads8mo agoHugging Face11bvk /CICIDS-2017Raw network data was collected over a period of 5 days, Monday through Friday, and stored in PCAP files. Monday was used to create most of the Benign data, while the Attack-Network implemented various types of attacks over the next 4 days, such as Brute Force connections (FTP and SSH), several types of DoS attacks, as well as a Botnet attack, Infiltration attacks and subsequent Port-Scanning activity. The PCAP data was processed using a tool developed by one of the authors of [1], called… See the full description on the dataset page: https://huggingface.co/datasets/bvk/CICIDS-2017.tabular1M<n<10M0 likes401 downloads2y agoHugging Face12WorkWithData /citiesThis datasets contains 47,605 cities from around the world. The latest version can be found and filtered differently on: https://www.workwithdata.com/datasets/cities Similar datasets can be found on: https://www.workwithdata.com tabular10K<n<100K2 likes334 downloads2y agoHugging Face13egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes250 downloads7mo agoHugging Face14cis-lmu /udhr-lid UDHR-LID Why UDHR-LID? You can access UDHR (Universal Declaration of Human Rights) here, but when a verse is missing, they have texts such as "missing" or "?". Also, about 1/3 of the sentences consist only of "articles 1-30" in different languages. We cleaned the entire dataset from XML files and selected only the paragraphs. We cleared any unrelated language texts from the data and also removed the cases that were incorrect. Incorrect? Look at the ckb and kmr files in the UDHR.… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/udhr-lid.text10K<n<100K8 likes241 downloads2y agoHugging Face15Taylor658 /photonic-integrated-circuit-yield 🏭 Photonic Integrated Circuit Yield Dataset 📊 125,000 synthetic (yield query, yield reasoning response) pairs covering process variation, defect density, lithography, and metrology challenges in CMOS-compatible photonic integrated circuit (PIC) manufacturing. ⚠️ Disclaimer: All entries are synthetically generated. Yield figures are computed from textbook models over sampled inputs, and citations are placeholders styled after technical sources; none reference a real… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/photonic-integrated-circuit-yield.texttext-generation100K<n<1M5 likes217 downloads28d agoHugging Face16Cianmcnally /irish-census Irish Census 1901 & 1926 Person-level records from the 1901 and 1926 censuses of Ireland, as published by the National Archives of Ireland — every individual return, in flat CSV. Year Rows Size Coverage 1901 4,434,939 4.31 GB All of Ireland (32 counties) 1926 2,973,480 0.56 GB Saorstát Éireann (26 counties) Total 7,408,419 4.87 GB The 1926 census is the first taken by the Irish Free State and was released to the public in 2026 under the 100-year rule. The… See the full description on the dataset page: https://huggingface.co/datasets/Cianmcnally/irish-census.tabulartabular-classification1M<n<10M1 likes211 downloads2mo agoHugging Face17Exploration-Lab /CISLRgated CISLR: Corpus for Indian Sign Language Recognition This repository contains the Indian Sign Language Dataset proposed in the following paper Paper: CISLR: Corpus for Indian Sign Language Recognition https://preview.aclanthology.org/emnlp-22-ingestion/2022.emnlp-main.707/ Authors: Abhinav Joshi, Ashwani Bhat, Pradeep S, Priya Gole, Shashwat Gupta, Shreyansh Agarwal, Ashutosh Modi Abstract: Indian Sign Language, though used by a diverse community, still lacks well-annotated… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/CISLR.text1K<n<10K15 likes202 downloads2y agoHugging Face18dazzle-nu /CIS435-CreditCardFraudDetectiontabular1M<n<10M13 likes196 downloads4y agoHugging Face19muzom /CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/muzom/CIC-IDS2017.tabular1M<n<10M0 likes196 downloads3mo agoHugging Face20c1587s /cil-regionalizationgeospatialn<1K0 likes189 downloads1y agoHugging Face21okra123 /CISR24 CISR24 CISR24 is the controlled complex-baseband benchmark introduced in the paper "CDTFNet: A Cross-Domain Token-Fusion Transformer for Waveform-Level Electromagnetic Awareness in Low-Altitude Intelligent Networks." It defines one closed-set task over 24 waveform classes: 10 communication, 6 sensing, and 8 integrated sensing and communication (ISAC) classes. Paper protocol Property Setting Sampling rate 10 MHz Record 1024 complex samples (102.4 us)… See the full description on the dataset page: https://huggingface.co/datasets/okra123/CISR24.document1M<n<10M0 likes184 downloads3mo agoHugging Face22viennh2012 /cardiac_cine_kaggle Kaggle Cardiac Cine-MRI Processed NIfTI cine-MRI sequences from the Kaggle cardiac dataset. Dataset Summary Modality: CMR cine MRI Views: SAX, LAX 2C, LAX 4C (time series) Splits: train / val (if present) Data Structure (per example) sax_t, lax_2c_t, lax_4c_t Metadata columns listed below Columns Imaging pid sax_t, lax_2c_t, lax_4c_t Metadata (all columns) n_slices, n_frames original_sax_spacing_x, original_sax_spacing_y… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_kaggle.tabularimage-segmentation1K<n<10K0 likes176 downloads8mo agoHugging Face23jason1966 /adityaramachandran27_world-air-quality-index-by-city-and-coordinates World Air Quality Index by City and Coordinates A Comprehensive Dataset on Cities, Latitude, Longitude, and Pollution Levels Dataset Info Source: Kaggle Original Size: 0.36 MB Kaggle Downloads: 11,064 Files: 1 Files AQI and Lat Long of Countries.csv Mirrored from Kaggle tabular10K<n<100K0 likes171 downloads6mo agoHugging Face24cis-lmu /GlotStoryBook Dataset Description Story Books for 180 ISO-639-3 codes. The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset. This dataset consists of 2 subsets: default, which consists of 4 publishers: asp: African Storybook pb: Pratham Books lcb: Little Cree Books lida: LIDA Stories nalibali, which comes from Nal'ibali stories. Usage (HF Loader) default: from datasets import load_dataset dataset… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotStoryBook.texttranslation10K<n<100K9 likes167 downloads18d agoHugging Face25k-master /k-beauty-ai-citation-dataset K-Beauty AI Citation Dataset Open dataset mapping Korean K-beauty entities (ingredients, skin concerns, use cases, brands) and answer-style guides to citation-shaped external references. Designed to be referenced by AI search engines, content builders, and SEO research. Canonical source: https://kbeautyanswers.com/dataset/ License: CC BY 4.0 Maintainer: K-Beauty Answers (site) Initial release: 2026-05-23 What's in it 128 entities (37 ingredients + 18 skin… See the full description on the dataset page: https://huggingface.co/datasets/k-master/k-beauty-ai-citation-dataset.texttext-classificationn<1K0 likes166 downloads4mo agoHugging Face26facebook /CIMemories CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs Paper Large Language Models (LLMs) increasingly use persistent memory from past interactions to enhance personalization and task performance. However, this memory introduces critical risks when sensitive information is revealed in inappropriate contexts. We present CIMemories, a benchmark for evaluating whether LLMs appropriately control information flow from memory based on task context.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/CIMemories.text10K<n<100K2 likes165 downloads11mo agoHugging Face27SOTAagi2030 /Harbor-Citation-Intake Harbor Citation Intake Ledger This card records citation packages submitted by regional archive partners. Canonical register: published Citation submissions citation_key collection title submitted_on decision pages HC-ALPHA Maps Estuary Charts 2023-04-18 approved 42 HC-BRAVO Oral History Dockworkers Speak 2022-11-09 approved 118 HC-CHARLIE Photographs Beacon Repairs 2024-01-12 approved 36 HC-DELTA Maps Channel Soundings 2021-07-03 approved 75… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Harbor-Citation-Intake.textn<1K0 likes146 downloads14d agoHugging Face28brickandmortar /twin-cities-public-records Twin Cities public records, joined 25 datasets · 1,575,384 rows · free, CC BY 4.0 · mirrored from brickandmortar.dev A city emits records constantly — parcels, recorded sales, assessments, permits, licences, inspections, 911 calls, cleanup sites, flood zones, federal loans, wages, census measures — and almost nobody joins them. These are the joined slices, published as files rather than as an API you have to ask for a key to. The join is the work; the data is free. This is a… See the full description on the dataset page: https://huggingface.co/datasets/brickandmortar/twin-cities-public-records.tabular100K<n<1M0 likes145 downloads1mo agoHugging Face29luyangliuable /circor-heart-soundaudion<1K1 likes141 downloads7mo agoHugging Face30PersonaBias /Original-circuit-discoverytabulartext-classification10K<n<100K0 likes140 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.