Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bencorn /CIC-IoT-2023tabular10M<n<100M0 likes13k downloads7mo agoHugging Face02Ciroc0 /dmi-aarhus-predictions DMI Aarhus Predictions Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0. Primary files File Purpose Produced by predictions_latest.parquet Current future + verified prediction store dmi-collector frontend_snapshot.json Primary integration contract for the Vercel frontend dmi-collector Compatibility files File Status Notes predictions.parquet Legacy Still read by compatibility… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-predictions.tabular1K<n<10K6 likes11k downloads51m agoHugging Face03google /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.tabulartext-classification1M<n<10M41 likes10k downloads3y agoHugging Face04somnath0100 /CICIoT2023Small CICIoT2023 This dataset provides a processed derivative of the CICIoT2023 traffic collection. The repository organizes truncated PCAP files and flow-based CSV extractions aligned to the original CICResearch folder hierarchy. Processing Workflow The processing pipeline follows four stages: Source acquisition from the CICResearch CICIoT2023 portal. Flow extraction from full PCAP files using TriFlowMeter. PCAP size reduction by truncating packet payloads to 128 bytes with… See the full description on the dataset page: https://huggingface.co/datasets/somnath0100/CICIoT2023Small.tabulartabular-classification1 likes3.3k downloads5mo agoHugging Face05evaluate /glue-ci Dataset Card for GLUE Dataset Summary GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems. Supported Tasks and Leaderboards The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks: ax A manually-curated evaluation dataset for fine-grained analysis of system… See the full description on the dataset page: https://huggingface.co/datasets/evaluate/glue-ci.tabulartext-classification1M<n<10M1 likes2.7k downloads1y agoHugging Face06Ciroc0 /dmi-aarhus-weather-data DMI Aarhus Weather Data Training data and model artifact dataset for the Aarhus weather pipeline. Maintained by Ciroc0. Primary files File Purpose Produced by training_matrix.parquet Current source of truth for training rows and causal observation context dmi-collector model_registry.json Active bucket registry per target dmi-ml-trainer model_meta.json Training timestamp, sample count and training window dmi-ml-trainer temperature_models.pkl… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-weather-data.tabular10K<n<100K1 likes2.4k downloads20h agoHugging Face07damo-da /ciaa-annual-reports CIAA Annual Reports — Nepali transcripts, ruled tables and chart data Machine-readable transcripts of the annual reports of Nepal's Commission for the Investigation of Abuse of Authority (अख्तियार दुरुपयोग अनुसन्धान आयोग, CIAA) — all 35 it has published to date. The 1st to 35th reports, fiscal years BS 2047/48 – 2081/82 (AD 1990–2025). The CIAA publishes these as PDFs whose text layer is, for several years, legacy pre-Unicode Devanagari that ordinary extractors turn into… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/ciaa-annual-reports.imagetext-retrieval100K<n<1M0 likes2.3k downloads2mo agoHugging Face08cis-lmu /GlotCC-V1 Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at https://github.com/cisnlp/GlotCC. List of Languages: See https://datasets-server.huggingface.co/splits?dataset=cis-lmu/GlotCC-V1 to get the list of splits available. Usage (Huggingface Hub -- Recommended)… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotCC-V1.tabular1B<n<10B61 likes2k downloads2y agoHugging Face09fahadhafeezofficial /cissp-llmbench CISSP-LLMBench tabulartext-generation10K<n<100K0 likes1.6k downloads3mo agoHugging Face10bvsam /cic-ids-2017 CIC-IDS-2017 Dataset This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use. Dataset Structure Configurations machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs). traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC. Raw Data The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/bvsam/cic-ids-2017.tabulartabular-classification1M<n<10M4 likes1.6k downloads10mo agoHugging Face11c01dsnap /CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.tabular1M<n<10M5 likes1.4k downloads3y agoHugging Face12viennh2012 /cardiac_cine_myops2020 MyoPS 2020 (Multi-Sequence CMR) Processed NIfTI multi-sequence CMR data (C0/DE/T2) for myocardial pathology segmentation. Dataset Summary Modality: CMR (C0, DE, T2) Task: Segmentation of myocardium/pathology Splits: train / test Data Structure (per example) c0, de, t2, label Metadata columns listed below Columns Imaging pid c0, de, t2, label Metadata (all columns) orig_spacing_x, orig_spacing_y, orig_spacing_z n_slices crop_lower_x… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_myops2020.tabularimage-segmentationn<1K0 likes1.4k downloads8mo agoHugging Face13CiferAI /Cifer-Fraud-Detection-Dataset-AF 📊 Cifer Fraud Detection Dataset 🧠 Overview The Cifer-Fraud-Detection-Dataset-AF is a high-fidelity, fully synthetic dataset created to support the development and benchmarking of privacy-preserving, federated, and decentralized machine learning systems in financial fraud detection. This dataset draws structural inspiration from the PaySim simulator, which was built using aggregated mobile money transaction data from a real financial provider operating in 14+ countries.… See the full description on the dataset page: https://huggingface.co/datasets/CiferAI/Cifer-Fraud-Detection-Dataset-AF.tabulartabular-classification10M<n<100M14 likes1.3k downloads1y agoHugging Face14allenai /asta-summary-citation-counts Dataset Summary This dataset tracks which scientific papers are most often cited by Asta, an agentic research platform that uses retrieval-augmented generation (RAG) to answer scientific questions. Each record is a paper cited by Asta's Summarize Literature tool, ranked by the number of times the system cited that paper. Across more than 113,000 user queries, we track 4M citations to over 2M distinct papers. By making this data public, we aim to create a transparent, trackable… See the full description on the dataset page: https://huggingface.co/datasets/allenai/asta-summary-citation-counts.tabular100M<n<1B11 likes1.2k downloads6d agoHugging Face15spaicom-lab /semasia-cifar100 Latents for cifar100 (timm) &nbsp;&nbsp;&nbsp; This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on cifar100, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability. Each config corresponds to a single model; only that model's Parquet files are read on load_dataset. Usage Load with datasets and… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-cifar100.tabularfeature-extraction100M<n<1B0 likes1.2k downloads3mo agoHugging Face16inference-optimization /speculators-ci-datasets speculator-tutorial Raw vs. on-policy regenerated conversation data for training speculative-decoding drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside so you can see exactly what regeneration changes and why it matters. Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B. Why regenerate at all? A speculative-decoding drafter is trained to predict what the verifier would say next. If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.tabulartext-generation1K<n<10K0 likes1.2k downloads2mo agoHugging Face17ciaochris /neuro2-neuroscience-datasets Neuro2 Neuroscience Dataset Atlas (unofficial mirror) Neuro2 is a catalog and 3D knowledge graph of open neuroscience datasets, created and maintained by Nataliya Kosmyna and Eugene Hauptmann. It indexes dataset records from OpenNeuro, Zenodo, DataCite, DANDI, OSF and dozens of other repositories and links them to authors, papers, tasks, institutions and funders. This repository is a dated snapshot of that public catalog, reshaped into six Parquet tables you can load with… See the full description on the dataset page: https://huggingface.co/datasets/ciaochris/neuro2-neuroscience-datasets.tabulartext-retrieval100K<n<1M0 likes997 downloads6d agoHugging Face18viennh2012 /cardiac_cine_acdc ACDC (Cardiac Cine-MRI) ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format. Dataset Summary Modality: Cardiac cine‑MRI (NIfTI) Task: Segmentation of LV, RV, and myocardium Frames: ED/ES + full SAX time series (sax_t) Labels: LV/RV cavities + myocardium Splits: train, test (as provided in processed output) Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.tabularimage-segmentationn<1K0 likes991 downloads8mo agoHugging Face19Shanmuk4622 /msc-cifar100 MSC — Minimum Sufficient Compute Artifacts for Is Compute Difficulty Architecture-Agnostic? Measuring and Distilling Per-Sample Minimum Sufficient Computation. Generated 2026-08-06T03:56:37Z by msc_lib v1.0.0. Repositories Shanmuk4622/msc-cifar100 — everything, one folder per run What MSC is The smallest cost-normalised configuration at which a network's decision has stably settled to its full-compute decision, defined uniformly over depth… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/msc-cifar100.tabular1M<n<10M0 likes967 downloads1mo agoHugging Face20JetBrains-Research /lca-ci-builds-repair 🏟️ Long Code Arena (CI builds repair) This is the benchmark for CI builds repair task as part of the 🏟️ Long Code Arena benchmark. 🛠️ Task. Given the logs of a failed GitHub Actions workflow and the corresponding repository snapshot, repair the repository contents in order to make the workflow pass. All the data is collected from repositories published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause, and BSD-2-Clause). The datapoints can be removed upon request. To… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-ci-builds-repair.tabularn<1K3 likes635 downloads2y agoHugging Face21CCB /cis5300-word-embeddings Word Embeddings and Semantic Similarity (CIS 5300) Dataset Description This dataset supports learning about word embeddings — dense vector representations that capture word meaning. It includes a standard similarity benchmark, a word sense disambiguation task, and a Shakespeare corpus for training custom embeddings. Configs SimLex-999: Word Similarity Benchmark SimLex-999 (Hill et al., 2015) is a gold-standard benchmark for evaluating word… See the full description on the dataset page: https://huggingface.co/datasets/CCB/cis5300-word-embeddings.tabularsentence-similarity1K<n<10K0 likes633 downloads5mo agoHugging Face22amanyagami /cifar100-224-adv-selfgentabularn<1K0 likes549 downloads1mo agoHugging Face23civbench /civbench-v1 CivBench A benchmark of LLM agents playing full games of Civilization VI through an MCP (Model Context Protocol) server. Each run captures the full per-turn game state, every tool call the agent issued, and the agent's own structured reflections. Configs tables — curated parquet tables, one row per logical record. Use this for analysis: datasets.load_dataset("civbench/civbench-v1", "tables", split="games"). raw — byte-identical mirror of the on-disk telemetry JSONL… See the full description on the dataset page: https://huggingface.co/datasets/civbench/civbench-v1.tabularother100K<n<1M1 likes506 downloads5mo agoHugging Face24viennh2012 /cardiac_cine_mnms2 M&Ms2 (Cardiac Cine-MRI, RV Focus) Processed NIfTI cine-MRI data derived from the M&Ms2 challenge. Dataset Summary Modality: CMR cine MRI Task: LV/RV/MYO segmentation (RV focus) Views: SAX + LAX 4C (LAX 2C if present) Splits: train / val / test Data Structure (per example) SAX: sax_ed, sax_ed_gt, sax_es, sax_es_gt LAX 4C: lax_4c_ed, lax_4c_ed_gt, lax_4c_es, lax_4c_es_gt LAX 2C (if present): lax_2c_ed, lax_2c_ed_gt, lax_2c_es, lax_2c_es_gt Metadata columns… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms2.tabularimage-segmentationn<1K0 likes496 downloads8mo agoHugging Face25Lystea /CICIOT2023-PARQUET CICIoT2023 — ipfixprobe flow records (Parquet) 1,479,074,715 bidirectional network flows re-exported from the raw PCAPs of CICIoT2023 (Canadian Institute for Cybersecurity, University of New Brunswick) with ipfixprobe 5.7.0, stored as 308 Parquet files (~19.4 GB) covering 33 attack classes + benign traffic from the 105-device IoT testbed. The original dataset ships ~587 GB of PCAPs and CSV features computed with a closed pipeline. This conversion provides an alternative… See the full description on the dataset page: https://huggingface.co/datasets/Lystea/CICIOT2023-PARQUET.tabulartabular-classification1B<n<10B0 likes495 downloads2mo agoHugging Face26circle-great /behavior-1k-2026-partial20-submission BEHAVIOR-1K 2026 Pi0.5 Partial-20 Self-Evaluation This repository contains an unedited partial self-evaluation for the 2026 BEHAVIOR Challenge. Scope BEHAVIOR-1K version: v3.9.1 Policy: Pi0.5 task-embedding policy derived from IliaLarchenko/behavior-1k-solution Robot: bundled R1Pro configuration Evaluation wrapper: omnigibson.eval.wrappers.DefaultWrapper Public instance indices: 0-9 (instance IDs 301-310) Rollouts per instance: 1 Evaluated tasks: 20 / 100… See the full description on the dataset page: https://huggingface.co/datasets/circle-great/behavior-1k-2026-partial20-submission.tabularn<1K0 likes465 downloads2mo agoHugging Face27viennh2012 /cardiac_cine_mnms M&Ms (Cardiac Cine-MRI) Processed NIfTI cine-MRI data derived from the M&Ms (Multi-Centre, Multi-Vendor & Multi-Disease) challenge. Dataset Summary Modality: CMR cine MRI Task: LV/RV/MYO segmentation Views: SAX (ED/ES) Splits: train / val / test Data Structure (per example) sax_ed, sax_ed_gt sax_es, sax_es_gt Optional: sax_t (if present) Metadata columns listed below Columns Imaging pid sax_ed, sax_ed_gt, sax_es, sax_es_gt sax_t (if present)… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_mnms.tabularimage-segmentationn<1K0 likes460 downloads8mo agoHugging Face28KerryMe /cinepile_10ktabular10K<n<100K0 likes459 downloads4mo agoHugging Face29thefcraft /civitai-stable-diffusion-337k How to Use from datasets import load_dataset dataset = load_dataset("thefcraft/civitai-stable-diffusion-337k") print(dataset['train'][0]) download images download zip files from images dir https://huggingface.co/datasets/thefcraft/civitai-stable-diffusion-337k/tree/main/images it contains some images with id from zipfile import ZipFile with ZipFile("filename.zip", 'r') as zObject: zObject.extractall() Dataset Summary GitHub URL:-… See the full description on the dataset page: https://huggingface.co/datasets/thefcraft/civitai-stable-diffusion-337k.image100K<n<1M45 likes426 downloads2y agoHugging Face30maedmatt /DREAM-pyramid-circlesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/maedmatt/DREAM-pyramid-circles.tabularrobotics100K<n<1M0 likes411 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.