datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ids-project-artifactscic-ids-2017
CIC-IDS-2017 Dataset
This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use.
Dataset Structure
Configurations
machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs).
traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC.
Raw Data
The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/bvsam/cic-ids-2017.CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles:
Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.CIC-IDS2017We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily.
NIDS Datasets
The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/rdpahalavan/CIC-IDS2017.the-stack-v2-train-full-ids
The Stack v2
The dataset consists of 4 versions:
bigcode/the-stack-v2: the full "The Stack v2" dataset
bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated
bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here
bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.ids-paragraph-commentary-corpusThe data was collected from GitHub using https://github.com/cheop-byeon/RFCRationaleBuilder.
This is our first version of data for RFC Rationale.
Dataset Citation
If you find this dataset useful and include it in your studies, please cite our paper:
@inproceedings{bian2024tell,
title={Tell Me Why: Language Models Help Explain the Rationale Behind Internet Protocol Design},
author={Bian, Jie and Welzl, Michael and Kutuzov, Andrey and Arefyev, Nikolay},
booktitle={2024 IEEE… See the full description on the dataset page: https://huggingface.co/datasets/jiebi/ids-paragraph-commentary-corpus.idsse-data
IDSSE — Sportec Open Tracking & Event Data
Mirror of the IDSSE dataset, originally hosted on Figshare, re-hosted here for stable and fast access via kloppy and related tools.
Contents
7 complete matches from the German Bundesliga season 2022/23 (1st and 2nd division), collected by Sportec Solutions using TRACAB optical tracking technology.
Each match includes:
Tracking data — position of all players and the ball at 25 fps
Event data — synchronized match events (passes… See the full description on the dataset page: https://huggingface.co/datasets/pysport/idsse-data.the-stack-v2-train-smol-ids
The Stack v2
The dataset consists of 4 versions:
bigcode/the-stack-v2: the full "The Stack v2" dataset
bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated
bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories.
bigcode/the-stack-v2-train-smol-ids: based on the bigcode/the-stack-v2-dedup… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids.4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids
4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task is not "write a thought". Instead, it is predict the next k=8 tokens
after the cut, and report the prediction as a ranked list of 3–6 candidate continuations,
each one being the tail copied verbatim plus the predicted continuation, written as
N. «tail + continuation».… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids.pcod_graph_ids32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-32B
with thinking mode off. Each thought is a few dense sentences of reasoning about the next
8 tokens after a cut, written from the document prefix alone — the generator never sees the
continuation. Stored thought_text includes the <thought>/</thought> wrapper.
The 4B parity counterpart is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.selfies-ids-cleaned4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the
exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next
8 tokens after a cut, written from the document prefix alone — the generator never sees the
continuation. Stored thought_text includes the <thought>/</thought> wrapper.
This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.dml-fl-iot-ids-dynamic3L-K3selfies-train-ids4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained
A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.
The source thoughts are two parts: a <think> block, then a single paragraph of dense
reasoning about the immediate continuation. Only the part after </think> — the paragraph —
becomes the VALUE. The reasoning inside the think block is dropped.
The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.zone-pretrain-ids-databelka-selfies-idsaflow_graph_ids4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids
Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by
Qwen/Qwen3-4B in thinking mode.
The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the
cut — and answer with a few dense, declarative sentences (about two to four) inside a
<thought>…</thought> block, focused on the exact state at the cut and what the local
grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.book_corpus-input_ids-valid-len256
Dataset Card for "book_corpus-input_ids-valid-len256"
More Information needed
dml-fl-iot-ids-dynamic3L-K5cic-ids-2017
CIC-IDS-2017 Dataset
This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use.
Dataset Structure
Configurations
machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs).
traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC.
Raw Data
The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/Ariasyah/cic-ids-2017.CIC-IDS2018LICENSE
You may redistribute, republish, and mirror the CSE-CIC-IDS2018 dataset in any form. However, any use or redistribution of the data must include a citation to the CSE-CIC-IDS2018 dataset and a link to this page in AWS.
Research paper outlining the details of analyzing the similar IDS/IPS dataset and related principles:
Iman Sharafaldin, Arash Habibi Lashkari, and Ali A. Ghorbani, “Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization”, 4th… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2018.cc-ids-to-title-multilingual
Common Crawl IDs to Titles
Per-config manifests mapping document IDs to their extracted titles.
Usage
from datasets import load_dataset
titles = load_dataset("orionweller/cc-ids-to-title-multilingual-2", "fw-edu")
dml-fl-iot-ids-dynamic3L-K10UltraData-Math-L2-preview-embedded-with-idsdml-fl-iot-ids-static3LCIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles:
Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/muzom/CIC-IDS2017.
