Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hossam3759180 /ids-project-artifactsimage10M<n<100M0 likes2.5k downloads21d agoHugging Face02bvsam /cic-ids-2017 CIC-IDS-2017 Dataset This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use. Dataset Structure Configurations machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs). traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC. Raw Data The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/bvsam/cic-ids-2017.tabulartabular-classification1M<n<10M4 likes1.6k downloads10mo agoHugging Face03c01dsnap /CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.tabular1M<n<10M5 likes1.4k downloads3y agoHugging Face04rdpahalavan /CIC-IDS2017We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily. NIDS Datasets The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/rdpahalavan/CIC-IDS2017.text-classification100M<n<1B4 likes1.2k downloads3y agoHugging Face05bigcode /the-stack-v2-train-full-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.tabulartext-generation10M<n<100M68 likes977 downloads2mo agoHugging Face06jiebi /ids-paragraph-commentary-corpusThe data was collected from GitHub using https://github.com/cheop-byeon/RFCRationaleBuilder. This is our first version of data for RFC Rationale. Dataset Citation If you find this dataset useful and include it in your studies, please cite our paper: @inproceedings{bian2024tell, title={Tell Me Why: Language Models Help Explain the Rationale Behind Internet Protocol Design}, author={Bian, Jie and Welzl, Michael and Kutuzov, Andrey and Arefyev, Nikolay}, booktitle={2024 IEEE… See the full description on the dataset page: https://huggingface.co/datasets/jiebi/ids-paragraph-commentary-corpus.text-classification0 likes943 downloads6mo agoHugging Face07pysport /idsse-data IDSSE — Sportec Open Tracking & Event Data Mirror of the IDSSE dataset, originally hosted on Figshare, re-hosted here for stable and fast access via kloppy and related tools. Contents 7 complete matches from the German Bundesliga season 2022/23 (1st and 2nd division), collected by Sportec Solutions using TRACAB optical tracking technology. Each match includes: Tracking data — position of all players and the ball at 25 fps Event data — synchronized match events (passes… See the full description on the dataset page: https://huggingface.co/datasets/pysport/idsse-data.geospatialother2 likes810 downloads7mo agoHugging Face08bigcode /the-stack-v2-train-smol-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. bigcode/the-stack-v2-train-smol-ids: based on the bigcode/the-stack-v2-dedup… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids.tabulartext-generation10M<n<100M63 likes809 downloads2mo agoHugging Face09JackHsieh /4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids 4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task is not "write a thought". Instead, it is predict the next k=8 tokens after the cut, and report the prediction as a ranked list of 3–6 candidate continuations, each one being the tail copied verbatim plus the predicted continuation, written as N. «tail + continuation».… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids.tabular1M<n<10M0 likes579 downloads1mo agoHugging Face10kamabata /pcod_graph_ids0 likes575 downloads1y agoHugging Face11JackHsieh /32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-32B with thinking mode off. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. The 4B parity counterpart is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes523 downloads1mo agoHugging Face12HoangHa /selfies-ids-cleanedtext100M<n<1B0 likes505 downloads2y agoHugging Face13JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes490 downloads1mo agoHugging Face14JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes461 downloads25d agoHugging Face15JabaleNurAdnan /dml-fl-iot-ids-dynamic3L-K30 likes373 downloads4mo agoHugging Face16HoangHa /selfies-train-idstext100M<n<1B0 likes346 downloads2y agoHugging Face17JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a single paragraph of dense reasoning about the immediate continuation. Only the part after </think> — the paragraph — becomes the VALUE. The reasoning inside the think block is dropped. The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes341 downloads1mo agoHugging Face18sullivan1502 /zone-pretrain-ids-data1M<n<10M0 likes335 downloads2mo agoHugging Face19HoangHa /belka-selfies-idstext10M<n<100M0 likes322 downloads2y agoHugging Face20kamabata /aflow_graph_ids0 likes299 downloads1y agoHugging Face21JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a few dense, declarative sentences (about two to four) inside a <thought>…</thought> block, focused on the exact state at the cut and what the local grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes253 downloads1mo agoHugging Face22tuanio /book_corpus-input_ids-valid-len256 Dataset Card for "book_corpus-input_ids-valid-len256" More Information needed 1M<n<10M0 likes220 downloads3y agoHugging Face23JabaleNurAdnan /dml-fl-iot-ids-dynamic3L-K50 likes220 downloads4mo agoHugging Face24Ariasyah /cic-ids-2017 CIC-IDS-2017 Dataset This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use. Dataset Structure Configurations machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs). traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC. Raw Data The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/Ariasyah/cic-ids-2017.tabulartabular-classification1M<n<10M0 likes219 downloads9mo agoHugging Face25c01dsnap /CIC-IDS2018LICENSE You may redistribute, republish, and mirror the CSE-CIC-IDS2018 dataset in any form. However, any use or redistribution of the data must include a citation to the CSE-CIC-IDS2018 dataset and a link to this page in AWS. Research paper outlining the details of analyzing the similar IDS/IPS dataset and related principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A. Ghorbani, “Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization”, 4th… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2018.0 likes215 downloads3y agoHugging Face26orionweller /cc-ids-to-title-multilingual Common Crawl IDs to Titles Per-config manifests mapping document IDs to their extracted titles. Usage from datasets import load_dataset titles = load_dataset("orionweller/cc-ids-to-title-multilingual-2", "fw-edu") text100M<n<1B0 likes213 downloads7mo agoHugging Face27JabaleNurAdnan /dml-fl-iot-ids-dynamic3L-K100 likes210 downloads4mo agoHugging Face28vibhuiitj /UltraData-Math-L2-preview-embedded-with-idstabular10M<n<100M0 likes208 downloads7mo agoHugging Face29JabaleNurAdnan /dml-fl-iot-ids-static3L0 likes203 downloads4mo agoHugging Face30muzom /CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/muzom/CIC-IDS2017.tabular1M<n<10M0 likes196 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.