Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hossam3759180 /ids-project-artifactsimage10M<n<100M0 likes2.5k downloads21d agoHugging Face02bvsam /cic-ids-2017 CIC-IDS-2017 Dataset This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use. Dataset Structure Configurations machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs). traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC. Raw Data The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/bvsam/cic-ids-2017.tabulartabular-classification1M<n<10M4 likes1.6k downloads10mo agoHugging Face03c01dsnap /CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/c01dsnap/CIC-IDS2017.tabular1M<n<10M5 likes1.4k downloads3y agoHugging Face04bigcode /the-stack-v2-train-full-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.tabulartext-generation10M<n<100M68 likes977 downloads2mo agoHugging Face05bigcode /the-stack-v2-train-smol-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. bigcode/the-stack-v2-train-smol-ids: based on the bigcode/the-stack-v2-dedup… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids.tabulartext-generation10M<n<100M63 likes809 downloads2mo agoHugging Face06JackHsieh /4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids 4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task is not "write a thought". Instead, it is predict the next k=8 tokens after the cut, and report the prediction as a ranked list of 3–6 candidate continuations, each one being the tail copied verbatim plus the predicted continuation, written as N. «tail + continuation».… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-ranked-v9b.stride-train2-test32.k-8.statml-arxiv.qwen3-ids.tabular1M<n<10M0 likes579 downloads1mo agoHugging Face07JackHsieh /32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-32B with thinking mode off. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. The 4B parity counterpart is… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/32B-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes523 downloads1mo agoHugging Face08JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes490 downloads1mo agoHugging Face09JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. Stored thought_text includes the <thought>/</thought> wrapper. This is the small-generator parity counterpart of… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes461 downloads25d agoHugging Face10JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a single paragraph of dense reasoning about the immediate continuation. Only the part after </think> — the paragraph — becomes the VALUE. The reasoning inside the think block is dropped. The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes341 downloads1mo agoHugging Face11JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a few dense, declarative sentences (about two to four) inside a <thought>…</thought> block, focused on the exact state at the cut and what the local grammar… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes253 downloads1mo agoHugging Face12Ariasyah /cic-ids-2017 CIC-IDS-2017 Dataset This repository contains the CIC-IDS-2017 dataset with the original PCAPs and the CSVs converted to Parquet format for easier use. Dataset Structure Configurations machine_learning: Contains the flow-based features used for ML training (Converted from MachineLearningCVE CSVs). traffic_labels: Contains the labelled flows (Converted from TrafficLabelling CSVs). Timestamps have been normalized to UTC. Raw Data The pcap/ folder… See the full description on the dataset page: https://huggingface.co/datasets/Ariasyah/cic-ids-2017.tabulartabular-classification1M<n<10M0 likes219 downloads9mo agoHugging Face13vibhuiitj /UltraData-Math-L2-preview-embedded-with-idstabular10M<n<100M0 likes208 downloads7mo agoHugging Face14muzom /CIC-IDS2017The CICIDS2017 dataset consists of labeled network flows, including full packet payloads in pcap format, the corresponding profiles and the labeled flows (GeneratedLabelledFlows.zip) and CSV files for machine and deep learning purpose (MachineLearningCSV.zip) are publicly available for researchers. If you are using our dataset, you should cite our related paper which outlining the details of the dataset and its underlying principles: Iman Sharafaldin, Arash Habibi Lashkari, and Ali A.… See the full description on the dataset page: https://huggingface.co/datasets/muzom/CIC-IDS2017.tabular1M<n<10M0 likes196 downloads3mo agoHugging Face15JackHsieh /4B-general-paragraph-nothink.k-8.statml-arxiv-iid.qwen3-ids.tags 4B-general-paragraph-nothink.k-8.statml-arxiv-iid.qwen3-ids.tags Tokenized, tag-wrapped form of JackHsieh/4B-general-paragraph-nothink.k-8.statml-arxiv-iid.qwen3-ids. Each thought is wrapped as <|note|><thought><|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base token ids (input_ids). Each thought is the note after the generator's empty think block. Notes cut off at the generator's 1024-token cap (train 182,597 rows, test 47,885; flagged truncated) are an… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-general-paragraph-nothink.k-8.statml-arxiv-iid.qwen3-ids.tags.tabulartext-generation10M<n<100M0 likes195 downloads4d agoHugging Face16thepowerfuldeez /the-stack-v2-train-smol-ids-updatedUpdate on The Stack V2 dataset: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids All repos from original dataset are parsed with Github API and re-downloaded, so respective updates are kept, metadata is updated. This took 10+ days to process due to GraphQL limits. Filtering rules Removed repos with no update in the last 6 years (no updates since September 2019) Removed files with a single line Removed repos with a single file Removed repos with more than 99%… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated.tabular100K<n<1M0 likes189 downloads1y agoHugging Face17JackHsieh /4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids 4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176 — Qwen3-4B-Instruct-2507 after RL against a frozen suffix conditional. Each thought is a few dense sentences of reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tabulartext-generation10M<n<100M0 likes177 downloads17d agoHugging Face18veera33 /CIC-IDS2017We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily. NIDS Datasets The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/veera33/CIC-IDS2017.tabulartext-classification100M<n<1B0 likes173 downloads8mo agoHugging Face19JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes161 downloads26d agoHugging Face20JackHsieh /4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a few dense sentences that the generator wrapped in a literal <thought>…</thought> block (plain text, not special tokens). Only the part after </think> becomes the VALUE, with every <thought> / </thought> tag removed… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-reason-only.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes156 downloads1mo agoHugging Face21JackHsieh /4B-general-paragraph-nothink.k-8.statml-arxiv-iid.qwen3-ids 4B-general-paragraph-nothink.k-8.statml-arxiv-iid.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-iid-160M-20M, generated by Qwen/Qwen3-4B with thinking off. Each thought is a note: a long, dense passage of plain prose reasoning about the next 8 tokens after a cut, written from the document prefix alone — the generator never sees the continuation. input_ids holds the 4-token empty-think prefill <think>\n\n</think>\n\n followed by the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-general-paragraph-nothink.k-8.statml-arxiv-iid.qwen3-ids.tabulartext-generation10M<n<100M0 likes146 downloads4d agoHugging Face22JackHsieh /4B-general-paragraph-nothink.k-8.statml-arxiv-iid.qwen3-ids.kv-tags-explained 4B-general-paragraph-nothink.k-8.statml-arxiv-iid.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-general-paragraph-nothink.k-8.statml-arxiv-iid.qwen3-ids. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note about what might come next. KEY: <last 8 prefix tokens> VALUE: <thought> <|/note|> and stored both as text (thought_text) and as… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-general-paragraph-nothink.k-8.statml-arxiv-iid.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes133 downloads4d agoHugging Face23JackHsieh /4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained Tokenized, tag-wrapped form of JackHsieh/4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids, the thoughts written by the prestar-RL policy Qwen3-4B-Instruct-2507.prestar-RL.reason-only.lr7e-7-kl0.step-2176. Each thought is wrapped as <|note|> This is a hint about a span that appears later in this document. KEY is the text immediately before that span; VALUE is a note… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-RL-step2176-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabulartext-generation10M<n<100M0 likes126 downloads16d agoHugging Face24hookprobe /edge-ids-threats HookProbe Edge IDS Threat Telemetry Real-world, anonymised threat verdicts from the HookProbe production edge intrusion-detection system. Unlike synthetic lab datasets (CICIDS2017, UNSW-NB15, Kitsune) this is what an actual edge sensor mesh observes on the open internet, labelled by the SENTINEL ensemble (isolation forest + calibrated naive-Bayes) that ships with HookProbe. Sensor: Raspberry Pi edge node + NAPSE AI-native flow classifier Enrichment: RDAP country + ASN lookups… See the full description on the dataset page: https://huggingface.co/datasets/hookprobe/edge-ids-threats.tabulartabular-classification1M<n<10M0 likes113 downloads10d agoHugging Face25simple-pretraining /bookcorpusopen_with_ids_chunked Dataset Card for "bookcorpusopen_with_ids_chunked" More Information needed tabular10M<n<100M0 likes100 downloads3y agoHugging Face26JackHsieh /4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tags 4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tags Tokenized, tag-wrapped form of JackHsieh/4B-reason-only.rule-r-1.0-k-8.L-512.statml-arxiv. Each thought is wrapped as <|note|><thought><|/note|> and stored both as text (thought_text) and as Qwen/Qwen3-4B-Base token ids (input_ids). Longest thought: 510 tokens — a training run's max_thought_length must be at least this. Delimiter ids: <|note|> = 151669, <|/note|> = 151670. A run must declare these under… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-Instruct-reason-only.stride-1.k-8.statml-arxiv.qwen3-ids.tags.tabulartext-generation10M<n<100M0 likes92 downloads29d agoHugging Face27JackHsieh /luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled Every non-first chunk of every document carries a thought: the gpt-5.6-luna reasoning thought where one was generated, and a content-free pause thought everywhere else. luna chunk <|reserved_special_token_1|> luna reasoning <|reserved_special_token_2|> filler chunk <|reserved_special_token_1|> 256x <|reserved_special_token_0|> <|reserved_special_token_2|> The filler is 258 tokens. Chunk 0 is excluded… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.k-8.statml-arxiv-llama32.llama32-ids.tags-pausefilled.tabulartext-generation1M<n<10M0 likes91 downloads2mo agoHugging Face28pcy12345BSU /CIC-IDS-2017We have developed a Python package as a wrapper around Hugging Face Hub and Hugging Face Datasets library to access this dataset easily. NIDS Datasets The nids-datasets package provides functionality to download and utilize specially curated and extracted datasets from the original UNSW-NB15 and CIC-IDS2017 datasets. These datasets, which initially were only flow datasets, have been enhanced to include packet-level information from the raw PCAP files. The dataset contains both… See the full description on the dataset page: https://huggingface.co/datasets/pcy12345BSU/CIC-IDS-2017.tabulartext-classification100M<n<1B1 likes89 downloads6mo agoHugging Face29deokhk /filtered_mgsm_with_ids Filtered MGSM with IDs This dataset is a filtered subset of [juletxara/mgsm] with an added integer id per language. English questions overlapping with a PolyMath English set were removed, and the same ids were excluded from all languages to avoid cross-language overlaps. Source: juletxara/mgsm (test split) Filtering date: 2025-09-15 Splits: each language is exposed as a split (en, bn, de, es, ja, sw, te, th). Fields id (int): per-language index assigned before… See the full description on the dataset page: https://huggingface.co/datasets/deokhk/filtered_mgsm_with_ids.tabularquestion-answering1K<n<10K0 likes81 downloads1y agoHugging Face30JackHsieh /luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by gpt-5.6-luna. Each thought is visible reasoning about the next 8 Qwen3 tokens after a cut, written without ever seeing that continuation. Intended to be spliced into the document before the chunk so a small model (Qwen3-4B-Base) can read the reasoning and predict the chunk. Thoughts are strings, not token ids;… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/luna-reason-only.stride-train8-test32.k-8.statml-arxiv.qwen3-ids.tabulartext-generation100K<n<1M0 likes77 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.