Team Ai
Datasetpublic

DollarSign/cicflow-ids-multiclass

CICFlow Multiclass Intrusion Detection Dataset Overview This dataset provides a multiclass network intrusion detection (IDS) benchmark derived from CICFlowMeter flow-level features. It is designed for attack-type classification, robust IDS research, and interpretable security modeling. Each network flow is labeled as either benign or one of nine attack categories, following a standard IDS taxonomy. The dataset is published in Hugging Face datasets format with… See the full description on the dataset page: https://huggingface.co/datasets/DollarSign/cicflow-ids-multiclass.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes20downloads
Dataset Card

CICFlow Multiclass Intrusion Detection Dataset

Overview

This dataset provides a multiclass network intrusion detection (IDS) benchmark derived from CICFlowMeter flow-level features. It is designed for attack-type classification, robust IDS research, and interpretable security modeling.

Each network flow is labeled as either benign or one of nine attack categories, following a standard IDS taxonomy.

The dataset is published in Hugging Face `datasets` format with explicit schemas, clean preprocessing, and reproducible splits.


Task

Multiclass classification

Given a network flow represented by CICFlowMeter features, predict the attack category.

Label Taxonomy

The dataset uses the following 10-class label space:

Label IDNameDescription
0benignNormal network traffic
1analysisPort scans, probing, and analysis activity
2backdoorBackdoor and remote access behavior
3dosDenial-of-Service attacks
4exploitsExploitation of vulnerabilities
5fuzzersFuzzing and malformed input attacks
6genericGeneric attack traffic
7reconnaissanceReconnaissance and information gathering
8shellcodeShellcode execution attempts
9wormsWorm propagation traffic

Additional label fields

  • —label → multiclass label (0–9)
  • —is_attack → binary indicator (1 if label ≠ 0, else 0)

This allows both multiclass IDS and binary detection experiments without reprocessing.


Dataset Structure

json
DatasetDict({
  train,
  validation,
  test
})

Each split contains records with the following schema:

json
flow_id: string
features: dict[str, float]
semantic_flags: dict[str, int]
label: ClassLabel (0–9)
is_attack: int (0/1)

Feature Description

1. Raw Numeric Features (features)

The features field contains CICFlowMeter-derived flow statistics, including:

Traffic volume & direction
  • —total_fwd_packets, total_bwd_packets
  • —total_length_of_fwd_packets
  • —total_length_of_bwd_packets
  • —down_up_ratio
Packet size statistics
  • —packet_length_min, packet_length_max
  • —packet_length_mean, packet_length_std, packet_length_variance
  • —Forward and backward packet length statistics
Timing & inter-arrival times
  • —flow_duration
  • —flow_iat_mean, flow_iat_std, flow_iat_max, flow_iat_min
  • —Forward and backward IAT statistics
  • —active_*, idle_* metrics
Rate-based features
  • —flow_bytes_per_s
  • —flow_packets_per_s
  • —fwd_packets_per_s, bwd_packets_per_s
TCP flag counters
  • —syn_flag_count, ack_flag_count, rst_flag_count
  • —fin_flag_count, psh_flag_count, urg_flag_count
  • —cwr_flag_count, ece_flag_count
Bulk and subflow statistics
  • —*_bulk_* features (conditionally emitted by CICFlowMeter)
  • —subflow_fwd_*, subflow_bwd_*
Note: Some CICFlowMeter features are conditionally emitted. Missing or undefined numeric values were filled with `0.0`, which semantically indicates absence of that behavior.

2. Semantic Flags (semantic_flags)

To support interpretable and rule-based IDS, each flow includes deterministic semantic indicators:

FlagMeaning
high_packet_rateVery high packets per second
one_way_trafficForward-only traffic (no response)
syn_ack_imbalanceLarge SYN/ACK imbalance
uniform_packet_sizeLow packet size variance
long_idle_c2Long idle periods (possible C2 behavior)

These flags are:

  • —Derived deterministically from numeric features
  • —Stable across dataset versions
  • —Suitable for rule-based or hybrid ML systems

Preprocessing Summary

The dataset was produced using a fully deterministic pipeline:

  • —Column name normalization
  • —Conversion of NaN / ±Infinity → 0.0
  • —Explicit label normalization (string → numeric)
  • —No row drops based on missing numeric values
  • —Stratified train/validation/test splits
  • —Explicit Hugging Face feature schemas

No synthetic balancing, oversampling, or augmentation was applied.


Class Distribution & Imbalance

The dataset is highly imbalanced, reflecting real-world network traffic:

  • —Benign traffic dominates
  • —Some attack classes are rare

This is intentional and realistic.

Users are strongly encouraged to:

  • —Use macro / weighted F1
  • —Inspect per-class recall
  • —Apply class weighting where appropriate

Accuracy alone is not a meaningful metric for this dataset.


Related Datasets

  • —A binary IDS version of this dataset is published separately for attack detection use cases.
  • —Both datasets share identical features and preprocessing logic.