Team Ai
Datasetpublic

skyylord/mitre-attack-ttp-labeled-instructions

MITRE ATT&CK TTP Mapping Dataset Training and evaluation data for mapping adversarial behavior descriptions (CTI reports, CTF writeups, CISA advisories) to MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs). Built as my individual contribution to a research project conducted at LORIA (supervised by Jean-Yves Marion). This dataset was developed and used to fine-tune skyylord/qwen3-emb-0.6b-ttp with CachedMultipleNegativesRankingLoss and ANCE-style hard negative re-mining.… See the full description on the dataset page: https://huggingface.co/datasets/skyylord/mitre-attack-ttp-labeled-instructions.

sourceHugging Facecc-by-nc-sa-4.0updated 20d agoView on Hugging Face
0likes60downloads
Dataset Card

MITRE ATT&CK TTP Mapping Dataset

Training and evaluation data for mapping adversarial behavior descriptions (CTI reports, CTF writeups, CISA advisories) to MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs).

Built as my individual contribution to a research project conducted at LORIA (supervised by Jean-Yves Marion). This dataset was developed and used to fine-tune skyylord/qwen3-emb-0.6b-ttp with CachedMultipleNegativesRankingLoss and ANCE-style hard negative re-mining.

Note: Due to their security nature, these datasets contain textual information about malware and other security aspects.

Dataset summary

Each example pairs a short instruction (a natural-language description of an adversarial action — "the attacker used X to achieve Y") with the ground-truth MITRE ATT&CK technique it corresponds to. The dataset is built for contrastive retrieval training: given an instruction, retrieve the correct technique out of the full ATT&CK Enterprise matrix (697 techniques/sub-techniques, v19).

Dataset Description

Expert, TRAM, Procedure, Derived Procedure

These four splits originate from tumeteor/Security-TTP-Mapping, built for the NCE paper (see Citation). Summarized from their dataset card:

  • —TRAM: sourced from CTID's TRAM project, then deduplicated, cleaned of short/noisy text and noisy labels, and remapped to ATT&CK by tumeteor.
  • —Procedure: procedure examples pulled from MITRE's CTI repo, with markup stripped.
  • —Derived Procedure: built by crawling the reference URLs cited by each procedure example and extracting the passages relevant to that example.
  • —Expert: paragraphs drawn from a broader pool of threat reports and annotated by security practitioners; multi-label, with the test split averaging ~4 labels per example.

For Laika, these splits were additionally remapped from the ATT&CK version tumeteor originally used to ATT&CK v19 Enterprise, and re-split/re-processed to match the rest of this dataset's format. All credit for original collection, cleaning, and the train/dev/test partitioning goes to tumeteor and the NCE paper authors.

Groups, Software, Campaigns

Collected and processed by me directly from MITRE ATT&CK's own Groups, Software, and Campaigns pages (not a pre-existing dataset — MITRE publishes these as reference pages, not training-ready pairs). In MITRE's own framing:

  • —Groups are adversary activity clusters tracked under a common name by the security community; each is mapped to its publicly-reported technique use, associated software, and campaigns.
  • —Software covers tools, malware, and utilities (commercial, open-source, or built-in) used to carry out adversary behavior, mapped to the techniques and groups they've been reported with.
  • —Campaigns are bounded-time intrusion activity sharing common targets/objectives, linked to the techniques, groups, and software involved where public reporting allows attribution. I converted these three pages into instruction/technique-label pairs for training.

CISA (held-out)

Scraped from published CISA cybersecurity advisories and processed into instruction/technique-label pairs. Used exclusively as a held-out evaluation set — never trained on.

CTF writeups (held-out)

A mix of my own CTF writeups and public writeups from 0xdf, processed by extracting the commands/actions used during each challenge and generating behavioral textual descriptions from them via an LLM, then labeling the result with the corresponding ATT&CK technique. The dataset contains only these derived behavioral descriptions, not the original writeup text. Held out — never trained on.

Configs / subsets

DatasetSourceRoleSize
expertNCE paper (Expert)train/val/test461/67/157
tramNCE paper (TRAM)train/val/test3346/582/703
procedureNCE paper (Procedure)train/val/test8323/1451/1735
derived_procedureNCE paper (Derived Procedure)train/val/test2463/482/507
campaignsMitre Campaignstrain/val/test621/115/116
softwareMitre Softwaretrain/val/test2910/645/655
groupsMitre Groupstrain/val/test1500/311/330
cisaCISA advisoriesheld-out (eval only)1535
ctfCustom CTF writeup corpusheld-out (eval only)102
mitre-enterprise-v19MITRE ATT&CK v19 Enterprisereference corpus697

Citation

If you use the expert, tram, procedure, or derived_procedure splits, please cite the original NCE paper, whose data these are built from:

bibtex
@inproceedings{nguyen-srndic-neth-ttpm,
    title     = "Noise Contrastive Estimation-based Matching Framework for Low-resource Security Attack Pattern Recognition",
    author    = "Nguyen, Tu and Srndic, Nedim and Neth, Alexander",
    booktitle = "Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics",
    month     = mar,
    year      = "2024",
    publisher = "Association for Computational Linguistics"
}

If you use this dataset as a whole (including the groups/software/campaigns/cisa/ctf_writeups splits and the ATT&CK-v19 remapping):

bibtex
@misc{laika-attack-ttp-dataset,
  author       = {Amine Akhlafa},
  title        = {Laika ATT&CK TTP Mapping Dataset},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/skyylord/mitre-attack-ttp-labeled-instructions}}
}

Related