skyylord/mitre-attack-ttp-labeled-instructions
MITRE ATT&CK TTP Mapping Dataset Training and evaluation data for mapping adversarial behavior descriptions (CTI reports, CTF writeups, CISA advisories) to MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs). Built as my individual contribution to a research project conducted at LORIA (supervised by Jean-Yves Marion). This dataset was developed and used to fine-tune skyylord/qwen3-emb-0.6b-ttp with CachedMultipleNegativesRankingLoss and ANCE-style hard negative re-mining.… See the full description on the dataset page: https://huggingface.co/datasets/skyylord/mitre-attack-ttp-labeled-instructions.
MITRE ATT&CK TTP Mapping Dataset
Training and evaluation data for mapping adversarial behavior descriptions (CTI reports, CTF writeups, CISA advisories) to MITRE ATT&CK Tactics, Techniques, and Procedures (TTPs).
Built as my individual contribution to a research project conducted at LORIA (supervised by Jean-Yves Marion). This dataset was developed and used to fine-tune skyylord/qwen3-emb-0.6b-ttp with CachedMultipleNegativesRankingLoss and ANCE-style hard negative re-mining.
Note: Due to their security nature, these datasets contain textual information about malware and other security aspects.
Dataset summary
Each example pairs a short instruction (a natural-language description of an adversarial action — "the attacker used X to achieve Y") with the ground-truth MITRE ATT&CK technique it corresponds to. The dataset is built for contrastive retrieval training: given an instruction, retrieve the correct technique out of the full ATT&CK Enterprise matrix (697 techniques/sub-techniques, v19).
Dataset Description
Expert, TRAM, Procedure, Derived Procedure
These four splits originate from tumeteor/Security-TTP-Mapping, built for the NCE paper (see Citation). Summarized from their dataset card:
- TRAM: sourced from CTID's TRAM project, then deduplicated, cleaned of short/noisy text and noisy labels, and remapped to ATT&CK by tumeteor.
- Procedure: procedure examples pulled from MITRE's CTI repo, with markup stripped.
- Derived Procedure: built by crawling the reference URLs cited by each procedure example and extracting the passages relevant to that example.
- Expert: paragraphs drawn from a broader pool of threat reports and annotated by security practitioners; multi-label, with the test split averaging ~4 labels per example.
For Laika, these splits were additionally remapped from the ATT&CK version tumeteor originally used to ATT&CK v19 Enterprise, and re-split/re-processed to match the rest of this dataset's format. All credit for original collection, cleaning, and the train/dev/test partitioning goes to tumeteor and the NCE paper authors.
Groups, Software, Campaigns
Collected and processed by me directly from MITRE ATT&CK's own Groups, Software, and Campaigns pages (not a pre-existing dataset — MITRE publishes these as reference pages, not training-ready pairs). In MITRE's own framing:
- Groups are adversary activity clusters tracked under a common name by the security community; each is mapped to its publicly-reported technique use, associated software, and campaigns.
- Software covers tools, malware, and utilities (commercial, open-source, or built-in) used to carry out adversary behavior, mapped to the techniques and groups they've been reported with.
- Campaigns are bounded-time intrusion activity sharing common targets/objectives, linked to the techniques, groups, and software involved where public reporting allows attribution. I converted these three pages into instruction/technique-label pairs for training.
CISA (held-out)
Scraped from published CISA cybersecurity advisories and processed into instruction/technique-label pairs. Used exclusively as a held-out evaluation set — never trained on.
CTF writeups (held-out)
A mix of my own CTF writeups and public writeups from 0xdf, processed by extracting the commands/actions used during each challenge and generating behavioral textual descriptions from them via an LLM, then labeling the result with the corresponding ATT&CK technique. The dataset contains only these derived behavioral descriptions, not the original writeup text. Held out — never trained on.
Configs / subsets
Citation
If you use the expert, tram, procedure, or derived_procedure splits, please cite the original NCE paper, whose data these are built from:
@inproceedings{nguyen-srndic-neth-ttpm,
title = "Noise Contrastive Estimation-based Matching Framework for Low-resource Security Attack Pattern Recognition",
author = "Nguyen, Tu and Srndic, Nedim and Neth, Alexander",
booktitle = "Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics",
month = mar,
year = "2024",
publisher = "Association for Computational Linguistics"
}If you use this dataset as a whole (including the groups/software/campaigns/cisa/ctf_writeups splits and the ATT&CK-v19 remapping):
@misc{laika-attack-ttp-dataset,
author = {Amine Akhlafa},
title = {Laika ATT&CK TTP Mapping Dataset},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/datasets/skyylord/mitre-attack-ttp-labeled-instructions}}
}