datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shell-attack-evolution-dataset
Shell Honeypot Attack Request–Response Dataset
A standardized, MITRE ATT&CK–annotated dataset of post-login shell
attacks captured by Cowrie SSH/Telnet
honeypots across two collection periods — 2021–2022 and 2024. It pairs
attacker shell commands with real captured system responses, enabling both
longitudinal threat analysis and the training/evaluation of AI-driven honeypots.
This is the open-source release accompanying the paper “Unveiling Evolving
Threats: A Data Analysis… See the full description on the dataset page: https://huggingface.co/datasets/zyw-286/shell-attack-evolution-dataset.douvras-algorithm-evolution-benchmark
Douvras Algorithm Evolution Benchmark v0.1
Synthetic candidate records with correctness, latency, memory and generation.
Candidates that fail correctness are invalid regardless of speed. It contains
48 records (32/8/8) across 12 workloads, split by workload.
Metrics are illustrative, not measured on real hardware. A real benchmark must
be run separately before claiming an optimization.
governed-skill-evolution
Governed Skill Evolution from Persistent Agent Experience
Prospective ablation and cross-model transfer study of three experience-retention conditions for governed Agent Skill evolution: no persistent history, flat chronological history, and a persistent Pattern Registry with a forward-chained Skill Impact Ledger.
Author: Song Luo
Version: 1.0.0
Source snapshot: d717c32396cfff1bef2800296541a70e9b4cabb8
Canonical repository: rrrrrredy/governed-skill-evolution
Zenodo:… See the full description on the dataset page: https://huggingface.co/datasets/RedinGhost/governed-skill-evolution.douvras-evolution-lab-constructions
Douvras Evolution Lab Constructions v0.1
Small, synthetic benchmark for the loop candidate → verifier → score.
Each problem is an eight-node cycle independent-set toy problem. The verifier
checks node range, uniqueness and absence of selected edges. The published
best_known_score is 3; a valid score of 4 is marked IMPROVED for this toy
family.
The dataset contains 54 records (36 train, 9 validation, 9 frozen test) across
six relabeled problem instances. Splitting is by… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-evolution-lab-constructions.ClaudioItaly__Evolutionstory-7B-v2.2-details
Dataset Card for Evaluation run of ClaudioItaly/Evolutionstory-7B-v2.2
Dataset automatically created during the evaluation run of model ClaudioItaly/Evolutionstory-7B-v2.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ClaudioItaly__Evolutionstory-7B-v2.2-details.
