VertexAGI-Archives/cyberdata-large
⚠️ THERE IS A NEWER VERSION This version performs poorly. Use it only for comparison and testing — not as a main dataset and not for enterprise use. View Vertex AGI for the latest version → CyberData Large 65,000 SFT examples combining verified agentic-coding trajectories, cybersecurity agent behavior, and structured vulnerability intelligence, built entirely from non-gated, redistributable sources. The CyberData family Size Repo Examples Train… See the full description on the dataset page: https://huggingface.co/datasets/VertexAGI-Archives/cyberdata-large.
<div align="center" style="border:6px solid #d00000;border-radius:18px;padding:34px 24px;margin:20px 0;background:#fff3f3;">
<div style="font-size:64px;line-height:1;">⚠️</div>
<div style="font-size:44px;font-weight:900;color:#b00000;margin:10px 0;">THERE IS A NEWER VERSION</div>
<div style="font-size:26px;font-weight:700;color:#222222;margin:10px 0 22px 0;"><b>This version performs poorly.</b> Use it <u>only for comparison and testing</u> — not as a main dataset and not for enterprise use.</div>
<a href="https://huggingface.co/VertexAGI" style="display:inline-block;padding:16px 34px;margin:8px;border-radius:12px;font-size:22px;font-weight:700;text-decoration:none;color:#ffffff;background:#111111;border:3px solid #111111;">View Vertex AGI for the latest version →</a>
</div>
<div align="center">
CyberData Large
65,000 SFT examples combining verified agentic-coding trajectories, cybersecurity agent behavior, and structured vulnerability intelligence, built entirely from non-gated, redistributable sources.
</div>
The CyberData family
The sizes are strictly nested (Small ⊂ Medium ⊂ Large ⊂ Full): every CyberData Small example is in Medium, Large and Full, and every Medium example is in Large and Full. Small's own validation examples stay validation examples in every size, and Medium's validation set is inside Large's, so you can train on a larger size and still evaluate on a smaller size's validation set without leakage.
Composition
How this size was built from Small
CyberData Small (15,000 examples, formerly cyberdata-1) already uses the entire AgentForge security split (576) and the entire UVID dataset (500), so those two cannot grow: their combined share is 1.7% here (7.2% in Small). The three sources that do have more supply, SWE-Hero, Open-SWE-Traces and CVE/CWE, grow in Small's own proportions, which is why the agentic-coding share is 80.1% here.
- New rows are chosen deterministically (md5 priorities), so the build is reproducible. Scripts are in
build_scripts/. - Open-SWE-Traces rows are
resolved=1only. The new rows are spread evenly across the ten harness/teacher/source-dataset combinations that contain resolved trajectories (three of the thirteen combinations contain none). Supply was not a limit at this size: 162,600 resolved trajectories exist upstream. - SWE-Hero rows are drawn across all 14 upstream shards.
- CVE/CWE rows use the identical deterministic template and the same shuffled order as Small, excluding rows already used. Every fact in an answer comes from a column in that row; blank fields are omitted, nothing is inferred or generated by another model.
- New examples are validation examples with probability 8% by id hash (the same share as Small's 1,200 / 15,000); Small's rows keep the split they had in Small.
Format
Each line is {"messages": [...]}, standard chat SFT format; some assistant turns carry tool_calls (the agentic-coding trajectories).
train-0000i-of-00016.jsonl: 59,745 training examples in 16 shards; every shard mixes all sources and is shuffled internallyvalid.jsonl: 5,255 validation examples (loads as thevalidationsplit)row_index.csv: one line per example:id,source,split,in_small(provenance, and the way to select exactly the Small subset)stats.json,build_scripts/
from datasets import load_dataset
ds = load_dataset("VertexAGI/cyberdata-large")
print(ds["train"][0]["messages"])Licensing
This is a mixture of five independently licensed sources: CC-BY-4.0 (nvidia/SWE-Hero-openhands-trajectories, nvidia/Open-SWE-Traces), Apache-2.0 (0xKitkat/AgentForge-1152), MIT (ismailtasdelen/unified-vulnerability-intelligence-dataset) and CC0-1.0 (stasvinokur/cve-and-cwe-dataset-1999-2025). No source in this mix is gated or carries a non-commercial restriction. If you redistribute this dataset, retain attribution to each upstream source above, per their licenses (required for the CC-BY-4.0 portions).
Quality checks performed
Run on the final files before release (build_scripts/verify.py):
- Exactly 65,000 examples; zero duplicate ids; every line parses and has at least two messages with valid roles and string content.
- All 15,000 CyberData Small ids are present.
- All 1,200 of Small's validation examples are in this set's
valid.jsonl(checked by content hash), and no Small training example became a validation example. - Validation share is 8.08% (target 8%).
Open-SWE-Tracesrows are restricted toresolved=1.
Limitations
This is a resampled mixture, not a from-scratch curated dataset: quality is bounded by the upstream sources. The CVE/CWE and UVID portions are single-turn Q&A synthesized from structured data via a template, not natural human-written conversation; they teach the answer format and style and are not a verified knowledge base, so a model fine-tuned on them can still state wrong CVE/CWE facts. The agentic-coding portion dominates (80.1%); the cybersecurity-agent and vulnerability-knowledge portions are comparatively small (1.7% combined) and cannot grow beyond Small's amounts.
