Team Ai
Datasetpublic

VertexAGI-Archives/cyberdata-large

⚠️ THERE IS A NEWER VERSION This version performs poorly. Use it only for comparison and testing — not as a main dataset and not for enterprise use. View Vertex AGI for the latest version → CyberData Large 65,000 SFT examples combining verified agentic-coding trajectories, cybersecurity agent behavior, and structured vulnerability intelligence, built entirely from non-gated, redistributable sources. The CyberData family Size Repo Examples Train… See the full description on the dataset page: https://huggingface.co/datasets/VertexAGI-Archives/cyberdata-large.

sourceHugging Faceotherupdated 5d agoView on Hugging Face
0likes743downloads
Dataset Card

<div align="center" style="border:6px solid #d00000;border-radius:18px;padding:34px 24px;margin:20px 0;background:#fff3f3;">

<div style="font-size:64px;line-height:1;">⚠️</div>

<div style="font-size:44px;font-weight:900;color:#b00000;margin:10px 0;">THERE IS A NEWER VERSION</div>

<div style="font-size:26px;font-weight:700;color:#222222;margin:10px 0 22px 0;"><b>This version performs poorly.</b> Use it <u>only for comparison and testing</u> — not as a main dataset and not for enterprise use.</div>

<a href="https://huggingface.co/VertexAGI" style="display:inline-block;padding:16px 34px;margin:8px;border-radius:12px;font-size:22px;font-weight:700;text-decoration:none;color:#ffffff;background:#111111;border:3px solid #111111;">View Vertex AGI for the latest version →</a>

</div>


<div align="center">

CyberData Large

65,000 SFT examples combining verified agentic-coding trajectories, cybersecurity agent behavior, and structured vulnerability intelligence, built entirely from non-gated, redistributable sources.

</div>


The CyberData family

SizeRepoExamplesTrainValid
SmallVertexAGI/cyberdata-small15,00013,8001,200
MediumVertexAGI/cyberdata-medium25,00023,0091,991
LargeVertexAGI/cyberdata-large65,00059,7455,255
FullVertexAGI/cyberdata-full479,214440,87338,341

The sizes are strictly nested (Small ⊂ Medium ⊂ Large ⊂ Full): every CyberData Small example is in Medium, Large and Full, and every Medium example is in Large and Full. Small's own validation examples stay validation examples in every size, and Medium's validation set is inside Large's, so you can train on a larger size and still evaluate on a smaller size's validation set without leakage.

Composition

SourceExamples%What it contributes
nvidia/SWE-Hero-openhands-trajectories27,54642.4%Execution-based software-engineering agent trajectories (OpenHands framework)
nvidia/Open-SWE-Traces24,51137.7%Multi-harness (OpenHands, SWE-agent, mini-swe-agent) coding-agent trajectories, resolved=1 only: the agent's patch was verified to fix the task
0xKitkat/AgentForge-11525760.9%Evidence-grounded defensive-cybersecurity agent trajectories: tool use, failed-check recovery, evidence-based findings (the whole security split)
ismailtasdelen/unified-vulnerability-intelligence-dataset5000.8%Structured vulnerability knowledge (CWE, CAPEC, MITRE ATT&CK, OWASP, CVSS, remediation, detection) templated into Q&A (the whole dataset)
stasvinokur/cve-and-cwe-dataset-1999-202511,86718.3%Real NVD CVE records (1999-2025) with severity, CVSS and CWE classification, templated into Q&A
Total65,000100%

How this size was built from Small

CyberData Small (15,000 examples, formerly cyberdata-1) already uses the entire AgentForge security split (576) and the entire UVID dataset (500), so those two cannot grow: their combined share is 1.7% here (7.2% in Small). The three sources that do have more supply, SWE-Hero, Open-SWE-Traces and CVE/CWE, grow in Small's own proportions, which is why the agentic-coding share is 80.1% here.

  • —New rows are chosen deterministically (md5 priorities), so the build is reproducible. Scripts are in build_scripts/.
  • —Open-SWE-Traces rows are resolved=1 only. The new rows are spread evenly across the ten harness/teacher/source-dataset combinations that contain resolved trajectories (three of the thirteen combinations contain none). Supply was not a limit at this size: 162,600 resolved trajectories exist upstream.
  • —SWE-Hero rows are drawn across all 14 upstream shards.
  • —CVE/CWE rows use the identical deterministic template and the same shuffled order as Small, excluding rows already used. Every fact in an answer comes from a column in that row; blank fields are omitted, nothing is inferred or generated by another model.
  • —New examples are validation examples with probability 8% by id hash (the same share as Small's 1,200 / 15,000); Small's rows keep the split they had in Small.

Format

Each line is {"messages": [...]}, standard chat SFT format; some assistant turns carry tool_calls (the agentic-coding trajectories).

  • —train-0000i-of-00016.jsonl: 59,745 training examples in 16 shards; every shard mixes all sources and is shuffled internally
  • —valid.jsonl: 5,255 validation examples (loads as the validation split)
  • —row_index.csv: one line per example: id, source, split, in_small (provenance, and the way to select exactly the Small subset)
  • —stats.json, build_scripts/
python
from datasets import load_dataset
ds = load_dataset("VertexAGI/cyberdata-large")
print(ds["train"][0]["messages"])

Licensing

This is a mixture of five independently licensed sources: CC-BY-4.0 (nvidia/SWE-Hero-openhands-trajectories, nvidia/Open-SWE-Traces), Apache-2.0 (0xKitkat/AgentForge-1152), MIT (ismailtasdelen/unified-vulnerability-intelligence-dataset) and CC0-1.0 (stasvinokur/cve-and-cwe-dataset-1999-2025). No source in this mix is gated or carries a non-commercial restriction. If you redistribute this dataset, retain attribution to each upstream source above, per their licenses (required for the CC-BY-4.0 portions).

Quality checks performed

Run on the final files before release (build_scripts/verify.py):

  • —Exactly 65,000 examples; zero duplicate ids; every line parses and has at least two messages with valid roles and string content.
  • —All 15,000 CyberData Small ids are present.
  • —All 1,200 of Small's validation examples are in this set's valid.jsonl (checked by content hash), and no Small training example became a validation example.
  • —Validation share is 8.08% (target 8%).
  • —Open-SWE-Traces rows are restricted to resolved=1.

Limitations

This is a resampled mixture, not a from-scratch curated dataset: quality is bounded by the upstream sources. The CVE/CWE and UVID portions are single-turn Q&A synthesized from structured data via a template, not natural human-written conversation; they teach the answer format and style and are not a verified knowledge base, so a model fine-tuned on them can still state wrong CVE/CWE facts. The agentic-coding portion dominates (80.1%); the cybersecurity-agent and vulnerability-knowledge portions are comparatively small (1.7% combined) and cannot grow beyond Small's amounts.