Team Ai
Datasetpublic

trapstreet/jev-class-cve-bench

Jev-class: name the weakness class You get the published description of one software vulnerability. You pick the weakness class it belongs to, from a list of 30 CWE identifiers. 30 options is a lot for this kind of benchmark, and the number of options is itself a limit. Some tools that do this job only accept short option lists — one published implementation refuses any list longer than 16 before it even runs the model. It does not score badly; it cannot answer at all. In… See the full description on the dataset page: https://huggingface.co/datasets/trapstreet/jev-class-cve-bench.

sourceHugging Faceotherupdated 7d agoView on Hugging Face
2likes98downloads
Dataset Card

Jev-class: name the weakness class

You get the published description of one software vulnerability. You pick the weakness class it belongs to, from a list of 30 CWE identifiers.

30 options is a lot for this kind of benchmark, and the number of options is itself a limit. Some tools that do this job only accept short option lists — one published implementation refuses any list longer than 16 before it even runs the model. It does not score badly; it cannot answer at all.

In linkloadgnssimage of linkdevice.c, there is a possible out-of-bounds write due to a missing bounds check. This could lead to local escalation of privilege with System execution privileges needed.

→ ANSWER: CWE-787

This is one task in [decision-layer-bench](https://github.com/trapstreet/decision-layer-bench), which tests Jev-class typed decision models: small models that take a list of options and return one of them plus a confidence, instead of calling an LLM. Jev 1.13.0 is on the leaderboard, along with open replicas and classifiers built on off-the-shelf models. We are not affiliated with TypeSafe — "Jev-class" is just what this category of model is called.

Two splits

casesanswerswhere you score it
dev150, 5 per classincludedyour own machine
test1,500, 50 per classnot includedthe leaderboard

No case appears in both, and neither overlaps the 500 cases used while designing the task.

judge.py is in this repository — the same file the grading server runs — so you can score dev yourself, and you can read exactly how a score is calculated.

The test answers are not here. If they were public, nobody could tell whether a high score meant the model classified well or just looked the answers up. The dev answers were never used for scoring, so publishing them costs nothing.

Do not compare dev scores. With 150 cases the margin of error is about ±0.036, and the top three entries on the leaderboard are within 0.006 of each other. Dev tells you your setup works. It cannot tell you who is better.

Running it

bash
uv tool install trap-cli
tp run <your-solution> --task cve-weakness-class     # local, no account, no network
tp auth login
tp submit <your-solution> --task cve-weakness-class

Anyone can run this task and put their result on the leaderboard: [trapstreet.run/tasks/cve-weakness-class](https://trapstreet.run/tasks/cve-weakness-class). Submissions are open. Every entry is scored by the same judge on the same cases, so a new result is directly comparable to the ones already there.

Output. Print one line. The identifier has to be one of the 30. If you print several, the last valid one is used.

ANSWER: CWE-787
CONFIDENCE: 0.85

The confidence line is optional and is never part of your score — a confident wrong answer and an unsure wrong answer both score zero. It is reported separately, as overconfidence: your average stated confidence minus your accuracy. An entry at +0.52 is saying it is about 79% sure and getting 27% right. A negative number means it is less confident than it should be.

Wrong answers are also split into two kinds: picking a parent or child of the correct class in the CWE hierarchy, versus picking something unrelated.

The floor

A plain string matcher, with no model at all, scores 0.478 here. That is because 392 of the 1,500 descriptions contain the name of their own class — "out-of-bounds write", "SQL injection" — so matching words gets them for free. 0.478 is the score you have to beat. Guessing randomly scores 0.033. (all baselines)

The descriptions are NVD's exact text, so anyone with a search engine can find the original record, which lists the answer. We cannot filter that out, so we are telling you instead. Descriptions that mention a CWE or CVE identifier are removed from the pool, since those give the answer away directly. Rebuilding the set each month keeps it newer than model training data, but it does not stop anyone looking things up.

What is in it

test1,500 cases — case_id, description
devdev/, 150 cases with answers
optionsoptions.json, the 30 CWE ids and their names
instructionstask.md
judgejudge.py and tools/cwe_tree.json (MITRE's CWE hierarchy)
sourceCVEs published 2026-07-01 to 2026-09-15

Every class has the same number of cases, which is not how CVEs are distributed in real life — a model that guessed the most common class would score well otherwise. Cases with identical description text are removed.

Each refresh adds a new config and never changes an old one. So build_2026_09_20 will always mean these same 1,500 cases, and results measured on it stay reproducible. See `REFRESH.md`.

Attribution

The vulnerability records come from the National Vulnerability Database (NIST), which is public domain. The weakness classes come from CWE™ (MITRE), released for free public use; CWE is a trademark of The MITRE Corporation, catalogue version 4.20. Neither NIST nor MITRE endorses this benchmark.

decision-layer-bench, task cve_weakness_class, build 2026_09_20.
Repo:  https://github.com/trapstreet/decision-layer-bench
Board: https://trapstreet.run/tasks/cve-weakness-class
Cases: https://huggingface.co/datasets/trapstreet/jev-class-cve-bench