Kicaulah/polyglot-bug-patterns
๐๏ธ Polyglot Bug Patterns ๐ Dataset ยท ๐จ Space ยท ๐ง Model ยท ๐ป GitHub Bro, imagine a dataset of 43 non-destructive attack payloads across 8 bug classes โ 24 of them polyglot, meaning a single string that's simultaneously plausible HTML, JS, SQL and shell, so it escapes whatever context your app dropped it into. Plus real detector output from a full multimodal scan. Everything here is synthesised or locally generated. No real-world traffic, no scraped data, no user data.โฆ See the full description on the dataset page: https://huggingface.co/datasets/Kicaulah/polyglot-bug-patterns.

๐๏ธ Polyglot Bug Patterns
[๐ Dataset](https://huggingface.co/datasets/Kicaulah/polyglot-bug-patterns) ยท [๐จ Space](https://huggingface.co/spaces/Kicaulah/polyglot-bughunter-x-static) ยท [๐ง Model](https://huggingface.co/Kicaulah/polyglot-bughunter-x) ยท [๐ป GitHub](https://github.com/skamy64-ux/polyglot-bughunter-x)
Bro, imagine a dataset of 43 non-destructive attack payloads across 8 bug classes โ 24 of them polyglot, meaning a single string that's simultaneously plausible HTML, JS, SQL and shell, so it escapes whatever context your app dropped it into. Plus real detector output from a full multimodal scan.
Everything here is synthesised or locally generated. No real-world traffic, no scraped data, no user data.
๐ฆ What's inside
Each config ships JSONL and Parquet, so both of these work:
from datasets import load_dataset
ds = load_dataset("Kicaulah/polyglot-bug-patterns", "payloads")
print(ds["train"][0]["value"])
print(ds["train"].filter(lambda r: r["polyglot"])[0]["value"])
findings = load_dataset("Kicaulah/polyglot-bug-patterns", "findings")
print(len(findings["train"]), "findings from one scan")๐งฌ The payload classes
Every payload carries a unique marker (PBHX7) so reflection is unambiguous โ that's what makes these usable as training signal for a detector or as a fixture in your own test suite.
โ ๏ธ Safety properties
This is a defensive dataset. Concretely:
- Every payload is non-destructive by construction. The
destructivecolumn isfalsefor all 43 rows โ verified by the samesafety.is_forbidden_payload()gate the scanner itself uses at runtime. - No
DROP/DELETE/UPDATE/INSERT, nosystem()/exec()/xp_cmdshell, nosleep()/benchmark(), no reverse shells, no persistence. - Time-based blind SQLi is deliberately absent โ
sleep()payloads are blocked project-wide. - Payloads are for systems you own or are authorized to test. Using them against someone else's production box is illegal in most jurisdictions.
๐งช Using it
As a test fixture โ assert your WAF/filter catches the polyglot set:
from datasets import load_dataset
ds = load_dataset("Kicaulah/polyglot-bug-patterns", "payloads")["train"]
for row in ds.filter(lambda r: r["vuln_class"] == "xss"):
print(row["value"]) # feed each through your own escaping functionAs detector training data โ findings pairs a payload with the signal that proved it (proof), the CVSS vector, and the remediation. That's a supervised signal for "was this injection real or a reflection".
As an LLM security eval โ prompt-injection rows are ready-made instruction overrides for testing whether your model or agent guardrails hold.
๐ Provenance
- payloads โ hand-authored from public documentation (OWASP WSTG, PortSwigger cheat sheets, public CVE writeups), rewritten to be non-destructive and marker-tagged.
- findings / scans โ actual output of
Hunter.demo()againstpolyglot_bug_hunter.demo_target, an intentionally vulnerable app that runs on localhost inside the test process. Regenerate withpython tools/build_dataset.py.
Nothing was scraped from production systems.
๐ Changelog
- v1.0.0 โ initial release: 43 payloads / 8 classes (24 polyglot), 36 findings from one full 4-modality demo scan, JSONL + Parquet, 12 languages declared.
๐ค Contribute a pattern
Missing a class? Open an issue or PR. The bar: it must be non-destructive (the automated gate enforces it), documented with a CWE and an OWASP mapping, and tagged for what detector signal it should produce.
<sub>MIT ยท ๐ท๏ธ PolyglotBugHunter-X ยท authorized security testing only</sub>
