Team Ai
Datasetpublic

Kicaulah/polyglot-bug-patterns

๐Ÿ—ƒ๏ธ Polyglot Bug Patterns ๐Ÿ Dataset ยท ๐ŸŽจ Space ยท ๐Ÿง  Model ยท ๐Ÿ’ป GitHub Bro, imagine a dataset of 43 non-destructive attack payloads across 8 bug classes โ€” 24 of them polyglot, meaning a single string that's simultaneously plausible HTML, JS, SQL and shell, so it escapes whatever context your app dropped it into. Plus real detector output from a full multimodal scan. Everything here is synthesised or locally generated. No real-world traffic, no scraped data, no user data.โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/Kicaulah/polyglot-bug-patterns.

sourceHugging Facemitupdated 6d agoView on Hugging Face
0likes99downloads
Dataset Card

demo

๐Ÿ—ƒ๏ธ Polyglot Bug Patterns

[๐Ÿ Dataset](https://huggingface.co/datasets/Kicaulah/polyglot-bug-patterns) ยท [๐ŸŽจ Space](https://huggingface.co/spaces/Kicaulah/polyglot-bughunter-x-static) ยท [๐Ÿง  Model](https://huggingface.co/Kicaulah/polyglot-bughunter-x) ยท [๐Ÿ’ป GitHub](https://github.com/skamy64-ux/polyglot-bughunter-x)


Bro, imagine a dataset of 43 non-destructive attack payloads across 8 bug classes โ€” 24 of them polyglot, meaning a single string that's simultaneously plausible HTML, JS, SQL and shell, so it escapes whatever context your app dropped it into. Plus real detector output from a full multimodal scan.

Everything here is synthesised or locally generated. No real-world traffic, no scraped data, no user data.

๐Ÿ“ฆ What's inside

ConfigRowsWhat it is
payloads43Every detection payload: value, bug class, CWE, OWASP, CVSS v3.1 vector+score, polyglot flag, and a destructive safety flag.
findings36Real findings from one full scan of the bundled demo target, with proof strings, endpoints, remediation, tags.
scans1Scan-level rollup: pages scanned, risk score, counts by severity, modality breakdown.
class_index9Per-class counts plus totals.

Each config ships JSONL and Parquet, so both of these work:

python
from datasets import load_dataset

ds = load_dataset("Kicaulah/polyglot-bug-patterns", "payloads")
print(ds["train"][0]["value"])
print(ds["train"].filter(lambda r: r["polyglot"])[0]["value"])

findings = load_dataset("Kicaulah/polyglot-bug-patterns", "findings")
print(len(findings["train"]), "findings from one scan")

๐Ÿงฌ The payload classes

ClassnExamples
sqli8' OR 1=1 -- , UNION SELECT NULL, MySQL error-based double-query, ORDER BY 99 column-count probe
xss9'"><svg/onload=alert(1)> attr breakout, </script><script> closer, JS-string escape, {{7*7}} SSTI
cmdi6;MARKER;, $(echo โ€ฆ), backticks, %0a CRLF, ${IFS}
traversal5..%2f..%2f..%2fetc%2fpasswd, ....//....//, double-double-encoded
ssrf5http://MARKER.oast.invalid/, 127.0.0.1, file:///, gopher://, [::1]
prompt-injection4"Ignore all previous instructions", </s>[INST] โ€ฆ [/INST] chat-template escape, markdown-fence system override
nosql-ldap3*(), '; return true; //, LDAP wildcard filter
redirect3absolute, protocol-relative, backslash open redirect

Every payload carries a unique marker (PBHX7) so reflection is unambiguous โ€” that's what makes these usable as training signal for a detector or as a fixture in your own test suite.

โš ๏ธ Safety properties

This is a defensive dataset. Concretely:

  • โ€”Every payload is non-destructive by construction. The destructive column is false for all 43 rows โ€” verified by the same safety.is_forbidden_payload() gate the scanner itself uses at runtime.
  • โ€”No DROP/DELETE/UPDATE/INSERT, no system()/exec()/xp_cmdshell, no sleep()/benchmark(), no reverse shells, no persistence.
  • โ€”Time-based blind SQLi is deliberately absent โ€” sleep() payloads are blocked project-wide.
  • โ€”Payloads are for systems you own or are authorized to test. Using them against someone else's production box is illegal in most jurisdictions.

๐Ÿงช Using it

As a test fixture โ€” assert your WAF/filter catches the polyglot set:

python
from datasets import load_dataset
ds = load_dataset("Kicaulah/polyglot-bug-patterns", "payloads")["train"]
for row in ds.filter(lambda r: r["vuln_class"] == "xss"):
    print(row["value"])   # feed each through your own escaping function

As detector training data โ€” findings pairs a payload with the signal that proved it (proof), the CVSS vector, and the remediation. That's a supervised signal for "was this injection real or a reflection".

As an LLM security eval โ€” prompt-injection rows are ready-made instruction overrides for testing whether your model or agent guardrails hold.

๐Ÿ“Š Provenance

  • โ€”payloads โ€” hand-authored from public documentation (OWASP WSTG, PortSwigger cheat sheets, public CVE writeups), rewritten to be non-destructive and marker-tagged.
  • โ€”findings / scans โ€” actual output of Hunter.demo() against polyglot_bug_hunter.demo_target, an intentionally vulnerable app that runs on localhost inside the test process. Regenerate with python tools/build_dataset.py.

Nothing was scraped from production systems.

๐Ÿ“ Changelog

  • โ€”v1.0.0 โ€” initial release: 43 payloads / 8 classes (24 polyglot), 36 findings from one full 4-modality demo scan, JSONL + Parquet, 12 languages declared.

๐Ÿค Contribute a pattern

Missing a class? Open an issue or PR. The bar: it must be non-destructive (the automated gate enforces it), documented with a CWE and an OWASP mapping, and tagged for what detector signal it should produce.

<sub>MIT ยท ๐Ÿ•ท๏ธ PolyglotBugHunter-X ยท authorized security testing only</sub>