nisaefendioglu/synthetic-sensitive-data-in-source-code-n300
Synthetic Sensitive Data in Source Code (N=300) Synthetic dataset of 300 source-code / config snippets containing hardcoded secrets and PII.Every sample includes at least one sensitive finding (no clean negatives in the main file). Designed for local masking, secret detection, and OWASP LLM02 — Sensitive Information Disclosure workflows. Version 1.3.6: README Files table documents split extension (split_pattern_custom_n300) and clean CSV. v1.3.4: split files + Dataset Viewer… See the full description on the dataset page: https://huggingface.co/datasets/nisaefendioglu/synthetic-sensitive-data-in-source-code-n300.
Synthetic Sensitive Data in Source Code (N=300)
Synthetic dataset of 300 source-code / config snippets containing hardcoded secrets and PII. Every sample includes at least one sensitive finding (no clean negatives in the main file).
Designed for local masking, secret detection, and OWASP LLM02 — Sensitive Information Disclosure workflows.
Version 1.3.6: README Files table documents split extension (split_pattern_custom_n300) and clean CSV. v1.3.4: split files + Dataset Viewer config split_pattern_n300. v1.3.2: clean extension (clean_n300).
All values are synthetic / fake. Do not treat them as real credentials.
Dataset Viewer
Use configs `sensitive_n300` (default), `clean_n300`, and `split_pattern_n300` for CSV previews. Download JSON for full sensitive_findings / custom_rules_recommended.
Files
Extension: clean negatives (N=300)
Extension: split / indirect secrets (N=300)
For evaluating built-in vs custom-regex masking on concatenated, folded, or comment-embedded secrets (thesis extension set).
Citation
Efendioğlu, N. N. (2026). Related thesis / dataset publication on Hugging Face (DOI: 10.57967/hf/9800).
