Team Ai
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nisaefendioglu /synthetic-sensitive-data-in-source-code-n300 Synthetic Sensitive Data in Source Code (N=300) Synthetic dataset of 300 source-code / config snippets containing hardcoded secrets and PII.Every sample includes at least one sensitive finding (no clean negatives in the main file). Designed for local masking, secret detection, and OWASP LLM02 — Sensitive Information Disclosure workflows. Version 1.3.6: README Files table documents split extension (split_pattern_custom_n300) and clean CSV. v1.3.4: split files + Dataset Viewer… See the full description on the dataset page: https://huggingface.co/datasets/nisaefendioglu/synthetic-sensitive-data-in-source-code-n300.texttext-classificationn<1K1 likes171 downloads12d agoHugging Face02trl-lab /contextual-sensitive-data Towards Contextual Sensitive Data Detection This dataset includes tables with sensitivity annotations that were used to train and evaluate methods for detecting contextual sensitive data. It accompanies the paper "Towards Contextual Sensitive Data Detection". Links: Paper: https://huggingface.co/papers/2512.04120 Code: https://github.com/trl-lab/sensitive-data-detection Sample Usage The GitHub repository provides scripts for running inference and fine-tuning using… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/contextual-sensitive-data.texttext-classification1K<n<10K2 likes85 downloads7mo agoHugging Face03Arthur-AI /arthur_sensitive_data_passwordtextn<1K1 likes21 downloads2y agoHugging Face04huytx267 /vietnamese_text_sensitive_dataset Vietnamese Text Sensitive Dataset Mô tả vietnamese_text_sensitive_dataset là bộ dữ liệu chứa các văn bản tiếng Việt nhạy cảm liên quan đến nội dung khiêu dâm, bạo lực, phân biệt đối xử, sai lệch chính trị và các chủ đề khác. Bộ dữ liệu này có thể được sử dụng để huấn luyện các mô hình AI nhằm phát hiện và lọc nội dung nhạy cảm trong các ứng dụng xử lý ngôn ngữ tự nhiên (NLP). Cấu trúc dữ liệu Bộ dữ liệu bao gồm các danh mục sau: Nội dung khiêu dâm, nhạy cảm… See the full description on the dataset page: https://huggingface.co/datasets/huytx267/vietnamese_text_sensitive_dataset.text10K<n<100K1 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.