Team Ai
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01filipwx /the-secrets-of-ceos-book-2k The-Secrets-Of-Ceos-Book-2k Made with ❤️ using 🦥 Unsloth Studio Beta2x was generated with Unsloth Recipe Studio. It contains 2,000 generated records. 🚀 Quick Start from datasets import load_dataset # Load the main dataset dataset = load_dataset("filipwx/the-secrets-of-ceos-book-2k", "data", split="train") df = dataset.to_pandas() 📊 Dataset Summary 📈 Records: 2,000 📋 Columns: 4 📋 Schema & Statistics Column Type Column Type Unique… See the full description on the dataset page: https://huggingface.co/datasets/filipwx/the-secrets-of-ceos-book-2k.text1K<n<10K0 likes847 downloads6mo agoHugging Face02Podric /prowl-secrets-corpus Prowl secrets corpus A labeled corpus for training and evaluating secret detectors (credentials, API keys, tokens, database URIs, private keys, and passwords) across code, Jira tickets, Confluence pages, chat, and logs, in multiple languages. It is the training data behind Prowl and its stage-3 encoder. 503,027 records: 181,530 positive, 321,497 negative. No live credentials. Every value is synthetic, format-preserving-obfuscated, or drawn from a public test fixture. The… See the full description on the dataset page: https://huggingface.co/datasets/Podric/prowl-secrets-corpus.texttext-classification100K<n<1M2 likes466 downloads3mo agoHugging Face03alea-institute /kl3m-data-dotgov-www.secretservice.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.secretservice.gov.textn<1K1 likes132 downloads2y agoHugging Face04RealEstateCRMBooks /Scott-Schmitz-Real-Estate-CRM-Secretstext0 likes120 downloads3mo agoHugging Face05Hodfa71 /pstu-synthetic-secrets PSTU Synthetic Secrets Dataset Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper: Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal Hoda Fakhar — ECML PKDD 2026 Dataset Description 175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric. All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.texttext-generationn<1K0 likes101 downloads7mo agoHugging Face06ram-lexsi /curatorkit-testrun-Secrets curatorkit-testrun-Secrets Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method curation Backend — Model — Formats alpaca Artifact dataset Published 2026-08-30 05:55 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-Secrets", "alpaca") texttext-generationn<1K0 likes100 downloads1mo agoHugging Face07ACCC1380 /TOP_SECRETSgated 绝密模型 TOP SECRET MODELS textn<1K0 likes87 downloads8d agoHugging Face08Naphula-Archives /Dark-Psychology-Secretstext1K<n<10K1 likes77 downloads8mo agoHugging Face09alea-institute /kl3m-filter-data-dotgov-www.secretservice.govtextn<1K0 likes31 downloads2y agoHugging Face10cudecanarim /hayleys-secretsimage10K<n<100K0 likes30 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.