Team Ai
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RedactionBench /RedactionBench Dataset Card for RedactionBench RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual. RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate. The above is… See the full description on the dataset page: https://huggingface.co/datasets/RedactionBench/RedactionBench.texttoken-classificationn<1K2 likes276 downloads5mo agoHugging Face02A10Networks /RedactionBench Dataset Card for RedactionBench RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual. RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate. The… See the full description on the dataset page: https://huggingface.co/datasets/A10Networks/RedactionBench.texttoken-classificationn<1K0 likes92 downloads4mo agoHugging Face03rmems /log-redaction-trajectories Log Redaction Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/log-redaction-trajectories.textn<1K0 likes57 downloads1mo agoHugging Face04aibotjock /RedactionBench Dataset Card for RedactionBench RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual. RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate. The… See the full description on the dataset page: https://huggingface.co/datasets/aibotjock/RedactionBench.texttoken-classificationn<1K0 likes42 downloads3mo agoHugging Face05nutrientdocs /DocPII-redaction-benchmark DocPII: Contextual Redaction Benchmark Dataset Dataset Description DocPII contains 1101 high-quality document samples enriched with embedded personally identifiable information (PII). Designed to evaluate context-aware redaction systems, it provides realistic, full-document contexts—a notable advancement over sentence-level datasets. All documents have been manually reviewed for accuracy, coherence, and redaction alignment, ensuring data quality for benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/DocPII-redaction-benchmark.texttext-generation1K<n<10K3 likes32 downloads1y agoHugging Face06MattStammers /Clinical_PII_Redaction_Testtexttoken-classificationn<1K0 likes23 downloads2y agoHugging Face07King-Harry /NinjaMasker-PII-Redactiontext10K<n<100K2 likes22 downloads3y agoHugging Face08ClarusC64 /legal-redaction-necessity-scope-consistency-coherence-v0.1What this dataset does You receive doc context sensitive elements proposed redactions justification tags duplicate signals over or under signals You decide coherent or incoherent Daily use redaction QC over redaction flag under redaction flag duplicate consistency check tabulartext-classificationn<1K0 likes22 downloads8mo agoHugging Face09King-Harry /NinjaMasker-PII-Redaction-Datasettext10K<n<100K1 likes20 downloads3y agoHugging Face10edithram23 /llama-redaction-instructiontext10K<n<100K0 likes19 downloads2y agoHugging Face11edithram23 /llama-redactiontext10K<n<100K0 likes18 downloads2y agoHugging Face12ClarusC64 /legal-disclosure-tagging-relevance-privilege-redaction-coherence-risk-v0.1What this dataset does You receive doc summary issue list tag privilege basis redaction rationale rule consistency notes You decide coherent or incoherent Daily use batch tagging QC privilege basis checking redaction logic consistency tabulartext-classificationn<1K0 likes18 downloads8mo agoHugging Face13aldersondev /phi-redaction-sfttext10K<n<100K0 likes12 downloads5mo agoHugging Face14edithram23 /redaction_keytext10K<n<100K0 likes9 downloads2y agoHugging Face15edithram23 /PII-redaction-berttext10K<n<100K0 likes8 downloads2y agoHugging Face16arkorlab /redaction-demotext1K<n<10K0 likes8 downloads5mo agoHugging Face17Toomate /Redaction-short📥 Télécharger le dataset complet 0 likes4 downloads5mo agoHugging Face18edithram23 /Redaction0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.