Team Ai
20 results

redaction

RedactionBench /RedactionBench Dataset Card for RedactionBench RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual. RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate. The above is… See the full description on the dataset page: https://huggingface.co/datasets/RedactionBench/RedactionBench.texttoken-classificationn<1K2 likes276 downloads5mo agoHugging FaceA10Networks /RedactionBench Dataset Card for RedactionBench RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual. RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate. The… See the full description on the dataset page: https://huggingface.co/datasets/A10Networks/RedactionBench.texttoken-classificationn<1K0 likes92 downloads4mo agoHugging Facermems /log-redaction-trajectories Log Redaction Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/rmems/log-redaction-trajectories.textn<1K0 likes57 downloads1mo agoHugging Faceaibotjock /RedactionBench Dataset Card for RedactionBench RedactionBench is an evaluation-only benchmark for character-level redaction across eleven document categories. Each of the 200 documents is manually-annotated with character spans that are either mandatory (must redact) or contextual. RedactionBench mixes 101 real-world documents manually sourced from the public web (transcribed, augmented) with 99 synthetic documents authored to fill categories where synthetic data is more appropriate. The… See the full description on the dataset page: https://huggingface.co/datasets/aibotjock/RedactionBench.texttoken-classificationn<1K0 likes42 downloads3mo agoHugging Facenutrientdocs /DocPII-redaction-benchmark DocPII: Contextual Redaction Benchmark Dataset Dataset Description DocPII contains 1101 high-quality document samples enriched with embedded personally identifiable information (PII). Designed to evaluate context-aware redaction systems, it provides realistic, full-document contexts—a notable advancement over sentence-level datasets. All documents have been manually reviewed for accuracy, coherence, and redaction alignment, ensuring data quality for benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/nutrientdocs/DocPII-redaction-benchmark.texttext-generation1K<n<10K3 likes32 downloads1y agoHugging FaceMattStammers /Clinical_PII_Redaction_Testtexttoken-classificationn<1K0 likes23 downloads2y agoHugging Face