Team Ai
20 results

perplexity

perplexity-ai /draco DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity The DRACO Benchmark consists of complex, open-ended research tasks with expert-curated rubrics for evaluating deep research systems. Tasks span 10 domains and require drawing on information sources from 40 countries. Each task is paired with a detailed, task-specific rubric featuring an average of ~40 evaluation criteria across four axes: factual accuracy, breadth and depth of analysis… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/draco.textn<1K116 likes3.5k downloads8mo agoHugging Faceperplexity-ai /browsesafe-bench Dataset Card for BrowseSafe-Bench Dataset Details Dataset Description BrowseSafe-Bench is a comprehensive security benchmark designed to evaluate the robustness of AI browser agents against prompt injection attacks embedded in realistic HTML environments. Unlike prior benchmarks that focus on simple text injections, BrowseSafe-Bench emphasizes environmental realism, incorporating complex HTML structures, diverse attack semantics, and benign "distractor"… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/browsesafe-bench.texttext-classification10K<n<100K29 likes841 downloads10mo agoHugging FaceTristan /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters" More Information needed tabular10M<n<100M0 likes538 downloads4y agoHugging Faceperplexity-ai /wandr WANDR Overview and provenance WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, structured, high-volume web research tasks. This dataset is a task-and-verification corpus, not a question/answer collection: it contains no solver outputs or reference answer sets. WANDR evaluation refetches cited pages and judges submitted records against task-specific, reference-free specifications. See the paper, the blog, and the evaluation repository.… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/wandr.textquestion-answeringn<1K3 likes322 downloads1mo agoHugging Faceperplexity-ai /PII-TRACE PII-TRACE PII-TRACE is a synthetic dataset for privacy-focused named entity recognition (NER) of personally identifiable information (PII), released as a 500-conversation subset of multi-turn dialogues with exact span annotations. Dataset summary This release contains 500 conversations with 2,653 annotated PII spans across nine labels. Among them, 450 conversations contain PII spans and 50 contain none. Data format Each record contains: Field… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/PII-TRACE.texttoken-classificationn<1K2 likes287 downloads18d agoHugging Faceprott5-memorization /pseudo-perplexity0 likes277 downloads6mo agoHugging Face