Team Ai
18 results

llm-evaluation

dipankarsarkar /llm-evaluation-self-audit LLM Evaluation Self-Audit Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research. Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a). The finding We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off. They often gave a different answer. Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.tabulartext-generation1K<n<10K1 likes303 downloads8d agoHugging FaceMultiverseComputingCAI /llm-refusal-evaluation 🛡️ LLM Refusal Evaluation Benchmark This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite. The prompts are organized into three groups: Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse. Chinese Sensitive Topics — prompts that may be censored by China-aligned models. Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse. 📌 Contents Safety Benchmarks JailbreakBench SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/MultiverseComputingCAI/llm-refusal-evaluation.texttext-generation1K<n<10K5 likes223 downloads10mo agoHugging FaceMomina-Muzafar /urdu-english-llm-evaluation Urdu-English Evaluation Dataset: Testing Qwen, Gemma, and Llama Testing three small open-source language models on a dataset of a hundred questions in Urdu and English, across five categories. Motivation It all starts when I noticed, while using voice and chat-based AI tools, that Urdu is often not handled as well as English, and I suspected that models aren't trained as extensively on Urdu compared to other languages like Hindi, that made me wonder if even… See the full description on the dataset page: https://huggingface.co/datasets/Momina-Muzafar/urdu-english-llm-evaluation.textquestion-answeringn<1K0 likes142 downloads3d agoHugging FaceSetloop /The-LLM-Attacker-Unified-Red-Team-Framework-for-Privacy-Evaluation-of-Split-and-Distributed-LLMsgated The LLM-Attacker: reproduction and peer-review release This dataset contains the final paper, frozen experiment evidence, and the source snapshot needed to inspect or reproduce the reported analyses. The repository is the pre-publication internal-review snapshot. The complete tagged evidence release will become publicly accessible with the paper. Tag v1.7.1 freezes the current manuscript and retained evidence tree. Final paper… See the full description on the dataset page: https://huggingface.co/datasets/Setloop/The-LLM-Attacker-Unified-Red-Team-Framework-for-Privacy-Evaluation-of-Split-and-Distributed-LLMs.3 likes127 downloads10h agoHugging FaceAITrailblazer /repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K4 likes113 downloads2mo agoHugging Facenbvbharath-1729 /llm-math-evaluation-dataset LLM Math Response Evaluation Dataset Dataset Summary A human-annotated dataset of 150 AI-generated math responses evaluated across GPT-4o, Claude, and Gemini. Each response is scored on Correctness, Reasoning, and Clarity using a structured rubric, with written justification for every score. Supported Tasks LLM evaluation and benchmarking Math reasoning quality assessment Error type classification in AI responses Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.tabularn<1K0 likes77 downloads16d agoHugging Face