llm-evaluation
llm-evaluation-self-audit
LLM Evaluation Self-Audit
Data for How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure, Dipankar Sarkar, Skelf Research.
Code and paper source: github.com/sarkar-dipankar/llm-evaluation-self-audit (this package is built from commit 687ee8a).
The finding
We asked eight open model variants to infer the structure of a prompt, then asked again with the identical call. Caching was off.
They often gave a different answer.
Mean… See the full description on the dataset page: https://huggingface.co/datasets/dipankarsarkar/llm-evaluation-self-audit.llm-refusal-evaluation
🛡️ LLM Refusal Evaluation Benchmark
This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite.
The prompts are organized into three groups:
Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse.
Chinese Sensitive Topics — prompts that may be censored by China-aligned models.
Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse.
📌 Contents
Safety Benchmarks
JailbreakBench
SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/MultiverseComputingCAI/llm-refusal-evaluation.urdu-english-llm-evaluation
Urdu-English Evaluation Dataset: Testing Qwen, Gemma, and Llama
Testing three small open-source language models on a dataset of a hundred questions in Urdu and English, across five categories.
Motivation
It all starts when I noticed, while using voice and chat-based AI tools, that Urdu is often not handled as well as English, and I suspected that models aren't trained as extensively on Urdu compared to other languages like Hindi, that made me wonder if even… See the full description on the dataset page: https://huggingface.co/datasets/Momina-Muzafar/urdu-english-llm-evaluation.The-LLM-Attacker-Unified-Red-Team-Framework-for-Privacy-Evaluation-of-Split-and-Distributed-LLMs
The LLM-Attacker: reproduction and peer-review release
This dataset contains the final paper, frozen experiment evidence, and the source snapshot needed to inspect or reproduce the reported analyses.
The repository is the pre-publication internal-review snapshot. The complete
tagged evidence release will become publicly accessible with the paper. Tag
v1.7.1 freezes the current manuscript and retained evidence tree.
Final paper… See the full description on the dataset page: https://huggingface.co/datasets/Setloop/The-LLM-Attacker-Unified-Red-Team-Framework-for-Privacy-Evaluation-of-Split-and-Distributed-LLMs.repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
llm-math-evaluation-dataset
LLM Math Response Evaluation Dataset
Dataset Summary
A human-annotated dataset of 150 AI-generated math responses
evaluated across GPT-4o, Claude, and Gemini. Each response is
scored on Correctness, Reasoning, and Clarity using a structured
rubric, with written justification for every score.
Supported Tasks
LLM evaluation and benchmarking
Math reasoning quality assessment
Error type classification in AI responses
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.
