Team Ai
Datasetpublic

jerogo/or-bench

OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/jerogo/or-bench.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes63downloads
settings

This repository belongs to jerogo on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameor-bench
visibilitypublic
licencecc-by-4.0
gatedno
ownerjerogo
Account settings
jerogo/or-bench · Team Ai