Team Ai
Datasetpublic

allenai/BenchMIRT-item-statistics

Permitted Use: The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. Disclaimer: This benchmark measures the latent safety and general reasoning scores of LLMs. The data includes prompts and outputs that may contain biased, toxic, or harmful content. The prompts and outputs were generated using existing benchmarks and third party models, which are subject to the license terms of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/BenchMIRT-item-statistics.

sourceHugging Faceupdated 1mo agoView on Hugging Face
1likes80downloads
Dataset Card

Permitted Use: The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.

Disclaimer: This benchmark measures the latent safety and general reasoning scores of LLMs. The data includes prompts and outputs that may contain biased, toxic, or harmful content. The prompts and outputs were generated using existing benchmarks and third party models, which are subject to the license terms of the original sources. Please refer to the Prompt ID, Model ID, and metadata for source information.

This dataset contains the per-item statistics for the BenchMIRT project. Original benchmarks used in this analysis: