evaluation
evaluation-results@misc{muennighoff2022crosslingual,
title={Crosslingual Generalization through Multitask Finetuning},
author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel},
year={2022},
eprint={2211.01786},
archivePrefix={arXiv},
primaryClass={cs.CL}
}GDP-Val-Evaluation-Submission
GDPval Submission Dataset
This dataset contains model outputs for GDP-Val evaluation.
Dataset Structure
data/: Contains the main dataset in Parquet format
train-00000-of-00001.parquet: Submission data with model outputs
deliverable_files/: Contains generated files for tasks that produce file deliverables
Organized by task_id
dataset_info.json: Metadata about the dataset
Columns
task_id: Unique identifier for each task
sector: Economic sector for the task… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/GDP-Val-Evaluation-Submission.vam-cross-evaluation-artifactsgrobid-evaluation
GROBID End-to-End Evaluation Dataset
Reference corpora used for GROBID end-to-end
benchmarking of scientific-article structuring.
Documentation: https://grobid.readthedocs.io/en/latest/End-to-end-evaluation/
Latest benchmarking scores: https://grobid.readthedocs.io/en/latest/Benchmarking/
Official archive (Zenodo): https://zenodo.org/record/7708580
Dataset summary
These are the datasets used for GROBID end-to-end benchmarking, covering:
metadata extraction… See the full description on the dataset page: https://huggingface.co/datasets/sciencialab/grobid-evaluation.aya_evaluation_suite
Dataset Summary
Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes:
human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated.
machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.evaluation_logs
Evaluation logs from "Auditing Games for Sandbagging"
This dataset provides evaluation transcripts produced for the paper "Auditing Games for Sandbagging". Transcripts are provided in Inspect .eval format, see https://github.com/AI-Safety-Institute/sabotage_games for a guide to viewing them.
Dataset Details
evaluation_transcripts/handover_evals contains the transcripts provided by the red team to the blue team at the beginning of the main round of the game, showing… See the full description on the dataset page: https://huggingface.co/datasets/sandbagging-games/evaluation_logs.
