Team Ai
20 results

evaluation

bigscience /evaluation-results@misc{muennighoff2022crosslingual, title={Crosslingual Generalization through Multitask Finetuning}, author={Niklas Muennighoff and Thomas Wang and Lintang Sutawika and Adam Roberts and Stella Biderman and Teven Le Scao and M Saiful Bari and Sheng Shen and Zheng-Xin Yong and Hailey Schoelkopf and Xiangru Tang and Dragomir Radev and Alham Fikri Aji and Khalid Almubarak and Samuel Albanie and Zaid Alyafeai and Albert Webson and Edward Raff and Colin Raffel}, year={2022}, eprint={2211.01786}, archivePrefix={arXiv}, primaryClass={cs.CL} }other100M<n<1B10 likes256k downloads3y agoHugging Facexiachongfeng /GDP-Val-Evaluation-Submission GDPval Submission Dataset This dataset contains model outputs for GDP-Val evaluation. Dataset Structure data/: Contains the main dataset in Parquet format train-00000-of-00001.parquet: Submission data with model outputs deliverable_files/: Contains generated files for tasks that produce file deliverables Organized by task_id dataset_info.json: Metadata about the dataset Columns task_id: Unique identifier for each task sector: Economic sector for the task… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/GDP-Val-Evaluation-Submission.textn<1K0 likes13k downloads1y agoHugging Facedreamdifferent /vam-cross-evaluation-artifacts0 likes8.5k downloads2m agoHugging Facesciencialab /grobid-evaluation GROBID End-to-End Evaluation Dataset Reference corpora used for GROBID end-to-end benchmarking of scientific-article structuring. Documentation: https://grobid.readthedocs.io/en/latest/End-to-end-evaluation/ Latest benchmarking scores: https://grobid.readthedocs.io/en/latest/Benchmarking/ Official archive (Zenodo): https://zenodo.org/record/7708580 Dataset summary These are the datasets used for GROBID end-to-end benchmarking, covering: metadata extraction… See the full description on the dataset page: https://huggingface.co/datasets/sciencialab/grobid-evaluation.document1K<n<10K1 likes8.2k downloads3mo agoHugging FaceCohereLabs /aya_evaluation_suite Dataset Summary Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes: human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated. machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.tabulartext-generation10K<n<100K55 likes7.1k downloads1y agoHugging Facesandbagging-games /evaluation_logs Evaluation logs from "Auditing Games for Sandbagging" This dataset provides evaluation transcripts produced for the paper "Auditing Games for Sandbagging". Transcripts are provided in Inspect .eval format, see https://github.com/AI-Safety-Institute/sabotage_games for a guide to viewing them. Dataset Details evaluation_transcripts/handover_evals contains the transcripts provided by the red team to the blue team at the beginning of the main round of the game, showing… See the full description on the dataset page: https://huggingface.co/datasets/sandbagging-games/evaluation_logs.3 likes5.1k downloads9mo agoHugging Face