Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tasksource /mmluMMLU (hendrycks_test on huggingface) without auxiliary train. It is much lighter (7MB vs 162MB) and faster than the original implementation, in which auxiliary train is loaded (+ duplicated!) by default for all the configs in the original version, making it quite heavy. We use this version in tasksource. Reference to original dataset: Measuring Massive Multitask Language Understanding - https://github.com/hendrycks/test @article{hendryckstest2021, title={Measuring Massive Multitask Language… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/mmlu.texttext-classification10K<n<100K36 likes36k downloads1y agoHugging Face02tasksource /bigbenchBIG-Bench but it doesn't require the hellish dependencies (tensorflow, pypi-bigbench, protobuf) of the official version. dataset = load_dataset("tasksource/bigbench",'movie_recommendation') Code to reproduce: https://colab.research.google.com/drive/1MKdLdF7oqrSQCeavAcsEnPdI85kD0LzU?usp=sharing Datasets are capped to 50k examples to keep things light. I also removed the default split when train was available also to save space, as default=train+val. @article{srivastava2022beyond… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/bigbench.textmultiple-choice100K<n<1M69 likes15k downloads1y agoHugging Face03tasksource /reclorhttps://whyu.me/reclor/ @inproceedings{yu2020reclor, author = {Yu, Weihao and Jiang, Zihang and Dong, Yanfei and Feng, Jiashi}, title = {ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning}, booktitle = {International Conference on Learning Representations (ICLR)}, month = {April}, year = {2020} } text1K<n<10K18 likes11k downloads3y agoHugging Face04tasksource /proofwriter Dataset Card for "proofwriter" More Information needed tabular100K<n<1M12 likes11k downloads3y agoHugging Face05tasksource /strategy-qatext1K<n<10K8 likes6k downloads4y agoHugging Face06tasksource /tasksource-jev-typed-decisions tasksource-jev-typed-decisions 2.5 million typed decisions (choices, ratings and probabilities) from 670 sources. Why use it Real supervision. Labels, ratings, and annotator votes come from established datasets, not a teacher model. Every row names its source. Breadth. Over 300 dataset families: NLI and reasoning, QA and commonsense, sentiment, intent and topic, toxicity and safety, preference pairs, fact checking, entity tagging, and dozens of languages. GLUE… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions.imagezero-shot-classification1M<n<10M19 likes5.2k downloads2h agoHugging Face07tasksource /babi_nli bAbi_nli bAbI tasks recasted as natural language inference. https://github.com/facebookarchive/bAbI-tasks tasksource recasting code: https://colab.research.google.com/drive/1J_RqDSw9iPxJSBvCJu-VRbjXnrEjKVvr?usp=sharing @article{weston2015towards, title={Towards ai-complete question answering: A set of prerequisite toy tasks}, author={Weston, Jason and Bordes, Antoine and Chopra, Sumit and Rush, Alexander M and Van Merri{\"e}nboer, Bart and Joulin, Armand and Mikolov, Tomas}… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/babi_nli.texttext-classification10K<n<100K3 likes4.7k downloads2y agoHugging Face08tasksource /lsat-lr Dataset Card for "lsat-lr" More Information needed text1K<n<10K0 likes4.2k downloads3y agoHugging Face09tasksource /lsat-rc Dataset Card for "lsat-rc" More Information needed text1K<n<10K0 likes3.8k downloads3y agoHugging Face10tasksource /esci Dataset Card for "esci" ESCI product search dataset https://github.com/amazon-science/esci-data/ Preprocessings: -joined the two relevant files -product_text aggregate all product text -mapped esci_label to full name @article{reddy2022shopping, title={Shopping Queries Dataset: A Large-Scale {ESCI} Benchmark for Improving Product Search}, author={Chandan K. Reddy and Lluís Màrquez and Fran Valero and Nikhil Rao and Hugo Zaragoza and Sambaran Bandyopadhyay and Arnab Biswas and Anlu… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/esci.tabulartext-classification1M<n<10M9 likes3.6k downloads3y agoHugging Face11tasksource /foliohttps://github.com/Yale-LILY/FOLIO @article{han2022folio, title={FOLIO: Natural Language Reasoning with First-Order Logic}, author = {Han, Simeng and Schoelkopf, Hailey and Zhao, Yilun and Qi, Zhenting and Riddell, Martin and Benson, Luke and Sun, Lucy and Zubova, Ekaterina and Qiao, Yujie and Burtell, Matthew and Peng, David and Fan, Jonathan and Liu, Yixin and Wong, Brian and Sailor, Malcolm and Ni, Ansong and Nan, Linyong and Kasai, Jungo and Yu, Tao and Zhang, Rui and Joty, Shafiq and… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/folio.tabulartext-classification1K<n<10K19 likes2.9k downloads3y agoHugging Face12tasksource /lsat-ar Dataset Card for "lsat-ar" More Information needed text1K<n<10K2 likes2.8k downloads3y agoHugging Face13tasksource /tasksource-instruct tasksource-instruct Instruction-tuning data recast from the ~480 English classification, multiple-choice and token-classification tasks of tasksource. Every example comes from a human-built dataset (NLI, logical reasoning, sentiment, hate speech, discourse, argumentation, ...), not from a teacher model. Each task is capped at 30k training examples, so no task dominates. Many tasks aren't in FLAN v2, for example DynaSent, DynaHate, discriminative bAbI, epistemic logic, RuleTaker… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-instruct.texttext-generation1M<n<10M24 likes2.7k downloads11d agoHugging Face14tasksource /procedural-typed-decisions procedural-typed-decisions Procedurally generated decision problems. Each row is one structured state (JSON, or a table, CSV, key=value lines, or prose for the arithmetic, retrieval, and aggregation configs) with several typed questions over that same state, following the Jev / System One request shape: choice (pick one criterion), noul (a number in [0, 1]; a probability or a yes/no), and score (an ordered rubric). Every answer is computed exactly from the state by rules that… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/procedural-typed-decisions.tabulartext-classification100K<n<1M4 likes1.6k downloads9d agoHugging Face15tasksource /ScienceQA_text_only Dataset Card for "scienceQA_text_only" ScienceQA text-only examples (examples where no image was initially present, which means they should be doable with text-only models.) @article{10.1007/s00799-022-00329-y, author = {Saikh, Tanik and Ghosal, Tirthankar and Mittal, Amish and Ekbal, Asif and Bhattacharyya, Pushpak}, title = {ScienceQA: A Novel Resource for Question Answering on Scholarly Articles}, year = {2022}, journal = {Int. J. Digit. Libr.}, month = {sep} } text10K<n<100K32 likes1.5k downloads3y agoHugging Face16tasksource /ruletaker Dataset Card for "ruletaker" https://github.com/allenai/ruletaker @inproceedings{ruletaker2020, title = {Transformers as Soft Reasoners over Language}, author = {Clark, Peter and Tafjord, Oyvind and Richardson, Kyle}, booktitle = {Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, {IJCAI-20}}, publisher = {International Joint Conferences on Artificial Intelligence Organization}, editor = {Christian Bessiere}… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/ruletaker.text100K<n<1M8 likes1.5k downloads3y agoHugging Face17tasksource /defeasible-nlihttps://github.com/rudinger/defeasible-nli @inproceedings{rudinger2020thinking, title={Thinking like a skeptic: feasible inference in natural language}, author={Rudinger, Rachel and Shwartz, Vered and Hwang, Jena D and Bhagavatula, Chandra and Forbes, Maxwell and Le Bras, Ronan and Smith, Noah A and Choi, Yejin}, booktitle={Findings of the Association for Computational Linguistics: EMNLP 2020}, pages={4661--4675}, year={2020} } texttext-classification100K<n<1M2 likes1.1k downloads2y agoHugging Face18tasksource /PRM800Khttps://github.com/openai/prm800k/tree/main 41 likes1.1k downloads3y agoHugging Face19tasksource /chaos-mnli-ambiguity chaos-mnli-ambiguity ChaosNLI, MNLI portion: 1,599 MNLI pairs relabeled by 100 annotators each (Nie et al., 2020). label_dist and label_count follow the entailment/neutral/contradiction order, and gini is the Gini coefficient of label_dist (0 = annotators evenly split, 1 = unanimous). Built from the jsonl first uploaded here, which flattens the ChaosNLI release (https://github.com/easonnie/ChaosNLI) and adds gini; the variable-key label_counter (a duplicate of label_count) is… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/chaos-mnli-ambiguity.tabular1K<n<10K0 likes1.1k downloads14d agoHugging Face20tasksource /logiqa-2.0-nlihttps://github.com/csitfun/LogiQA2.0 Temporary citation: @article{liu2020logiqa, title={Logiqa: A challenge dataset for machine reading comprehension with logical reasoning}, author={Liu, Jian and Cui, Leyang and Liu, Hanmeng and Huang, Dandan and Wang, Yile and Zhang, Yue}, journal={arXiv preprint arXiv:2007.08124}, year={2020} } text10K<n<100K5 likes926 downloads3y agoHugging Face21tasksource /Boardgame-QAhttps://arxiv.org/pdf/2306.07934.pdf text10K<n<100K8 likes804 downloads3y agoHugging Face22tasksource /commonsense_qa_2.0https://github.com/allenai/csqa2 @article{talmor2022commonsenseqa, title={CommonsenseQA 2.0: Exposing the limits of AI through gamification}, author={Talmor, Alon and Yoran, Ori and Bras, Ronan Le and Bhagavatula, Chandra and Goldberg, Yoav and Choi, Yejin and Berant, Jonathan}, journal={arXiv preprint arXiv:2201.05320}, year={2022} } textquestion-answering10K<n<100K4 likes793 downloads3y agoHugging Face23tasksource /FOL-nli Dataset Card for "FOL-nli" https://github.com/sileod/unigram/ https://arxiv.org/abs/2406.11035 Citation: @article{sileo2024scaling, title={Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars}, author={Sileo, Damien}, journal={arXiv preprint arXiv:2406.11035}, year={2024} } texttext-classification100K<n<1M3 likes739 downloads9mo agoHugging Face24tasksource /jigsaw_toxicitytabular100K<n<1M2 likes733 downloads3y agoHugging Face25tasksource /doc-nli Dataset Card for "doc-nli" https://github.com/salesforce/DocNLI/tree/main @inproceedings{yin-etal-2021-docnli, title = "{D}oc{NLI}: A Large-scale Dataset for Document-level Natural Language Inference", author = "Yin, Wenpeng and Radev, Dragomir and Xiong, Caiming", editor = "Zong, Chengqing and Xia, Fei and Li, Wenjie and Navigli, Roberto", booktitle = "Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021"… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/doc-nli.texttext-classification1M<n<10M1 likes668 downloads2y agoHugging Face26tasksource /ecqa Dataset Card for "ecqa" https://github.com/dair-iitd/ECQA-Dataset @inproceedings{aggarwaletal2021ecqa, title={{E}xplanations for {C}ommonsense{QA}: {N}ew {D}ataset and {M}odels}, author={Shourya Aggarwal and Divyanshu Mandowara and Vishwajeet Agrawal and Dinesh Khandelwal and Parag Singla and Dinesh Garg}, booktitle="Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/ecqa.textquestion-answering10K<n<100K0 likes611 downloads3y agoHugging Face27tasksource /dolci-instructtext1M<n<10M0 likes595 downloads10mo agoHugging Face28tasksource /scruplestext10K<n<100K1 likes529 downloads4y agoHugging Face29tasksource /planbench Dataset Card for "planbench" https://arxiv.org/abs/2206.10498 @article{valmeekam2024planbench, title={Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change}, author={Valmeekam, Karthik and Marquez, Matthew and Olmo, Alberto and Sreedharan, Sarath and Kambhampati, Subbarao}, journal={Advances in Neural Information Processing Systems}, volume={36}, year={2024} } tabular10K<n<100K12 likes509 downloads2y agoHugging Face30tasksource /prontoqahttps://github.com/asaparov/prontoqa/ @article{saparov2022language, title={Language models are greedy reasoners: A systematic formal analysis of chain-of-thought}, author={Saparov, Abulhair and He, He}, journal={arXiv preprint arXiv:2210.01240}, year={2022} } question-answering1 likes480 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.