Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01genbio-ai /rna-downstream-tasks GB.RNA Benchmark Datasets mRNA related tasks Translation efficiency prediction from Chu et al.(2024) [1] 3 cell lines: Muscle, pc3, HEK input sequence: 5'UTR 10-fold cross-validation split mRNA expression level prediction from Chu et al.(2024) [1] 3 cell lines: Muscle, pc3, HEK input sequence: 5'UTR 10-fold cross-validation split Mean ribosome load prediction from Sample et al. (2019) [2] input sequence: 5'UTR ouput: mean ribosome load the original data… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks.tabular1M<n<10M0 likes1.8k downloads1mo agoHugging Face02tasksource /jigsaw_toxicitytabular100K<n<1M2 likes778 downloads3y agoHugging Face03tasksource /blog_authorship_corpustabular100K<n<1M2 likes391 downloads2y agoHugging Face04tasksource /social-chemestry-101tabular100K<n<1M4 likes354 downloads4y agoHugging Face05tasksource /simlextabularn<1K0 likes261 downloads3y agoHugging Face06tasksource /acceptability-prediction@inproceedings{lau-etal-2015-unsupervised, title = "Unsupervised Prediction of Acceptability Judgements", author = "Lau, Jey Han and Clark, Alexander and Lappin, Shalom", booktitle = "Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)", month = jul, year = "2015", address = "Beijing, China", publisher = "Association for… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/acceptability-prediction.tabulartext-classification1K<n<10K1 likes252 downloads4y agoHugging Face07sabir15 /osworld_tasks_filesdocumentn<1K0 likes240 downloads9mo agoHugging Face081-800-SHARED-TASKS /COLING-2025-CHIPSAL Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-CHIPSAL.tabular100K<n<1M2 likes131 downloads2y agoHugging Face09tasksource /winowhyhttps://github.com/HKUST-KnowComp/WinoWhy @inproceedings{zhang2020WinoWhy, author = {Hongming Zhang and Xinran Zhao and Yangqiu Song}, title = {WinoWhy: A Deep Diagnosis of Essential Commonsense Knowledge for Answering Winograd Schema Challenge}, booktitle = {Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL) 2020}, year = {2020} } tabular1K<n<10K2 likes114 downloads3y agoHugging Face10tasksource /nli-veridicality-transitivity@inproceedings{yanaka-etal-2021-exploring, title = "Exploring Transitivity in Neural {NLI} Models through Veridicality", author = "Yanaka, Hitomi and Mineshima, Koji and Inui, Kentaro", booktitle = "Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume", year = "2021", pages = "920--934", } tabulartext-classification100K<n<1M1 likes113 downloads4y agoHugging Face11tasksource /paradehttps://github.com/heyunh2015/PARADE_dataset @inproceedings{he-etal-2020-parade, title = "{PARADE}: {A} {N}ew {D}ataset for {P}araphrase {I}dentification {R}equiring {C}omputer {S}cience {D}omain {K}nowledge", author = "He, Yun and Wang, Zhuoer and Zhang, Yin and Huang, Ruihong and Caverlee, James", booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)", month = nov, year = "2020", address… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/parade.tabularsentence-similarity10K<n<100K0 likes103 downloads3y agoHugging Face12tasksource /english-gradinghttps://www.kaggle.com/competitions/feedback-prize-english-language-learning tabular1K<n<10K4 likes98 downloads3y agoHugging Face13tasksource /sts-companionhttps://ixa2.si.ehu.eus/stswiki/index.php/STSbenchmark The companion datasets to the STS Benchmark comprise the rest of the English datasets used in the STS tasks organized by us in the context of SemEval between 2012 and 2017. Authors collated two datasets, one with pairs of sentences related to machine translation evaluation. Another one with the rest of datasets, which can be used for domain adaptation studies. @inproceedings{cer-etal-2017-semeval, title = "{S}em{E}val-2017 Task 1:… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/sts-companion.tabularsentence-similarity1K<n<10K3 likes95 downloads4y agoHugging Face14tasksource /monotonicity-entailment@inproceedings{yanaka-etal-2019-neural, title = "Can Neural Networks Understand Monotonicity Reasoning?", author = "Yanaka, Hitomi and Mineshima, Koji and Bekki, Daisuke and Inui, Kentaro and Sekine, Satoshi and Abzianidze, Lasha and Bos, Johan", booktitle = "Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP", year = "2019", pages = "31--40", } tabular1K<n<10K0 likes95 downloads4y agoHugging Face15tasksource /offensive-humor@article{tang2022naughtyformer, title={The Naughtyformer: A Transformer Understands Offensive Humor}, author={Tang, Leonard and Cai, Alexander and Li, Steve and Wang, Jason}, journal={arXiv preprint arXiv:2211.14369}, year={2022} } tabular100K<n<1M9 likes93 downloads4y agoHugging Face16vitor-cirilo-santos /osworld_tasks_filestabularn<1K0 likes89 downloads1y agoHugging Face17tasksource /wouldyourathertabular1K<n<10K0 likes80 downloads4y agoHugging Face18ehyo /GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks GPUMemNet and GPUUtilNet Dataset This dataset accompanies the paper “GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations.” It contains synthetic deep learning training configurations and their measured GPU memory consumption and utilization characteristics. Dataset configurations The dataset is divided into separate configurations because MLP, CNN, and Transformer workloads use different feature schemas.… See the full description on the dataset page: https://huggingface.co/datasets/ehyo/GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks.tabulartabular-regression10K<n<100K0 likes47 downloads4mo agoHugging Face19lgy0404 /mobileforge-generated-tasks MobileForge Generated Tasks This dataset contains the consolidated task pool generated by MobileGym-Curriculum from target-app exploration trajectories. These tasks are used by MobileForge for rollout collection and annotation-free adaptation. Dataset summary File Rows Apps Size Description generated_tasks_26020301-all.csv 3,249 20 1.93 MB Consolidated AndroidWorld-side MobileForge task pool. The task pool is generated from real target-app… See the full description on the dataset page: https://huggingface.co/datasets/lgy0404/mobileforge-generated-tasks.tabulartext-generation1K<n<10K0 likes35 downloads4mo agoHugging Face20tasksource /semantic-feature-production-normstabular1K<n<10K0 likes33 downloads4y agoHugging Face21tasksource /context_toxicityhttps://github.com/ipavlopoulos/context_toxicity/ @inproceedings{xenos-etal-2021-context, title = "Context Sensitivity Estimation in Toxicity Detection", author = "Xenos, Alexandros and Pavlopoulos, John and Androutsopoulos, Ion", booktitle = "Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021)", month = aug, year = "2021", address = "Online", publisher = "Association for Computational Linguistics", url =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/context_toxicity.tabular1K<n<10K3 likes25 downloads3y agoHugging Face22jhonatan-ospina /osworld_tasks_filestabularn<1K0 likes23 downloads11mo agoHugging Face23Nagepalli /osworld_tasks_filestabularn<1K0 likes22 downloads11mo agoHugging Face24saweerawal /osworld_tasks_filestabularn<1K0 likes22 downloads11mo agoHugging Face25saaduuu /osworld_tasks_filestabularn<1K0 likes22 downloads11mo agoHugging Face26tasksource /rankme-nlg-acceptability@inproceedings{novikova-etal-2018-rankme, title = "RankME: Reliable Human Ratings for Natural Language Generation", author = "Novikova, Jekaterina and Duvsek, Ondvrej and Rieser, Verena", booktitle = "Proceedings of the NAACL2018", month = jun, year = "2018", address = "New Orleans, Louisiana", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/N18-2012", doi = "10.18653/v1/N18-2012", pages = "72--78"… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/rankme-nlg-acceptability.tabulartext-classification1K<n<10K0 likes21 downloads4y agoHugging Face27usamamukhtar328 /osworld_tasks_filestabularn<1K0 likes20 downloads11mo agoHugging Face28jeremiahnthiwa /osworld_tasks_filestabularn<1K0 likes19 downloads11mo agoHugging Face29zaidturing /osworld_tasks_filestabularn<1K0 likes18 downloads11mo agoHugging Face30biruklturing /osworld_tasks_filestabularn<1K0 likes18 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.