datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rna-downstream-tasks
GB.RNA Benchmark Datasets
mRNA related tasks
Translation efficiency prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
mRNA expression level prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
Mean ribosome load prediction from Sample et al. (2019) [2]
input sequence: 5'UTR
ouput: mean ribosome load
the original data… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks.ELYZA-tasks-100
ELYZA-tasks-100: 日本語instructionモデル評価データセット
Data Description
本データセットはinstruction-tuningを行ったモデルの評価用データセットです。詳細は リリースのnote記事 を参照してください。
特徴:
複雑な指示・タスクを含む100件の日本語データです。
役に立つAIアシスタントとして、丁寧な出力が求められます。
全てのデータに対して評価観点がアノテーションされており、評価の揺らぎを抑えることが期待されます。
具体的には以下のようなタスクを含みます。
要約を修正し、修正箇所を説明するタスク
具体的なエピソードから抽象的な教訓を述べるタスク
ユーザーの意図を汲み役に立つAIアシスタントとして振る舞うタスク
場合分けを必要とする複雑な算数のタスク
未知の言語からパターンを抽出し日本語訳する高度な推論を必要とするタスク
複数の指示を踏まえた上でyoutubeの対話を生成するタスク
架空の生き物や熟語に関する生成・大喜利などの想像力が求められるタスク… See the full description on the dataset page: https://huggingface.co/datasets/elyza/ELYZA-tasks-100.jigsaw_toxicitygeneb-tasks
GENEB — Genomic Embedding Benchmark (task data)
Task-level sequence classification data for GENEB, a multi-task benchmark for DNA sequence encoders introduced in the paper: GENEB: Why Genomic Models Are Hard to Compare.
Paper: https://huggingface.co/papers/2606.04525
Source code: GitHub - darlednik/GENEB
Leaderboard: Hugging Face Space
GENEB evaluates frozen representations from 40 genomic foundation models across 100 tasks in 13 functional categories using a unified… See the full description on the dataset page: https://huggingface.co/datasets/darlednik/geneb-tasks.terminal-tasksblog_authorship_corpussocial-chemestry-101arct
The Argument Reasoning Comprehension Task: Identification and Reconstruction of Implicit Warrants
https://github.com/UKPLab/argument-reasoning-comprehension-task
@InProceedings{Habernal.et.al.2018.NAACL.ARCT,
title = {The Argument Reasoning Comprehension Task: Identification
and Reconstruction of Implicit Warrants},
author = {Habernal, Ivan and Wachsmuth, Henning and
Gurevych, Iryna and Stein, Benno},
publisher = {Association for… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/arct.simlexcounterfactually-augmented-imdb@article{kaushik2020learning,
title={Learning the Difference that Makes a Difference with Counterfactually Augmented Data},
author={Kaushik, Divyansh and Hovy, Eduard and Lipton, Zachary C},
journal={International Conference on Learning Representations (ICLR)},
year={2020}
}
acceptability-prediction@inproceedings{lau-etal-2015-unsupervised,
title = "Unsupervised Prediction of Acceptability Judgements",
author = "Lau, Jey Han and
Clark, Alexander and
Lappin, Shalom",
booktitle = "Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)",
month = jul,
year = "2015",
address = "Beijing, China",
publisher = "Association for… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/acceptability-prediction.traciehttps://github.com/allenai/aristo-leaderboard/tree/master/tracie/data
@inproceedings{ZRNKSR21,
author = {Ben Zhou and Kyle Richardson and Qiang Ning and Tushar Khot and Ashish Sabharwal and Dan Roth},
title = {Temporal Reasoning on Implicit Events from Distant Supervision},
booktitle = {NAACL},
year = {2021},
}
osworld_tasks_filestomi-nlitomi dataset (theory of mind question answering) recasted as natural language inference
https://colab.research.google.com/drive/1J_RqDSw9iPxJSBvCJu-VRbjXnrEjKVvr?usp=sharing
@article{sileo2023tasksource,
title={tasksource: Structured Dataset Preprocessing Annotations for Frictionless Extreme Multi-Task Learning and Evaluation},
author={Sileo, Damien},
url= {https://arxiv.org/abs/2301.05948},
journal={arXiv preprint arXiv:2301.05948},
year={2023}
}… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tomi-nli.implicit-hate-stg1https://github.com/SALT-NLP/implicit-hate
@inproceedings{elsherief-etal-2021-latent,
title = "Latent Hatred: A Benchmark for Understanding Implicit Hate Speech",
author = "ElSherief, Mai and
Ziems, Caleb and
Muchlinski, David and
Anupindi, Vaishnavi and
Seybolt, Jordyn and
De Choudhury, Munmun and
Yang, Diyi",
booktitle = "Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing",
month = nov,
year =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/implicit-hate-stg1.AES2-essay-scoringhttps://www.kaggle.com/competitions/learning-agency-lab-automated-essay-scoring-2/data
temporal-nli@inproceedings{thukral-etal-2021-probing,
title = "Probing Language Models for Understanding of Temporal Expressions",
author = "Thukral, Shivin and
Kukreja, Kunal and
Kavouras, Christian",
booktitle = "Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP",
month = nov,
year = "2021",
address = "Punta Cana, Dominican Republic",
publisher = "Association for Computational Linguistics",
url =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/temporal-nli.implicaturesImplicature corpus
@article{george2020conversational,
title={Conversational implicatures in English dialogue: Annotated dataset},
author={George, Elizabeth Jasmi and Mamidi, Radhika},
journal={Procedia Computer Science},
volume={171},
pages={2316--2323},
year={2020},
publisher={Elsevier}
}
Augmented with generated distractors https://colab.research.google.com/drive/1ix0FgwzPAjQkIQA2E3ctlylvcmya7vGy?usp=sharing, for tasksource
@article{sileo2023tasksource,
title={tasksource:… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/implicatures.counterfactually-augmented-snli@article{kaushik2020learning,
title={Learning the Difference that Makes a Difference with Counterfactually Augmented Data},
author={Kaushik, Divyansh and Hovy, Eduard and Lipton, Zachary C},
journal={International Conference on Learning Representations (ICLR)},
year={2020}
}
COLING-2025-CHIPSAL
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/COLING-2025-CHIPSAL.subjectivity@misc{antici2023corpus,
title={A Corpus for Sentence-level Subjectivity Detection on English News Articles},
author={Francesco Antici and Andrea Galassi and Federico Ruggeri and Katerina Korre and Arianna Muti and Alessandra Bardi and Alice Fedotova and Alberto Barrón-Cedeño},
year={2023},
eprint={2305.18034},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
datasheet:… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/subjectivity.I2D2code:
https://i2d2.allen.ai/
https://arxiv.org/abs/2212.09246
@inproceedings{Bhagavatula2022GenGen,
title={Generating Generics: Knowledge Induction with NeuroLogic and Self-Imitation},
author={Chandra Bhagavatula, Jena D. Hwang, Doug Downey, Ronan Le Bras, Ximing Lu, Lianhui Qin, Keisuke Sakaguchi, Swabha Swayamdipta, Peter West, Yejin Choi},
booktitle={arXiv},
year={2022}
}
help-nlihttps://github.com/verypluming/HELP
@InProceedings{yanaka-EtAl:2019:starsem,
author = {Yanaka, Hitomi and Mineshima, Koji and Bekki, Daisuke and Inui, Kentaro and Sekine, Satoshi and Abzianidze, Lasha and Bos, Johan},
title = {HELP: A Dataset for Identifying Shortcomings of Neural Models in Monotonicity Reasoning},
booktitle = {Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM2019)},
year = {2019},
}
syntactic-augmentation-nlihttps://github.com/Aatlantise/syntactic-augmentation-nli/tree/master/datasets
@inproceedings{min-etal-2020-syntactic,
title = "Syntactic Data Augmentation Increases Robustness to Inference Heuristics",
author = "Min, Junghyun and
McCoy, R. Thomas and
Das, Dipanjan and
Pitler, Emily and
Linzen, Tal",
booktitle = "Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics",
month = jul,
year = "2020",
address =… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/syntactic-augmentation-nli.clcd-english@article{salvatore2019logical,
title={A logical-based corpus for cross-lingual evaluation},
author={Salvatore, Felipe and Finger, Marcelo and Hirata Jr, Roberto},
journal={arXiv preprint arXiv:1905.05704},
year={2019}
}
winowhyhttps://github.com/HKUST-KnowComp/WinoWhy
@inproceedings{zhang2020WinoWhy,
author = {Hongming Zhang and Xinran Zhao and Yangqiu Song},
title = {WinoWhy: A Deep Diagnosis of Essential Commonsense Knowledge for Answering Winograd Schema Challenge},
booktitle = {Proceedings of Annual Meeting of the Association for Computational Linguistics (ACL) 2020},
year = {2020}
}
nli-veridicality-transitivity@inproceedings{yanaka-etal-2021-exploring,
title = "Exploring Transitivity in Neural {NLI} Models through Veridicality",
author = "Yanaka, Hitomi and
Mineshima, Koji and
Inui, Kentaro",
booktitle = "Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume",
year = "2021",
pages = "920--934",
}
SDOH-NLISDOH-NLI is a natural language inference dataset containing ~30k premise-hypothesis pairs with binary entailment labels in the domain of social and behavioral determinants of health.
@misc{lelkes2023sdohnli,
title={SDOH-NLI: a Dataset for Inferring Social Determinants of Health from Clinical Notes},
author={Adam D. Lelkes and Eric Loreaux and Tal Schuster and Ming-Jun Chen and Alvin Rajkomar},
year={2023},
eprint={2310.18431},
archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/SDOH-NLI.paradehttps://github.com/heyunh2015/PARADE_dataset
@inproceedings{he-etal-2020-parade,
title = "{PARADE}: {A} {N}ew {D}ataset for {P}araphrase {I}dentification {R}equiring {C}omputer {S}cience {D}omain {K}nowledge",
author = "He, Yun and
Wang, Zhuoer and
Zhang, Yin and
Huang, Ruihong and
Caverlee, James",
booktitle = "Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)",
month = nov,
year = "2020",
address… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/parade.english-gradinghttps://www.kaggle.com/competitions/feedback-prize-english-language-learning
