Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tamdd18 /CEH_question_answertextn<1K0 likes1.7k downloads2y agoHugging Face02Malikeh1375 /medical-question-answering-datasetstextquestion-answering1M<n<10M86 likes1.3k downloads6mo agoHugging Face03aisingapore /NLU-Question-Answeringgated SEA Question Answering SEA Question Answering evaluates a model's ability to predict a contiguous span of characters that answers the question about a given passage. It is sampled from TyDi QA-GoldP for Indonesian, IndicQA for Tamil, and XQuaD for Thai and Vietnamese. Supported Tasks and Leaderboards SEA Question Answering is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Question-Answering.texttext-generation1K<n<10K0 likes1.2k downloads9mo agoHugging Face04xwjzds /extractive_qa_question_answering_hr Dataset Card HR-Multiwoz is a fully-labeled dataset of 5980 extractive qa spanning 10 HR domains to evaluate LLM Agent. It is the first labeled open-sourced conversation dataset in the HR domain for NLP research. Please refer to HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent for details about the dataset construction. Dataset Sources Repository: xwjzds/extractive_qa_question_answering_hr Paper: HR-MultiWOZ: A Task Oriented Dialogue (TOD)… See the full description on the dataset page: https://huggingface.co/datasets/xwjzds/extractive_qa_question_answering_hr.text1K<n<10K12 likes892 downloads3y agoHugging Face05OpenFinAL /Financial_Question_Answeringtext1K<n<10K2 likes525 downloads11mo agoHugging Face06mariiazhiv /cybersecurity_full_question_answerstext1K<n<10K0 likes459 downloads11mo agoHugging Face07nreimers /reddit_question_best_answersQuestion & question body together with the best answers to that question from Reddit. The score for the question / answer is the upvote count (i.e. positive-negative upvotes). Only questions / answers that have these properties were extracted: min_score = 3 min_title_len = 20 min_body_len = 100 text1M<n<10M17 likes419 downloads4y agoHugging Face08PrimeIntellect /stackexchange-question-answering SYNTHETIC-1 This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here text100K<n<1M17 likes296 downloads2y agoHugging Face09petkopetkov /medical-question-answering-splittext100K<n<1M0 likes269 downloads2y agoHugging Face10addy88 /nq-question-answeronlytext100K<n<1M1 likes263 downloads5y agoHugging Face11Lots-of-LoRAs /task290_tellmewhy_question_answerability Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task290_tellmewhy_question_answerability Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task290_tellmewhy_question_answerability.texttext-generation1K<n<10K0 likes241 downloads2y agoHugging Face12AswiN037 /tamil-question-answering-datasetthis dataset contains 5 columns context, question, answer_start, answer_text, source Column Description context A general small paragraph in tamil language question question framed form the context answer_text text span that extracted from context answer_start index of answer_text source who framed this context, question, answer pair source team KBA => (Karthi, Balaji, Azeez) these people manually created CHAII =>a kaggle competition XQA => multilingual QA… See the full description on the dataset page: https://huggingface.co/datasets/AswiN037/tamil-question-answering-dataset.text1K<n<10K8 likes226 downloads4y agoHugging Face13AnonymousSub /MedQuAD_47441_Question_Answer_Pairs Dataset Card for "MedQuAD_47441_Question_Answer_Pairs" More Information needed text10K<n<100K13 likes221 downloads4y agoHugging Face14nezahatkorkmaz /Turkish-medical-visual-question-answering-LLaVa-dataset Türkçe Radyoloji Görüntüleme Veri Seti - data_RAD data_RAD veri seti, radyoloji görüntüleri üzerinde görsel soru-cevaplama (VQA) araştırmaları yapmak amacıyla Türkçeye çevrilmiş ve LLaVa mimarisiyle uyumlu hale getirilmiştir. Bu veri seti, tıbbi görüntü analizi ve yapay zeka destekli radyoloji uygulamalarını geliştirmek için kullanılabilir. Veri Seti İçeriği Toplam Görüntü Sayısı: 316 Veri Yapısı: DatasetDict({ train: Dataset({ features: ['image'], num_rows: 316 }) }) Özellikler:… See the full description on the dataset page: https://huggingface.co/datasets/nezahatkorkmaz/Turkish-medical-visual-question-answering-LLaVa-dataset.imagequestion-answeringn<1K13 likes177 downloads2y agoHugging Face15OneEyeDJ /Art-Vision-Question-Answering-Dataset Art Vision Question Answering Dataset 🎨 A curated dataset for training AI models on digital artwork analysis and visual question answering. Dataset Overview This dataset contains 577 question-answer pairs extracted from artwork conversations, designed for training multimodal AI models on art analysis tasks. ✨ Key Features 🖼️ Visual Thumbnails: Artwork images displayed directly in the dataset viewer 💬 Rich Q&A: Expert-level questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/OneEyeDJ/Art-Vision-Question-Answering-Dataset.imageimage-to-textn<1K2 likes173 downloads1y agoHugging Face16Gliscor /Kafka-Donusum-Question-Answer Dataset Card for "Kafka-Donusum-FineTuning" More Information needed 0 likes160 downloads1y agoHugging Face17nirantk /chaii-hindi-and-tamil-question-answeringtextquestion-answering1K<n<10K0 likes156 downloads3y agoHugging Face18kurehamnm /Chinese_Question_Answering_Datasettextquestion-answering1M<n<10M6 likes148 downloads2y agoHugging Face19RUCAIBox /Question-AnsweringThis is the question answering datasets collected by TextBox, including: SQuAD (squad) CoQA (coqa) Natural Questions (nq) TriviaQA (tqa) WebQuestions (webq) NarrativeQA (nqa) MS MARCO (marco) NewsQA (newsqa) HotpotQA (hotpotqa) MSQG (msqg) QuAC (quac). The detail and leaderboard of each dataset can be found in TextBox page. question-answering1 likes142 downloads4y agoHugging Face20shahrukh95 /OWASP-question-answer-datasettextn<1K0 likes139 downloads3y agoHugging Face21taesiri /video-game-question-answeringimage10K<n<100K3 likes136 downloads3y agoHugging Face22BoltMonkey /psychology-question-answerA JSON formatted dataset comprising 197,180 question and answer pairs covering a wide range of topics encountered in a Bachelor level psychology course. I have included a broad range of question types, topics, and answer styles. The dataset was created using personal notes and several LLMs (such as GPT4) and manually assessed for veracity and completeness of response. Despite this, the size of the dataset prohibits me from ensuring every single answer is 100% accurate and up-to-date. As such… See the full description on the dataset page: https://huggingface.co/datasets/BoltMonkey/psychology-question-answer.textquestion-answering100K<n<1M11 likes127 downloads2y agoHugging Face23Lots-of-LoRAs /task865_mawps_addsub_question_answering Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task865_mawps_addsub_question_answering Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task865_mawps_addsub_question_answering.texttext-generation1K<n<10K0 likes127 downloads2y agoHugging Face24open-source-metrics /visual-question-answering-checkpoint-downloadstabularn<1K7 likes126 downloads4y agoHugging Face25CrossNow /medical-question-answering-datasetstextquestion-answering1M<n<10M0 likes126 downloads5mo agoHugging Face26fawern /visual-question-answering-cocoimagen<1K11 likes125 downloads2y agoHugging Face27ZackZhu00 /CFQA_Chinese_Finance_Question_Answering Citation For the complete project, please check Here If you use CFQA in your research, experiments, benchmarks, or publications, please cite the accompanying paper: @inproceedings{zhu2026cfqa, title = {CFQA: A Chinese Financial Question Answering Benchmark From Corporate Annual Reports}, author = {Tianning Zhu and Mo Liu and Murathan Kurfali}, booktitle = {Proceedings of The 7th Financial Narrative Processing Workshop (FNP 2026)}, year = {2026}, address =… See the full description on the dataset page: https://huggingface.co/datasets/ZackZhu00/CFQA_Chinese_Finance_Question_Answering.textn<1K0 likes117 downloads2mo agoHugging Face28toughdata /quora-question-answer-datasetQuora Question Answer Dataset (Quora-QuAD) contains 56,402 question-answer pairs scraped from Quora. Usage: For instructions on fine-tuning a model (Flan-T5) with this dataset, please check out the article: https://www.toughdata.net/blog/post/finetune-flan-t5-question-answer-quora-dataset textquestion-answering10K<n<100K20 likes110 downloads3y agoHugging Face29nogyxo /question-answering-ukrainian-json-answerstext100K<n<1M8 likes105 downloads3y agoHugging Face30mou3az /Question-Answering-Generation-Choices The dataset is a merged compilation of QuAIL, RACE, and Cosmos QA datasets, having undergone preprocessing. textquestion-answering10K<n<100K10 likes105 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.