Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nreimers /reddit_question_best_answersQuestion & question body together with the best answers to that question from Reddit. The score for the question / answer is the upvote count (i.e. positive-negative upvotes). Only questions / answers that have these properties were extracted: min_score = 3 min_title_len = 20 min_body_len = 100 text1M<n<10M17 likes377 downloads4y agoHugging Face02mariiazhiv /cybersecurity_full_question_answerstext1K<n<10K0 likes287 downloads1y agoHugging Face03nirantk /chaii-hindi-and-tamil-question-answeringtextquestion-answering1K<n<10K0 likes142 downloads3y agoHugging Face04ZackZhu00 /CFQA_Chinese_Finance_Question_Answering Citation For the complete project, please check Here If you use CFQA in your research, experiments, benchmarks, or publications, please cite the accompanying paper: @inproceedings{zhu2026cfqa, title = {CFQA: A Chinese Financial Question Answering Benchmark From Corporate Annual Reports}, author = {Tianning Zhu and Mo Liu and Murathan Kurfali}, booktitle = {Proceedings of The 7th Financial Narrative Processing Workshop (FNP 2026)}, year = {2026}, address =… See the full description on the dataset page: https://huggingface.co/datasets/ZackZhu00/CFQA_Chinese_Finance_Question_Answering.textn<1K0 likes130 downloads2mo agoHugging Face05BoltMonkey /psychology-question-answerA JSON formatted dataset comprising 197,180 question and answer pairs covering a wide range of topics encountered in a Bachelor level psychology course. I have included a broad range of question types, topics, and answer styles. The dataset was created using personal notes and several LLMs (such as GPT4) and manually assessed for veracity and completeness of response. Despite this, the size of the dataset prohibits me from ensuring every single answer is 100% accurate and up-to-date. As such… See the full description on the dataset page: https://huggingface.co/datasets/BoltMonkey/psychology-question-answer.textquestion-answering100K<n<1M11 likes121 downloads2y agoHugging Face06LangChainDatasets /question-answering-paul-grahamtextn<1K9 likes105 downloads4y agoHugging Face07LangChainDatasets /question-answering-state-of-the-uniontextn<1K6 likes101 downloads4y agoHugging Face08toughdata /quora-question-answer-datasetQuora Question Answer Dataset (Quora-QuAD) contains 56,402 question-answer pairs scraped from Quora. Usage: For instructions on fine-tuning a model (Flan-T5) with this dataset, please check out the article: https://www.toughdata.net/blog/post/finetune-flan-t5-question-answer-quora-dataset textquestion-answering10K<n<100K20 likes101 downloads3y agoHugging Face09naklecha /minecraft-question-answer-700k minecraft-question-answer-700k Introducing the largest synthetic Minecraft Q&A dataset, covering every topic, game mechanic, item and craft in Minecraft. The dataset was generated by extracting over 18,000 Minecraft wiki pages, and using glaive.ai's synthetic data generation pipeline. about the dataset rows - 694,814 tokens - 47,133,624 source - https://minecraft.wiki/ Hit me up on twitter if you see a bug or need a synthetic dataset for your company:… See the full description on the dataset page: https://huggingface.co/datasets/naklecha/minecraft-question-answer-700k.textquestion-answering100K<n<1M46 likes86 downloads2y agoHugging Face10LLaMAX /BenchMAX_Question_Answering Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Question_Answering is a dataset of BenchMAX for evaluating the long-context capability of LLMs in multilingual scenarios. The subtasks are similar to the subtasks in RULER. The data is sourcing from UN Parallel Corpus and xquad. The haystacks… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Question_Answering.texttext-generationn<1K0 likes86 downloads2y agoHugging Face11nreimers /reddit_question_best_answers_langstext1M<n<10M1 likes83 downloads4y agoHugging Face12Aixr /Math-Question-Answertexttext-generation1K<n<10K3 likes82 downloads2y agoHugging Face13MrBananaHuman /kor_ethical_question_answertext10K<n<100K10 likes64 downloads3y agoHugging Face14sabin1234 /Dengue_Surveillance_Data_Question_Answering_Dataset Nepali ShareGPT Clean Final Dataset Comprehensive Documentation & Analysis Report Dengue Surveillance Data - Question Answering Dataset 📋 Dataset Overview This dataset is a curated collection of 256 question-answer pairs focused on Dengue Surveillance in Nepal. It contains data-grounded questions in Nepali language paired with factual, statistical answers sourced from the Department of Health Services (DoHS), Nepal. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Dengue_Surveillance_Data_Question_Answering_Dataset.textn<1K0 likes46 downloads1mo agoHugging Face15kurumikz /Question-Answering_Kazakh 🇰🇿 Question-Answering_Kazakh A comprehensive Kazakh-language question-answer dataset for fine-tuning and training language models.Created and maintained by Kurumikz. Free to use with attribution. 📌 Overview Question-Answering_Kazakh is an open-domain QA dataset written entirely in the Kazakh language (kk). It covers a wide range of topics — from the history and geography of Kazakhstan to Kazakh grammar, culture, economy, and language learning (Kazakh ↔ English).… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-Answering_Kazakh.textquestion-answering1K<n<10K1 likes43 downloads6mo agoHugging Face16Mwnthai /bodo-legal-question-answering-ai4bharat Bodo Legal Question Answering Dataset Overview This dataset is a Bodo-language legal Question Answering (QA) resource created for research in low-resource Natural Language Processing (NLP) and legal language processing. The supplied source files contain legal judgment contexts together with multiple questions and answers. For Hugging Face compatibility and question-answering model training, each question-answer pair has been flattened into a separate JSONL example… See the full description on the dataset page: https://huggingface.co/datasets/Mwnthai/bodo-legal-question-answering-ai4bharat.textquestion-answering10K<n<100K0 likes36 downloads1mo agoHugging Face17Bytte-AI /Pidgin_Question-English_Answer_Dataset Pidgin Question - English Answer Dataset (Sample) Data Card v1.0 Dataset Name: Pidgin Question - English Answer Dataset (Sample)Dataset Type: Sample DatasetVersion: 1.0Release Date: 2026Organization: Bytte AILicense: CC-BY-4.0Contact: contact@bytteai.xyzWebsite: https://www.bytte.xyz/ Note: This is a sample dataset containing 331 cross-lingual question-answer pairs (Pidgin questions → English answers). Generated through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin_Question-English_Answer_Dataset.texttext-classificationn<1K0 likes34 downloads8mo agoHugging Face18sabin1234 /Nepali_Finance_QA_Romanized_Question_Answer_Pairs Nepali Finance QA — 3,000 Romanized Question–Answer Pairs A single-turn question-answering dataset about personal finance, banking, insurance and capital-market investing, written in romanized Nepali (Nepali in Latin letters). Every row is one question and one answer, and every row sits at a unique point of a fixed grid: 30 subdomains × 10 topics × 10 answer aspects = 3,000 rows. This README documents everything about the dataset: its structure, every field, the full taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepali_Finance_QA_Romanized_Question_Answer_Pairs.text1K<n<10K0 likes34 downloads10d agoHugging Face19restack /conversational-question-answer-wikipedia-v1.0 Dataset Information A dataset containing questions and conversational answers, based on sections of Wikipedia articles from wikipedia-en-chunks. This is a synthetic dataset created with the help of gemini-2.0-flash-001. Dataset Structure The dataset consists of a single JSON file with the following structure: [ { "messages": [ { "role": "system", "content": "You are a helpful assistant. You answer questions in a… See the full description on the dataset page: https://huggingface.co/datasets/restack/conversational-question-answer-wikipedia-v1.0.text10K<n<100K1 likes30 downloads1y agoHugging Face20kurumikz /Question-answeringsmall-ru Dataset Card for Question Answering Russian Dataset 🧠 Quick Summary Небольшой, чистый и тестовый датасет, созданный энтузиастом.Содержит базовые фундаментальные знания по математике, странам и тюркским народам.Подходит для обучения и тестирования моделей в образовательных и исследовательских целях. 📚 Dataset Details Curated by: @kurumikz Language(s): Russian (ru) License: CC-BY 4.0 — свободное использование с обязательным указанием автора Size Category:… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-answeringsmall-ru.textquestion-answering1K<n<10K1 likes30 downloads1y agoHugging Face21mkly /crypto-sales-question-answersA dataset consisting of questions, answers, and cryptocurrency descriptions textquestion-answeringn<1K3 likes29 downloads3y agoHugging Face22Mwnthai /bodo-legal-question-answering-iiith Bodo Legal Question Answering Dataset — IIITH Translation Overview A Bodo-language legal Question Answering (QA) resource derived from English legal judgments. Each example contains a judgment context, a question, and its corresponding answer. Data Provenance Original Legal Source The underlying English legal judgments were extracted from the publicly accessible Gauhati High Court judgment repository:… See the full description on the dataset page: https://huggingface.co/datasets/Mwnthai/bodo-legal-question-answering-iiith.textquestion-answering10K<n<100K0 likes29 downloads1mo agoHugging Face23Khyatimirani /pcos-patient-assist-question-and-answer Dataset Card for PCOS Patient Assist Question and Answer Dataset Dataset Details Dataset Description The PCOS Patient Assist Question and Answer Dataset is a curated dataset of question–answer pairs designed to represent common questions asked by patients diagnosed with or concerned about Polycystic Ovary Syndrome (PCOS). The dataset is structured to simulate real patient queries that arise during different stages of the PCOS journey, including diagnosis… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos-patient-assist-question-and-answer.textquestion-answeringn<1K0 likes27 downloads7mo agoHugging Face24sk75 /Turkish_Cyber_Security_CVE_Question_AnswerBu veri seti Türkçe siber güvenlik verisi oluşması adına tarafımca üretilmiştir. Son 5 yıldaki (2021-2026) CVE zaafiyetleri araştırılıp açıklamaları bulunup Ollama üzerinden Gpt-Oss:120b modeli ile Türkçeleştirilip soru-cevap haline getirilmiştir. Veri CVE adı kullanıcı mesajı modelin thinking süreci ve model cevabını içermektedir. text1K<n<10K0 likes27 downloads3mo agoHugging Face25Khyatimirani /pcos_question_answer_hindi PCOS Hindi Lifestyle & Clinical Q&A Dataset Dataset Details Dataset Description This dataset contains patient-facing conversational question–answer pairs in Hindi (Devanagari script) focused on Polycystic Ovary Syndrome (PCOS/PCOD). The dataset is designed to support training and evaluation of healthcare conversational AI systems that provide lifestyle and general clinical guidance for women diagnosed with PCOS. All conversations are structured in a chat format… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos_question_answer_hindi.textquestion-answeringn<1K0 likes25 downloads8mo agoHugging Face26kimleang123 /khmer_question_answergatedThe data collected from https://www.khsearch.com/ related to the general question-answering examination. It used to train fine-tuned models from many LLMs, including LlaMa, Qwen, Mistral, and Gemma. Under the research title "Fine-tuning for Question Answering in Low-Resource Languages: A Case Study on Khmer" conducted at ViLa Lab, Institute of Technology of Cambodia, Phnom Penh. Lab Info: https://www.facebook.com/vilalabitc Paper:… See the full description on the dataset page: https://huggingface.co/datasets/kimleang123/khmer_question_answer.textquestion-answering10K<n<100K4 likes24 downloads1y agoHugging Face27azizmatin /question_answering Dataset Information This Question Answering dataset is a reading comprehension resource derived from Persian Wikipedia. This crowd-sourced dataset contains over 9,000 entries, each of which can either be an unanswerable question or a question with one or more answers based on the provided context. Similar to the SQuAD2.0 dataset, the inclusion of unanswerable questions allows for the development of systems that "know they don't know the answer." Additionally, the dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/azizmatin/question_answering.textquestion-answeringn<1K0 likes22 downloads2y agoHugging Face28RomainPct /steve-jobs-question-and-answerstexttext-generationn<1K0 likes22 downloads2y agoHugging Face29Saleh11623 /questionanswering-datasettextquestion-answeringn<1K0 likes21 downloads2y agoHugging Face30kth8 /1B-question-answerLlama-3.2-1B-Instruct and gemma-3-1b-it responses to MuskumPillerum/General-Knowledge dataset. text10K<n<100K0 likes21 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.