datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NLU-Question-Answering
SEA Question Answering
SEA Question Answering evaluates a model's ability to predict a contiguous span of characters that answers the question about a given passage. It is sampled from TyDi QA-GoldP for Indonesian, IndicQA for Tamil, and XQuaD for Thai and Vietnamese.
Supported Tasks and Leaderboards
SEA Question Answering is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Question-Answering.task290_tellmewhy_question_answerability
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task290_tellmewhy_question_answerability
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task290_tellmewhy_question_answerability.psychology-question-answerA JSON formatted dataset comprising 197,180 question and answer pairs covering a wide range of topics encountered in a Bachelor level psychology course. I have included a broad range of question types, topics, and answer styles.
The dataset was created using personal notes and several LLMs (such as GPT4) and manually assessed for veracity and completeness of response. Despite this, the size of the dataset prohibits me from ensuring every single answer is 100% accurate and up-to-date. As such… See the full description on the dataset page: https://huggingface.co/datasets/BoltMonkey/psychology-question-answer.task865_mawps_addsub_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task865_mawps_addsub_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task865_mawps_addsub_question_answering.Question-Answering-Generation-Choices
The dataset is a merged compilation of QuAIL, RACE, and Cosmos QA datasets,
having undergone preprocessing.
task867_mawps_multiop_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task867_mawps_multiop_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task867_mawps_multiop_question_answering.task1731_quartz_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1731_quartz_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1731_quartz_question_answering.minecraft-question-answer-700k
minecraft-question-answer-700k
Introducing the largest synthetic Minecraft Q&A dataset, covering every topic, game mechanic, item and craft in Minecraft. The dataset was generated by extracting over 18,000 Minecraft wiki pages, and using glaive.ai's synthetic data generation pipeline.
about the dataset
rows - 694,814
tokens - 47,133,624
source - https://minecraft.wiki/
Hit me up on twitter if you see a bug or need a synthetic dataset for your company:… See the full description on the dataset page: https://huggingface.co/datasets/naklecha/minecraft-question-answer-700k.BenchMAX_Question_Answering
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Question_Answering is a dataset of BenchMAX for evaluating the long-context capability of LLMs in multilingual scenarios.
The subtasks are similar to the subtasks in RULER.
The data is sourcing from UN Parallel Corpus and xquad.
The haystacks… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Question_Answering.task861_asdiv_addsub_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task861_asdiv_addsub_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task861_asdiv_addsub_question_answering.Math-Question-Answertask144_subjqa_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task144_subjqa_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task144_subjqa_question_answering.task751_svamp_subtraction_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task751_svamp_subtraction_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task751_svamp_subtraction_question_answering.task1326_qa_zre_question_generation_from_answer
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1326_qa_zre_question_generation_from_answer
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1326_qa_zre_question_generation_from_answer.task868_mawps_singleop_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task868_mawps_singleop_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task868_mawps_singleop_question_answering.Dermatology-Question-Answer-Dataset-For-Fine-Tuning
Dataset Details
The data set has about 1 Million Tokens for Training and about 1500 question answers.
Dataset Description
This dataset is a comprehensive compilation of questions related to dermatology, spanning inquiries about various skin diseases, their symptoms, recommended medications, and available treatment modalities. Each question is paired with a concise and informative response, making it an ideal resource for training and fine-tuning language models in the… See the full description on the dataset page: https://huggingface.co/datasets/Mreeb/Dermatology-Question-Answer-Dataset-For-Fine-Tuning.obekt-question-answer-reasoning-micro-v0.1
Obekt Micro Reasoning Dataset (v0.1)
Dataset Description
This is a "micro" dataset containing questions, answers, and reasoning traces. It is generated using the Xiaomi MiMo V2 Flash LLM and is intended for experimental purposes, quick prototyping, and fine-tuning trials where reasoning capability is a focus.
Source Model: xiaomi/mimo-v2-flash
Contains
obekt-question-answer-reasoning-micro-v0.1.csv: The main data file.
Columns:
question: The input query.… See the full description on the dataset page: https://huggingface.co/datasets/obekt/obekt-question-answer-reasoning-micro-v0.1.task753_svamp_addition_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task753_svamp_addition_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task753_svamp_addition_question_answering.task754_svamp_common-division_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task754_svamp_common-division_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task754_svamp_common-division_question_answering.civil-human-rights-question-answering
Dataset Card for rag-prompt
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/rag-prompt/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/civil-human-rights-question-answering.arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.task178_quartz_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task178_quartz_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task178_quartz_question_answering.task864_asdiv_singleop_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task864_asdiv_singleop_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task864_asdiv_singleop_question_answering.task1286_openbookqa_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1286_openbookqa_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1286_openbookqa_question_answering.question-answer-Subject-Finance-InstructQuestion-Answering_Kazakh
🇰🇿 Question-Answering_Kazakh
A comprehensive Kazakh-language question-answer dataset for fine-tuning
and training language models.Created and maintained by Kurumikz. Free to use with attribution.
📌 Overview
Question-Answering_Kazakh is an open-domain QA dataset written entirely
in the Kazakh language (kk). It covers a wide range of topics — from the
history and geography of Kazakhstan to Kazakh grammar, culture, economy, and
language learning (Kazakh ↔ English).… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-Answering_Kazakh.task1594_yahoo_answers_topics_question_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1594_yahoo_answers_topics_question_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1594_yahoo_answers_topics_question_generation.task866_mawps_multidiv_question_answering
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task866_mawps_multidiv_question_answering
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task866_mawps_multidiv_question_answering.Pidgin_Question-English_Answer_Dataset
Pidgin Question - English Answer Dataset (Sample)
Data Card v1.0
Dataset Name: Pidgin Question - English Answer Dataset (Sample)Dataset Type: Sample DatasetVersion: 1.0Release Date: 2026Organization: Bytte AILicense: CC-BY-4.0Contact: contact@bytteai.xyzWebsite: https://www.bytte.xyz/
Note: This is a sample dataset containing 331 cross-lingual question-answer pairs (Pidgin questions → English answers). Generated through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin_Question-English_Answer_Dataset.QuestionAnswer_MCQ
