datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Customer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/MakTek/Customer_support_faqs_dataset.faq-bacen
FaqBacenRetrieval
Retrieve the correct answer to a citizen question about Brazilian financial/banking regulation, from the Banco Central do Brasil (BACEN) public FAQ. 373 test questions over a pool of 1673 unique regulatory answers. Native PT-BR; financial/government domain.
Part of MTEB-BR — the native Brazilian-Portuguese MTEB sub-benchmark. Task type: Retrieval · Language: Brazilian Portuguese (mined from real-world sources) · Domains: Financial, Government, Written.… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/faq-bacen.faquad-ir
FaQuADIR
FaQuAD reformulated as PT-BR academic retrieval: given a question about Brazilian higher education, retrieve the source paragraph containing the answer. 900 questions over 244 unique paragraphs, drawn from 18 official documents of a federal-university CS program plus 21 Wikipedia articles about Brazil's higher-education system. Complements legal/web retrieval domains with academic/educational content.
Part of MTEB-BR — the native Brazilian-Portuguese MTEB sub-benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/faquad-ir.faq
QA4FAQ @ EVALITA 2016
Original dataset information available here
Data format
The data has been converted to be used as a questin answering task.
There are two splits, test-1 and test-2, each containing the same data processed in slightly different ways.
test-1
The data is in jsonl format, where each line is a json object with the following fields:
id: a unique identifier for the question
question: the question
A, B, C, D: the possible answers to the question… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/faq.bitcoin-wallet-recovery-faq
Bitcoin Wallet Recovery FAQ Dataset v1.0
A high-quality Question & Answer dataset focused exclusively on Bitcoin wallet recovery and self-custody best practices. It is designed for training, fine-tuning, and evaluating LLMs and retrieval-augmented generation (RAG) systems in the domain of bitcoin security, seed backup, device loss, and fund recovery.
Dataset Summary
Total records: 500
Language: English
Answer length: 150–300 words per record
Categories: 39… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-recovery-faq.Mental_Health_FAQ
License & Attribution
MTEB-format derivative of tolu07/Mental_Health_FAQ. Licensed under MIT (same as source). Text encoding repaired with ftfy.
sap_faqhc3-faq-fr-gouv
License & Attribution
MTEB-format derivative of the faq_fr_gouv subset of almanach/hc3_french_ood (French government FAQ). Query = question; corpus = answer. Licensed under CC-BY-SA-4.0 (same as source).
georgian-faqFrequently asked questions (FAQs) and answers mined from Georgian websites via Common Crawl
faquad-nli
Dataset Card for FaQuAD-NLI
Dataset Summary
FaQuAD is a Portuguese reading comprehension dataset that follows the format of the Stanford Question Answering Dataset (SQuAD). It is a pioneer Portuguese reading comprehension dataset using the challenging format of SQuAD. The dataset aims to address the problem of abundant questions sent by academics whose answers are found in available institutional documents in the Brazilian higher education system. It consists of 900… See the full description on the dataset page: https://huggingface.co/datasets/ruanchaves/faquad-nli.faq
Council of AI — FAQ door
FAQ door. Living GET only. No 13-of-14. No 24/7. No “all 313 verify.”
Live: https://councilof.ai/faq
Council OS: https://councilof.ai/os
Council Space: https://councilof.ai/gspc-arena
Measurement, not certification. Empty slots are not for sale. No scores on this card.
Jail is a measured floor, not a 16th pane.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is a printer of… See the full description on the dataset page: https://huggingface.co/datasets/csoai/faq.Ecommerce_FAQEcommerce FAQ Chatbot Dataset
Overview
The Ecommerce FAQ Chatbot Dataset is a valuable collection of questions and corresponding answers, meticulously curated for training and evaluating chatbot models in the context of an Ecommerce environment. This dataset is designed to assist developers, researchers, and data scientists in building effective chatbots that can handle customer inquiries related to an Ecommerce platform.
Contents
The dataset comprises a total of 79 question-answer pairs… See the full description on the dataset page: https://huggingface.co/datasets/Andyrasika/Ecommerce_FAQ.synthetic-persian-chatbot-rag-faq-retrieval
Dataset Summary
Synthetic Persian Chatbot RAG FAQ Retrieval (SynPerChatbotRAGFAQRetrieval) is a Persian (Farsi) dataset built for the Retrieval task in Retrieval-Augmented Generation (RAG)-based chatbot systems. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini. The dataset is designed to evaluate how well models retrieve relevant FAQ entries based on a user's message and prior conversation context.
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-rag-faq-retrieval.FDA_Pharmaceuticals_FAQ
FDA Pharmaceutical Q&A Dataset
Description
This dataset contains a collection of question-and-answer pairs related to pharmaceutical regulatory compliance provided by the Food and Drug Administration (FDA). It is designed to support research and development in the field of natural language processing, particularly for tasks involving information retrieval, question answering, and conversational agents within the pharmaceutical domain.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/Jaymax/FDA_Pharmaceuticals_FAQ.faquadAcademic secretaries and faculty members of higher education institutions face a common problem:
the abundance of questions sent by academics
whose answers are found in available institutional documents.
The official documents produced by Brazilian public universities are vast and disperse,
which discourage students to further search for answers in such sources.
In order to lessen this problem, we present FaQuAD:
a novel machine reading comprehension dataset
in the domain of Brazilian higher education institutions.
FaQuAD follows the format of SQuAD (Stanford Question Answering Dataset) [Rajpurkar et al. 2016].
It comprises 900 questions about 249 reading passages (paragraphs),
which were taken from 18 official documents of a computer science college
from a Brazilian federal university
and 21 Wikipedia articles related to Brazilian higher education system.
As far as we know, this is the first Portuguese reading comprehension dataset in this format.synthetic-persian-chatbot-rag-faq-pair-classification
Dataset Summary
Synthetic Persian Chatbot RAG FAQ Pair Classification (SynPerChatbotRAGFAQPC) is a Persian (Farsi) dataset for the Pair Classification task, specifically designed for evaluating Retrieval-Augmented Generation (RAG) chatbot systems. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini.
The dataset measures a model’s ability to assess whether a given FAQ (question–answer pair) is relevant to a new user… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-rag-faq-pair-classification.embedded_faqs_medicaredomeggook_faqfiltered-finephrase-faqFAQ_BACENThis dataset was used in the article: https://arxiv.org/abs/2311.11331
Trade_Balance_Forest_Products_Comparative_FAQ_Dataset
Trade Balance & Forest Products Comparative FAQ Dataset (Nepali)
A Nepali-language (Devanagari script) instruction-following dataset built from two real Nepali government/official statistical publications: Nepal's Trade Balance by Partner Countries (FY 2074/75) and Forest Production Data Nepal (published by the Department of National Parks and Wildlife Conservation). Unlike simple single-value statistic lookups, almost every record in this dataset asks the model to compare two… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Trade_Balance_Forest_Products_Comparative_FAQ_Dataset.ShareHub_FAQ_Nepali_Dataset
ShareHub FAQ Nepali Dataset
A Nepali-language, instruction-following (question–answer) dataset of 2,448 real frequently-asked-questions about securities listed on the Nepal Stock Exchange (NEPSE) — covering common stocks, debentures, and mutual funds — sourced from ShareHub. Every record is a single-turn human → gpt conversation in Devanagari script, and every record carries rich provenance and generation metadata.
File analyzed: sharehub_faq.jsonl
1. Quick Facts… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/ShareHub_FAQ_Nepali_Dataset.ShareHub_NEPSE_Securities_FAQ_Dataset_Romanized
ShareHub NEPSE Securities FAQ Dataset (Romanized)
File: sharehub_faq_romanized.jsonl
Total records: 2,448
Format: JSON Lines (.jsonl) — one JSON object per line
Language: Nepali (ne / ISO 639-3 npi) — content is written in romanized Nepali (Latin script) in both the question and the answer, despite the metadata declaring script: "Deva" (see Section 16 for this discrepancy)
Domain: Financial services — real-time-style market data for securities listed on the Nepal Stock Exchange… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/ShareHub_NEPSE_Securities_FAQ_Dataset_Romanized.romanian-legal-faq-2026
Dataset: Romanian Legal FAQ 2026 (Coltuc Legal Knowledge Base)
Descriere
Set de date structurat în limba română conținând instrucțiuni, întrebări frecvente și soluții procedurale din dreptul civil, drept bancar (clauze abuzive, executări silite), dreptul muncii și dreptul pensiilor.
Dataset-ul este optimizat pentru fine-tuning LLM, sisteme RAG (Retrieval-Augmented Generation) și modele de asistență juridică automată.
Autor și Proprietate Intelectuală… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/romanian-legal-faq-2026.retail-faqbanking-faq-hi-en-speech
Banking FAQ Hindi-English Speech Dataset
A multilingual banking FAQ dataset that combines an English FAQ source with Hindi translation and synthetic speech generation.
Overview
This project starts from a public English banking FAQ dataset and converts it into a Hindi-English speech dataset for research and experimentation in conversational AI and multilingual speech systems.
Source dataset
Kaggle:… See the full description on the dataset page: https://huggingface.co/datasets/ashirbadsahu/banking-faq-hi-en-speech.NRB_Inflation_Annual_Report_FAQ_Dataset
NRB Inflation & Annual Report FAQ Dataset (Nepali)
A Nepali-language (Devanagari script) instruction-following dataset built from Nepal Rastra Bank (NRB) official statistical publications — specifically Annual Report economic-indicator datasets and the Inflation Indicators report for Fiscal Year 2079/80 (2022/23). Each record is a single-turn human → gpt FAQ pair where the question asks for a specific statistic (e.g., a growth rate, an index value, an inflation rate, an… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/NRB_Inflation_Annual_Report_FAQ_Dataset.e-faqiran-legal-Faq-dataset
📚 Persian Legal Q&A Dataset | دیتاست سوال و جوابهای حقوقی فارسی
This dataset contains a curated collection of 170 Persian legal question–answer pairs categorized into common legal domains such as:
Family Law (حقوق خانواده)
Criminal Law (حقوق کیفری)
Contract Law (قراردادها)
Labor Law (روابط کارگر و کارفرما)
Property & Real Estate (دعاوی ملکی)
Inheritance Law (ارث و وصیت)
Civil Procedure (آیین دادرسی)
Enforcement (اجرای احکام)
Public Law (حقوق عمومی)
Financial Claims (مطالبات مالی)… See the full description on the dataset page: https://huggingface.co/datasets/sasanbarok/iran-legal-Faq-dataset.Nepal_Economic_Statistics_FAQ
🇳🇵 Nepal Economic Statistics FAQ (Nepali) — Merged Dataset README
A merged, sequentially re-indexed dataset of 645 Nepali-language factual Q&A pairs covering three related economic domains: commodity/SITC trade, GDP by ISIC sector, and customs import/export by checkpoint. Built by combining three source datasets into a single JSONL file.
🔖 TL;DR
What
Value
Total records
645
Output file
merged_nepal_economic_stats.jsonl
File size
~1.21 MB… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Nepal_Economic_Statistics_FAQ.
