Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01xwjzds /extractive_qa_question_answering_hr Dataset Card HR-Multiwoz is a fully-labeled dataset of 5980 extractive qa spanning 10 HR domains to evaluate LLM Agent. It is the first labeled open-sourced conversation dataset in the HR domain for NLP research. Please refer to HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent for details about the dataset construction. Dataset Sources Repository: xwjzds/extractive_qa_question_answering_hr Paper: HR-MultiWOZ: A Task Oriented Dialogue (TOD)… See the full description on the dataset page: https://huggingface.co/datasets/xwjzds/extractive_qa_question_answering_hr.text1K<n<10K12 likes892 downloads3y agoHugging Face02OpenFinAL /Financial_Question_Answeringtext1K<n<10K2 likes525 downloads11mo agoHugging Face03AswiN037 /tamil-question-answering-datasetthis dataset contains 5 columns context, question, answer_start, answer_text, source Column Description context A general small paragraph in tamil language question question framed form the context answer_text text span that extracted from context answer_start index of answer_text source who framed this context, question, answer pair source team KBA => (Karthi, Balaji, Azeez) these people manually created CHAII =>a kaggle competition XQA => multilingual QA… See the full description on the dataset page: https://huggingface.co/datasets/AswiN037/tamil-question-answering-dataset.text1K<n<10K8 likes226 downloads4y agoHugging Face04shahrukh95 /OWASP-question-answer-datasettextn<1K0 likes139 downloads3y agoHugging Face05nogyxo /question-answering-ukrainian-json-answerstext100K<n<1M8 likes105 downloads3y agoHugging Face06mou3az /Question-Answering-Generation-Choices The dataset is a merged compilation of QuAIL, RACE, and Cosmos QA datasets, having undergone preprocessing. textquestion-answering10K<n<100K10 likes105 downloads3y agoHugging Face07shahrukh95 /OWASP-and-NVD-question-answer-datasettext10K<n<100K2 likes96 downloads3y agoHugging Face08nogyxo /question-answering-ukrainiantabular100K<n<1M8 likes92 downloads3y agoHugging Face09code-switching /question-answertext1K<n<10K0 likes86 downloads1mo agoHugging Face10Mreeb /Dermatology-Question-Answer-Dataset-For-Fine-Tuning Dataset Details The data set has about 1 Million Tokens for Training and about 1500 question answers. Dataset Description This dataset is a comprehensive compilation of questions related to dermatology, spanning inquiries about various skin diseases, their symptoms, recommended medications, and available treatment modalities. Each question is paired with a concise and informative response, making it an ideal resource for training and fine-tuning language models in the… See the full description on the dataset page: https://huggingface.co/datasets/Mreeb/Dermatology-Question-Answer-Dataset-For-Fine-Tuning.tabulartext-generation1K<n<10K7 likes67 downloads3y agoHugging Face11its-myrto /fitness-question-answersA total of 965 q&a pairs i gathered from the web related to physical activity and fitness. textquestion-answeringn<1K9 likes55 downloads2y agoHugging Face12CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes49 downloads2y agoHugging Face13Kubermatic /cncf-question-and-answer-dataset-for-llm-training CNCF QA Dataset for LLM Tuning Description This dataset, named cncf-qa-dataset-for-llm-tuning, is designed for fine-tuning large language models (LLMs) and is formatted in a question-answer (QA) style. The data is sourced from PDF and markdown (MD) files extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. These files were processed and converted into a QA format to be fed into the LLM model. The dataset includes the… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/cncf-question-and-answer-dataset-for-llm-training.text10K<n<100K3 likes40 downloads2y agoHugging Face14aisquared /dais-question-answers DAIS-Question-Answers Dataset This dataset contains question-answer pairs created using ChatGPT using text data scraped from the Databricks Data and AI Summit 2023 (DAIS 2023) homepage as well as text from any public page that is linked in that page or is a two-hop linked page. We have used this dataset to fine-tune our DAIS DLite model, along with our dataset of webpage texts. Feel free to check them out! Note that, due to the use of ChatGPT to curate these question-answer pairs… See the full description on the dataset page: https://huggingface.co/datasets/aisquared/dais-question-answers.text1K<n<10K1 likes36 downloads3y agoHugging Face15Subi1152 /tamil-question-answering-datasetthis dataset contains 5 columns context, question, answer_start, answer_text, source Column Description context A general small paragraph in tamil language question question framed form the context answer_text text span that extracted from context answer_start index of answer_text source who framed this context, question, answer pair source team KBA => (Karthi, Balaji, Azeez) these people manually created CHAII =>a kaggle competition XQA => multilingual QA… See the full description on the dataset page: https://huggingface.co/datasets/Subi1152/tamil-question-answering-dataset.text1K<n<10K0 likes34 downloads7mo agoHugging Face16GokulWork /QuestionAnswer_MCQtexttext-generationn<1K1 likes31 downloads3y agoHugging Face17nharshavardhana /Santali-Ol-Chiki-Agriculture_Question-Answer_DatasetSantali (Ol Chiki) Agriculture Question-Answer Dataset is a curated collection of question–answer pairs focused on agricultural knowledge in the Santali language, written in the Ol Chiki script. The dataset consists of question-answer pairs in the Santali language focusing on agriculture, animal husbandry, and rural health topics. The content covers crop diseases, soil management, livestock care, and farming techniques tailored for tribal communities. This dataset is designed to support… See the full description on the dataset page: https://huggingface.co/datasets/nharshavardhana/Santali-Ol-Chiki-Agriculture_Question-Answer_Dataset.textn<1K2 likes25 downloads6mo agoHugging Face18StaAhmed /Football_Question_Answerstext1K<n<10K2 likes24 downloads3y agoHugging Face19MasterControlAIML /question_answer_finetuning_embeddings.csvtext100K<n<1M0 likes24 downloads2y agoHugging Face20junaid20 /question_answertextn<1K1 likes20 downloads3y agoHugging Face21HanHan055 /quora_question_answer_pairtext100K<n<1M0 likes18 downloads4y agoHugging Face22chenle015 /OpenMP_Question_Answering OpenMP Question Answering Dataset OpenMP Question Answering Dataset is a new OpenMP question answering introduced in paper "LM4HPC: Towards Effective Language Model Application in High-Performance Computing". It is designed to probe the capabilities of language models in single-turn interactions with users. Similar to other QA datasets, we include some request-response pairs which are not strictly question-answering pairs. The categories and examples of questions in the OMPQA… See the full description on the dataset page: https://huggingface.co/datasets/chenle015/OpenMP_Question_Answering.textn<1K1 likes18 downloads3y agoHugging Face23RahulS3 /siddha_vaithiyam_question_answering_chatbot Medical Home Remedy Chatbot Dataset Overview This dataset is designed for a chatbot that answers questions related to medical problems with simple home remedies. The information in this dataset has been sourced from old books containing traditional remedies used in the past. Contents Dataset Files: dataset.csv : The main dataset file containing questions and corresponding home remedy answers. Data Structure: Each row in the CSV file… See the full description on the dataset page: https://huggingface.co/datasets/RahulS3/siddha_vaithiyam_question_answering_chatbot.texttable-question-answering1K<n<10K2 likes18 downloads3y agoHugging Face24shahrukh95 /NVD-question-answer-datasettext10K<n<100K0 likes14 downloads3y agoHugging Face25hmmamalrjoub /Islam_Question_and_Answer1tabularquestion-answeringn<1K1 likes11 downloads2y agoHugging Face26geekyfreak /question_answertextn<1K0 likes8 downloads2y agoHugging Face27yamini0506 /extractive_multi_turn_question_answeringtext1K<n<10K1 likes8 downloads2y agoHugging Face28gaydmi /question-without-answerstext1K<n<10K0 likes8 downloads2y agoHugging Face29alucent /mirror-cncf-question-and-answer-dataset-for-llm-traininggated CNCF QA Dataset for LLM Tuning Description This dataset, named cncf-qa-dataset-for-llm-tuning, is designed for fine-tuning large language models (LLMs) and is formatted in a question-answer (QA) style. The data is sourced from PDF and markdown (MD) files extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. These files were processed and converted into a QA format to be fed into the LLM model. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-cncf-question-and-answer-dataset-for-llm-training.text10K<n<100K0 likes8 downloads3mo agoHugging Face30pdp19 /context_question_answertext1K<n<10K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.