Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MatinaAI /peka_persian_knowledge_assessmentgated PeKA (Persian Knowledge Assessment) PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics. For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper. This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.tabularquestion-answering1K<n<10K3 likes1.5k downloads1y agoHugging Face02kalpesh77 /vedic-knowledge-base Vedic Knowledge Base (Shriyantra OS Architecture) हा डेटासेट वैदिक ज्ञान, चक्र, मंत्र, पवित्र भूमिती (Sacred Geometry) आणि आधुनिक न्यूरल नेटवर्क आर्किटेक्चर (Artificial Neural Networks) यांना एकत्र जोडणारे एक नाविन्यपूर्ण मॉडेल आहे. रचना (Structure) मल्टी-कॉन्फिगरेशन सेटिंग्समुळे तुम्ही प्रत्येक फाईल आता स्वतंत्रपणे लोड करू शकता. tabulartext-classificationn<1K0 likes150 downloads3mo agoHugging Face03Knowledge-aware-AI /GPTKB_v1.5This hosts the GPTKB v1.5 dataset. Visit https://gptkb.org to browse GPTKB and for further information. Papers: GPTKB methodology: https://arxiv.org/pdf/2411.04920 GPTKB v1.5: https://arxiv.org/pdf/2507.05740 Citations: @InProceedings{GPTKB, title={Enabling LLM Knowledge Analysis via Extensive Materialization}, author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon}, year={2025}, booktitle={ACL}, } @article{GPTKB15, title={GPTKB v1.5: A Massive… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/GPTKB_v1.5.texttext-generation100M<n<1B1 likes124 downloads11mo agoHugging Face04Knowledge-aware-AI /GPTKB_v1This is the GPTKB dataset from the ACL 2025 paper: @InProceedings{GPTKB, title={Enabling LLM Knowledge Analysis via Extensive Materialization}, author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon}, year={2025}, booktitle={ACL}, } Preprint: https://arxiv.org/pdf/2411.04920 Web interface for browsing GPTKB: https://gptkb.org texttext-generation100M<n<1B0 likes101 downloads1y agoHugging Face05ktiyab /cooking-knowledge-basics Comprehensive Cooking Knowledge Q&A Dataset This dataset (cooking_knowledge.csv) contains a rich collection of synthetically generated Question-Answer (Q&A) pairs covering diverse aspects of cooking knowledge, with particular emphasis on food chemistry, flavor pairing, cooking techniques, dietary accommodations, and culinary traditions. The data was created using a large language model with advanced reasoning capabilities, prompted with various grounded contexts and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/cooking-knowledge-basics.textquestion-answering1K<n<10K6 likes95 downloads2y agoHugging Face06QCRI /islamic-knowledge-ai-survey Advances in AI Systems on Islamic Knowledge Capabilities: A Critical Survey A comprehensive systematic survey of 160+ papers (2016–2026) examining how AI systems operationalize Islamic knowledge, spanning NLP, information retrieval, speech processing, multimodal learning, educational technology, and LLM alignment. Abstract AI systems are increasingly mediating how Islamic communities access, study, and apply Islamic sources; still, research on Islamic-knowledge… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/islamic-knowledge-ai-survey.textn<1K1 likes55 downloads8mo agoHugging Face07rakhasetiawan /herbal-knowledge-embedding-wikipedia Herbs Domain Knowledge Safety Wikipedia Embeddings Pre-computed vector embeddings from Wikipedia articles covering Assorted Herbs, spices, and other botanical items used for alternative medicine — ready to drop into your RAG pipeline without any embedding overhead. Dataset Details Property Details Embedding Model nomic-ai/nomic-embed-text-v1.5 (135M) Vector Dimensions 768 Source Wikipedia Topics herbs, plants, spices, and other botanical items for… See the full description on the dataset page: https://huggingface.co/datasets/rakhasetiawan/herbal-knowledge-embedding-wikipedia.text1K<n<10K0 likes55 downloads8mo agoHugging Face082npc /KnowledgeBerg KnowledgeBerg KnowledgeBerg is a multilingual benchmark for knowledge-grounded reasoning over enumerated answer sets. The default configuration is English. Use the dataset viewer's configuration selector to open each language separately. Data Each language is stored as one CSV file under data/<Language>/. Columns: id original_id domain original_question original_answer question_type implicit_question options correct_option reasoning_depth knowledge_width validation_cot tabular10K<n<100K0 likes46 downloads6mo agoHugging Face09TPelc /Current_Trivia_Knowledge-benchmark Current Trivia Knowledge RAG Benchmark Short Summary: A 140-QA pair (70 train, 70 test) dataset for real-world RAG evaluation. It features current knowledge questions unavailable to LLMs trained before 2024 (e.g., GPT-4o) across diverse domains, and includes human feedback for the training set, enabling robust assessment of contextual information's critical impact on LLM accuracy. Introduction & Motivation: This dataset addresses the critical need for a dynamic… See the full description on the dataset page: https://huggingface.co/datasets/TPelc/Current_Trivia_Knowledge-benchmark.textquestion-answeringn<1K0 likes44 downloads1y agoHugging Face10rakhasetiawan /agricultural-knowledge-embedding-wikipedia Agriculture Domain Knowledge Safety Wikipedia Embeddings Pre-computed vector embeddings from Wikipedia articles covering Agriculture, Planting Methods, and related topics — ready to drop into your RAG pipeline without any embedding overhead. Dataset Details Property Details Embedding Model nomic-ai/nomic-embed-text-v1.5 (135M) Vector Dimensions 768 Source Wikipedia Topics Agriculture, various plants, planting methods and more Format CSV (vectors +… See the full description on the dataset page: https://huggingface.co/datasets/rakhasetiawan/agricultural-knowledge-embedding-wikipedia.text1K<n<10K0 likes41 downloads8mo agoHugging Face11harmesh95 /healthcare-disease-knowledge Disease Symptoms & Treatment Dataset This dataset contains structured information about 687 diseases and their associated details.It is intended for research, educational, and prototyping purposes in healthcare-related ML/NLP tasks. Contents Each row corresponds to one disease, with 17 columns: disease — Name of the disease main_link — Reference link Diagnosis_treatment_link — Link to diagnosis/treatment page Doctors_departments_link — Relevant medical… See the full description on the dataset page: https://huggingface.co/datasets/harmesh95/healthcare-disease-knowledge.texttext-retrievaln<1K1 likes38 downloads11mo agoHugging Face12aashrithareddyy /NeuroExplain-Neurological-Knowledge NeuroExplain: Neurological Knowledge Dataset About This dataset is created for the NeuroExplain: LLM-Based Neurological Report Interpretation and Patient Education System project. The purpose of this dataset is to provide structured neurological information that can be used to develop an AI system capable of explaining neurological conditions in simple and understandable language. Dataset Contents The dataset currently contains information about… See the full description on the dataset page: https://huggingface.co/datasets/aashrithareddyy/NeuroExplain-Neurological-Knowledge.textn<1K1 likes38 downloads27d agoHugging Face13Qulture /qulture-general-knowledge-dataset Qulture General Knowledge Question Dataset An open dataset containing 16 families of general knowledge questions, for a total of 64 records. tabularquestion-answeringn<1K0 likes35 downloads2mo agoHugging Face14rakhasetiawan /food-knowledge-embedding-wikipedia Food Domain Knowledge Safety Wikipedia Embeddings Pre-computed vector embeddings from Wikipedia articles covering Food, Cooking Methods, and other items related to food preparation — ready to drop into your RAG pipeline without any embedding overhead. Dataset Details Property Details Embedding Model nomic-ai/nomic-embed-text-v1.5 (135M) Vector Dimensions 768 Source Wikipedia Topics Food, Cooking methods, and other items related to food preparation… See the full description on the dataset page: https://huggingface.co/datasets/rakhasetiawan/food-knowledge-embedding-wikipedia.text10K<n<100K0 likes33 downloads8mo agoHugging Face15vishnun /NLP-KnowledgeGraph Dataset Card for Dataset Name Dataset Summary KG dataset created by using spaCy PoS and Dependency parser. Supported Tasks and Leaderboards Can be leveraged for token classification for detection of knowledge graph entities and relations. Languages English Dataset Structure Data Instances [More Information Needed] Data Fields Important fields for the token classification task are tokens - tokenized text tags - Tags… See the full description on the dataset page: https://huggingface.co/datasets/vishnun/NLP-KnowledgeGraph.texttoken-classification10K<n<100K17 likes32 downloads3y agoHugging Face16DropTheHQ /dropthe-knowledge-graph DropThe Entity Relationship Graph A large-scale entity relationship dataset containing 2.9 million typed, directional connections between 1.8 million entities spanning entertainment, media, finance, and technology. Extracted from the DropThe knowledge graph. Dataset Description While most open knowledge graphs focus on encyclopedic facts (Wikidata) or narrow domains (MovieLens for ratings), this dataset captures operational relationships -- the connections that actually… See the full description on the dataset page: https://huggingface.co/datasets/DropTheHQ/dropthe-knowledge-graph.textgraph-mln<1K0 likes32 downloads7mo agoHugging Face17dhgottesman /keen_estimating_knowledge_in_llmstext1K<n<10K0 likes29 downloads11mo agoHugging Face18owkin /medical_knowledge_from_extractsThis dataset is used to train LLMs for medical knowledge extraction tasks textsummarization1K<n<10K17 likes26 downloads3y agoHugging Face19Aditya210 /knowledge-graph Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset can be used for training the SVO Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources… See the full description on the dataset page: https://huggingface.co/datasets/Aditya210/knowledge-graph.text10K<n<100K5 likes24 downloads2y agoHugging Face20s-emanuilov /rivers-knowledge-base Rivers Knowledge Base Dataset Full structured knowledge base combining DBpedia extractions with LLM-augmented data for U.S. rivers. This repository contains raw DBpedia SPARQL query results, augmented hydrological measurements, alternative river names, and geographic and administrative metadata. The knowledge base serves as the source data for knowledge graph construction used in the Licensing Oracle experiments. Citation @article{ackermann2025stemming, title={Stemming… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/rivers-knowledge-base.text10K<n<100K0 likes24 downloads11mo agoHugging Face21MagicHub /railway-knowledge-QAtextn<1K1 likes23 downloads3y agoHugging Face22ClarusC64 /legal-file-handover-knowledge-action-deadline-coherence-risk-v0.1What this dataset does You receive handover summary key facts strategy tasks deadlines new actions You decide coherent or incoherent Daily use transfer QC missed deadline prevention strategy drift detection tabulartext-classificationn<1K0 likes23 downloads8mo agoHugging Face23TuringsSolutions /Universal-Knowledge-GPTtext1K<n<10K0 likes22 downloads2y agoHugging Face24kunu5402 /Diverse-Knowledge Everything Data This data is synthetically generated by a ton of open and closed source models. This is basically a parsed version of yearly log form a small dialouge based testing to anylyze model's response on it then perform human evals on it. The data contains information about everything from every domain, most of the pairs included in this data are preferred by humans as the model's response. It can be used for topic modeling, or human preference evals etc. Rest anyone can do… See the full description on the dataset page: https://huggingface.co/datasets/kunu5402/Diverse-Knowledge.texttext-generation10K<n<100K0 likes22 downloads2y agoHugging Face25Adapting /Knowledge-Driven-Dialoguestext10K<n<100K1 likes20 downloads4y agoHugging Face26chriswikoff /Beginner-Golf-Knowledgetextn<1K0 likes20 downloads10mo agoHugging Face27rakhasetiawan /fire-knowledge-embedding-wikipedia Fire Domain Knowledge Safety Wikipedia Embeddings Pre-computed vector embeddings from Wikipedia articles covering fire hazard safety, environmental effects, and related topics — ready to drop into your RAG pipeline without any embedding overhead. Dataset Details Property Details Embedding Model nomic-ai/nomic-embed-text-v1.5 (135M) Vector Dimensions 768 Source Wikipedia Topics Fire safety, hazard prevention, environmental impact, and more Format CSV… See the full description on the dataset page: https://huggingface.co/datasets/rakhasetiawan/fire-knowledge-embedding-wikipedia.text1K<n<10K0 likes20 downloads8mo agoHugging Face28oopere /knowledge_transfer_1500_base Dataset Summary knowledge_transfer_1500_base is a dataset consisting of 1500 prompts distributed across 10 different topics. The prompts are designed to facilitate sentence completion by models, with half of the sentences containing fewer than 5 words and the other half containing between 5 and 10 words. The dataset is intended to serve as a foundation for knowledge transfer tasks, specifically for creating response datasets that can be useful for fine-tuning models to mimic the… See the full description on the dataset page: https://huggingface.co/datasets/oopere/knowledge_transfer_1500_base.text1K<n<10K2 likes19 downloads2y agoHugging Face29elvincth /durecdial-knowledgetext10K<n<100K0 likes17 downloads2y agoHugging Face30sagepond /Self-knowledge-datasettextn<1K0 likes17 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.