Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MatinaAI /peka_persian_knowledge_assessmentgated PeKA (Persian Knowledge Assessment) PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics. For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper. This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.tabularquestion-answering1K<n<10K3 likes1.5k downloads1y agoHugging Face02MuskumPillerum /General-Knowledge Dataset Card for Dataset Name Dataset Summary The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'. It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself. Distribution The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.texttext-classification10K<n<100K51 likes467 downloads10mo agoHugging Face03FreedomIntelligence /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M52 likes372 downloads3y agoHugging Face04chembricks /chemistry-knowledge ChemBricks Knowledge Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule? These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer. Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.tabularquestion-answering10K<n<100K1 likes357 downloads23d agoHugging Face05snuh /specialist-level_medical_knowledge_dataset_sft specialist-level_medical_knowledge_dataset_sft Dataset Summary specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K2 likes232 downloads2mo agoHugging Face06wuwu616 /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/wuwu616/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M0 likes222 downloads23d agoHugging Face07emgena /omnimcp_graphrag_knowledge_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_knowledge_teaser.texttext-generationn<1K0 likes163 downloads24d agoHugging Face08jackliu2006 /car_knowledge car_knowledge This dataset contains car knowledge instruction-output pairs generated for LLM fine-tuning. Dataset Description Each record contains: instruction: The input question or task about car knowledge. gpt_output: The response generated by GPT-5. gemini_output: The response generated by Gemini. Dataset Statistics Total records: 3027 Files: 4 parquet file(s) in data/, up to 1000 records each. Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/jackliu2006/car_knowledge.tabulartext-generation1K<n<10K1 likes160 downloads7mo agoHugging Face09ranjithraj /cancer-knowledge-base Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation The only open CC-BY-4.0 oncology knowledge base that combines: 110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this. A provable 152-question MCQ benchmark — every answer derives from this KB's own structured data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.tabularquestion-answering10K<n<100K0 likes142 downloads2mo agoHugging Face10Yxanul /Mephisto-Knowledge_538k Mephisto-Knowledge_538k 538,861 English knowledge SFT examples generated by Qwen/Qwen3.5-4B in non-thinking (Instruct) mode on the Knowledge prompts of openbmb/UltraData-SFT-2605. Responses contain no chain-of-thought — thinking was disabled at generation time, so each assistant turn is a direct answer, usually with a short justification. Companion dataset: Mephisto-IF_172k (instruction-following, same teacher and pipeline). Read this before training: ref_agrees… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-Knowledge_538k.textquestion-answering100K<n<1M2 likes142 downloads2mo agoHugging Face11sosa123454321 /greenpars-knowledge GreenPars knowledge base (RAG) Trilingual reports (FA/EN/TR) of the GreenPars plastic-recycling business plan: business plan, financial model, funding, scopes, Turkish partners, sanctions critique, WtE & sponsor assessment. chunks.json = 130 pre-chunked passages used by the GreenPars AI assistant. textquestion-answeringn<1K0 likes130 downloads12d agoHugging Face12Knowledge-aware-AI /GPTKB_v1.5This hosts the GPTKB v1.5 dataset. Visit https://gptkb.org to browse GPTKB and for further information. Papers: GPTKB methodology: https://arxiv.org/pdf/2411.04920 GPTKB v1.5: https://arxiv.org/pdf/2507.05740 Citations: @InProceedings{GPTKB, title={Enabling LLM Knowledge Analysis via Extensive Materialization}, author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon}, year={2025}, booktitle={ACL}, } @article{GPTKB15, title={GPTKB v1.5: A Massive… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/GPTKB_v1.5.texttext-generation100M<n<1B1 likes124 downloads11mo agoHugging Face13jiosephlee /auxiliary-views-knowledge-acquisition Auxiliary Views Knowledge Acquisition This repository contains the cleaned source documents and evaluation probes used in Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views (arXiv:2609.04180). News August 21, 2026: Our paper was accepted to Findings of EMNLP 2026. Configurations Configuration Split Rows documents train 30 factual_cloze test 6,435 factual_mcqa_5shot test 4,515… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/auxiliary-views-knowledge-acquisition.texttext-generation10K<n<100K1 likes124 downloads1mo agoHugging Face14Knowledge-aware-AI /GPTKB_v1This is the GPTKB dataset from the ACL 2025 paper: @InProceedings{GPTKB, title={Enabling LLM Knowledge Analysis via Extensive Materialization}, author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon}, year={2025}, booktitle={ACL}, } Preprint: https://arxiv.org/pdf/2411.04920 Web interface for browsing GPTKB: https://gptkb.org texttext-generation100M<n<1B0 likes101 downloads1y agoHugging Face15k-mktr /latest-news-knowledge-qa-sep2026-pilot Latest-News Knowledge QA Dataset — PILOT (September 2026) Status: pilot / proof-of-concept. This is the first experimental release of a self-hosted news→QA pipeline, covering a single window: September 2026 (ISO weeks 38–40). It is a time-boxed snapshot, not an ongoing collection — it will not be updated with newer news in this version. If the pilot proves out, follow-up releases will cover later windows. A synthetic fine-tuning dataset of multi-turn question–answer threads… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/latest-news-knowledge-qa-sep2026-pilot.textquestion-answering1K<n<10K0 likes101 downloads10d agoHugging Face16ktiyab /cooking-knowledge-basics Comprehensive Cooking Knowledge Q&A Dataset This dataset (cooking_knowledge.csv) contains a rich collection of synthetically generated Question-Answer (Q&A) pairs covering diverse aspects of cooking knowledge, with particular emphasis on food chemistry, flavor pairing, cooking techniques, dietary accommodations, and culinary traditions. The data was created using a large language model with advanced reasoning capabilities, prompted with various grounded contexts and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/cooking-knowledge-basics.textquestion-answering1K<n<10K6 likes95 downloads2y agoHugging Face17Lots-of-LoRAs /task685_mmmlu_answer_generation_clinical_knowledge Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task685_mmmlu_answer_generation_clinical_knowledge Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task685_mmmlu_answer_generation_clinical_knowledge.texttext-generationn<1K0 likes84 downloads2y agoHugging Face18Uunan /turkish-knowledge-sft Turkish Knowledge SFT Turkish Knowledge SFT is a large-scale synthetic instruction-following dataset designed to improve the factual knowledge, explanation quality, and instructional capabilities of Turkish Large Language Models (LLMs). The dataset is designed for Supervised Fine-Tuning (SFT) and follows a conversation-oriented format compatible with modern chat models. Features 🇹🇷 Entirely in Turkish 🤖 Synthetic instruction-following dataset 📚… See the full description on the dataset page: https://huggingface.co/datasets/Uunan/turkish-knowledge-sft.texttext-generation100K<n<1M0 likes80 downloads3mo agoHugging Face19snuh /essential-level_medical_knowledge_dataset_sft essential-level_medical_knowledge_dataset_sft Dataset Summary essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K0 likes80 downloads19d agoHugging Face20public-knowledge-project /ref-annotation-benchmark RenoBench: A Citation Parsing Benchmark RenoBench (Reference Annotation Benchmark) is a standardized evaluation benchmark for citation parsing—the task of annotating plain-text bibliographic references with structured components following the JATS (Journal Article Tag Suite) standard. Dataset Description RenoBench contains 10,000 plain-text citations paired with their corresponding JATS XML annotations. The dataset was assembled by extracting plain-text references from… See the full description on the dataset page: https://huggingface.co/datasets/public-knowledge-project/ref-annotation-benchmark.texttoken-classification10K<n<100K1 likes79 downloads9mo agoHugging Face21nuhmanpk /dev-knowledge-base Dev Knowledge Base (Programming Documentation Dataset) A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems. Do Follow me on Github: https://github.com/nuhmanpk Overview This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as: Programming languages Frameworks (frontend, backend) DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.tabularquestion-answering100K<n<1M1 likes76 downloads7mo agoHugging Face22MaatAI /african-history-knowledge-merged-sft-cleaned African History Knowledge Merged SFT — Cleaned A reproducible, format-cleaned version of MaatAI/african-history-knowledge-merged-sft, pinned to source commit 0a40eb041d85d59b86219641de0fd87786ee0f77. Split Rows train 28,585 validation 1,589 test 1,589 Total 31,763 Cleaning performed Quarantined 13 training records: 12 have no final answer after a closing thinking tag, and one has ambiguous repeated closing tags. Their original text and… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/african-history-knowledge-merged-sft-cleaned.texttext-generation10K<n<100K0 likes75 downloads1mo agoHugging Face23sthanika-ai /Bharat-Knowledge-Probe-Benchmarkgated BKP-500 — Bharat Knowledge Probe Does your model know where it is? BKP-500 is a benchmark of things every Indian knows and frontier LLMs routinely fumble — lakh/crore arithmetic, Indian digit grouping, state-specific land units (bigha, katha, guntha...), traditional mass units, the Indian fiscal year, agricultural crop seasons, government schemes, and structural identifiers (PAN, GSTIN, IFSC, PIN codes). The evaluation harness that runs a model against this dataset and… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Bharat-Knowledge-Probe-Benchmark.textquestion-answeringn<1K1 likes71 downloads9d agoHugging Face24Pinkstackorg /HQ-knowledgedistills-1.2M-magpieThis dataset is.an exact mix of 900k general qwen conversation with general questions, math, code and another 300k of Gemma 2 27B generations, for creative writing. The dataset was made for "healing" pruned LLM's, especially ones based off of qwen2.5 series, as some conversations include the models saying who they are. Unlike the previous 900K version, we also mixed in Gemma generations, to add more creative writing examples. Many thanks to the magpie project for making this possible, this… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstackorg/HQ-knowledgedistills-1.2M-magpie.texttext-generation1M<n<10M1 likes64 downloads2y agoHugging Face25Knowledge-aware-AI /LLMpedia LLMpedia Encyclopedic articles generated entirely from the parametric memory of large language models — no retrieval — released as a benchmark for studying LLM factuality, unverifiability, and subject-choice behavior at scale. This dataset accompanies the paper "LLMpedia: A Transparent Framework to Materialize an LLM's Encyclopedic Knowledge at Scale" (Saeed & Razniewski, 2026), arXiv:2603.24080. Motivation Benchmarks like MMLU suggest frontier models are near… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/LLMpedia.tabulartext-generation100K<n<1M0 likes63 downloads3mo agoHugging Face26Nekochu /Tree-of-Web-KnowledgeInspired by Tree of Knowledge (ToK), now remade as Proof of Concept: Tree-of-Web-Knowledge aka ToWK. Alpaca Dataset created using llama2, Code, Cleaned using score of llm-blender/PairRM and dedup. Possible improvement: - custom Web search instead of JSON obj by VinciGit00/Scrapegraph-ai. 🔍 .hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .img-lbl { position: relative; display: inline-block; cursor: pointer; } .hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .pv { width: 500px; height: auto;… See the full description on the dataset page: https://huggingface.co/datasets/Nekochu/Tree-of-Web-Knowledge.textquestion-answering1K<n<10K0 likes62 downloads3mo agoHugging Face27knowledge-distillation /openthoughts3_math OpenThoughts3 Math This dataset contains the math-only, complete-solution subset used for supervised fine-tuning in LLM-Fusion experiments. It was derived from open-thoughts/OpenThoughts3-1.2M. Dataset summary 103,760 training rows 32,193 unique math questions Up to four solutions per question, selected deterministically with seed 20260910 All rows have domain = "math" and source = "ai2-adapt-dev/openmath-2-math" Solutions are retained only when the assistant… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-distillation/openthoughts3_math.texttext-generation100K<n<1M0 likes61 downloads24d agoHugging Face28sujitpandey /k-10-general-knowledge-50k k-10-general-knowledge-50k Synthetic, original expository text aligned to the K-10 (CBSE/NCERT-style) curriculum. Rows: 50,000 Total words: 37,314,580 Subject(s): General Knowledge Rows per grade: 1: 3,112, 10: 6,930, 2: 3,112, 3: 5,712, 4: 4,672, 5: 4,328, 6: 5,222, 7: 5,334, 8: 5,789, 9: 5,789 Fields Field Type Description text string The generated passage subject string Subject name grade int Grade level word_count int Number of words in text… See the full description on the dataset page: https://huggingface.co/datasets/sujitpandey/k-10-general-knowledge-50k.tabulartext-generation10K<n<100K0 likes56 downloads2d agoHugging Face29metehan777 /global-seo-knowledgetexttext-generation1K<n<10K3 likes55 downloads2y agoHugging Face30ReactiveAI /AI-Knowledge-Chat-SMAT Dataset Card for ReactiveAI/AI-Knowledge-Chat-SMAT Conversational dataset for Supervised Memory Aware Training (SMAT) of Reactive Language Models (RxLM), containing dialogues with AI/Data Science knowledge. DOCS IN PROGRESS Dataset Details Dataset Description Curated by: Adam Filipek / Reactive AI Language(s) (NLP): English-only License: Apache-2.0 Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/AI-Knowledge-Chat-SMAT.textquestion-answering10K<n<100K2 likes52 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.