datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
peka_persian_knowledge_assessment
PeKA (Persian Knowledge Assessment)
PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics.
For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper.
This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.chemistry-knowledge
ChemBricks Knowledge
Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule?
These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer.
Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.specialist-level_medical_knowledge_dataset_sft
specialist-level_medical_knowledge_dataset_sft
Dataset Summary
specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/wuwu616/huatuo_knowledge_graph_qa.omnimcp_graphrag_knowledge_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_knowledge_teaser.car_knowledge
car_knowledge
This dataset contains car knowledge instruction-output pairs generated for LLM fine-tuning.
Dataset Description
Each record contains:
instruction: The input question or task about car knowledge.
gpt_output: The response generated by GPT-5.
gemini_output: The response generated by Gemini.
Dataset Statistics
Total records: 3027
Files: 4 parquet file(s) in data/, up to 1000 records each.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/jackliu2006/car_knowledge.cancer-knowledge-base
Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation
The only open CC-BY-4.0 oncology knowledge base that combines:
110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis
rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this.
A provable 152-question MCQ benchmark — every answer derives from this KB's own structured
data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.Mephisto-Knowledge_538k
Mephisto-Knowledge_538k
538,861 English knowledge SFT examples generated by
Qwen/Qwen3.5-4B in non-thinking
(Instruct) mode on the Knowledge prompts of
openbmb/UltraData-SFT-2605.
Responses contain no chain-of-thought — thinking was disabled at generation
time, so each assistant turn is a direct answer, usually with a short
justification.
Companion dataset: Mephisto-IF_172k
(instruction-following, same teacher and pipeline).
Read this before training: ref_agrees… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/Mephisto-Knowledge_538k.greenpars-knowledge
GreenPars knowledge base (RAG)
Trilingual reports (FA/EN/TR) of the GreenPars plastic-recycling business plan: business plan, financial model, funding, scopes, Turkish partners, sanctions critique, WtE & sponsor assessment. chunks.json = 130 pre-chunked passages used by the GreenPars AI assistant.
GPTKB_v1.5This hosts the GPTKB v1.5 dataset. Visit https://gptkb.org to browse GPTKB and for further information.
Papers:
GPTKB methodology: https://arxiv.org/pdf/2411.04920
GPTKB v1.5: https://arxiv.org/pdf/2507.05740
Citations:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
@article{GPTKB15,
title={GPTKB v1.5: A Massive… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/GPTKB_v1.5.auxiliary-views-knowledge-acquisition
Auxiliary Views Knowledge Acquisition
This repository contains the cleaned source documents and evaluation
probes used in Knowledge Acquisition During Pre-training? Large Language Models
Learn Better With Auxiliary Views (arXiv:2609.04180).
News
August 21, 2026: Our paper was accepted to Findings of EMNLP 2026.
Configurations
Configuration
Split
Rows
documents
train
30
factual_cloze
test
6,435
factual_mcqa_5shot
test
4,515… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/auxiliary-views-knowledge-acquisition.GPTKB_v1This is the GPTKB dataset from the ACL 2025 paper:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
Preprint: https://arxiv.org/pdf/2411.04920
Web interface for browsing GPTKB: https://gptkb.org
latest-news-knowledge-qa-sep2026-pilot
Latest-News Knowledge QA Dataset — PILOT (September 2026)
Status: pilot / proof-of-concept. This is the first experimental release
of a self-hosted news→QA pipeline, covering a single window: September
2026 (ISO weeks 38–40). It is a time-boxed snapshot, not an ongoing
collection — it will not be updated with newer news in this version. If the
pilot proves out, follow-up releases will cover later windows.
A synthetic fine-tuning dataset of multi-turn question–answer threads… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/latest-news-knowledge-qa-sep2026-pilot.cooking-knowledge-basics
Comprehensive Cooking Knowledge Q&A Dataset
This dataset (cooking_knowledge.csv) contains a rich collection of synthetically generated Question-Answer (Q&A) pairs covering diverse aspects of cooking knowledge, with particular emphasis on food chemistry, flavor pairing, cooking techniques, dietary accommodations, and culinary traditions. The data was created using a large language model with advanced reasoning capabilities, prompted with various grounded contexts and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/cooking-knowledge-basics.task685_mmmlu_answer_generation_clinical_knowledge
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task685_mmmlu_answer_generation_clinical_knowledge
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task685_mmmlu_answer_generation_clinical_knowledge.turkish-knowledge-sft
Turkish Knowledge SFT
Turkish Knowledge SFT is a large-scale synthetic instruction-following dataset designed to improve the factual knowledge, explanation quality, and instructional capabilities of Turkish Large Language Models (LLMs).
The dataset is designed for Supervised Fine-Tuning (SFT) and follows a conversation-oriented format compatible with modern chat models.
Features
🇹🇷 Entirely in Turkish
🤖 Synthetic instruction-following dataset
📚… See the full description on the dataset page: https://huggingface.co/datasets/Uunan/turkish-knowledge-sft.essential-level_medical_knowledge_dataset_sft
essential-level_medical_knowledge_dataset_sft
Dataset Summary
essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.ref-annotation-benchmark
RenoBench: A Citation Parsing Benchmark
RenoBench (Reference Annotation Benchmark) is a standardized evaluation benchmark for citation parsing—the task of annotating plain-text bibliographic references with structured components following the JATS (Journal Article Tag Suite) standard.
Dataset Description
RenoBench contains 10,000 plain-text citations paired with their corresponding JATS XML annotations. The dataset was assembled by extracting plain-text references from… See the full description on the dataset page: https://huggingface.co/datasets/public-knowledge-project/ref-annotation-benchmark.dev-knowledge-base
Dev Knowledge Base (Programming Documentation Dataset)
A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems.
Do Follow me on Github: https://github.com/nuhmanpk
Overview
This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as:
Programming languages
Frameworks (frontend, backend)
DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.african-history-knowledge-merged-sft-cleaned
African History Knowledge Merged SFT — Cleaned
A reproducible, format-cleaned version of MaatAI/african-history-knowledge-merged-sft, pinned to source commit 0a40eb041d85d59b86219641de0fd87786ee0f77.
Split
Rows
train
28,585
validation
1,589
test
1,589
Total
31,763
Cleaning performed
Quarantined 13 training records: 12 have no final answer after a closing thinking tag, and one has ambiguous repeated closing tags. Their original text and… See the full description on the dataset page: https://huggingface.co/datasets/MaatAI/african-history-knowledge-merged-sft-cleaned.Bharat-Knowledge-Probe-Benchmark
BKP-500 — Bharat Knowledge Probe
Does your model know where it is?
BKP-500 is a benchmark of things every Indian knows and frontier LLMs routinely fumble — lakh/crore
arithmetic, Indian digit grouping, state-specific land units (bigha, katha, guntha...), traditional
mass units, the Indian fiscal year, agricultural crop seasons, government schemes, and structural
identifiers (PAN, GSTIN, IFSC, PIN codes).
The evaluation harness that runs a model against this dataset and… See the full description on the dataset page: https://huggingface.co/datasets/sthanika-ai/Bharat-Knowledge-Probe-Benchmark.HQ-knowledgedistills-1.2M-magpieThis dataset is.an exact mix of 900k general qwen conversation with general questions, math, code and another 300k of Gemma 2 27B generations, for creative writing.
The dataset was made for "healing" pruned LLM's, especially ones based off of qwen2.5 series, as some conversations include the models saying who they are.
Unlike the previous 900K version, we also mixed in Gemma generations, to add more creative writing examples.
Many thanks to the magpie project for making this possible, this… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstackorg/HQ-knowledgedistills-1.2M-magpie.LLMpedia
LLMpedia
Encyclopedic articles generated entirely from the parametric memory of large
language models — no retrieval — released as a benchmark for studying LLM
factuality, unverifiability, and subject-choice behavior at scale.
This dataset accompanies the paper
"LLMpedia: A Transparent Framework to Materialize an LLM's Encyclopedic
Knowledge at Scale" (Saeed & Razniewski, 2026), arXiv:2603.24080.
Motivation
Benchmarks like MMLU suggest frontier models are near… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/LLMpedia.Tree-of-Web-KnowledgeInspired by Tree of Knowledge (ToK), now remade as Proof of Concept: Tree-of-Web-Knowledge aka ToWK.
Alpaca Dataset created using llama2, Code, Cleaned using score of llm-blender/PairRM and dedup.
Possible improvement: - custom Web search instead of JSON obj by VinciGit00/Scrapegraph-ai.
🔍
.hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .img-lbl { position: relative; display: inline-block; cursor: pointer; }
.hf-sanitized.hf-sanitized-UDgbtn3GgVkKb3cKXMTHL .pv { width: 500px; height: auto;… See the full description on the dataset page: https://huggingface.co/datasets/Nekochu/Tree-of-Web-Knowledge.openthoughts3_math
OpenThoughts3 Math
This dataset contains the math-only, complete-solution subset used for supervised fine-tuning in LLM-Fusion experiments. It was derived from open-thoughts/OpenThoughts3-1.2M.
Dataset summary
103,760 training rows
32,193 unique math questions
Up to four solutions per question, selected deterministically with seed 20260910
All rows have domain = "math" and source = "ai2-adapt-dev/openmath-2-math"
Solutions are retained only when the assistant… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-distillation/openthoughts3_math.k-10-general-knowledge-50k
k-10-general-knowledge-50k
Synthetic, original expository text aligned to the K-10 (CBSE/NCERT-style) curriculum.
Rows: 50,000
Total words: 37,314,580
Subject(s): General Knowledge
Rows per grade: 1: 3,112, 10: 6,930, 2: 3,112, 3: 5,712, 4: 4,672, 5: 4,328, 6: 5,222, 7: 5,334, 8: 5,789, 9: 5,789
Fields
Field
Type
Description
text
string
The generated passage
subject
string
Subject name
grade
int
Grade level
word_count
int
Number of words in text… See the full description on the dataset page: https://huggingface.co/datasets/sujitpandey/k-10-general-knowledge-50k.global-seo-knowledgeAI-Knowledge-Chat-SMAT
Dataset Card for ReactiveAI/AI-Knowledge-Chat-SMAT
Conversational dataset for Supervised Memory Aware Training (SMAT) of Reactive Language Models (RxLM), containing dialogues
with AI/Data Science knowledge.
DOCS IN PROGRESS
Dataset Details
Dataset Description
Curated by: Adam Filipek / Reactive AI
Language(s) (NLP): English-only
License: Apache-2.0
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/AI-Knowledge-Chat-SMAT.
