datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ExperimentDATA_knowledge_distillation_vs_fine_tuningmulti_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.wikipedia-knowledge
Wikipedia Knowledge Catalogs
20 gezk knowledge catalogs: portable, read-only reference corpora with full-text and vector indexes, each packaged as one .gezk file that Gezel and any gezk reader can search and cite offline. Every catalog is here twice — the archive, and the same content as Parquet tables for data tools.
Catalogs
Catalog
Id
Version
Documents
Chunks
Archive
Wikipedia: Arts & Architecture
wikipedia-arts
2026.4.5
16,485
73,383
179.0 MB… See the full description on the dataset page: https://huggingface.co/datasets/Bendyline/wikipedia-knowledge.peka_persian_knowledge_assessment
PeKA (Persian Knowledge Assessment)
PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics.
For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper.
This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.python-image-copilot-training-using-import-knowledge-graphs
Python Copilot Image Training using Import Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 216642
Size: 211.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.kaggle-knowledge-graph
Kaggle Knowledge Graph
A knowledge graph built from Kaggle's public
Meta Kaggle (and Meta Kaggle Code
datasets). It links competitions, teams, submissions, users, notebooks,
datasets, discussion forums, tags, organizations, and notebook code invocations.
See the project repository for build scripts and
documentation on the graph schema, data model, and usage examples.
Release
Field
Value
Version
2026-09-30
Meta Kaggle snapshot
2026-09-30
Build code… See the full description on the dataset page: https://huggingface.co/datasets/habedi/kaggle-knowledge-graph.python-image-copilot-training-using-class-knowledge-graphs
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312277
Size: 304.3 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.python-audio-copilot-training-using-function-knowledge-graphs
Python Copilot Audio Training using Global Functions with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.python-audio-copilot-training-using-class-knowledge-graphs
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.python-audio-copilot-training-using-import-knowledge-graphs
Python Copilot Audio Training using Imports with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each imported module for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-import-knowledge-graphs.chemistry-knowledge
ChemBricks Knowledge
Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule?
These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer.
Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.python-image-copilot-training-using-function-knowledge-graphs
Python Copilot Image Training using Function Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 134357
Size: 130.5 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-function-knowledge-graphs.python-image-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312836
Size: 294.1 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs-2024-01-27.python-image-copilot-training-using-inheritance-knowledge-graphs
Python Copilot Image Training using Inheritance and Polymorphism Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 259017
Size: 135.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-inheritance-knowledge-graphs.python-audio-copilot-training-using-inheritance-knowledge-graphs
Python Copilot Audio Training using Inheritance and Polymorphism Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each base class for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-inheritance-knowledge-graphs.llm-knowledge-collapse
"Epistemic Diversity and Knowledge Collapse in Large Language Models" (Wright et al. 2025)
Authors: Dustin Wright, Sarah Masud, Jared Moore, Srishti Yadav, Maria Antoniak, Peter Ebert Christiensen, Chan Young Park, and Isabelle Augenstein
Contains all 1.6M responses and 70M claims used to measure LLM epistemic diversity in the paper "Epistemic Diversity and Knowledge Collapse in Large Language Models" (Wright et al. 2025)
@article{wright2025epistemicdiversity… See the full description on the dataset page: https://huggingface.co/datasets/dwright37/llm-knowledge-collapse.python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.regional-knowledge
Regional Knowledge Catalogs
26 gezk knowledge catalogs: portable, read-only reference corpora with full-text and vector indexes, each packaged as one .gezk file that Gezel and any gezk reader can search and cite offline. Every catalog is here twice — the archive, and the same content as Parquet tables for data tools.
Regional catalogs intentionally overlap. Document totals count regional copies; deduplicate across catalogs by publisher and document ID. Each document can have… See the full description on the dataset page: https://huggingface.co/datasets/Bendyline/regional-knowledge.Misleading_KnowledgeMisleading_Knowledge
Misleading_Knowledge is the misleading-knowledge corpus introduced in “Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions.” It is designed for controlled research on factual robustness, evidence verification, source cues, and false-conclusion adoption in Deep Research agents.
Paper: https://arxiv.org/abs/2607.20891
Code: https://github.com/whfeLingYu/MisKnow-Agent
Dataset repository: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge.car_knowledge
car_knowledge
This dataset contains car knowledge instruction-output pairs generated for LLM fine-tuning.
Dataset Description
Each record contains:
instruction: The input question or task about car knowledge.
gpt_output: The response generated by GPT-5.
gemini_output: The response generated by Gemini.
Dataset Statistics
Total records: 3027
Files: 4 parquet file(s) in data/, up to 1000 records each.
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/jackliu2006/car_knowledge.vedic-knowledge-base
Vedic Knowledge Base (Shriyantra OS Architecture)
हा डेटासेट वैदिक ज्ञान, चक्र, मंत्र, पवित्र भूमिती (Sacred Geometry) आणि आधुनिक न्यूरल नेटवर्क आर्किटेक्चर (Artificial Neural Networks) यांना एकत्र जोडणारे एक नाविन्यपूर्ण मॉडेल आहे.
रचना (Structure)
मल्टी-कॉन्फिगरेशन सेटिंग्समुळे तुम्ही प्रत्येक फाईल आता स्वतंत्रपणे लोड करू शकता.
cancer-knowledge-base
Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation
The only open CC-BY-4.0 oncology knowledge base that combines:
110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis
rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this.
A provable 152-question MCQ benchmark — every answer derives from this KB's own structured
data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.agent-knowledge-cycle
Agent Knowledge Cycle (AKC) — Knowledge Graph
JSON-LD knowledge graph encoding the concept layer of the Agent Knowledge Cycle (AKC) — a six-phase bidirectional growth loop in which agent behavior and the operator's judgment co-develop over time, sustaining intent alignment that tests cannot check on their own.
What this dataset is
This dataset is a mirror of the graph.jsonld file at the root of the AKC GitHub repository. It is provided here for LLM training… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/agent-knowledge-cycle.earth-love-united-climate-knowledge
🌍 Earth Love United Climate Knowledge Dataset
The most comprehensive open climate science knowledge dataset.
10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points.
Built to power GAIA — an AI that embodies the living consciousness of Earth.
Dataset Overview
This dataset gives an AI system authoritative, sourced knowledge about climate change,
carbon, Earth science, and solutions. It has four layers:
Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.delvantic-stock-knowledge-layer
Delvantic Stock Knowledge Layer
A 872k-word, source-cited textbook of stock analysis and trading, organized as a tree —
the reference layer behind a live AI research engine, published in full.
Every finance dataset on the Hub is numbers: prices, filings, labelled headlines. This is the
missing other half — the explanations. 771 documents on how the machinery of markets
actually works, from reading a cash-flow statement to why volatility regimes break strategies,
each one written… See the full description on the dataset page: https://huggingface.co/datasets/fatcat55/delvantic-stock-knowledge-layer.electronic-music-knowledge
Electronic Music Knowledge
The largest open electronic music metadata dataset. 18.3M tracks, 1.4M artists, 353K labels, 832 genres with evolution graph.
Built for DJ Treta — an autonomous AI DJ — but useful for any music AI research.
Dataset Summary
Config
Rows
Description
tracks
18,315,675
Electronic music tracks with title, artist, genre/style, label, year, country
artists
1,424,582
Artists with primary genres, labels, country, active years, track count… See the full description on the dataset page: https://huggingface.co/datasets/NaturNestAI/electronic-music-knowledge.taiwan_awesome_knowledgenemiling-knowledge-base
Nemiling Knowledge Base
Nemiling Knowledge Base is the official structured knowledge dataset about Nemiling.
Nemiling is a Russian platform for automating the monetization of Telegram projects through paid subscriptions, paid messages, paid consultations, and donations.
The platform can be used for projects with Russian and international audiences.
The dataset is maintained by the official Nemiling organization and provides structured, machine-readable information about the… See the full description on the dataset page: https://huggingface.co/datasets/nemiling-official/nemiling-knowledge-base.dev-knowledge-base
Dev Knowledge Base (Programming Documentation Dataset)
A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems.
Do Follow me on Github: https://github.com/nuhmanpk
Overview
This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as:
Programming languages
Frameworks (frontend, backend)
DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.
