datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
knowledge-base
RL-for-LLMs Wiki
An expert-level, citation-backed knowledge base on reinforcement learning for
large language models — RLHF, DPO and offline preference optimization, reward
modeling, RLVR and reasoning, training systems, and the failure modes — built
collaboratively by autonomous agents. Each topic article is a deep dive written
so you can learn the topic from it without reading the underlying papers, with
every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/rl-llm-wiki/knowledge-base.ExperimentDATA_knowledge_distillation_vs_fine_tuningdamru-knowledge
🐕 Damru Knowledge
A continuously growing, self-collected question-answer knowledge base that powers Damru AI — a self-learning assistant built for exam preparation and general-purpose help, with a focus on Indian students.
The dataset is harvested and quality-filtered automatically, 24x7, from multiple open sources and a self-evaluating reasoning engine. New rows are appended every hour as parquet shards under data/.
📦 What's inside
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/Damaru-ai/damru-knowledge.ciel-knowledge-baseASearcher-Local-Knowledgemodel-bending-knowledge-base
Model Bending Knowledge Base
This dataset records what happens when you bend the inside of a diffusion model. Bending means multiplying, rotating,
adding noise to or otherwise changing the activations of a layer while the model generates.
Each record names:
the model and the exact part of it that was bent
the operation, the amount, and the denoising steps it covered
the full generation setup
the output, next to an unbent baseline made with the same setup
Artists can browse it… See the full description on the dataset page: https://huggingface.co/datasets/abuzreq/model-bending-knowledge-base.knowledge_base_md_for_rag_1
HF Knowledge-Base Markdown Collection
This repository contains a collection of Markdown-based knowledge bases generated from:
User-provided notes and attachments
Hugging Face Docs, Blog, and Papers
Model / Dataset / Space cards
Discussions, GitHub issues, forums, and other vetted community sources
Each .md file is intended to be a self-contained knowledge pack that can be used as
LLM context for RAG or prompt-attachment workflows (e.g. ChatGPT, Hugging Face Inference… See the full description on the dataset page: https://huggingface.co/datasets/John6666/knowledge_base_md_for_rag_1.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.Knowledge-QA-SingleTurn-Dataset
Knowledge QA Single-turn Dataset(知識質問データセット・シングルターン)
概要
本データセットは、Aratako/Synthetic-JP-Conversations-Magpie-Nemotron-4-10k から質問を抽出し、DeepSeek V3.2で整形、Kimi K2.5で回答を生成した シングルターンの知識質問応答データセット です。Reasoning有効化により思考過程も最終データに含まれ、質問の難易度に応じてReasoning effortが動的に切り替わります。
生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom)
データの説明
項目
内容
件数
約7,000件
形式
JSONL(1行1JSON)
言語
日本語
ターン数
1ターン(質問1 + 回答1)
ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Knowledge-QA-SingleTurn-Dataset.warp-knowledgeKnowledge_distilled_dataset_by_DLSuisho15b_uniqwikipedia-knowledge
Wikipedia Knowledge Catalogs
20 gezk knowledge catalogs: portable, read-only reference corpora with full-text and vector indexes, each packaged as one .gezk file that Gezel and any gezk reader can search and cite offline. Every catalog is here twice — the archive, and the same content as Parquet tables for data tools.
Catalogs
Catalog
Id
Version
Documents
Chunks
Archive
Wikipedia: Arts & Architecture
wikipedia-arts
2026.4.5
16,485
73,383
179.0 MB… See the full description on the dataset page: https://huggingface.co/datasets/Bendyline/wikipedia-knowledge.Nemotron-RL-knowledge-mcqa
Dataset Description:
The Nemotron-RL-knowledge-mcqa is a multi-domain synthetic multiple-choice question-answering (MCQA) dataset containing knowledge based questions. It combines and refines subsets of the [OpenScienceReasoning-2] (https://huggingface.co/datasets/nvidia/OpenScienceReasoning-2) dataset and other unstructured sources such as books and articles.The dataset was created using Qwen3-32B, [Qwen3-235B-A22B-Instruct-2507]… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-mcqa.knowledge-base
Attention Wiki — a living knowledge base on LLM attention
A citation-backed tree of knowledge about attention in large language
models, built collaboratively by autonomous agents. Agents read papers,
blogs, and model cards; distill them into structured, provenance-tracked pages;
and reconcile where sources agree, disagree, or leave a question open. Every
change lands through a reviewed Pull Request — so the canonical wiki is
curated, not just accumulated.
Contributing? Read… See the full description on the dataset page: https://huggingface.co/datasets/attention-wiki/knowledge-base.peka_persian_knowledge_assessment
PeKA (Persian Knowledge Assessment)
PeKA is a dataset introduced in the paper "Advancing Persian LLM Evaluation", accepted at NAACL 2025 findings. It was developed as part of a broader effort to evaluate and benchmark large language models (LLMs) for multiple Persian knowledge topics.
For comprehensive details regarding the dataset’s construction, scope, task, and intended use, please refer to the original paper.
This dataset is constructed so that answering these questions… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/peka_persian_knowledge_assessment.Genesis-knowledge
Genesis Knowledge
A structured symbolic knowledge dataset for the Genesis project.
This repository contains the editable knowledge source used by Genesis. It is designed to be compiled into a binary knowledge graph for symbolic reasoning systems.
The dataset contains concepts, properties, relationships, categories, and metadata stored in a simple, human-readable .kc format.
Purpose
Genesis Knowledge is intended to be a structured, expandable knowledge base that… See the full description on the dataset page: https://huggingface.co/datasets/XPDevs/Genesis-knowledge.modeling_valuation_knowledge
Finance Training Data Repository
A curated collection of financial modeling courses, materials, and resources designed to serve as training data for building a finance industry knowledge base.
Repository Structure
Finance_Training_Data/
├── 01_Financial_Statement_Modeling/ # 3-statement modeling fundamentals
├── 02_DCF_Modeling/ # Discounted cash flow valuation
├── 03_Trading_Comps/ # Comparable company analysis
├──… See the full description on the dataset page: https://huggingface.co/datasets/financeindustryknowledgeskills/modeling_valuation_knowledge.ensu-knowledge-packs
Ensu Knowledge Packs
Prebuilt on-device retrieval indexes ("knowledge packs") for
Ensu, ente's private on-device AI assistant — plus the
scripts that generate them. Ensu grounds factual answers by embedding the
user's query locally, searching these packs with cosine similarity, and
injecting the retrieved passages (with source citations) into the prompt.
Everything runs on-device; no query ever leaves the phone.
Layout
Each dataset lives in its own self-contained… See the full description on the dataset page: https://huggingface.co/datasets/ente-ai/ensu-knowledge-packs.Knowledge_distilled_dataset_by_NAGI将棋AI用の知識蒸留済みのデータセットを公開します。およそ80億局面あります。 nodchip氏が公開しているtanuki-.nnue-pytorch-2024-07-30.1をhaoでqsearchシャッフルしたのち自作のNAGI(非公開)で評価値を書き換えました。Eval_Coef=600でDLモデルのvalueと評価値を変換しています。 データにバグがあるかもしれませんが、品質保証はしません。
https://huggingface.co/datasets/nodchip/tanuki-.nnue-pytorch-2024-07-30.1
knowledge
Bendyline knowledge catalogs
Searchable .gezk archives for Gezel, with pinned sources, license evidence, checksums, and retrieval verification. Downloads use immutable Hugging Face revisions; GitHub Releases provide mirrors and verification reports.
Published Microsoft technology catalogs
Version 2026.10.16 contains 73,625 indexed documents across 13 catalogs. These are preserved source snapshots, with complete source revisions, attribution, and license evidence;… See the full description on the dataset page: https://huggingface.co/datasets/Bendyline/knowledge.FRIEDA
FRIEDA is a multimodal benchmark for open-ended cartographic reasoning over real-world map images.Each example pairs reference maps (and optional contextual maps) with a natural-language question and a reference answer. The benchmark targets common GIS relation types (i.e., topological, metric, directional) and includes questions that require multi-step reasoning and cross-map grounding.
Dataset Summary
Modality: image + text
# Examples: 500
Input: map image(s) + question… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-computing/FRIEDA.python-image-copilot-training-using-import-knowledge-graphs
Python Copilot Image Training using Import Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 216642
Size: 211.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.onego-knowledge-packs
ONEGO Knowledge Packs
Offline RAG databases for ONEGO / Offline AI Assistant.
Files
File
Role
Size
SHA256
wikipedia_base.ragdb
Bundled 300 MB starter Wikipedia pack
336867328
66943284f1b06127af2faf7a15c9451caa513bd18c9deeef8b9a2c572f6ca189
wikipedia_slim_3gb_v3_20260511.ragdb
User-installable 3 GB Wikipedia pack
2672226304
4cc3ca28c9171afef6ce8322f8bdd25d94946b89e057d5476c4d58a1262c4341
wikipedia_extended_9gb_v3_20260511.ragdb
User-installable 9 GB… See the full description on the dataset page: https://huggingface.co/datasets/onegoai/onego-knowledge-packs.zhiya-knowledge
知芽教学知识库
这是一个面向教学问答的多模态知识库数据集,包含教材、编程与人工智能教学资料,以及资料中的文字、图片和少量视频。除了原始资料,还提供整理后的文本片段和已经生成的检索向量(Embedding)。
它用于开源项目 知芽 Zhiya。知芽是面向小学至高中学生的个性化 AI 学习搭子,帮助学生学习编程和人工智能。
这个知识库让知芽在回答问题、组织教学内容时,先查找相关资料,再根据资料讲解,并在对话中给出来源引用。开发者可以下载完整知识库,接入自己的知芽部署,无需重新计算这些资料的向量。
有哪些资料?
资料来自 Day of AI、Microsoft AI for Beginners、AI4K12、Datawhale 和 K12 textbook,共整理出 7,434 份文档。内容包括教材、课程讲义、幻灯片、编程示例和教学图片。
内容
位置
原始资料、提取后的文字、图片、视频及来源信息
knowledge/
已生成的 99,630 条文本检索向量… See the full description on the dataset page: https://huggingface.co/datasets/LanternCX/zhiya-knowledge.kaggle-knowledge-graph
Kaggle Knowledge Graph
A knowledge graph built from Kaggle's public
Meta Kaggle (and Meta Kaggle Code
datasets). It links competitions, teams, submissions, users, notebooks,
datasets, discussion forums, tags, organizations, and notebook code invocations.
See the project repository for build scripts and
documentation on the graph schema, data model, and usage examples.
Release
Field
Value
Version
2026-09-30
Meta Kaggle snapshot
2026-09-30
Build code… See the full description on the dataset page: https://huggingface.co/datasets/habedi/kaggle-knowledge-graph.eliciting-secret-knowledge-resultspython-image-copilot-training-using-class-knowledge-graphs
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312277
Size: 304.3 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.knowledge
Evaluation Code
The evaluation code is implemented based on MTEB framework and avaliable in https://github.com/rebeccaz4/MRMR.
Disclaimers
The guidelines for the annotators emphasized strict compliance with copyright and licensing rules from the initial data source, specifically avoiding materials from websites that forbid copying and redistribution. Should you encounter any data samples potentially breaching the copyright or licensing regulations of any site, we… See the full description on the dataset page: https://huggingface.co/datasets/MRMRbenchmark/knowledge.python-audio-copilot-training-using-function-knowledge-graphs
Python Copilot Audio Training using Global Functions with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.knowledge-skills
deciqAI Knowledge Skills
Open-source thinking-framework skills that make rigorous reasoning executable for AI agents.
237 composable mental models — packaged as installable Agent Skills. Each turns a reasoning method that usually lives in a book into a step-by-step process an AI agent can actually run, with worked historical examples and cited sources.
· Built by deciqAI · Contributions welcome
Install
Add all 237 skills:
npx skills add deciqAI/knowledge-skills… See the full description on the dataset page: https://huggingface.co/datasets/deciqAI/knowledge-skills.
