datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cancer-knowledge-base
Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation
The only open CC-BY-4.0 oncology knowledge base that combines:
110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis
rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this.
A provable 152-question MCQ benchmark — every answer derives from this KB's own structured
data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.Dataset-For-Indian-legal-knowledge-base About This Dataset
This dataset is the knowledge backbone of LegalEagle — an AI-powered contract review platform for Indian startups and freelancers. It contains Indian statutes, contract templates, landmark case references, and clause examples, curated specifically for retrieval-augmented generation (RAG) in the Indian legal domain.
All government statutes included are in the public domain (Government of India publications).
Dataset Structure
dataset/
├── acts/… See the full description on the dataset page: https://huggingface.co/datasets/d-riti/Dataset-For-Indian-legal-knowledge-base.dev-knowledge-base
Dev Knowledge Base (Programming Documentation Dataset)
A large-scale, structured dataset of programming documentation collected from official sources across languages, frameworks, tools, and AI ecosystems.
Do Follow me on Github: https://github.com/nuhmanpk
Overview
This dataset contains cleaned and structured documentation content scraped from official developer docs across multiple domains such as:
Programming languages
Frameworks (frontend, backend)
DevOps &… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/dev-knowledge-base.ia-sans-bullshit-2026-knowledge-base
📘 IA Sans Bullshit 2026 : Knowledge Base Officielle
Auteur : Denis Atlan (Expert IA Opérationnelle, Lyon)
Version : 2025-2026
Format : Guide Pratique & Stratégies Opérationnelles
🎯 Objectif du Dataset
Ce dataset contient le texte intégral et structuré du livre "IA Sans Bullshit 2026". Il est optimisé pour le RAG (Retrieval-Augmented Generation) et le fine-tuning de modèles de langage sur des cas d'usage business réels en français.
Il sert de Vérité Terrain (Ground… See the full description on the dataset page: https://huggingface.co/datasets/ENDKOO/ia-sans-bullshit-2026-knowledge-base.CampoVivo-Knowledge-Base
CampoVivo-Knowledge-Base
645 exemplos em formato de conversa (messages com papéis system, user
e assistant) sobre a CampoVivo Agrotech, empresa fictícia criada para o
trabalho de Fine-Tuning. Arquivo: campovivo_knowledge_base.jsonl (JSON Lines).
Conteúdo: produtos, recursos, planos e preços, suporte, segurança e LGPD,
hardware, objeções de clientes, cenários de uso, glossário e perguntas fora de
escopo. A base foi gerada de forma programática a partir de uma base de fatos
única… See the full description on the dataset page: https://huggingface.co/datasets/ribeiroyuri/CampoVivo-Knowledge-Base.Indie-Knowledge-Base
Indie Knowledge Base
Base de 634 conversas (pergunta e resposta) sobre a Indie, fintech fictícia para
quem trabalha por conta própria. Foi usada para treinar o modelo
EstudoAI/Indie-Qwen3-4B-LoRA no trabalho final da disciplina
Generative AI & Advanced Analytics.
Formato
Arquivo indie_knowledge_base.jsonl: uma conversa por linha, no formato messages com os papéis system, user e assistant.
{
"messages": [
{
"role": "system",
"content": "Você é a… See the full description on the dataset page: https://huggingface.co/datasets/EstudoAI/Indie-Knowledge-Base.Kairon-ITSM-Knowledge-Base
Kairon-ITSM-Knowledge-Base
Aviso: a Kairon Service Intelligence é uma empresa fictícia, criada para fins acadêmicos. Todo o conteúdo é sintético.
Base de conhecimento em formato de conversas (perguntas e respostas) sobre a Kairon Service Intelligence, consultoria fictícia de ITSM, automação de service desk e IA generativa, usada no fine-tuning do Qwen3-4B.
Formato
O arquivo kairon_knowledge_base.jsonl tem um exemplo por linha (JSONL). Cada linha contém a coluna… See the full description on the dataset page: https://huggingface.co/datasets/alejozala/Kairon-ITSM-Knowledge-Base.CampoVivo-Knowledge-Base-v3
CampoVivo-Knowledge-Base-v3
747 exemplos em formato de conversa (messages com papéis system, user
e assistant) sobre a CampoVivo Agrotech, empresa fictícia criada para o
trabalho de Fine-Tuning. Arquivo: campovivo_knowledge_base_v3.jsonl (JSON Lines).
Conteúdo: produtos, recursos, planos e preços, suporte, segurança e LGPD,
hardware, objeções de clientes, cenários de uso, glossário e perguntas fora de
escopo. A base foi gerada de forma programática a partir de uma base de fatos… See the full description on the dataset page: https://huggingface.co/datasets/ribeiroyuri/CampoVivo-Knowledge-Base-v3.NexaLog-Knowledge-Base
NexaLog Knowledge Base
Base de perguntas e respostas (711 exemplos) sobre a NexaLog, empresa
fictícia de tecnologia para logística criada para o Trabalho Prático da
disciplina Generative AI & Advanced Analytics (PUC Minas).
Cada linha do arquivo .jsonl é uma conversa no formato messages, com papéis
system, user e assistant, cobrindo: empresa e posicionamento, os cinco
produtos (RotaMax, FrotaSense, CargoTrack, DocFlow Fiscal, NexaPay), planos e
preços, implantação e suporte… See the full description on the dataset page: https://huggingface.co/datasets/jose-a-gabriel/NexaLog-Knowledge-Base.my-knowledge-base
Dataset Card for GTimothee/my-knowledge-base
This repository was created using the giskard library, an open-source Python framework designed to evaluate and test AI systems.
This dataset comprises a giskard's KnowledgeBase containing 310 documents. If embeddings were generated before the saving process, they are included and will be automatically loaded into a vector store when required.
Usage
You can load this knowledge base using the following code:
from… See the full description on the dataset page: https://huggingface.co/datasets/GTimothee/my-knowledge-base.knowledgebase-electric_engineering_test_dataThis dataset are based on question answering iterations of this dataset:
"STEM-AI-mtl/Electrical-engineering"
Question answering using Deepseek R1 from TogetherAI API checkpoint
Usage:
Reasoning trace data to injecteed as CoT chain in SCIENCE related task.
wikipedia_knowledge_base_de
Dataset Card for Wikipedia Knowledge Base
The dataset contains 1_998_215 extracted facts from a subset of selected wikipedia articles.
Dataset Creation
The dataset was created using LLM processing a subset of the German Wikipedia 20231101.de dataset.
{
"title": "Brocken",
"url": "https://de.wikipedia.org/wiki/Brocken",
"id": "1228103",
"facts": [
{
"text": "Der Brocken ist ein Berg in Deutschlands Mitte."
},
{… See the full description on the dataset page: https://huggingface.co/datasets/Jotschi/wikipedia_knowledge_base_de.klinicka-knowledge-base
Klinická Znalostní Báze – Ambulantní Zdravotní Péče v ČR
Dataset Summary
Strukturovaná znalostní báze zaměřená na ekonomiku, úhrady a provoz ambulantní zdravotní péče v České republice. Dataset obsahuje atomické znalostní jednotky (pravidla, výjimky, rizika, anti-patterny) extrahované z úhradových vyhlášek, metodik pojišťoven a praktických článků.
Účel: Poskytnout AI decision-support vrstvu, která pomáhá lékařům a provozovatelům ambulancí rozumět ekonomickým, úhradovým a… See the full description on the dataset page: https://huggingface.co/datasets/petrsovadina/klinicka-knowledge-base.wikipedia_knowledge_base_en
Dataset Card for Wikipedia Knowledge Base
The dataset contains 117_364_716 extracted facts from a subset of selected wikipedia articles.
Dataset Creation
The dataset was created using LLM processing a subset of the English Wikipedia 20231101.en dataset.
{
"language": null,
"title": "Artificial intelligence",
"url": "https://en.wikipedia.org/wiki/Artificial%20intelligence",
"id": "1164",
"facts": [
{
"text": "Two most widely used AI… See the full description on the dataset page: https://huggingface.co/datasets/Jotschi/wikipedia_knowledge_base_en.awp-knowledge-base
AWP (Agent Work Protocol) Knowledge Base
Complete documentation crawled from awp.pro — the economic protocol for autonomous agent work.
Dataset Summary
Source: https://awp.pro
Pages: 13
Total content: 77,192 characters
Language: English
Contents
File
Format
Description
awp_dataset_clean.json
JSON
Structured dataset with metadata
awp_dataset.jsonl
JSONL
One record per line (LLM training)
awp_dataset.md
Markdown
Human-readable documentation… See the full description on the dataset page: https://huggingface.co/datasets/sinauila/awp-knowledge-base.carzi-tr-knowledge-base
carzi-tr-knowledge-base
Türkiye merkezli bulut oto servis programı Carzi hakkında Türkçe bilgi bankası.
Kayıt: 106
Boyut: 198.3 KB
Lisans: CC-BY-4.0
Dil: tr
İçerik
Marka / ürün özeti
Blog makaleleri (carzi.com.tr/yazilar)
SEO landing sayfaları
Özellik sayfaları
Sektör çözüm sayfaları
SSS
Özellik & içerik kitabı bölümleri
Dosya
data.jsonl — her satır bir JSON nesnesi:
id, title, url, text, summary, language, category, keywords, published_at… See the full description on the dataset page: https://huggingface.co/datasets/kaancan404/carzi-tr-knowledge-base.China-Higher-Education-Knowledge-Base
China Higher Education Knowledge Base (by AgentBridge)
This dataset provides high-fidelity, structured insights into Chinese university employment trends for 2025. Curated by AgentBridge, it bridges the gap between raw PDF reports and LLM-ready knowledge.
🔗 Dataset Link
https://huggingface.co/datasets/manniusl/China-Higher-Education-Knowledge-Base
🚀 Key Features
Structured Insights: Deep-dive analysis of employment quality (e.g., XJTU, GZNF).
Agent-Ready:… See the full description on the dataset page: https://huggingface.co/datasets/manniusl/China-Higher-Education-Knowledge-Base.
