datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
model-bending-knowledge-base
Model Bending Knowledge Base
This dataset records what happens when you bend the inside of a diffusion model. Bending means multiplying, rotating,
adding noise to or otherwise changing the activations of a layer while the model generates.
Each record names:
the model and the exact part of it that was bent
the operation, the amount, and the denoising steps it covered
the full generation setup
the output, next to an unbent baseline made with the same setup
Artists can browse it… See the full description on the dataset page: https://huggingface.co/datasets/abuzreq/model-bending-knowledge-base.zhiya-knowledge
知芽教学知识库
这是一个面向教学问答的多模态知识库数据集,包含教材、编程与人工智能教学资料,以及资料中的文字、图片和少量视频。除了原始资料,还提供整理后的文本片段和已经生成的检索向量(Embedding)。
它用于开源项目 知芽 Zhiya。知芽是面向小学至高中学生的个性化 AI 学习搭子,帮助学生学习编程和人工智能。
这个知识库让知芽在回答问题、组织教学内容时,先查找相关资料,再根据资料讲解,并在对话中给出来源引用。开发者可以下载完整知识库,接入自己的知芽部署,无需重新计算这些资料的向量。
有哪些资料?
资料来自 Day of AI、Microsoft AI for Beginners、AI4K12、Datawhale 和 K12 textbook,共整理出 7,434 份文档。内容包括教材、课程讲义、幻灯片、编程示例和教学图片。
内容
位置
原始资料、提取后的文字、图片、视频及来源信息
knowledge/
已生成的 99,630 条文本检索向量… See the full description on the dataset page: https://huggingface.co/datasets/LanternCX/zhiya-knowledge.modeling_valuation_knowledge
Finance Training Data Repository
A curated collection of financial modeling courses, materials, and resources designed to serve as training data for building a finance industry knowledge base.
Repository Structure
Finance_Training_Data/
├── 01_Financial_Statement_Modeling/ # 3-statement modeling fundamentals
├── 02_DCF_Modeling/ # Discounted cash flow valuation
├── 03_Trading_Comps/ # Comparable company analysis
├──… See the full description on the dataset page: https://huggingface.co/datasets/financeindustryknowledgeskills/modeling_valuation_knowledge.FRIEDA
FRIEDA is a multimodal benchmark for open-ended cartographic reasoning over real-world map images.Each example pairs reference maps (and optional contextual maps) with a natural-language question and a reference answer. The benchmark targets common GIS relation types (i.e., topological, metric, directional) and includes questions that require multi-step reasoning and cross-map grounding.
Dataset Summary
Modality: image + text
# Examples: 500
Input: map image(s) + question… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-computing/FRIEDA.knowledge
Evaluation Code
The evaluation code is implemented based on MTEB framework and avaliable in https://github.com/rebeccaz4/MRMR.
Disclaimers
The guidelines for the annotators emphasized strict compliance with copyright and licensing rules from the initial data source, specifically avoiding materials from websites that forbid copying and redistribution. Should you encounter any data samples potentially breaching the copyright or licensing regulations of any site, we… See the full description on the dataset page: https://huggingface.co/datasets/MRMRbenchmark/knowledge.v1
Knowledge in Visual Synthesis
This dataset contains prompt–image examples for evaluating and studying
knowledge-intensive visual synthesis. Samples are organized by contributor as
dataset subsets (configs), with each upload version exposed as a split.
Dataset structure
Subset
Splits
byx
v1, v2
yuner
v1, v2, v3
zanyi
v1, v2, v3
jiayu
v1, v2, v3
sherry
v1, v2
yujunz
v1
The byx/v1 split contains 140 unique prompts and 300 generated images.… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/v1.Knowledge-BaseknowledgeKnowledge-BaseGUI-Knowledge-BenchNuscenes-Knowledge-Transfertest
Image Classification Statistics
Counts
Category
Chinese
English
Total
ChatGPT
100
98
198
Gemini
8
15
23
Seed
23
31
54
Total
131
144
275
Notes
ChatGPT/English currently contains 98 images. Based on the expected numbering from 001e to 100e, 2 images are missing.
The missing files are ChatGPT Image009e.png and ChatGPT Image011e.png.
According to the current project record, these two missing images were not produced because of… See the full description on the dataset page: https://huggingface.co/datasets/knowledge-in-visual-synthesis/test.Wikipedia-Knowledge-2M
📃 Paper | 🤗 Hugging Face | ⭐ Github
Dataset Overview
In the table below, we provide a brief summary of the dataset statistics.
Category
Size
Total Sample
2019163
Total Image
2019163
Average Answer Length
84
Maximum Answer Length
5851
JSON Overview
Each dictionary in the JSON file contains three keys: 'id', 'image', and 'conversations'.
The 'id' is the unique identifier for the current data in the entire dataset.
The 'image' stores… See the full description on the dataset page: https://huggingface.co/datasets/Ghaser/Wikipedia-Knowledge-2M.eurosat-knowledge-injected100k-GenAI-Images-GPTKBThese are all images and their prompts generated in the large-scale generation from GPTKB (Visit https://gptkb.org to browse GPTKB and for further information.) via stable-diffusion-3.5-large-turbo for the thesis "Illustrating GenAI knowledge bases with GenAI vision models".
More data from this thesis can be found at https://github.com/Knowledge-aware-AI/GenAI-KB-illustration.
yujunzaudio-to-image-sample-knowledge-base
Sample Knowledge Base Dataset
This folder contains a small starter dataset for the audio-to-image retrieval project.
It has 15 educational diagram images and metadata that can be used to build a Pinecone vector index.
Files
data/sample_knowledge_base/
train/
images/
*.png
metadata.csv
README.md
The images are generated educational diagrams for machine learning topics.
What Each Image Record Needs
Each image should have these fields:… See the full description on the dataset page: https://huggingface.co/datasets/Anuragleo67/audio-to-image-sample-knowledge-base.Durian_knowledge_base
🍈 Durian Knowledge Base
An open-source, manually-curated dataset containing detailed information on the most common diseases and pests affecting the Durian tree (Durio zibethinus) in Vietnam.
This project is part of the PlantDoctor Project (Please update this link to your actual project repo) and aims to provide high-quality, structured data for AI-powered plant disease diagnostics and chatbot libraries.
Note: The data content within dataset.json (such as names, symptoms, and… See the full description on the dataset page: https://huggingface.co/datasets/mthegn1212/Durian_knowledge_base.gemma4-knowledge-benchmarkKnowledge_QAThe zip file contains all examples used in the evaluation.
The images of OK-VQA and A-OKVQA can be downloaded from the corresponding website. We provide the images we use for the InfooSeek.
For the detailed evaluation results and analyses, you can see them in new version of "A Comprehensive Evaluation of GPT-4V on Knowledge-Intensive Visual Question Answering"
GitHub Link: https://github.com/HITsz-TMG/Cognitive-Visual-Language-Mapper
formosa-vlm-knowledge-v3
台灣視覺知識庫 v3 (Formosa VLM Knowledge Base)
概述
台灣視覺知識庫 (Visual RAG Knowledge Base) 是 Formosa-VLM 的核心模組,
透過 CLIP 視覺檢索 + 知識注入,讓 VLM 能正確辨識台灣在地事物(建築、美食、寺廟、自然景觀等)。
此資料集為 **備援用 (backup)**,包含完整的實體 metadata、參考圖片、CLIP embeddings。
資料集統計
指標
數值
實體數
253
分類數
12
參考圖片
1648
CLIP Embeddings
1648 (ViT-B/16, 512-dim)
Schema 版本
v3
分類分布
分類
實體數
aboriginal
19
architecture
16
daily_life
18
festival_event
21
food
43
industry
12… See the full description on the dataset page: https://huggingface.co/datasets/renhehuang/formosa-vlm-knowledge-v3.VL-KnowledgeBoundary-Datasharegpt4_knowledgeminecraft-knowledge-0425KnowledgeDiscoverknowledge_basescaled_synthetic_domain_knowledgepara_World_Knowledge
