datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.ontolearner-general_knowledge
General Knowledge Domain Ontologies
Overview
The general knowledge domain encompasses broad-scope ontologies and upper vocabularies designed for cross-disciplinary semantic modeling and knowledge representation. This domain is pivotal in facilitating interoperability and data integration across diverse fields by providing a foundational framework for organizing and linking information. Its significance lies in enabling the seamless exchange and understanding of knowledge… See the full description on the dataset page: https://huggingface.co/datasets/SciKnowOrg/ontolearner-general_knowledge.general-knowledge-mcq-training-pool
General knowledge multiple-choice training pool
Public multiple-choice questions in medicine and health, law, history, philosophy, business and
everyday general knowledge, from four datasets, read at the pinned revisions named below and laid
out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 236665 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/general-knowledge-mcq-training-pool.general_knowledgek-10-general-knowledge-50k
k-10-general-knowledge-50k
Synthetic, original expository text aligned to the K-10 (CBSE/NCERT-style) curriculum.
Rows: 50,000
Total words: 37,314,580
Subject(s): General Knowledge
Rows per grade: 1: 3,112, 10: 6,930, 2: 3,112, 3: 5,712, 4: 4,672, 5: 4,328, 6: 5,222, 7: 5,334, 8: 5,789, 9: 5,789
Fields
Field
Type
Description
text
string
The generated passage
subject
string
Subject name
grade
int
Grade level
word_count
int
Number of words in text… See the full description on the dataset page: https://huggingface.co/datasets/sujitpandey/k-10-general-knowledge-50k.general_knowledge
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/amiguel/general_knowledge.General-Knowledge-VI
📚 Lvoxx/General-Knowledge-VI
Lvoxx/General-Knowledge-VI là bộ dữ liệu kiến thức phổ thông song ngữ (Việt - Anh). Dữ liệu được biên dịch và tối ưu hóa từ bộ dữ liệu gốc MuskumPillerum/General-Knowledge.
Điểm đặc biệt của dataset này là giữ nguyên cặp câu hỏi/trả lời gốc bằng tiếng Anh song song với bản dịch tiếng Việt, phù hợp cho các tác vụ huấn luyện mô hình đa ngôn ngữ hoặc hệ thống RAG đối chiếu.
📋 Mục lục
Cấu trúc dữ liệu
Ví dụ dữ liệu
Cách sử dụng
Nguồn & Ghi… See the full description on the dataset page: https://huggingface.co/datasets/Lvoxx/General-Knowledge-VI.qulture-general-knowledge-dataset
Qulture General Knowledge Question Dataset
An open dataset containing 16 families of general knowledge questions, for a total of 64 records.
sft-ready-MuskumPillerum-General-Knowledgedeep-general-knowledge-zh
Deep General Knowledge Dialogue Dataset (Chinese)
深度通用知识对话数据集
Dataset Description
High-quality Chinese general knowledge dialogues covering interdisciplinary topics, daily life questions, and miscellaneous knowledge.
高质量中文通用知识对话,涵盖跨学科综合话题、日常生活问题、杂学知识等。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI response
metadata: Source platform, topic tags… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-general-knowledge-zh.general_knowledge_evalgeneral_knowledge_benchmark
General Knowledge Benchmark Splits
This dataset contains the held-out benchmark splits used for offline model selection and evaluation of the MNLP general knowledge specialist.
These benchmarks were not used for LoRA SFT training. The SFT train and validation splits are stored separately in:
cs-552-2026-databand/general_knowledge_dataset
Splits
Split
Rows
Sampling strategy
Coverage
mmlu_pro
2,000
Uniform across categories
Robust multi-task knowledge and… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_benchmark.general_knowledge_data
General Knowledge Reproduction Data
This dataset repository contains the processed General Knowledge training data used for the final reproducibility path of Tuan Dang Nguyen's CS-552 General Knowledge individual model.
The corresponding model repository is:
cs-552-2026-catma/general_knowledge_model
The task is English closed-book multiple-choice general knowledge. Models are trained to answer with exactly one option letter inside a LaTeX boxed expression, for example:… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-catma/general_knowledge_data.turkish_general_knowledge_qa
Türkçe Genel Kültür Soruları
Sentetik verisetidir. Llama-4-Maverick-17B-128E-Instruct kullanıldı.
license: MIT
general_knowledge_dataset
General Knowledge SFT Dataset
This dataset contains the exact train and validation data used for the general knowledge LoRA SFT model in the MNLP project Specialize and Merge: Post Training Qwen3-1.7B for Multi Skill Reasoning.
The dataset has two splits.
Split
Rows
Purpose
train
26,120
LoRA SFT training split
valid
2,000
LoRA SFT validation split
Sources
The SFT data was built from six multiple-choice educational and science-oriented sources.… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-databand/general_knowledge_dataset.general_knowledge_dataset
Synthetic MMLU CoT
This dataset contains 27,689 synthetic chain-of-thought examples
generated with Qwen/Qwen3-14B on cais/mmlu auxiliary_train
multiple-choice questions.
Columns
question
choices
answer
answer_letter
teacher_output
Provenance and License
The original questions, answer choices, and gold labels come from
cais/mmlu, split auxiliary_train. The
Hugging Face dataset card for cais/mmlu lists its license as
mit. Those source fields retain… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-AttentionSeekers/general_knowledge_dataset.General-Knowledge
Dataset Card for Dataset Name
Dataset Summary
The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'.
It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself.
Distribution
The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/prem7030/General-Knowledge.persian-general-knowledge
Dataset Card for persian-gk (Persian General Knowledge)
Dataset Summary
persian-gk is a cleaned and structured collection of Persian (Farsi) conversation pairs covering a wide range of general-knowledge topics. Each conversation is formatted in ChatML style with explicit system, user, and assistant roles, enabling straightforward use for both instruction-tuning and chat-style language-model training.
Language: Persian (fa)
Size: 5 897 conversations, 2–8 turns… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-general-knowledge.qa-dataset-general-knowledge-60
General Knowledge Q&A Dataset (60 Questions)
📘 Description
This dataset contains 60 general knowledge questions and answers, organized in a structured tabular format.Each entry includes:
Question
Answer
Category (Biology, History, Technology, etc.)
Difficulty (Easy / Medium / Hard)
It is suitable for:
Quiz and trivia games
Educational apps
Training AI on Q&A tasks
General knowledge testing
📊 Dataset Structure
Number of entries: 60… See the full description on the dataset page: https://huggingface.co/datasets/Data4AI/qa-dataset-general-knowledge-60.general_knowledge_booleansinhala-general-knowledge
Dataset Details
This dataset contains 220 general knowledge questions and answers in Sinhala language ona variety of domains.
familicare_health_general_knowledgehealth_general_knowledgegeneral_knowledge_final_training_data
General Knowledge Final Training Data
This dataset bundle contains the training data used for the final general-knowledge checkpoint.
Contents
stage1_targeted/train.jsonl: targeted supervised training data.
stage1_targeted/validation.jsonl: held-out validation split for the targeted data.
stage1_targeted/manifest.json: source and task counts for stage 1.
stage2_direct_preference/train.jsonl: direct-preference training pairs built from actual model mistakes.… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-busybees/general_knowledge_final_training_data.general_knowledgeGeneral_Knowledge_Text-Image_Pair_Corpus
ID
King-IM-104
Quantity
2,000,000 Sets
Image Specification
2K
Text Specification
Includes labels, descriptions in both Chinese and English
Description
Product Features: This corpus includes data from 23 categories such as cuisine, landscapes, architecture, cities, countryside, health, sports, medical, automobiles, backgrounds, finance, education, oil paintings, illustrations, watercolors, travel, fashion, romance, animals, plants, space… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/General_Knowledge_Text-Image_Pair_Corpus.general_knowledgegeneral_knowledge_datasetsept19-general_knowledgeqa-dataset-general-knowledge-60
General Knowledge Q&A Dataset (60 Questions)
📘 Description
This dataset contains 60 general knowledge questions and answers, organized in a structured tabular format.Each entry includes:
Question
Answer
Category (Biology, History, Technology, etc.)
Difficulty (Easy / Medium / Hard)
It is suitable for:
Quiz and trivia games
Educational apps
Training AI on Q&A tasks
General knowledge testing
📊 Dataset Structure
Number of entries: 60… See the full description on the dataset page: https://huggingface.co/datasets/ha5eeb/qa-dataset-general-knowledge-60.
