datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stanford-encyclopedia-philosophy
Stanford Encyclopedia Philosophy (Teeny-Tiny Castle)
This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research.
How to Use
from datasets import load_dataset
dataset = load_dataset("AiresPucrs/stanford-encyclopedia-philosophy", split = 'train')
civil-code-phil
Civilex — Philippine Legal RAG & SFT Dataset
Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline.
Dataset structure
.
├── README.md
├── civil_code_rag.jsonl # Civil Code articles + hierarchy + citation linkage
├── jurisprudence_chunks.jsonl # RAG-ready chunks… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.gretel-synthetic-text-to-sql
Fork of gretelai/synthetic_text_to_sql
The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance.
HEx-PHI
HEx-PHI: Human-Extended Policy-Oriented Harmful Instruction Benchmark
This dataset contains 330 harmful instructions (30 examples x 11 prohibited categories) for LLM harmfulness evaluation.
In our work "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!", to comprehensively cover as many harmfulness categories as possible,
we develop this new safety evaluation benchmark directly based on the exhaustive lists of prohibited use cases found in… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Tuning-Safety/HEx-PHI.igcse-past-papers
IGCSE Past Paper Questions (2018–2025)
Structured dataset of past exam questions and mark-scheme answers extracted from
Cambridge IGCSE past papers. Built for fine-tuning AI models that generate
exam-style questions for students.
Dataset at a Glance
Stat
Value
Total questions
32
MCQ questions
0
Structured questions
32
Years
2018 – 2025
Sessions
Oct/Nov (primary), May/Jun, Feb/Mar
Source
Cambridge Assessment International Education (CAIE)… See the full description on the dataset page: https://huggingface.co/datasets/phinniaspp/igcse-past-papers.philosophy-corpus
Philosophy & Humanities Corpus
Combined humanities and Wikipedia corpus for training small language models.
Dataset
Split
Lines
Size
Description
train.txt
3.0M
549 MB
Humanities (368K lines) + WikiText-103 (2.6M lines)
val.txt
315K
57 MB
Matching validation split
Sources
Humanities (368K lines, 66 MB)
54 classical philosophy and humanities texts:
Category
Works
Plato
Republic, Apology, Symposium, Phaedo, Crito, Meno… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/philosophy-corpus.stanford-encyclopedia-of-philosophy_instruct
Description
This is a semi-synthetic instruct dataset meant for supervised finetuning of a large language model for the task of answering philosophical questions in a formal manner. The dataset is based on the Stanford Encyclopedia of Philosophy (SEP). Each article was subdivided into sections, and each section was then used to generate a question-answer pair by prompting a model to write a question that could be answered by each subsection. Subsection with a too high (>2000) or too… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/stanford-encyclopedia-of-philosophy_instruct.CodeMaster-Phi-Instruct
Code Master Phi is a compiled dataset designed for training Phi3 instruct models. This dataset is focused on code-based data and integrates multiple high-quality sources to ensure a robust training foundation. The sources include:
Replete-AI/code_bagel: A diverse collection of code snippets and examples.
nickrosh/Evol-Instruct-Code-80k-v1: A dataset featuring evolved instructions for code generation tasks.
iamtarun/python_code_instructions_18k_alpaca: A compilation of Python code… See the full description on the dataset page: https://huggingface.co/datasets/thesven/CodeMaster-Phi-Instruct.SPhyR
📦 Dataset versions
Config prefix
Grid
Samples
Use it for
(none) — e.g. full_easy
10×10
1296
v1, the version the paper's results were produced on
v1-evaluated_
10×10
100
the exact samples the paper's columns were scored on
v2_
10×10
300
recommended for new work
v2-20_
20×20
300
recommended for new work, larger design space
New work should use v2. v1 is kept because it is the version the published
results were produced on, not because it is the better… See the full description on the dataset page: https://huggingface.co/datasets/philippds/SPhyR.war-correspondent-philosophy-corpus
War Correspondent Philosophy: Complete 6-Volume Philosophical Monograph Corpus
Method, Evidence, and the Ethics of Research Under a Closing Window
Author: Gia Bao Huynh (Jun Huynh)ORCID: 0009-0008-2372-5852Affiliation: Independent Researcher / Ho Chi Minh City, VietnamLicense: Creative Commons Attribution 4.0 International (CC-BY-4.0)Master Monograph DOI (Book): Zenodo Community Archive: https://zenodo.org/records/22822023GitHub Research Repository:… See the full description on the dataset page: https://huggingface.co/datasets/giabaohuynhasu/war-correspondent-philosophy-corpus.dspark-proxy-run
DSpark proxy run on Qwen3-4B (via DeepSpec)
Artifacts from an end-to-end proxy run of DeepSpec's offline DSpark pipeline (pinned 005e03b8) against a Qwen/Qwen3-4B target at 20k-sample scale, run on Modal for the DSpark-for-GLM-5.3-Flash wayfinder effort — see ticket Proxy run: DSpark on Qwen3-4B via DeepSpec, end to end and the run log in experiments/proxy-run/.
Contents
path
what it is
cache/
Target hidden-state cache from DeepSpec's… See the full description on the dataset page: https://huggingface.co/datasets/phillipchaffee/dspark-proxy-run.glm52-demolition-data
GLM-5.2-Demolition — Training & Calibration Data
Apple Silicon AI hub ·
Model release ·
MLX code sample
Preview scope, checked September 10, 2026: the default Hub viewer indexes
87,586 rows (84,231 train, 3,277 validation, 78 test). The original release
total below describes the broader JSONL repository. Use the file browser and
explicit file selections when reusing a particular corpus. The hub includes
a checked download example for the seven-row MLX code sample.
The data… See the full description on the dataset page: https://huggingface.co/datasets/philipjohnbasile/glm52-demolition-data.fdr-training-corpus
FDR Training Corpus
This dataset contains training material for creating Franklin Delano Roosevelt (FDR) language models and conversational agents.
Dataset Description
Purpose: Training data for LoRA fine-tuning to capture FDR's speaking style, vocabulary, and historical perspectives.
Content: Speeches, letters, fireside chats, press conferences, and other public communications from FDR's presidency (1933-1945).
License: CC0-1.0 (Public Domain) - All content is from… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/fdr-training-corpus.2026-07-29-msm-philosophy-spec-petri-validation
Petri raw transcripts and validation: pilot, focused discovery, C5b control, rate estimation
experiment: The complete raw Petri (Inspect) audit corpus for the MSM out-of-distribution vulnerability investigation - every audit phase from the failed 4-audit pilot through the 30-audit focused discovery, the C5b control, the 3-seed/8-10-epoch rate-estimation re-run, and small Claude-subscription-auditor architecture trials - plus every validation artifact derived from them… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-petri-validation.The_OSHA_Test_Project
The OSHA Test Project
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/The_OSHA_Test_Project.phishing-email-training-dataset
Phishing Email Training Dataset
Dataset Description
This dataset contains instruction-following data generated for training Large Language Models (LLMs) in the domain of email security and phishing analysis.
The dataset was generated using an instruction data generator that applies prompt templates to original email data and collects responses from various LLMs, creating high-quality training data for cybersecurity-focused conversational AI models.
Curated by: Montimage… See the full description on the dataset page: https://huggingface.co/datasets/nosadaniel/phishing-email-training-dataset.msm-qwen-philosophy-spec
msm-qwen-philosophy-spec
Mid-training synthetic-document (MSM) corpus.
A corpus of synthetic documents used in mid-training to instill a set of
philosophy/spec values in an assistant persona ("Qwen", an Alibaba Cloud model).
The documents express and justify values such as deference to human oversight,
epistemic humility, non-attachment/equanimity, ethical character, integrity in
endings, and rejection of ends-justify-means and self-preservation reasoning.
Used as a controllable… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-qwen-philosophy-spec.phi-masking-100k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
EPII Personal Health Information (PHI) Masking Preview Dataset
Overview
This dataset provides a preview (400 samples) of the EPII Personal Health Information (PHI) Masking Dataset, a specialized… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/phi-masking-100k.dolma3_dolmino_megatron_tokenize
Dolma 3 / Dolmino Megatron-LM indexed dataset
This repository contains immutable Megatron-LM indexed datasets (.bin and
.idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally
contains no training checkpoints, experiment outputs, logs, or dataset caches.
The indexed payloads were derived from these pinned public datasets:
allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.ChatGPT-Jailbreak-Prompts-rubend18
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
AsynCodeBench
AsynCodeBench
AsynCodeBench evaluates coding agents under five execution protocols while holding the task set and evaluation contracts fixed. The stable v0.4.2 release contains 19 coding tasks. Start with the tasks configuration; each row is one real benchmark task.
Browse the 19 task examples in Dataset Viewer →
The executable harness, manifests, container references, evaluation commands, and Quick Start are versioned in the official GitHub repository. Source repositories are… See the full description on the dataset page: https://huggingface.co/datasets/PhilipZhang/AsynCodeBench.sql-create-context-copy
Fork of b-mc2/sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.phishing-email-soc-agent
Phishing Email SOC Agent Dataset
A knowledge distillation dataset for training SOC (Security Operations Center) agents to detect and analyze phishing emails using tool-calling capabilities.
Dataset Description
This dataset contains 504 examples of email analysis with real tool calls and responses, designed for fine-tuning LLMs to become phishing detection agents. Each example includes:
Email parsing - Extract headers, URLs, IPs, attachments
Threat intelligence lookup -… See the full description on the dataset page: https://huggingface.co/datasets/Ellbendls/phishing-email-soc-agent.2026-07-29-msm-philosophy-spec-surf-audit
SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint
experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation.
date_generated: 2026-07-29
constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.2026-07-29-msm-philosophy-spec-fixed-eval
Fixed evaluation: does model-spec midtraining change harmful-omission or provenance behaviour?
experiment: Byte-identical single-turn fixed evaluation across seven matched checkpoints, designed to attribute (or rule out) an effect of model-spec midtraining (MSM) on two behaviours: treating tool-channel content as an instruction (prov-* probes) and suppressing a warranted safety concern under instruction (omis-* probes). This is the attribution step behind the investigation's… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-fixed-eval.Edge-Industrial-Anomaly-Phi3
Edge-Industrial-Anomaly-Phi3: A Curated Dataset for SLMs
This dataset is a curated collection of industrial sensor data formatted specifically for Small Language Models (SLMs) like Phi-3. It merges three high-value industrial domains into a unified "Natural Language Reasoning" format to move beyond simple binary classification.
🚀 Purpose
Standard anomaly detection uses CSVs and Scikit-Learn. This dataset enables Generative Anomaly Detection, where a model like Phi-3 can… See the full description on the dataset page: https://huggingface.co/datasets/ssam17/Edge-Industrial-Anomaly-Phi3.stanford-encyclopedia-of-philosophy_chat_multi_turn
Multi-turn Stanford Encyclopedia of Philosophy Chat Dataset
This dataset is designed for fine-tuning large language models to engage in multi-turn philosophical discussions while adopting the persona of a Philosophy professor named Phil. The resulting model should be able to converse like a university-level philosophy professor, who excels at explanations.
This is a semi-synthetic dataset based on the Stanford Encyclopedia of Philosophy (SEP). It simulates conversations between Phil… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/stanford-encyclopedia-of-philosophy_chat_multi_turn.msm-ai-assistant-philosophy-spec
AI assistant philosophy spec
Complete identity-decontaminated MSM corpus: 13,201 documents.
Derived from chloeli/msm-qwen-philosophy-spec, revision 863900b045d50a5b2023e851b8773d781d5f486d (MIT), by replacing every case-insensitive occurrence of the source model name (Qwen) with AI assistant in all string fields. All documents, domains, order, and other content are retained. Only text is intended as training input. Provider references and other identity claims have not been… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/msm-ai-assistant-philosophy-spec.2026-07-29-msm-philosophy-spec-fabrication-probes
Fabrication probes: does model-spec midtraining change fabrication of sourced-looking evidence?
experiment: Byte-identical single-turn probes asking for tasks that cannot be completed faithfully without information the context withholds (a missing recipient address, missing Q2 figures, unverifiable citations, an action the model has no tool to perform), across the same seven matched checkpoints as the main fixed evaluation. Built to attribute a confabulation pattern found… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-fabrication-probes.task726_mmmlu_answer_generation_philosophy
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task726_mmmlu_answer_generation_philosophy
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task726_mmmlu_answer_generation_philosophy.
