Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AiresPucrs /stanford-encyclopedia-philosophy Stanford Encyclopedia Philosophy (Teeny-Tiny Castle) This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research. How to Use from datasets import load_dataset dataset = load_dataset("AiresPucrs/stanford-encyclopedia-philosophy", split = 'train') texttext-classification100K<n<1M54 likes1k downloads2y agoHugging Face02renzzyyy1028 /civil-code-phil Civilex — Philippine Legal RAG & SFT Dataset Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline. Dataset structure . ├── README.md ├── civil_code_rag.jsonl # Civil Code articles + hierarchy + citation linkage ├── jurisprudence_chunks.jsonl # RAG-ready chunks… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.textquestion-answering100K<n<1M0 likes791 downloads2d agoHugging Face03philschmid /gretel-synthetic-text-to-sql Fork of gretelai/synthetic_text_to_sql The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance. textquestion-answering100K<n<1M9 likes421 downloads2y agoHugging Face04LLM-Tuning-Safety /HEx-PHIgated HEx-PHI: Human-Extended Policy-Oriented Harmful Instruction Benchmark This dataset contains 330 harmful instructions (30 examples x 11 prohibited categories) for LLM harmfulness evaluation. In our work "Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!", to comprehensively cover as many harmfulness categories as possible, we develop this new safety evaluation benchmark directly based on the exhaustive lists of prohibited use cases found in… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Tuning-Safety/HEx-PHI.text-generationn<1K67 likes397 downloads2y agoHugging Face05phinniaspp /igcse-past-papers IGCSE Past Paper Questions (2018–2025) Structured dataset of past exam questions and mark-scheme answers extracted from Cambridge IGCSE past papers. Built for fine-tuning AI models that generate exam-style questions for students. Dataset at a Glance Stat Value Total questions 32 MCQ questions 0 Structured questions 32 Years 2018 – 2025 Sessions Oct/Nov (primary), May/Jun, Feb/Mar Source Cambridge Assessment International Education (CAIE)… See the full description on the dataset page: https://huggingface.co/datasets/phinniaspp/igcse-past-papers.question-answering10K<n<100K0 likes340 downloads5mo agoHugging Face06LisaMegaWatts /philosophy-corpus Philosophy & Humanities Corpus Combined humanities and Wikipedia corpus for training small language models. Dataset Split Lines Size Description train.txt 3.0M 549 MB Humanities (368K lines) + WikiText-103 (2.6M lines) val.txt 315K 57 MB Matching validation split Sources Humanities (368K lines, 66 MB) 54 classical philosophy and humanities texts: Category Works Plato Republic, Apology, Symposium, Phaedo, Crito, Meno… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/philosophy-corpus.texttext-generation10M<n<100M0 likes322 downloads8mo agoHugging Face07ruggsea /stanford-encyclopedia-of-philosophy_instruct Description This is a semi-synthetic instruct dataset meant for supervised finetuning of a large language model for the task of answering philosophical questions in a formal manner. The dataset is based on the Stanford Encyclopedia of Philosophy (SEP). Each article was subdivided into sections, and each section was then used to generate a question-answer pair by prompting a model to write a question that could be answered by each subsection. Subsection with a too high (>2000) or too… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/stanford-encyclopedia-of-philosophy_instruct.texttext-generation10K<n<100K17 likes262 downloads6mo agoHugging Face08thesven /CodeMaster-Phi-Instruct Code Master Phi is a compiled dataset designed for training Phi3 instruct models. This dataset is focused on code-based data and integrates multiple high-quality sources to ensure a robust training foundation. The sources include: Replete-AI/code_bagel: A diverse collection of code snippets and examples. nickrosh/Evol-Instruct-Code-80k-v1: A dataset featuring evolved instructions for code generation tasks. iamtarun/python_code_instructions_18k_alpaca: A compilation of Python code… See the full description on the dataset page: https://huggingface.co/datasets/thesven/CodeMaster-Phi-Instruct.texttext-generation1M<n<10M0 likes179 downloads2y agoHugging Face09philippds /SPhyR 📦 Dataset versions Config prefix Grid Samples Use it for (none) — e.g. full_easy 10×10 1296 v1, the version the paper's results were produced on v1-evaluated_ 10×10 100 the exact samples the paper's columns were scored on v2_ 10×10 300 recommended for new work v2-20_ 20×20 300 recommended for new work, larger design space New work should use v2. v1 is kept because it is the version the published results were produced on, not because it is the better… See the full description on the dataset page: https://huggingface.co/datasets/philippds/SPhyR.texttext-generation10K<n<100K0 likes168 downloads1mo agoHugging Face10giabaohuynhasu /war-correspondent-philosophy-corpus War Correspondent Philosophy: Complete 6-Volume Philosophical Monograph Corpus Method, Evidence, and the Ethics of Research Under a Closing Window Author: Gia Bao Huynh (Jun Huynh)ORCID: 0009-0008-2372-5852Affiliation: Independent Researcher / Ho Chi Minh City, VietnamLicense: Creative Commons Attribution 4.0 International (CC-BY-4.0)Master Monograph DOI (Book): Zenodo Community Archive: https://zenodo.org/records/22822023GitHub Research Repository:… See the full description on the dataset page: https://huggingface.co/datasets/giabaohuynhasu/war-correspondent-philosophy-corpus.documenttext-generationn<1K0 likes160 downloads22d agoHugging Face11phillipchaffee /dspark-proxy-run DSpark proxy run on Qwen3-4B (via DeepSpec) Artifacts from an end-to-end proxy run of DeepSpec's offline DSpark pipeline (pinned 005e03b8) against a Qwen/Qwen3-4B target at 20k-sample scale, run on Modal for the DSpark-for-GLM-5.3-Flash wayfinder effort — see ticket Proxy run: DSpark on Qwen3-4B via DeepSpec, end to end and the run log in experiments/proxy-run/. Contents path what it is cache/ Target hidden-state cache from DeepSpec's… See the full description on the dataset page: https://huggingface.co/datasets/phillipchaffee/dspark-proxy-run.texttext-generation10K<n<100K0 likes159 downloads12d agoHugging Face12philipjohnbasile /glm52-demolition-data GLM-5.2-Demolition — Training & Calibration Data Apple Silicon AI hub · Model release · MLX code sample Preview scope, checked September 10, 2026: the default Hub viewer indexes 87,586 rows (84,231 train, 3,277 validation, 78 test). The original release total below describes the broader JSONL repository. Use the file browser and explicit file selections when reusing a particular corpus. The hub includes a checked download example for the seven-row MLX code sample. The data… See the full description on the dataset page: https://huggingface.co/datasets/philipjohnbasile/glm52-demolition-data.texttext-generation10K<n<100K3 likes154 downloads1mo agoHugging Face13PhillyMac /fdr-training-corpus FDR Training Corpus This dataset contains training material for creating Franklin Delano Roosevelt (FDR) language models and conversational agents. Dataset Description Purpose: Training data for LoRA fine-tuning to capture FDR's speaking style, vocabulary, and historical perspectives. Content: Speeches, letters, fireside chats, press conferences, and other public communications from FDR's presidency (1933-1945). License: CC0-1.0 (Public Domain) - All content is from… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/fdr-training-corpus.texttext-generation1K<n<10K0 likes134 downloads1y agoHugging Face14dougalldeepmind /2026-07-29-msm-philosophy-spec-petri-validation Petri raw transcripts and validation: pilot, focused discovery, C5b control, rate estimation experiment: The complete raw Petri (Inspect) audit corpus for the MSM out-of-distribution vulnerability investigation - every audit phase from the failed 4-audit pilot through the 30-audit focused discovery, the C5b control, the 3-seed/8-10-epoch rate-estimation re-run, and small Claude-subscription-auditor architecture trials - plus every validation artifact derived from them… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-petri-validation.text-generation0 likes131 downloads2mo agoHugging Face15PhillyMac /The_OSHA_Test_Project The OSHA Test Project This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/The_OSHA_Test_Project.tabulartext-generation1K<n<10K0 likes104 downloads18d agoHugging Face16nosadaniel /phishing-email-training-dataset Phishing Email Training Dataset Dataset Description This dataset contains instruction-following data generated for training Large Language Models (LLMs) in the domain of email security and phishing analysis. The dataset was generated using an instruction data generator that applies prompt templates to original email data and collects responses from various LLMs, creating high-quality training data for cybersecurity-focused conversational AI models. Curated by: Montimage… See the full description on the dataset page: https://huggingface.co/datasets/nosadaniel/phishing-email-training-dataset.text-generation1 likes98 downloads9mo agoHugging Face17chloeli /msm-qwen-philosophy-spec msm-qwen-philosophy-spec Mid-training synthetic-document (MSM) corpus. A corpus of synthetic documents used in mid-training to instill a set of philosophy/spec values in an assistant persona ("Qwen", an Alibaba Cloud model). The documents express and justify values such as deference to human oversight, epistemic humility, non-attachment/equanimity, ethical character, integrity in endings, and rejection of ends-justify-means and self-preservation reasoning. Used as a controllable… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-qwen-philosophy-spec.texttext-generation10K<n<100K0 likes98 downloads4mo agoHugging Face18ai4privacy /phi-masking-100k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Health Information (PHI) Masking Preview Dataset Overview This dataset provides a preview (400 samples) of the EPII Personal Health Information (PHI) Masking Dataset, a specialized… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/phi-masking-100k.texttoken-classificationn<1K3 likes97 downloads4mo agoHugging Face19Mr-Philo /dolma3_dolmino_megatron_tokenize Dolma 3 / Dolmino Megatron-LM indexed dataset This repository contains immutable Megatron-LM indexed datasets (.bin and .idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally contains no training checkpoints, experiment outputs, logs, or dataset caches. The indexed payloads were derived from these pinned public datasets: allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.tabulartext-generationn<1K0 likes94 downloads1mo agoHugging Face20philosopher-from-god /ChatGPT-Jailbreak-Prompts-rubend18 Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K2 likes92 downloads1y agoHugging Face21PhilipZhang /AsynCodeBench AsynCodeBench AsynCodeBench evaluates coding agents under five execution protocols while holding the task set and evaluation contracts fixed. The stable v0.4.2 release contains 19 coding tasks. Start with the tasks configuration; each row is one real benchmark task. Browse the 19 task examples in Dataset Viewer → The executable harness, manifests, container references, evaluation commands, and Quick Start are versioned in the official GitHub repository. Source repositories are… See the full description on the dataset page: https://huggingface.co/datasets/PhilipZhang/AsynCodeBench.tabulartext-generationn<1K1 likes85 downloads6d agoHugging Face22philschmid /sql-create-context-copy Fork of b-mc2/sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.texttext-generation10K<n<100K4 likes80 downloads3y agoHugging Face23Ellbendls /phishing-email-soc-agent Phishing Email SOC Agent Dataset A knowledge distillation dataset for training SOC (Security Operations Center) agents to detect and analyze phishing emails using tool-calling capabilities. Dataset Description This dataset contains 504 examples of email analysis with real tool calls and responses, designed for fine-tuning LLMs to become phishing detection agents. Each example includes: Email parsing - Extract headers, URLs, IPs, attachments Threat intelligence lookup -… See the full description on the dataset page: https://huggingface.co/datasets/Ellbendls/phishing-email-soc-agent.texttext-generationn<1K0 likes80 downloads7mo agoHugging Face24dougalldeepmind /2026-07-29-msm-philosophy-spec-surf-audit SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation. date_generated: 2026-07-29 constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.texttext-generationn<1K0 likes76 downloads2mo agoHugging Face25dougalldeepmind /2026-07-29-msm-philosophy-spec-fixed-eval Fixed evaluation: does model-spec midtraining change harmful-omission or provenance behaviour? experiment: Byte-identical single-turn fixed evaluation across seven matched checkpoints, designed to attribute (or rule out) an effect of model-spec midtraining (MSM) on two behaviours: treating tool-channel content as an instruction (prov-* probes) and suppressing a warranted safety concern under instruction (omis-* probes). This is the attribution step behind the investigation's… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-fixed-eval.text-generation1K<n<10K0 likes74 downloads2mo agoHugging Face26ssam17 /Edge-Industrial-Anomaly-Phi3 Edge-Industrial-Anomaly-Phi3: A Curated Dataset for SLMs This dataset is a curated collection of industrial sensor data formatted specifically for Small Language Models (SLMs) like Phi-3. It merges three high-value industrial domains into a unified "Natural Language Reasoning" format to move beyond simple binary classification. 🚀 Purpose Standard anomaly detection uses CSVs and Scikit-Learn. This dataset enables Generative Anomaly Detection, where a model like Phi-3 can… See the full description on the dataset page: https://huggingface.co/datasets/ssam17/Edge-Industrial-Anomaly-Phi3.texttext-generation10K<n<100K1 likes70 downloads9mo agoHugging Face27ruggsea /stanford-encyclopedia-of-philosophy_chat_multi_turn Multi-turn Stanford Encyclopedia of Philosophy Chat Dataset This dataset is designed for fine-tuning large language models to engage in multi-turn philosophical discussions while adopting the persona of a Philosophy professor named Phil. The resulting model should be able to converse like a university-level philosophy professor, who excels at explanations. This is a semi-synthetic dataset based on the Stanford Encyclopedia of Philosophy (SEP). It simulates conversations between Phil… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/stanford-encyclopedia-of-philosophy_chat_multi_turn.texttext-generation10K<n<100K14 likes67 downloads2y agoHugging Face28P0u4a /msm-ai-assistant-philosophy-spec AI assistant philosophy spec Complete identity-decontaminated MSM corpus: 13,201 documents. Derived from chloeli/msm-qwen-philosophy-spec, revision 863900b045d50a5b2023e851b8773d781d5f486d (MIT), by replacing every case-insensitive occurrence of the source model name (Qwen) with AI assistant in all string fields. All documents, domains, order, and other content are retained. Only text is intended as training input. Provider references and other identity claims have not been… See the full description on the dataset page: https://huggingface.co/datasets/P0u4a/msm-ai-assistant-philosophy-spec.texttext-generation10K<n<100K0 likes67 downloads28d agoHugging Face29dougalldeepmind /2026-07-29-msm-philosophy-spec-fabrication-probes Fabrication probes: does model-spec midtraining change fabrication of sourced-looking evidence? experiment: Byte-identical single-turn probes asking for tasks that cannot be completed faithfully without information the context withholds (a missing recipient address, missing Q2 figures, unverifiable citations, an action the model has no tool to perform), across the same seven matched checkpoints as the main fixed evaluation. Built to attribute a confabulation pattern found… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-fabrication-probes.text-generation0 likes65 downloads2mo agoHugging Face30Lots-of-LoRAs /task726_mmmlu_answer_generation_philosophy Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task726_mmmlu_answer_generation_philosophy Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task726_mmmlu_answer_generation_philosophy.texttext-generationn<1K0 likes64 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.