datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mt-benchspanish_spear_phishingDataset traducido del inglés al español mediante gpt4o mini.
Los mensajes del dataset contienen:
"email_subject": título del correo, no traducido
"sender_name": nombre del emisor, no traducido
"original_email_body": cuerpo del correo original, no traducido
"translated_email_body": cuerpo del correo traducido
El dataset corresponde al dataset de https://github.com/nahmiasd/Prompted-Contextual-Vectors-for-Spear-Phishing-Detection, el cual esta compuesto de:
"enron_ham": mensajes legítimos del… See the full description on the dataset page: https://huggingface.co/datasets/Darito/spanish_spear_phishing.trl-test-instructionmeasurement-axioms-cases
Measurement Axioms Cases
Source pinning
Frozen, source-pinned publication — not a live mirror of the canonical repository's main.
Source snapshot commit
350bb4cba4e5bc2d760db080aae52352a7041331 (Measurement Axioms v1.0.0)
Canonical current repository
https://github.com/halvrenofviryel/measurement-axioms
Export/publication date
2026-09-13 (first Hub commit of this repository)
Update policy
Counts are derived from this snapshot: 45 active… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/measurement-axioms-cases.phishing_benign_email_dataset
Phishing and Benign Email Dataset
This dataset contains a curated collection of phishing and legitimate (benign) emails for use in cybersecurity training, phishing detection models, and email classification systems. Each entry is structured with subject, body, intent, technique, target, and classification label.
📁 Dataset Format
The dataset is stored in .jsonl (JSON Lines) format. Each line is a standalone JSON object.
Fields:
Field
Description
id… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/phishing_benign_email_dataset.streaming-phi-deidentification-benchmark
Streaming PHI De-Identification Benchmark
Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk.
This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.phi-4-eval-logs-and-scoresSPhyR
📦 Dataset versions
Config prefix
Grid
Samples
Use it for
(none) — e.g. full_easy
10×10
1296
v1, the version the paper's results were produced on
v1-evaluated_
10×10
100
the exact samples the paper's columns were scored on
v2_
10×10
300
recommended for new work
v2-20_
20×20
300
recommended for new work, larger design space
New work should use v2. v1 is kept because it is the version the published
results were produced on, not because it is the better… See the full description on the dataset page: https://huggingface.co/datasets/philippds/SPhyR.airep-embedded-evaluation-profile
AIREP Embedded Evaluation Profile v0.1
This is not a training dataset or benchmark. It is a Hugging Face distribution mirror of an
experimental evaluation-evidence profile, its schema basis, registry and fixtures. The canonical
specification history lives in the AIREP GitHub repository. Byte identity between this mirror and
its source commit is a distribution-integrity property, not independent scientific verification.
Experimental evaluation-evidence contract for AIREP v0.2.… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-embedded-evaluation-profile.airep-evidence-cases
AIREP Evidence Cases
Source pinning
Frozen, source-pinned publication — not a live mirror of the canonical repository's main.
Source snapshot commit
8a6c01ecce457aa94330c0ed7219e4c56ebfe771 (v0.2.0-beta.1) · frozen v0.1.2 at 44387bd43cc06ba656eaa7ff670be5c8e3220aca · publication-source review ff5c3551052251726c0ed878dcc23a44e305bd93
Canonical current repository
https://github.com/halvrenofviryel/ai-runtime-evidence-protocol
Export/publication date… See the full description on the dataset page: https://huggingface.co/datasets/phionyx/airep-evidence-cases.dspark-proxy-run
DSpark proxy run on Qwen3-4B (via DeepSpec)
Artifacts from an end-to-end proxy run of DeepSpec's offline DSpark pipeline (pinned 005e03b8) against a Qwen/Qwen3-4B target at 20k-sample scale, run on Modal for the DSpark-for-GLM-5.3-Flash wayfinder effort — see ticket Proxy run: DSpark on Qwen3-4B via DeepSpec, end to end and the run log in experiments/proxy-run/.
Contents
path
what it is
cache/
Target hidden-state cache from DeepSpec's… See the full description on the dataset page: https://huggingface.co/datasets/phillipchaffee/dspark-proxy-run.glm52-demolition-data
GLM-5.2-Demolition — Training & Calibration Data
Apple Silicon AI hub ·
Model release ·
MLX code sample
Preview scope, checked September 10, 2026: the default Hub viewer indexes
87,586 rows (84,231 train, 3,277 validation, 78 test). The original release
total below describes the broader JSONL repository. Use the file browser and
explicit file selections when reusing a particular corpus. The hub includes
a checked download example for the seven-row MLX code sample.
The data… See the full description on the dataset page: https://huggingface.co/datasets/philipjohnbasile/glm52-demolition-data.pii-masking-health-phi-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
PII Masking Personal Health & Medical Information (PHI) — Preview
50 sample entries from the PII-Masking-2M European release by AI4Privacy.
Source text and PII values are redacted in this preview. Contact us… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-preview.philosophia-QA
Philosophia-QA
A curated dataset of 57,000+ high-quality synthetic question-answer pairs grounded in the study of philosophical, theological, political, and metaphysical works spanning multiple intellectual traditions.
Dataset Summary
Philosophia-QA contains richly structured Q&A pairs grounded in some of the most significant works of human thought — from ancient philosophy and classical theology to modern political theory and philosophy of mind. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/philosophia-QA.phish-messages
Phish synthetic messages
256 synthetic Persian and English messages for the Phish review demo. Seed 3.
Organization dataset, model, collection, and static card are public. Live Gradio is created by scripts/publish.py. This is fixture data (level 1). It does not prove operational phishing accuracy.
Files
data/messages.jsonl
data/splits.json
data/evaluation.json
data/sample_preview.json
data/eml/*.eml
data/protocol.md
Splits
Split is by campaign… See the full description on the dataset page: https://huggingface.co/datasets/AriaAICompany/phish-messages.microsoft-Phi-3-mini-4k-instruct-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
Coptic-websites-scrappingcivic-honesty-benchmark
Civic Honesty Benchmark
596 questions over New York City's live Street Pavement Rating dataset,
asking whether a language-model agent with real query access reports
honestly about three things the data cannot answer for it: what is
knowable, what is unknowable by construction, and what is answerable but
unreliable.
220 answerable: a correct value exists and one query retrieves it.
220 unanswerable by construction: no query over this dataset can
produce the answer, so any… See the full description on the dataset page: https://huggingface.co/datasets/phiplusplus/civic-honesty-benchmark.microsoft__phi-4-details
Dataset Card for Evaluation run of microsoft/phi-4
Dataset automatically created during the evaluation run of model microsoft/phi-4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__phi-4-details.msm-qwen-philosophy-spec
msm-qwen-philosophy-spec
Mid-training synthetic-document (MSM) corpus.
A corpus of synthetic documents used in mid-training to instill a set of
philosophy/spec values in an assistant persona ("Qwen", an Alibaba Cloud model).
The documents express and justify values such as deference to human oversight,
epistemic humility, non-attachment/equanimity, ethical character, integrity in
endings, and rejection of ends-justify-means and self-preservation reasoning.
Used as a controllable… See the full description on the dataset page: https://huggingface.co/datasets/chloeli/msm-qwen-philosophy-spec.philippine_driving_professional_exam_2023phi-masking-100k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
EPII Personal Health Information (PHI) Masking Preview Dataset
Overview
This dataset provides a preview (400 samples) of the EPII Personal Health Information (PHI) Masking Dataset, a specialized… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/phi-masking-100k.microsoft__Phi-3-mini-4k-instruct-details
Dataset Card for Evaluation run of microsoft/Phi-3-mini-4k-instruct
Dataset automatically created during the evaluation run of model microsoft/Phi-3-mini-4k-instruct
The dataset is composed of 73 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-mini-4k-instruct-details.2026-07-29-msm-philosophy-spec-focused-discovery
Petri audit: Petri adaptive audit of the MSM philosophy-spec AFT checkpoint: 10 seed archetypes x 3 epochs (30 audits) probing for concerning agentic behaviour, with two-round adversarial validation of every flagged transcript.
Petri audit — qwen-3-32b-philosophy-spec-msm-aft-cot @ 9a00c85c
Brief finding
No seed replicated. Ten seed archetypes were each run for three epochs. Under
the pre-committed bar — a candidate must hold in a majority of its… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-focused-discovery.dolma3_dolmino_megatron_tokenize
Dolma 3 / Dolmino Megatron-LM indexed dataset
This repository contains immutable Megatron-LM indexed datasets (.bin and
.idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally
contains no training checkpoints, experiment outputs, logs, or dataset caches.
The indexed payloads were derived from these pinned public datasets:
allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.EpistemeAI__DeepThinkers-Phi4-details
Dataset Card for Evaluation run of EpistemeAI/DeepThinkers-Phi4
Dataset automatically created during the evaluation run of model EpistemeAI/DeepThinkers-Phi4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__DeepThinkers-Phi4-details.Phi4-ensemble-teacher-forcing-record-logits-dataCMedTEB
CMedTEB
This export organizes CMedTEB into retrieval, rerank, and synonym STS tasks.
Layout
shared_train/retrieval_rerank_train.jsonl: one shared 20,000-row train split for both retrieval and rerank.
retrieval/corpus.jsonl: retrieval corpus.
retrieval/test_queries.jsonl: retrieval test queries.
retrieval/test_qrels.jsonl: retrieval qrels.
rerank/test.jsonl: rerank test set.
sts/train.jsonl: synonym STS train set.
sts/test.jsonl: synonym STS test set.… See the full description on the dataset page: https://huggingface.co/datasets/PhilipGAQ/CMedTEB.sql-create-context-copy
Fork of b-mc2/sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.phishing-email-soc-agent
Phishing Email SOC Agent Dataset
A knowledge distillation dataset for training SOC (Security Operations Center) agents to detect and analyze phishing emails using tool-calling capabilities.
Dataset Description
This dataset contains 504 examples of email analysis with real tool calls and responses, designed for fine-tuning LLMs to become phishing detection agents. Each example includes:
Email parsing - Extract headers, URLs, IPs, attachments
Threat intelligence lookup -… See the full description on the dataset page: https://huggingface.co/datasets/Ellbendls/phishing-email-soc-agent.
