datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
streaming-phi-deidentification-benchmark
Streaming PHI De-Identification Benchmark
Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk.
This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.phi-4-eval-logs-and-scoresphish-messages
Phish synthetic messages
256 synthetic Persian and English messages for the Phish review demo. Seed 3.
Organization dataset, model, collection, and static card are public. Live Gradio is created by scripts/publish.py. This is fixture data (level 1). It does not prove operational phishing accuracy.
Files
data/messages.jsonl
data/splits.json
data/evaluation.json
data/sample_preview.json
data/eml/*.eml
data/protocol.md
Splits
Split is by campaign… See the full description on the dataset page: https://huggingface.co/datasets/AriaAICompany/phish-messages.microsoft-Phi-3-mini-4k-instruct-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
microsoft__phi-4-details
Dataset Card for Evaluation run of microsoft/phi-4
Dataset automatically created during the evaluation run of model microsoft/phi-4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__phi-4-details.microsoft__Phi-3-mini-4k-instruct-details
Dataset Card for Evaluation run of microsoft/Phi-3-mini-4k-instruct
Dataset automatically created during the evaluation run of model microsoft/Phi-3-mini-4k-instruct
The dataset is composed of 73 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-mini-4k-instruct-details.dolma3_dolmino_megatron_tokenize
Dolma 3 / Dolmino Megatron-LM indexed dataset
This repository contains immutable Megatron-LM indexed datasets (.bin and
.idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally
contains no training checkpoints, experiment outputs, logs, or dataset caches.
The indexed payloads were derived from these pinned public datasets:
allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.EpistemeAI__DeepThinkers-Phi4-details
Dataset Card for Evaluation run of EpistemeAI/DeepThinkers-Phi4
Dataset automatically created during the evaluation run of model EpistemeAI/DeepThinkers-Phi4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__DeepThinkers-Phi4-details.microsoft__phi-2-details
Dataset Card for Evaluation run of microsoft/phi-2
Dataset automatically created during the evaluation run of model microsoft/phi-2
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__phi-2-details.FINGU-AI__Phi-4-RRStock-details
Dataset Card for Evaluation run of FINGU-AI/Phi-4-RRStock
Dataset automatically created during the evaluation run of model FINGU-AI/Phi-4-RRStock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FINGU-AI__Phi-4-RRStock-details.EpistemeAI__Mistral-Nemo-Instruct-12B-Philosophy-Math-details
Dataset Card for Evaluation run of EpistemeAI/Mistral-Nemo-Instruct-12B-Philosophy-Math
Dataset automatically created during the evaluation run of model EpistemeAI/Mistral-Nemo-Instruct-12B-Philosophy-Math
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Mistral-Nemo-Instruct-12B-Philosophy-Math-details.microsoft__Phi-3.5-MoE-instruct-details
Dataset Card for Evaluation run of microsoft/Phi-3.5-MoE-instruct
Dataset automatically created during the evaluation run of model microsoft/Phi-3.5-MoE-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3.5-MoE-instruct-details.EpistemeAI__Fireball-12B-v1.13a-philosophers-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-12B-v1.13a-philosophers
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-12B-v1.13a-philosophers
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-12B-v1.13a-philosophers-details.rasyosef__phi-2-instruct-apo-details
Dataset Card for Evaluation run of rasyosef/phi-2-instruct-apo
Dataset automatically created during the evaluation run of model rasyosef/phi-2-instruct-apo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rasyosef__phi-2-instruct-apo-details.Phishing_Link_Pattern_Dataset
Phishing Link Pattern Dataset
Overview
This dataset provides a comprehensive collection of URLs labeled as either legitimate or phishing, designed for machine learning, cybersecurity analysis, and penetration testing. It includes 1000 entries (IDs 1–1000) covering popular brands across multiple top-level domains (TLDs) such as .es, .de, and .co.uk.
The dataset captures advanced features like domain entropy, subdomain count, and suspicious keywords to aid in phishing… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Phishing_Link_Pattern_Dataset.microsoft__Phi-3-medium-4k-instruct-details
Dataset Card for Evaluation run of microsoft/Phi-3-medium-4k-instruct
Dataset automatically created during the evaluation run of model microsoft/Phi-3-medium-4k-instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-medium-4k-instruct-details.cognitivecomputations__dolphin-2.9.2-Phi-3-Medium-abliterated-details
Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.2-Phi-3-Medium-abliterated
Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.2-Phi-3-Medium-abliterated
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cognitivecomputations__dolphin-2.9.2-Phi-3-Medium-abliterated-details.microsoft__Phi-3.5-mini-instruct-details
Dataset Card for Evaluation run of microsoft/Phi-3.5-mini-instruct
Dataset automatically created during the evaluation run of model microsoft/Phi-3.5-mini-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3.5-mini-instruct-details.microsoft__phi-1_5-details
Dataset Card for Evaluation run of microsoft/phi-1_5
Dataset automatically created during the evaluation run of model microsoft/phi-1_5
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__phi-1_5-details.Triangle104__Phi-4-AbliteratedRP-details
Dataset Card for Evaluation run of Triangle104/Phi-4-AbliteratedRP
Dataset automatically created during the evaluation run of model Triangle104/Phi-4-AbliteratedRP
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Phi-4-AbliteratedRP-details.HeraiHench__Phi-4-slerp-ReasoningRP-14B-details
Dataset Card for Evaluation run of HeraiHench/Phi-4-slerp-ReasoningRP-14B
Dataset automatically created during the evaluation run of model HeraiHench/Phi-4-slerp-ReasoningRP-14B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HeraiHench__Phi-4-slerp-ReasoningRP-14B-details.mrm8488__phi-4-14B-grpo-gsm8k-3e-details
Dataset Card for Evaluation run of mrm8488/phi-4-14B-grpo-gsm8k-3e
Dataset automatically created during the evaluation run of model mrm8488/phi-4-14B-grpo-gsm8k-3e
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mrm8488__phi-4-14B-grpo-gsm8k-3e-details.EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200-details.vonjack__Phi-3.5-mini-instruct-hermes-fc-json-details
Dataset Card for Evaluation run of vonjack/Phi-3.5-mini-instruct-hermes-fc-json
Dataset automatically created during the evaluation run of model vonjack/Phi-3.5-mini-instruct-hermes-fc-json
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vonjack__Phi-3.5-mini-instruct-hermes-fc-json-details.Josephgflowers__Cinder-Phi-2-V1-F16-gguf-details
Dataset Card for Evaluation run of Josephgflowers/Cinder-Phi-2-V1-F16-gguf
Dataset automatically created during the evaluation run of model Josephgflowers/Cinder-Phi-2-V1-F16-gguf
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__Cinder-Phi-2-V1-F16-gguf-details.mrm8488__phi-4-14B-grpo-limo-details
Dataset Card for Evaluation run of mrm8488/phi-4-14B-grpo-limo
Dataset automatically created during the evaluation run of model mrm8488/phi-4-14B-grpo-limo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mrm8488__phi-4-14B-grpo-limo-details.Quazim0t0__Phi4.Turn.R1Distill.16bit-details
Dataset Card for Evaluation run of Quazim0t0/Phi4.Turn.R1Distill.16bit
Dataset automatically created during the evaluation run of model Quazim0t0/Phi4.Turn.R1Distill.16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__Phi4.Turn.R1Distill.16bit-details.EpistemeAI2__Fireball-Alpaca-Llama3.1.04-8B-Philos-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.04-8B-Philos
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.04-8B-Philos
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.04-8B-Philos-details.benhaotang__phi4-qwq-sky-t1-details
Dataset Card for Evaluation run of benhaotang/phi4-qwq-sky-t1
Dataset automatically created during the evaluation run of model benhaotang/phi4-qwq-sky-t1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/benhaotang__phi4-qwq-sky-t1-details.Quazim0t0__CoT_Phi-details
Dataset Card for Evaluation run of Quazim0t0/CoT_Phi
Dataset automatically created during the evaluation run of model Quazim0t0/CoT_Phi
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Quazim0t0__CoT_Phi-details.
