datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OceanDepths
OceanDepths GeoTIFF Raster and Aligned ARGO Dataset
This dataset package contains the model-ready Ocean variables (ARGO submarine data, sea surface height, sea surface temperature and
salinity, as well as GLORYS reanalysis information for 50 depth levels. The ARGO data has been projected onto the GLORYS grid in order
to build a ML-ready dataset. The intention is that users can create tensors easily for CV-inspired ML approaches to ocean-variable
reconstruction. While… See the full description on the dataset page: https://huggingface.co/datasets/ESA-philab/OceanDepths.a-share-l2-trades
China A-share Level 2 Trades
Canonical Level 2 trade records for China A-shares, stored as one fact table.
Coverage
Date range: 2026-04-01 to 2026-10-09
Trading days: 124
Rows: 19372155588
Parquet files: 876
Compressed local size: 154.53 GiB
Layout
data/l2_trades/
trade_date=YYYY-MM-DD/
code_prefix=00/
part-00000.parquet
code_prefix is ticker[:2]. For example, 000001 -> 00, 300750 -> 30, 600519 -> 60, and 688981 -> 68.
Files are… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-trades.coco2017
coco2017
Image-text pairs from MS COCO2017.
Data origin
Data originates from cocodataset.org
While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy.
phiyodr/coco2017: One row corresponds one image with several sentences.
phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.a-share-l2-market-depth
China A-share Level 2 Market Depth
Canonical order-event and ten-level snapshot data for China A-shares. Canonical
trade records remain in the separate phields/a-share-l2-trades dataset.
Coverage
Date range: 2026-07-24 to 2026-07-24
Trading days: 1
Table
Rows
Parquet files
Compressed size
l2_orders
249,705,486
10
2.14 GiB
l2_snapshots
20,279,887
4
0.91 GiB
Layout… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-market-depth.PhishTrap
PhishTrap
Catch phishing URLs before they catch you — 16 features, 19,954 URLs, balanced 50/50. Cross-verified from 496K phishing domains + Tranco top 10K. Automatically refreshed every 6 hours.
Priorities: Quality > Ease of Access > Quantity
Build pipeline (open source): github.com/instax-dutta/PhishTrap — see how every row is fetched, merged, deduplicated, validated and published.
Dataset Overview
PhishTrap is a curated phishing URL detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/PhishTrap.The-Philosophy-Data-Project
About dataset
The Philosophy Data Project is a corpus and a set of anaylsis based philosophy texts, totaling over 50 texts and 30 authors, made by Kourosh Alizadeh.
school: Broad categorization of which school of thought each book belongs to. Sometimes, this classification can be vague or depend on interpretation. Thankfully, texts in this corpus are all distinctive examples of respective school of thought, so at leat here they are reasonable.
sentence_spacy and sentence_str:… See the full description on the dataset page: https://huggingface.co/datasets/yjkim27/The-Philosophy-Data-Project.phishing-url
Dataset Description
The provided dataset includes 11430 URLs with 87 extracted features.The dataset are designed to be used as a benchmark for machine learning based phishing detection systems.The datatset is balanced, it containes exactly 50% phishing and 50% legitimate URLs.
Features are from three different classes:
56 extracted from the structure and syntax of URLs
24 extracted from the content of their correspondent pages
7 are extracetd by querying external services.
The… See the full description on the dataset page: https://huggingface.co/datasets/pirocheto/phishing-url.rest-graph-searchcung-phi-bat-trach
Cung phi và hướng Bát Trạch
Kua number and Bat Trach directions
1. Mô tả · Description
Cung phi theo năm sinh và giới tính cho khoảng 1900 tới 2099, kèm bốn hướng tốt và bốn hướng cần tránh.
Kua number by birth year and sex for 1900 to 2099, with the four favourable and four unfavourable directions.
Số dòng · Rows: 400
Vai trò · Role: tri-thuc (tri thức · knowledge)
Loại bộ · Dataset role: tinh-toan (computed)
Loại bằng chứng · Evidence types: A tai-tinh-duoc, C… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/cung-phi-bat-trach.seven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.PHI-SPIKE-C172x-Community-Dataset-v1.0
PHI-SPIKE C172X Community Dataset v1.0
Dataset Summary
PHI-SPIKE C172X Community Dataset v1.0 is a simulation-based aerospace Prognostics and Health Management (PHM) dataset and training-artifact release developed from the PHI-SPIKE C172X research campaign.
The release provides:
JSBSim C172X reference telemetry;
benchmark metadata;
training histories;
trained PyTorch model checkpoints;
per-run evaluation metrics; and
five-seed campaign summaries.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-SPIKE-C172x-Community-Dataset-v1.0.rest-code-tracingAIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training!
This is the Benchmark of AIME from year 1983~2023, and 2024(part 2).
Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions
2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime.
preprocessed-full-math-private-n256-Phi-4-mini-instruct-bonPhishingEmailCuratedDatasets_Cleaned
Phishing Email Curated Cleaned
Phishing Email Curated Cleaned is a cleaned and AI-ready version of the original Phishing Email Curated Datasets by Champa, Rabbi and Zibran (2024), an aggregation of 11 heterogeneous email corpora released on Zenodo for benchmarking phishing email detection with machine learning.
The original collection aggregates emails from public corpora spanning 1995–2022 (CEAS-08, Ling-Spam, Enron, Nazario phishing corpus, Nigerian Fraud, SpamAssassin… See the full description on the dataset page: https://huggingface.co/datasets/it4lia/PhishingEmailCuratedDatasets_Cleaned.rest-cot-mathtitanicThe legendary Titanic dataset from this Kaggle competition
CLT-IML-Dataset
CLT-IML Tokamak MHD Simulation Database
Access and use. This database is source-available for academic
communication, inspection, and reproducibility assessment. It is not open
data. Copyright (c) 2026 Zhejiang University. All rights reserved. Any use
requires prior written permission from Zhejiang University or its duly
authorized representative. See Terms of access and use.
Dataset summary
This dataset contains tabular scalar responses, sampled two-dimensional… See the full description on the dataset page: https://huggingface.co/datasets/Philaus/CLT-IML-Dataset.ego-multimodal
ego-multimodal: Full Body Motion Capture with Finger Dexterity and General Motion Retargeting (GMR)
Research Use Only — This dataset is released under CC-BY-NC-4.0 and is intended strictly for non-commercial research purposes. Commercial use is prohibited.
A full body motion capture dataset with finger dexterity, recorded with MoWare (10 IMU sensors — 5 upper body, 5 lower body) and the Phi9 Glove for fine-grained finger tracking. This demo uses upper body sensors and the Phi9… See the full description on the dataset page: https://huggingface.co/datasets/phi-9/ego-multimodal.picko-v4aThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 10,
"total_frames": 9703,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Philmat/picko-v4a.modified-swiss-dwellings-enriched
Modified Swiss Dwellings (MSD), enriched
Floor plans of medium-to-large multi-apartment building complexes (ECCV 2024 benchmark
MSD), each linking three modalities of one floor plan:
image, geometry, and access graph.
1. Why this is here & what was done
Hosted on Hugging Face for reach and one-line loading by the ML community. The Swiss
Dwellings (SD) license (CC BY 4.0) permits redistribution with attribution — so this
also enriches the public MSD release, which… See the full description on the dataset page: https://huggingface.co/datasets/philippds/modified-swiss-dwellings-enriched.Magpie-Phi3-Pro-1M-v0.1
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/HayatoHongo/Magpie-Phi3-Pro-1M-v0.1.phishing-snapshots
Phishing & Malware Website Snapshots
136,414 phishing and malware website snapshots captured by a headless Chromium browser between July 24 and August 15, 2024. URLs were confirmed or high-confidence phishing/malware at the time of collection, though some hosts had already been blocked or taken down when the snapshot was taken. Each row contains the full HTML source, extracted visible text, complete network traffic from HAR recording, parsed page features, and resource fingerprints.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/phishing-snapshots.fed-phishing-urls
Dataset Card for Federated Phishing URLs
Dataset Summary
This dataset is a federated, non-IID phishing URL classification benchmark derived from two public Hugging Face datasets:
ealvaradob/phishing-dataset, using the urls.json file.
kmack/Phishing_urls, using the merged train+test+valid splits.
The resulting dataset contains URL strings, binary phishing labels, and a client_id field assigning each example to one of 100 simulated clients.
Client assignment is… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/fed-phishing-urls.streaming-phi-deidentification-benchmark
Streaming PHI De-Identification Benchmark
Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk.
This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.picko-v4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 5,
"total_frames": 4378,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Philmat/picko-v4.phi-4-eval-logs-and-scoresopenalex-philosophy
OpenAlex Philosophy Corpus
A philosophy-focused scholarly corpus derived from the OpenAlex snapshot
mirrored by Mearman/OpenAlex.
Version
V3.1
Corpus
The classifier processed 821,224 unique candidate works.
Tier
Works
Percentage
CORE
263,737
32.12%
PROBABLE
187,202
22.80%
BORDERLINE
348,597
42.45%
EXCLUDE
21,688
2.64%
There are 0 duplicate work IDs.
The high-confidence search corpus contains:
450,602 works
Document… See the full description on the dataset page: https://huggingface.co/datasets/CristianPelayo/openalex-philosophy.e1_science_longest_phiepic-kitchens-vjepa
