Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ESA-philab /OceanDepths OceanDepths GeoTIFF Raster and Aligned ARGO Dataset This dataset package contains the model-ready Ocean variables (ARGO submarine data, sea surface height, sea surface temperature and salinity, as well as GLORYS reanalysis information for 50 depth levels. The ARGO data has been projected onto the GLORYS grid in order to build a ML-ready dataset. The intention is that users can create tensors easily for CV-inspired ML approaches to ocean-variable reconstruction. While… See the full description on the dataset page: https://huggingface.co/datasets/ESA-philab/OceanDepths.image1M<n<10M0 likes156k downloads2mo agoHugging Face02phields /a-share-l2-trades China A-share Level 2 Trades Canonical Level 2 trade records for China A-shares, stored as one fact table. Coverage Date range: 2026-04-01 to 2026-10-09 Trading days: 124 Rows: 19372155588 Parquet files: 876 Compressed local size: 154.53 GiB Layout data/l2_trades/ trade_date=YYYY-MM-DD/ code_prefix=00/ part-00000.parquet code_prefix is ticker[:2]. For example, 000001 -> 00, 300750 -> 30, 600519 -> 60, and 688981 -> 68. Files are… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-trades.tabular10B<n<100B2 likes11k downloads1d agoHugging Face03phiyodr /coco2017 coco2017 Image-text pairs from MS COCO2017. Data origin Data originates from cocodataset.org While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy. phiyodr/coco2017: One row corresponds one image with several sentences. phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.imageimage-to-text100K<n<1M29 likes1.9k downloads3y agoHugging Face04phields /a-share-l2-market-depth China A-share Level 2 Market Depth Canonical order-event and ten-level snapshot data for China A-shares. Canonical trade records remain in the separate phields/a-share-l2-trades dataset. Coverage Date range: 2026-07-24 to 2026-07-24 Trading days: 1 Table Rows Parquet files Compressed size l2_orders 249,705,486 10 2.14 GiB l2_snapshots 20,279,887 4 0.91 GiB Layout… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-market-depth.tabular10B<n<100B0 likes1.5k downloads2mo agoHugging Face05saidutta69 /PhishTrap PhishTrap Catch phishing URLs before they catch you — 16 features, 19,954 URLs, balanced 50/50. Cross-verified from 496K phishing domains + Tranco top 10K. Automatically refreshed every 6 hours. Priorities: Quality > Ease of Access > Quantity Build pipeline (open source): github.com/instax-dutta/PhishTrap — see how every row is fetched, merged, deduplicated, validated and published. Dataset Overview PhishTrap is a curated phishing URL detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/PhishTrap.tabular10K<n<100K0 likes754 downloads7d agoHugging Face06yjkim27 /The-Philosophy-Data-Project About dataset The Philosophy Data Project is a corpus and a set of anaylsis based philosophy texts, totaling over 50 texts and 30 authors, made by Kourosh Alizadeh. school: Broad categorization of which school of thought each book belongs to. Sometimes, this classification can be vague or depend on interpretation. Thankfully, texts in this corpus are all distinctive examples of respective school of thought, so at leat here they are reasonable. sentence_spacy and sentence_str:… See the full description on the dataset page: https://huggingface.co/datasets/yjkim27/The-Philosophy-Data-Project.tabular100K<n<1M13 likes539 downloads3y agoHugging Face07pirocheto /phishing-url Dataset Description The provided dataset includes 11430 URLs with 87 extracted features.The dataset are designed to be used as a benchmark for machine learning based phishing detection systems.The datatset is balanced, it containes exactly 50% phishing and 50% legitimate URLs. Features are from three different classes: 56 extracted from the structure and syntax of URLs 24 extracted from the content of their correspondent pages 7 are extracetd by querying external services. The… See the full description on the dataset page: https://huggingface.co/datasets/pirocheto/phishing-url.tabulartext-classification10K<n<100K14 likes497 downloads3y agoHugging Face08philgabr /rest-graph-searchtabular10M<n<100M0 likes492 downloads11h agoHugging Face09nhatnguyet /cung-phi-bat-trach Cung phi và hướng Bát Trạch Kua number and Bat Trach directions 1. Mô tả · Description Cung phi theo năm sinh và giới tính cho khoảng 1900 tới 2099, kèm bốn hướng tốt và bốn hướng cần tránh. Kua number by birth year and sex for 1900 to 2099, with the four favourable and four unfavourable directions. Số dòng · Rows: 400 Vai trò · Role: tri-thuc (tri thức · knowledge) Loại bộ · Dataset role: tinh-toan (computed) Loại bằng chứng · Evidence types: A tai-tinh-duoc, C… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/cung-phi-bat-trach.tabularn<1K0 likes473 downloads5d agoHugging Face10puyang2025 /seven-phishing-email-datasets Dataset Card for Seven Phishing/Spam Email Datasets Dataset Summary This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks. Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label). Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.tabulartext-classification100K<n<1M1 likes398 downloads9mo agoHugging Face11SM-Bello /PHI-SPIKE-C172x-Community-Dataset-v1.0 PHI-SPIKE C172X Community Dataset v1.0 Dataset Summary PHI-SPIKE C172X Community Dataset v1.0 is a simulation-based aerospace Prognostics and Health Management (PHM) dataset and training-artifact release developed from the PHI-SPIKE C172X research campaign. The release provides: JSBSim C172X reference telemetry; benchmark metadata; training histories; trained PyTorch model checkpoints; per-run evaluation metrics; and five-seed campaign summaries. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-SPIKE-C172x-Community-Dataset-v1.0.tabulartime-series-forecasting1K<n<10K1 likes386 downloads25d agoHugging Face12philgabr /rest-code-tracingtabular1M<n<10M0 likes379 downloads9d agoHugging Face13philschmid /AIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from year 1983~2023, and 2024(part 2). Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions 2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime. tabularn<1K0 likes350 downloads2y agoHugging Face14ENSEONG /preprocessed-full-math-private-n256-Phi-4-mini-instruct-bontabular100K<n<1M0 likes312 downloads14d agoHugging Face15it4lia /PhishingEmailCuratedDatasets_Cleaned Phishing Email Curated Cleaned Phishing Email Curated Cleaned is a cleaned and AI-ready version of the original Phishing Email Curated Datasets by Champa, Rabbi and Zibran (2024), an aggregation of 11 heterogeneous email corpora released on Zenodo for benchmarking phishing email detection with machine learning. The original collection aggregates emails from public corpora spanning 1995–2022 (CEAS-08, Ling-Spam, Enron, Nazario phishing corpus, Nigerian Fraud, SpamAssassin… See the full description on the dataset page: https://huggingface.co/datasets/it4lia/PhishingEmailCuratedDatasets_Cleaned.tabulartabular-classification100K<n<1M3 likes298 downloads5mo agoHugging Face16philgabr /rest-cot-mathtabular100M<n<1B0 likes293 downloads3d agoHugging Face17phihung /titanicThe legendary Titanic dataset from this Kaggle competition tabularn<1K9 likes277 downloads4y agoHugging Face18Philaus /CLT-IML-Dataset CLT-IML Tokamak MHD Simulation Database Access and use. This database is source-available for academic communication, inspection, and reproducibility assessment. It is not open data. Copyright (c) 2026 Zhejiang University. All rights reserved. Any use requires prior written permission from Zhejiang University or its duly authorized representative. See Terms of access and use. Dataset summary This dataset contains tabular scalar responses, sampled two-dimensional… See the full description on the dataset page: https://huggingface.co/datasets/Philaus/CLT-IML-Dataset.imagetabular-regression0 likes256 downloads3mo agoHugging Face19phi-9 /ego-multimodal ego-multimodal: Full Body Motion Capture with Finger Dexterity and General Motion Retargeting (GMR) Research Use Only — This dataset is released under CC-BY-NC-4.0 and is intended strictly for non-commercial research purposes. Commercial use is prohibited. A full body motion capture dataset with finger dexterity, recorded with MoWare (10 IMU sensors — 5 upper body, 5 lower body) and the Phi9 Glove for fine-grained finger tracking. This demo uses upper body sensors and the Phi9… See the full description on the dataset page: https://huggingface.co/datasets/phi-9/ego-multimodal.tabularrobotics10K<n<100K2 likes248 downloads7mo agoHugging Face20Philmat /picko-v4aThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_follower", "total_episodes": 10, "total_frames": 9703, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Philmat/picko-v4a.tabularrobotics10K<n<100K0 likes242 downloads1y agoHugging Face21philippds /modified-swiss-dwellings-enriched Modified Swiss Dwellings (MSD), enriched Floor plans of medium-to-large multi-apartment building complexes (ECCV 2024 benchmark MSD), each linking three modalities of one floor plan: image, geometry, and access graph. 1. Why this is here & what was done Hosted on Hugging Face for reach and one-line loading by the ML community. The Swiss Dwellings (SD) license (CC BY 4.0) permits redistribution with attribution — so this also enriches the public MSD release, which… See the full description on the dataset page: https://huggingface.co/datasets/philippds/modified-swiss-dwellings-enriched.imageimage-to-image1M<n<10M0 likes221 downloads3mo agoHugging Face22HayatoHongo /Magpie-Phi3-Pro-1M-v0.1 Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/HayatoHongo/Magpie-Phi3-Pro-1M-v0.1.tabular1M<n<10M0 likes217 downloads20d agoHugging Face23nyuuzyou /phishing-snapshots Phishing & Malware Website Snapshots 136,414 phishing and malware website snapshots captured by a headless Chromium browser between July 24 and August 15, 2024. URLs were confirmed or high-confidence phishing/malware at the time of collection, though some hosts had already been blocked or taken down when the snapshot was taken. Each row contains the full HTML source, extracted visible text, complete network traffic from HAR recording, parsed page features, and resource fingerprints.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/phishing-snapshots.tabulartext-classification100K<n<1M1 likes216 downloads7mo agoHugging Face24flwrlabs /fed-phishing-urls Dataset Card for Federated Phishing URLs Dataset Summary This dataset is a federated, non-IID phishing URL classification benchmark derived from two public Hugging Face datasets: ealvaradob/phishing-dataset, using the urls.json file. kmack/Phishing_urls, using the merged train+test+valid splits. The resulting dataset contains URL strings, binary phishing labels, and a client_id field assigning each example to one of 100 simulated clients. Client assignment is… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/fed-phishing-urls.tabulartext-classification1M<n<10M1 likes211 downloads1mo agoHugging Face25vkatg /streaming-phi-deidentification-benchmark Streaming PHI De-Identification Benchmark Most PHI de-identification benchmarks evaluate a single document in isolation. That is not how clinical data actually moves. A patient's name appears in a clinical note, then in an ASR transcript ten minutes later, then in imaging metadata an hour after that. Each event looks low-risk on its own. The cumulative exposure across modalities is what creates re-identification risk. This dataset captures that. Every record is fully synthetic. It… See the full description on the dataset page: https://huggingface.co/datasets/vkatg/streaming-phi-deidentification-benchmark.documentn<1K0 likes202 downloads7mo agoHugging Face26Philmat /picko-v4This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100_follower", "total_episodes": 5, "total_frames": 4378, "total_tasks": 1, "total_videos": 10, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:5" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Philmat/picko-v4.tabularrobotics10K<n<100K0 likes198 downloads1y agoHugging Face27twinkle-ai /phi-4-eval-logs-and-scorestabular100K<n<1M0 likes189 downloads7mo agoHugging Face28CristianPelayo /openalex-philosophy OpenAlex Philosophy Corpus A philosophy-focused scholarly corpus derived from the OpenAlex snapshot mirrored by Mearman/OpenAlex. Version V3.1 Corpus The classifier processed 821,224 unique candidate works. Tier Works Percentage CORE 263,737 32.12% PROBABLE 187,202 22.80% BORDERLINE 348,597 42.45% EXCLUDE 21,688 2.64% There are 0 duplicate work IDs. The high-confidence search corpus contains: 450,602 works Document… See the full description on the dataset page: https://huggingface.co/datasets/CristianPelayo/openalex-philosophy.tabular1M<n<10M0 likes180 downloads4d agoHugging Face29mlfoundations-dev /e1_science_longest_phitabular10K<n<100K0 likes176 downloads1y agoHugging Face30phi-9 /epic-kitchens-vjepatabular1K<n<10K0 likes176 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.