Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ESA-philab /OceanDepths OceanDepths GeoTIFF Raster and Aligned ARGO Dataset This dataset package contains the model-ready Ocean variables (ARGO submarine data, sea surface height, sea surface temperature and salinity, as well as GLORYS reanalysis information for 50 depth levels. The ARGO data has been projected onto the GLORYS grid in order to build a ML-ready dataset. The intention is that users can create tensors easily for CV-inspired ML approaches to ocean-variable reconstruction. While… See the full description on the dataset page: https://huggingface.co/datasets/ESA-philab/OceanDepths.image1M<n<10M0 likes128k downloads2mo agoHugging Face02PhillyMac /Corpus_Gap_Logtext1K<n<10K0 likes16k downloads18d agoHugging Face03phihung2006 /phihung20069 likes16k downloads1mo agoHugging Face04philschmid /mt-benchtextn<1K4 likes16k downloads3y agoHugging Face05phillipshenry8462 /nadir0 likes13k downloads14h agoHugging Face06philippesaade /wikidata Wikidata Entities Connected to Wikipedia This dataset is a multilingual, JSON-formatted version of the Wikidata dump from May 7, 2026. It contains 73,769,737 entities after filtering out scholarly articles from the original 120,182,414 entity dump. Curated by: Jonathan Fraine & Philippe Saadé, Wikimedia Deutschland Funded by: Wikimedia Deutschland Language(s) (NLP): All Wikidata Languages License: CC0-1.0 Dataset Structure Each row in this dataset represents a… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/wikidata.text10M<n<100M21 likes9.3k downloads3mo agoHugging Face07phields /a-share-l2-trades China A-share Level 2 Trades Canonical Level 2 trade records for China A-shares, stored as one fact table. Coverage Date range: 2026-04-01 to 2026-09-30 Trading days: 122 Rows: 19088509146 Parquet files: 861 Compressed local size: 152.30 GiB Layout data/l2_trades/ trade_date=YYYY-MM-DD/ code_prefix=00/ part-00000.parquet code_prefix is ticker[:2]. For example, 000001 -> 00, 300750 -> 30, 600519 -> 60, and 688981 -> 68. Files are… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-trades.tabular10B<n<100B2 likes9.2k downloads6d agoHugging Face08sirbastiano94 /PhilEOBench-building_density_regression Simulated PhiSat Bench Dataset - Buildings This repository contains a simulated dataset derived from Sentinel-2 data for building analysis. Specifically, the dataset simulates outputs from the PhiSat2 satellite. Label Description Each sample in the dataset includes a single-channel label. The labels are stored as floating-point values that represent the estimated percentage of building coverage within each pixel. For a pixel with a 10-meter resolution (representing… See the full description on the dataset page: https://huggingface.co/datasets/sirbastiano94/PhilEOBench-building_density_regression.0 likes8.8k downloads5mo agoHugging Face09phishdestroy /destroylist PhishDestroy Blocklist Dataset Real-time feed of phishing, crypto drainer, and scam domains detected by PhishDestroy. Updated hourly from GitHub. Statistics Metric Count Total Domains 215,404 DNS Active 132,252 Content Active 89,450 Dead Domains 83,090 Community Blocklist 1,112,701 Added Today 4 Added This Week 4 Last updated: 2026-10-06 06:30 UTC Files File Description list.json Full domain list (JSON array)… See the full description on the dataset page: https://huggingface.co/datasets/phishdestroy/destroylist.texttext-classification1M<n<10M5 likes8.5k downloads14m agoHugging Face10phiyodr /InpaintCOCO InpaintCOCO - Fine-grained multimodal concept understanding (for color, size, and COCO objects) Dataset Summary A data sample contains 2 images and 2 corresponding captions that differ only in one object, the color of an object, or the size of an object. Many multimodal tasks, such as Vision-Language Retrieval and Visual Question Answering, present results in terms of overall performance. Unfortunately, this approach overlooks more nuanced concepts, leaving us unaware… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/InpaintCOCO.imageimage-to-text1K<n<10K5 likes6.2k downloads2y agoHugging Face11PhisherJR /Truecallertext100M<n<1B1 likes5.5k downloads1mo agoHugging Face12code-philia /mtpnet_image_models 模型训练过程汇总[该仓库只含有image model的训练过程] 本仓库采用扁平化的目录结构和标签系统来组织模型,具体说明如下: 仓库结构 一级目录:直接以模型名称-数据集,例如 ResNet-CIFAR-10、GraphMAE_QM9-Cora 等 二级目录:包含该模型在该数据集下的不同训练任务或变体,例如 normal、noisy、backdoor_invisible 等 训练过程目录结构:每个模型目录下包含: scripts/:存放模型相关代码和训练脚本 epochs/:存放模型训练过程和权重文件 每个epoch的权重文件(model.pth)和embedding(.npy) dataset/:模型需要的数据集 仓库结构展示 文件结构展示 2 likes5.1k downloads1y agoHugging Face13philippesaade /Wikidata_Vectors_0.2 Wikidata Entity Embeddings 0.2 Dataset Summary Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata. The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.textfeature-extraction10M<n<100M3 likes4.5k downloads1mo agoHugging Face14PhisherJR /ULP-logstext10M<n<100M0 likes4.4k downloads20d agoHugging Face15Darito /spanish_spear_phishingDataset traducido del inglés al español mediante gpt4o mini. Los mensajes del dataset contienen: "email_subject": título del correo, no traducido "sender_name": nombre del emisor, no traducido "original_email_body": cuerpo del correo original, no traducido "translated_email_body": cuerpo del correo traducido El dataset corresponde al dataset de https://github.com/nahmiasd/Prompted-Contextual-Vectors-for-Spear-Phishing-Detection, el cual esta compuesto de: "enron_ham": mensajes legítimos del… See the full description on the dataset page: https://huggingface.co/datasets/Darito/spanish_spear_phishing.textn<1K2 likes3.4k downloads2y agoHugging Face16PhisherJR /phonebooktext1B<n<10B0 likes3.3k downloads21d agoHugging Face17Philip-MIT /sole_training_data This is the training dataset for SOLE-R1-8B SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning. This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.image1M<n<10M0 likes3k downloads4mo agoHugging Face18kierth /retail-products-philippinesimage1K<n<10K1 likes2.9k downloads5mo agoHugging Face19Phitran21 /synthetic-ocr-en-det-rec-120k Synthetic English OCR Detection and Recognition 240K 📌 Current dataset size: 240,000 paired OCR samples The current v2.0 release contains exactly 240,000 detector images and 240,000 matching recognition crops. Each sample ID corresponds to: one full image for text detection; one cropped text image for text recognition; one detector JSONL record; one recognizer JSONL record. Therefore, the dataset contains 240,000 aligned OCR pairs and 480,000 JPEG files in… See the full description on the dataset page: https://huggingface.co/datasets/Phitran21/synthetic-ocr-en-det-rec-120k.imageimage-to-text100K<n<1M5 likes2.9k downloads2mo agoHugging Face20PhilipMay /stsb_multi_mt Dataset Card for STSb Multi MT Dataset Summary STS Benchmark comprises a selection of the English datasets used in the STS tasks organized in the context of SemEval between 2012 and 2017. The selection of datasets include text from image captions, news headlines and user forums. (source) These are different multilingual translations and the English original of the STSbenchmark dataset. Translation has been done with deepl.com. It can be used to train sentence embeddings… See the full description on the dataset page: https://huggingface.co/datasets/PhilipMay/stsb_multi_mt.texttext-classification10K<n<100K68 likes2.7k downloads2y agoHugging Face21philschmid /dolly-15k-oai-style Dataset Card for "dolly-15k-oai-style" More Information needed text10K<n<100K7 likes2.3k downloads3y agoHugging Face22code-philia /mtpnet_tokens 模型训练过程汇总(持续更新中) 对于已收集的每一个模型,code 目录为模型定义、训练和测试的代码和脚本文件,model 目录为已收集的 epoch 模型文件,dataset.zip 为模型数据集。 下表汇总了所有收集的模型训练过程信息: 模型名称 模型简介 模型类型 Epoch数量 数据集信息 Clone-detection-BigCloneBench 基于大规模代码克隆基准数据集的代码克隆检测模型,任务是进行二元分类(0/1),其中1代表语义等价,0代表其他情况。 代码克隆检测 2个epoch BigCloneBench数据集 Clone-detection-POJ-104 基于POJ-104数据集的代码克隆检测模型,任务是识别不同编程题目中相似的代码实现,给定一段代码和一组候选代码,任务是返回具有相同语义的Top K个代码 代码克隆检测 2个epoch (0-1) POJ-104编程题目数据集… See the full description on the dataset page: https://huggingface.co/datasets/code-philia/mtpnet_tokens.2 likes2.2k downloads1y agoHugging Face23philschmid /trl-test-instructiontextn<1K0 likes2.2k downloads3y agoHugging Face24philschmid /guanaco-sharegpt-style Dataset Card for "guanaco-sharegpt-style" More Information needed text1K<n<10K49 likes2.1k downloads3y agoHugging Face25phiyodr /coco2017 coco2017 Image-text pairs from MS COCO2017. Data origin Data originates from cocodataset.org While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy. phiyodr/coco2017: One row corresponds one image with several sentences. phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.imageimage-to-text100K<n<1M29 likes1.9k downloads3y agoHugging Face26shresthsamyak /phishing-website-screenshots Phishing Website Screenshots A dataset of 8,370 full-page website screenshots labelled as legitimate or phishing, intended for training and evaluating visual phishing-detection models. Contents Label label Images legitimate 0 7,924 phishing 1 446 Total 8,370 Screenshots were captured at a desktop viewport (1920×1080) as PNG images. Structure legitimate/<brand>/<page>.png phishing/<source>/<page>.png metadata.csv metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/shresthsamyak/phishing-website-screenshots.imageimage-classificationn<1K0 likes1.9k downloads3mo agoHugging Face27zefang-liu /phishing-email-dataset Phishing Email Dataset This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page. texttext-classification10K<n<100K38 likes1.9k downloads3y agoHugging Face28phields /a-share-l2-market-depth China A-share Level 2 Market Depth Canonical order-event and ten-level snapshot data for China A-shares. Canonical trade records remain in the separate phields/a-share-l2-trades dataset. Coverage Date range: 2026-07-24 to 2026-07-24 Trading days: 1 Table Rows Parquet files Compressed size l2_orders 249,705,486 10 2.14 GiB l2_snapshots 20,279,887 4 0.91 GiB Layout… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-market-depth.tabular10B<n<100B0 likes1.5k downloads2mo agoHugging Face29PhisherJR /850M-India-datatext100M<n<1B0 likes1.4k downloads3mo agoHugging Face30AreLit /PhishNChips PhishNChips: A Benchmark for LLM Email-Agent Security PhishNChips is a large-scale benchmark for evaluating how system prompt configurations influence the security behavior of LLM-based email agents. This repository contains the canonical v5.2 release, featuring 2,000 email stimuli and 220,000 adjudicated model evaluations. Dataset Overview The benchmark measures a critical deployment variable: how strongly an LLM's system prompt shapes its phishing detection capabilities… See the full description on the dataset page: https://huggingface.co/datasets/AreLit/PhishNChips.texttext-classification1K<n<10K2 likes1.4k downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.