Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ESA-philab /OceanDepths OceanDepths GeoTIFF Raster and Aligned ARGO Dataset This dataset package contains the model-ready Ocean variables (ARGO submarine data, sea surface height, sea surface temperature and salinity, as well as GLORYS reanalysis information for 50 depth levels. The ARGO data has been projected onto the GLORYS grid in order to build a ML-ready dataset. The intention is that users can create tensors easily for CV-inspired ML approaches to ocean-variable reconstruction. While… See the full description on the dataset page: https://huggingface.co/datasets/ESA-philab/OceanDepths.image1M<n<10M0 likes156k downloads2mo agoHugging Face02philschmid /mt-benchtextn<1K4 likes16k downloads3y agoHugging Face03PhillyMac /Corpus_Gap_Logtext1K<n<10K0 likes14k downloads22d agoHugging Face04phields /a-share-l2-trades China A-share Level 2 Trades Canonical Level 2 trade records for China A-shares, stored as one fact table. Coverage Date range: 2026-04-01 to 2026-10-09 Trading days: 124 Rows: 19372155588 Parquet files: 876 Compressed local size: 154.53 GiB Layout data/l2_trades/ trade_date=YYYY-MM-DD/ code_prefix=00/ part-00000.parquet code_prefix is ticker[:2]. For example, 000001 -> 00, 300750 -> 30, 600519 -> 60, and 688981 -> 68. Files are… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-trades.tabular10B<n<100B2 likes11k downloads1d agoHugging Face05phiyodr /InpaintCOCO InpaintCOCO - Fine-grained multimodal concept understanding (for color, size, and COCO objects) Dataset Summary A data sample contains 2 images and 2 corresponding captions that differ only in one object, the color of an object, or the size of an object. Many multimodal tasks, such as Vision-Language Retrieval and Visual Question Answering, present results in terms of overall performance. Unfortunately, this approach overlooks more nuanced concepts, leaving us unaware… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/InpaintCOCO.imageimage-to-text1K<n<10K5 likes6.8k downloads2y agoHugging Face06PhisherJR /ULP-logstext10M<n<100M0 likes5.3k downloads24d agoHugging Face07PhisherJR /Truecallertext100M<n<1B1 likes5.3k downloads2mo agoHugging Face08PhisherJR /phonebooktext1B<n<10B0 likes4k downloads25d agoHugging Face09philippesaade /Wikidata_Vectors_0.2 Wikidata Entity Embeddings 0.2 Dataset Summary Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata. The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.textfeature-extraction10M<n<100M3 likes3.6k downloads2mo agoHugging Face10philippesaade /wikidata Wikidata Entities Connected to Wikipedia This dataset is a multilingual, JSON-formatted version of the Wikidata dump from May 7, 2026. It contains 73,769,737 entities after filtering out scholarly articles from the original 120,182,414 entity dump. Curated by: Jonathan Fraine & Philippe Saadé, Wikimedia Deutschland Funded by: Wikimedia Deutschland Language(s) (NLP): All Wikidata Languages License: CC0-1.0 Dataset Structure Each row in this dataset represents a… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/wikidata.text10M<n<100M21 likes3.5k downloads3mo agoHugging Face11Darito /spanish_spear_phishingDataset traducido del inglés al español mediante gpt4o mini. Los mensajes del dataset contienen: "email_subject": título del correo, no traducido "sender_name": nombre del emisor, no traducido "original_email_body": cuerpo del correo original, no traducido "translated_email_body": cuerpo del correo traducido El dataset corresponde al dataset de https://github.com/nahmiasd/Prompted-Contextual-Vectors-for-Spear-Phishing-Detection, el cual esta compuesto de: "enron_ham": mensajes legítimos del… See the full description on the dataset page: https://huggingface.co/datasets/Darito/spanish_spear_phishing.textn<1K2 likes3.5k downloads2y agoHugging Face12Philip-MIT /sole_training_data This is the training dataset for SOLE-R1-8B SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning. This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.image1M<n<10M0 likes3.2k downloads4mo agoHugging Face13PhilipMay /stsb_multi_mt Dataset Card for STSb Multi MT Dataset Summary STS Benchmark comprises a selection of the English datasets used in the STS tasks organized in the context of SemEval between 2012 and 2017. The selection of datasets include text from image captions, news headlines and user forums. (source) These are different multilingual translations and the English original of the STSbenchmark dataset. Translation has been done with deepl.com. It can be used to train sentence embeddings… See the full description on the dataset page: https://huggingface.co/datasets/PhilipMay/stsb_multi_mt.texttext-classification10K<n<100K68 likes2.6k downloads2y agoHugging Face14philschmid /trl-test-instructiontextn<1K0 likes2.2k downloads3y agoHugging Face15philschmid /dolly-15k-oai-style Dataset Card for "dolly-15k-oai-style" More Information needed text10K<n<100K7 likes2.2k downloads3y agoHugging Face16zefang-liu /phishing-email-dataset Phishing Email Dataset This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page. texttext-classification10K<n<100K38 likes2.1k downloads3y agoHugging Face17philschmid /guanaco-sharegpt-style Dataset Card for "guanaco-sharegpt-style" More Information needed text1K<n<10K49 likes2k downloads3y agoHugging Face18phiyodr /coco2017 coco2017 Image-text pairs from MS COCO2017. Data origin Data originates from cocodataset.org While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy. phiyodr/coco2017: One row corresponds one image with several sentences. phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.imageimage-to-text100K<n<1M29 likes1.9k downloads3y agoHugging Face19AreLit /PhishNChips PhishNChips: A Benchmark for LLM Email-Agent Security PhishNChips is a large-scale benchmark for evaluating how system prompt configurations influence the security behavior of LLM-based email agents. This repository contains the canonical v5.2 release, featuring 2,000 email stimuli and 220,000 adjudicated model evaluations. Dataset Overview The benchmark measures a critical deployment variable: how strongly an LLM's system prompt shapes its phishing detection capabilities… See the full description on the dataset page: https://huggingface.co/datasets/AreLit/PhishNChips.texttext-classification1K<n<10K2 likes1.8k downloads6mo agoHugging Face20phields /a-share-l2-market-depth China A-share Level 2 Market Depth Canonical order-event and ten-level snapshot data for China A-shares. Canonical trade records remain in the separate phields/a-share-l2-trades dataset. Coverage Date range: 2026-07-24 to 2026-07-24 Trading days: 1 Table Rows Parquet files Compressed size l2_orders 249,705,486 10 2.14 GiB l2_snapshots 20,279,887 4 0.91 GiB Layout… See the full description on the dataset page: https://huggingface.co/datasets/phields/a-share-l2-market-depth.tabular10B<n<100B0 likes1.5k downloads2mo agoHugging Face21PhisherJR /850M-India-datatext100M<n<1B1 likes1.5k downloads3mo agoHugging Face22kierth /retail-products-philippinesimage1K<n<10K1 likes1.5k downloads5mo agoHugging Face23PhisherJR /GlobalTGtext100K<n<1M0 likes1.2k downloads2d agoHugging Face24AiresPucrs /stanford-encyclopedia-philosophy Stanford Encyclopedia Philosophy (Teeny-Tiny Castle) This dataset is part of a tutorial tied to the Teeny-Tiny Castle, an open-source repository containing educational tools for AI Ethics and Safety research. How to Use from datasets import load_dataset dataset = load_dataset("AiresPucrs/stanford-encyclopedia-philosophy", split = 'train') texttext-classification100K<n<1M54 likes1k downloads2y agoHugging Face25ealvaradob /phishing-datasetDataset designed for phishing classification tasks in various data types.texttext-classification10K<n<100K64 likes978 downloads3y agoHugging Face26philgzl /ears EARS: Expressive Anechoic Recordings of Speech This is a mirror of the Expressive Anechoic Recordings of Speech (EARS) dataset. The original files were converted from WAV to Opus to reduce the size and accelerate streaming. Sampling rate: 48 kHz Channels: 1 Format: Opus Splits: Train: 92 hours, 15939 utterances, speakers p001 to p099 Validation: 2 hours, 322 utterances, speakers p100 and p101 Test: 6 hours, 966 utterances, speakers p102 to p107 License: CC BY-NC 4.0 Source:… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/ears.audio10K<n<100K0 likes892 downloads1y agoHugging Face27philgzl /fsd50k FSD50K: An open dataset of human-labeled sound events This is a mirror of the FSD50K sound event dataset. The original files were converted from WAV to Opus to reduce the size and accelerate streaming. Sampling rate: 48 kHz Channels: 1 Format: Opus Splits: Dev: 80 hours, 40966 clips. Eval: 28 hours, 10231 clips. License: FSD50K is released under CC-BY. However, each clip has its own licence. Clip licenses include CC0, CC-BY, CC-BY-NC and CC Sampling+. Clip licenses are specified… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/fsd50k.audio10K<n<100K0 likes871 downloads1y agoHugging Face28renzzyyy1028 /civil-code-phil Civilex — Philippine Legal RAG & SFT Dataset Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline. Dataset structure . ├── README.md ├── civil_code_rag.jsonl # Civil Code articles + hierarchy + citation linkage ├── jurisprudence_chunks.jsonl # RAG-ready chunks… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.textquestion-answering100K<n<1M0 likes791 downloads2d agoHugging Face29if001 /hle_math_category_phi4textn<1K0 likes770 downloads1y agoHugging Face30philipp-zettl /inaturalist-enriched Enriched iNaturalist dataset from 2026-03-27. This dataset is based on philipp-zettl/inaturalist-s3-massive. The data was enriched using the ./enrich.py script inside the repository. It contains the following features photo_id: The original ID of the photo inside the inaturalist dataset observation_uuid: The observation's UUID image: The actual image content taxon_id: The ID of the taxonomy species_name: The name of the species inside the image taxonomic_rank: The type of taxonomic rank… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/inaturalist-enriched.text1M<n<10M1 likes756 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.