Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zefang-liu /phishing-email-dataset Phishing Email Dataset This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page. texttext-classification10K<n<100K38 likes1.9k downloads3y agoHugging Face02AreLit /PhishNChips PhishNChips: A Benchmark for LLM Email-Agent Security PhishNChips is a large-scale benchmark for evaluating how system prompt configurations influence the security behavior of LLM-based email agents. This repository contains the canonical v5.2 release, featuring 2,000 email stimuli and 220,000 adjudicated model evaluations. Dataset Overview The benchmark measures a critical deployment variable: how strongly an LLM's system prompt shapes its phishing detection capabilities… See the full description on the dataset page: https://huggingface.co/datasets/AreLit/PhishNChips.texttext-classification1K<n<10K2 likes1.4k downloads6mo agoHugging Face03alexkstern /phishing_urlstext100K<n<1M4 likes523 downloads3y agoHugging Face04Mitake /PhishingURLsANDBenignURLstext100K<n<1M3 likes517 downloads4y agoHugging Face05yjkim27 /The-Philosophy-Data-Project About dataset The Philosophy Data Project is a corpus and a set of anaylsis based philosophy texts, totaling over 50 texts and 30 authors, made by Kourosh Alizadeh. school: Broad categorization of which school of thought each book belongs to. Sometimes, this classification can be vague or depend on interpretation. Thankfully, texts in this corpus are all distinctive examples of respective school of thought, so at leat here they are reasonable. sentence_spacy and sentence_str:… See the full description on the dataset page: https://huggingface.co/datasets/yjkim27/The-Philosophy-Data-Project.tabular100K<n<1M13 likes514 downloads3y agoHugging Face06philschmid /AIME_1983_2024Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from year 1983~2023, and 2024(part 2). Original: https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions 2024(part 1) can be find at https://huggingface.co/datasets/AI-MO/aimo-validation-aime. tabularn<1K0 likes420 downloads2y agoHugging Face07SM-Bello /PHI-SPIKE-C172x-Community-Dataset-v1.0 PHI-SPIKE C172X Community Dataset v1.0 Dataset Summary PHI-SPIKE C172X Community Dataset v1.0 is a simulation-based aerospace Prognostics and Health Management (PHM) dataset and training-artifact release developed from the PHI-SPIKE C172X research campaign. The release provides: JSBSim C172X reference telemetry; benchmark metadata; training histories; trained PyTorch model checkpoints; per-run evaluation metrics; and five-seed campaign summaries. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-SPIKE-C172x-Community-Dataset-v1.0.tabulartime-series-forecasting1K<n<10K1 likes373 downloads21d agoHugging Face08nhatnguyet /cung-phi-bat-trach Cung phi và hướng Bát Trạch Kua number and Bat Trach directions 1. Mô tả · Description Cung phi theo năm sinh và giới tính cho khoảng 1900 tới 2099, kèm bốn hướng tốt và bốn hướng cần tránh. Kua number by birth year and sex for 1900 to 2099, with the four favourable and four unfavourable directions. Số dòng · Rows: 400 Vai trò · Role: tri-thuc (tri thức · knowledge) Loại bộ · Dataset role: tinh-toan (computed) Loại bằng chứng · Evidence types: A tai-tinh-duoc, C… See the full description on the dataset page: https://huggingface.co/datasets/nhatnguyet/cung-phi-bat-trach.tabularn<1K0 likes319 downloads15h agoHugging Face09LingoIITGN /PHINCAbstract Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities, it is a very frequently observed pattern of communication on social media platforms. Flexibility to use multiple languages in one text message might help to communicate efficiently with the target audience. But, the noisy user-generated code-mixed text adds to the challenge of processing and understanding natural language to a much larger extent. Machine translation from… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/PHINC.texttranslation10K<n<100K1 likes262 downloads2y agoHugging Face10imanoop7 /phishing_url_classification Phishing URL Classification Dataset This dataset contains URLs labeled as 'Safe' (0) or 'Not Safe' (1) for phishing detection tasks. Dataset Summary This dataset contains URLs labeled for phishing detection tasks. It's designed to help train and evaluate models that can identify potentially malicious URLs. Dataset Creation The dataset was synthetically generated using a custom script that creates both legitimate and potentially phishing URLs. This approach… See the full description on the dataset page: https://huggingface.co/datasets/imanoop7/phishing_url_classification.texttext-classification100K<n<1M5 likes258 downloads2y agoHugging Face11phihung /titanicThe legendary Titanic dataset from this Kaggle competition tabularn<1K9 likes250 downloads4y agoHugging Face12phil329 /OpenVid-1M-mapping Summary This is the extent dataset proposed in the paper "OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation". OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets. New Feature: Video-ZIP mapping files now available for efficient video lookup (see Dataset… See the full description on the dataset page: https://huggingface.co/datasets/phil329/OpenVid-1M-mapping.texttext-to-video1M<n<10M0 likes222 downloads2y agoHugging Face13locuoco /the-biggest-spam-ham-phish-email-dataset-300000 The Biggest Spam Ham Phish Email Dataset (250000+) This dataset is a large-scale, unified, and deduplicated collection of text messages and emails created for spam, ham, and phishing detection. It has been constructed by combining multiple publicly available and open-source datasets into a single standardized format, making it suitable for machine learning, deep learning, and NLP-based projects. The dataset contains approximately unique 250,000+ samples, covering a diverse range of… See the full description on the dataset page: https://huggingface.co/datasets/locuoco/the-biggest-spam-ham-phish-email-dataset-300000.texttext-classification100K<n<1M0 likes182 downloads8mo agoHugging Face14jhonrayo99 /phishing-email-balanced-6000 Balanced Phishing Email Detection Subset This dataset is a derived, randomly sampled subset of Cyber Cop's Phishing Email Detection dataset on Kaggle. The original dataset is distributed under the GNU Lesser General Public License 3.0. Dataset structure The file phishing_email_subset.csv contains 6,000 English email examples: text: email text. label: 0 for a safe email and 1 for a phishing email. Label Class Examples 0 Safe email 3,000 1 Phishing… See the full description on the dataset page: https://huggingface.co/datasets/jhonrayo99/phishing-email-balanced-6000.texttext-classification1K<n<10K0 likes168 downloads2mo agoHugging Face15Dizzzy0x00 /LLMGen-Phishing-Email-Dataset LLM-Generated Phishing Email Dataset Dataset Description This dataset comprises a collection of phishing and legitimate emails generated using Large Language Models (LLMs), specifically DeepSeek for Chinese emails and OpenAI models for English emails. The primary purpose of this dataset is to facilitate research and development in phishing email detection and classification. The dataset is structured with two key columns: content: The full text content of the email.… See the full description on the dataset page: https://huggingface.co/datasets/Dizzzy0x00/LLMGen-Phishing-Email-Dataset.texttext-classification1K<n<10K2 likes140 downloads10mo agoHugging Face16sayhan /strix-philosophy-qa Strix 134k question-answer pairs based on AiresPucrs' stanford-encyclopedia-philosophy dataset. textquestion-answering100K<n<1M29 likes107 downloads3y agoHugging Face17PhilippXXY /AudibleLight_Eigenmike32-5_DCASE-STARSS23_Dataset AudibleLight Eigenmike32-5 DCASE-STARSS23 Dataset This dataset contains 121 synthetic spatial audio scenes — 111 for training and 10 for evaluation — of 60 seconds each, generated with the AudibleLight dataset generator (DOI). Each scene is rendered as five independent simulated Eigenmike32 captures, with 32 channels per capture, resulting in 570 minutes of multichannel audio in total at 24 kHz. Foreground Audio Foreground events are sampled from ESC-50: Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/PhilippXXY/AudibleLight_Eigenmike32-5_DCASE-STARSS23_Dataset.audion<1K0 likes94 downloads6mo agoHugging Face18datastax /philosopher-quotes450 quotes by 9 philosophers (50 quotes each), labeled with the author and with a variable number of topic tags. The quotes originally come from https://www.kaggle.com/datasets/mertbozkurt5/quotes-by-philosophers (CC BY-NC-SA 4.0). The text of each quote has been cleaned of soft-hyphens (\xad) and other weird characters. The topic labeling has been done with a default HuggingFace zero-shot classifier pipeline with multi_labels. textn<1K9 likes88 downloads3y agoHugging Face19philosopher-from-god /ChatGPT-Jailbreak-Prompts-rubend18 Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K2 likes87 downloads1y agoHugging Face20Rishik001 /PhishingURLDatasetstabular10K<n<100K1 likes85 downloads1y agoHugging Face21Febriyansyah /phishing-emails-multilingual Phishing Emails Multilingual (ID/EN) — Synthetic Dataset sintetis & edukatif 600 email dwibahasa Indonesia 🇮🇩 & English 🇺🇸 untuk riset deteksi phishing — oleh Febriyansyah. ⚠️ Synthetic & edu-defense-only — dibuat untuk pembelajaran defensive security, bukan untuk kampanye nyata. Jangan gunakan untuk aktivitas ilegal. Ringkasan 600 baris — 300 phishing / 300 benign (seimbang), 321 EN / 279 ID Kolom: id (int), language (id/en), text (string, badan email)… See the full description on the dataset page: https://huggingface.co/datasets/Febriyansyah/phishing-emails-multilingual.tabulartext-classificationn<1K0 likes79 downloads28d agoHugging Face22jmLuis /MediaFrameCorpus-PhilippineFrameCorpus-CombinedThis training and validation dataset is a combination of Media Frame Corpus and Philippine Frame Corpus, labeled using the Policy Issue Frames Codebook. Train-test split of 80-20. Code_frames column contains annotations following the PolicyIssue Frames Codebook (1-15), wherein at least two(2) annotators agree with the label. The text column contains sentences/phrases from online news articles. The label column is the 0th index code_frames used for training. tabulartext-classification10K<n<100K0 likes78 downloads3y agoHugging Face23mfgiguere /erudit-french-philosophy Dataset Card for Dataset Name Dataset Description Dataset Summary This dataset contains all french philosophy that has been published on erudit.org. It has been generated using a Bs4 web parser that you can find in this repo: https://github.com/MFGiguere/french-philosophy-generator. Supported Tasks and Leaderboards This dataset could be useful for this (non-exhaustive) set of tasks: detect if a text is philosophical or not, generate philosophical… See the full description on the dataset page: https://huggingface.co/datasets/mfgiguere/erudit-french-philosophy.tabular100K<n<1M2 likes69 downloads3y agoHugging Face24bgspaditya /phishing-datasettext100K<n<1M5 likes69 downloads3y agoHugging Face25kjhq /Philippines-Stock-Symbols-and-Metadata Philippines Stock Symbols & Company Metadata This dataset contains stock symbols and basic company metadata for all listed companies in Philippines.It is updated weekly if new changes are there. 📊 Dataset Contents The dataset is provided as a CSV file with the following columns: Column Description name Full company name ticker Stock ticker symbol (e.g., AAPL, MSFT) market The exchange/market where the stock is listed sector The primary business sector… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Philippines-Stock-Symbols-and-Metadata.textn<1K0 likes65 downloads1y agoHugging Face26gelcloudy /philippine-elections-2025 Philippine Elections 2025 Dataset (COMELEC) This dataset contains Philippine election results from COMELEC, including both local and overseas voting data. Dataset Structure The dataset is provided in two formats: 1. Combined Datasets combined_local.csv — consolidated local election data combined_overseas.csv — consolidated overseas election data These files merge all regions into a single dataset for easier analysis. 2. Per-Region Datasets… See the full description on the dataset page: https://huggingface.co/datasets/gelcloudy/philippine-elections-2025.tabular10M<n<100M1 likes65 downloads6mo agoHugging Face27SM-Bello /PHI-CTRL-F16-Fault-Recovery-Telemetry PHI-CTRL F-16 Actuator Fault Recovery Dataset High-Fidelity JSBSim 6-DOF Telemetry for Physics-Hybrid Self-Healing Flight Control Official verification artifacts of the PHI-CTRL (Physics-Hybrid Integrity Control) architecture — a digital-twin-driven, self-healing flight control framework that actively compensates actuator degradation in real time. Author: Mohammed Bello Sani (SM-Bello) Affiliation: Air Force Institute of Technology (AFIT), Kaduna · Penelope Inc. / PHI Lab… See the full description on the dataset page: https://huggingface.co/datasets/SM-Bello/PHI-CTRL-F16-Fault-Recovery-Telemetry.tabulartime-series-forecasting10K<n<100K0 likes62 downloads28d agoHugging Face28Parv-09 /phi_so101_8bin_v1_trim phi_so101_8bin_v1 — opening-pause trim table Companion to BrutalCaesar/phi_so101_8bin_v1. This is not a dataset. It is a 119-row table plus the script that produced it. The original dataset is unmodified and remains authoritative. Applying this table excludes each episode's pre-teleop dead air as a chunk start point, without deleting a single frame from disk. Why Every episode begins with the arm sitting still while the operator has not yet moved the leader.… See the full description on the dataset page: https://huggingface.co/datasets/Parv-09/phi_so101_8bin_v1_trim.tabularn<1K0 likes57 downloads2mo agoHugging Face29mdsajjadullah /bangla-phishing-detection-2026 Bangla Phishing Detection Dataset (SMS, Email, URLs) 2026 Synthetic dataset (~2000 rows) of phishing and legitimate messages in Bangla (Bengali) + some English, simulating common Bangladesh scams (bKash, Nagad, Daraz, Eid offers, job fraud, account lock alerts, etc.). Research Motivation Phishing/smishing attacks are rising in Bangladesh and South Asia, often in Bangla using local services. Most phishing datasets are English-only and miss these patterns.This dataset fills… See the full description on the dataset page: https://huggingface.co/datasets/mdsajjadullah/bangla-phishing-detection-2026.tabulartext-classification1K<n<10K0 likes56 downloads7mo agoHugging Face30marcuscedricridia /philippines-typhoon-tracks Philippines Typhoon Tracks IBTrACS v04r01 Western Pacific basin tracks joined with intensity classification. Fields storm_id — IBTrACS Storm ID time — ISO timestamp (3-hourly) lat, lon — position (decimal degrees) wind_kt — max sustained wind (knots) WMO_PRES — min central pressure (mb, may be NaN) intensity — Saffir-Simpson class (0=TD/TS, 1=Cat 1-2, 2=Cat 3-4, 3=Cat 5) Stats 26,786 rows 895 storms 1980–2025 Source IBTrACS v04r01… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/philippines-typhoon-tracks.tabulartime-series-forecasting10K<n<100K0 likes56 downloads23d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.