Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ise-uiuc /Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies. tabulartext-generation10K<n<100K173 likes47k downloads3y agoHugging Face02ise-uiuc /Magicoder-Evol-Instruct-110KA decontaminated version of evol-codealpaca-v1. Decontamination is done in the same way as StarCoder (bigcode decontamination process). texttext-generation100K<n<1M190 likes30k downloads3y agoHugging Face03LiamLian0727 /UIISThis dataset is proposed by the ICCV 2023 paper "WaterMask: Instance Segmentation for Underwater Imagery", specific parameters about the dataset can be viewed in the paper The Underwater Image Instance Segmentation (UIIS) dataset contains 4,628 images with pixel-level annotations in seven categories used for the underwater instance segmentation task. The dataset is organized in MS COCO format and the annotation files and images for training and testing are in UDW files. Updatae:… See the full description on the dataset page: https://huggingface.co/datasets/LiamLian0727/UIIS.imageimage-segmentation1K<n<10K1 likes4.1k downloads1y agoHugging Face04uilab /BLEnD BLEnD This is the official repository of BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages (Submitted to NeurIPS 2024 Datasets and Benchmarks Track). 24/12/05: Updated translation errors25/05/02: Updated multiple choice questions file (v1.1)26/09/15: Added new data collected for SemEval-2026 Task 7, covering 17 additional language-culture pairs (semeval-annotations, semeval-questions, and semeval split of multiple-choice-questions)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/BLEnD.textquestion-answering100K<n<1M16 likes2.5k downloads25d agoHugging Face05UII-AI /MedVidBench MedVidBench: A Benchmark for Medical Video Understanding Introduced in the paper: MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding (CVPR 2026). 📄 Paper: arxiv.org/abs/2512.06581 🌐 Project Page: uii-ai.github.io/MedGRPO 💻 Code: UII-AI/MedGRPO-Code 🤗 Model: UII-AI/uAI-NEXUS-MedVLM-1.0a-7B-RL 🎮 Demo: UII-AI/MedGRPO-Demo 📊 Leaderboard: UII-AI/MedVidBench-Leaderboard Dataset Description MedVidBench is a test benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/UII-AI/MedVidBench.imagevideo-classification1K<n<10K14 likes1.4k downloads5mo agoHugging Face06agentsea /wave-uiLICENSE image10K<n<100K27 likes1.3k downloads2y agoHugging Face07uisp /pali-commentary-thai-script-siamrath-version Multi-File CSV Dataset คำอธิบาย อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์ 01/010001.csv: เล่ม 1 หน้า 1 01/010002.csv: เล่ม 1 หน้า 2 ... 02/020001.csv: เล่ม 2 หน้า 1 คำอธิบายของแต่ละเล่ม เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑) เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒) เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓) เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑) เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒) เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.tabular100K<n<1M3 likes1.2k downloads2y agoHugging Face08uisp /tripitaka-siamrath Multi-File CSV Dataset คำอธิบาย พระไตรปิฎกภาษาไทยฉบับสยามรัฏฐ จำนวน 45 เล่ม ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์ 01/010001.csv: เล่ม 1 หน้า 1 01/010002.csv: เล่ม 1 หน้า 2 ... 02/020001.csv: เล่ม 2 หน้า 1 คำอธิบายของแต่ละเล่ม เล่ม 1 (754 หน้า): พระวินัยปิฎก เล่ม ๑ มหาวิภังค์ ปฐมภาค เล่ม 2 (717 หน้า): พระวินัยปิฎก เล่ม ๒ มหาวิภังค์ ทุติภาค เล่ม 3 (328 หน้า): พระวินัยปิฎก เล่ม ๓ ภิกขุณี วิภังค์ เล่ม 4 (304 หน้า): พระวินัยปิฎก เล่ม ๔ มหาวรรคภาค ๑ เล่ม 5 (278… See the full description on the dataset page: https://huggingface.co/datasets/uisp/tripitaka-siamrath.tabular100K<n<1M10 likes952 downloads2y agoHugging Face09yyyang /UI-Grounding-Benchmarks UI-Grounding-Benchmarks This is a collection of UI grounding benchmarks: ScreenSpot ScreenSpot-V2 ScreenSpot-Pro OS-World-G UI-Vision Thanks for their great work! This benchmark collection is used in the paper: FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection 🖼️ Project Page: https://showlab.github.io/FocusUI/ 🏠 Github Repo: https://github.com/showlab/FocusUI 📝 Paper: https://arxiv.org/pdf/2601.03928 Model Zoo Model Backbone 🤗… See the full description on the dataset page: https://huggingface.co/datasets/yyyang/UI-Grounding-Benchmarks.image1K<n<10K2 likes821 downloads8mo agoHugging Face10taidng /UIT-ViQuAD2.0 Vietnamese Question Answering Dataset Dataset Card for UIT-ViQuAD2.0 Dataset Summary The HF version for Vietnamese QA dataset created by Nguyen et al. (2020) and released in the shared task. The original UIT-ViQuAD contains over 23,000 QA pairs based on 174 Vietnamese Wikipedia articles. UIT-ViQuAD2.0 adds over 12K unanswerable questions for the same passage. The dataset has been processed to remove a few duplicated questions and answers. Version 2.0 contains… See the full description on the dataset page: https://huggingface.co/datasets/taidng/UIT-ViQuAD2.0.textquestion-answering10K<n<100K21 likes779 downloads2y agoHugging Face11uitnlp /vietnamese_students_feedbackStudents’ feedback is a vital resource for the interdisciplinary research involving the combining of two different research fields between sentiment analysis and education. Vietnamese Students’ Feedback Corpus (UIT-VSFC) is the resource consists of over 16,000 sentences which are human-annotated with two different tasks: sentiment-based and topic-based classifications. To assess the quality of our corpus, we measure the annotator agreements and classification evaluation on the UIT-VSFC corpus. As a result, we obtained the inter-annotator agreement of sentiments and topics with more than over 91% and 71% respectively. In addition, we built the baseline model with the Maximum Entropy classifier and achieved approximately 88% of the sentiment F1-score and over 84% of the topic F1-score.texttext-classification10K<n<100K32 likes748 downloads4y agoHugging Face12agentsea /wave-ui-25k WaveUI-25k This dataset contains 25k examples of labeled UI elements. It is a subset of a collection of ~80k preprocessed examples assembled from the following sources: WebUI RoboFlow GroundUI-18K These datasets were preprocessed to have matching schemas and to filter out unwanted examples, such as duplicated, overlapping and low-quality datapoints. We also filtered out many text elements which were not in the main scope of this work. The WaveUI-25k dataset includes the original… See the full description on the dataset page: https://huggingface.co/datasets/agentsea/wave-ui-25k.image10K<n<100K40 likes741 downloads2y agoHugging Face13DUDE-Framework /Real-UI-Clickboxes RUC: Real UI Clickboxes Click carefully, even when the page is trying to trick you! 👀 Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents. ACL Anthology: https://aclanthology.org/2026.acl-long.310/ PDF: https://aclanthology.org/2026.acl-long.310.pdf DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.imageimage-text-to-text1K<n<10K1 likes683 downloads3mo agoHugging Face14ijlewis /ui-navigation-corpus User Interface (Navigation) Corpus Overview This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them. Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent. Dataset Structure The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/ijlewis/ui-navigation-corpus.imageimage-segmentation1M<n<10M0 likes619 downloads10mo agoHugging Face15mlfoundations-cua-dev /easyr1-103k-4MP-jedi-ui-vision-gta1-data easyr1-103k-4MP-jedi-ui-vision-gta1-data Merged dataset composed of the following sources: datasets/easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP (63031 samples in split train) datasets/easyr1-grounding-gta1-4MP-easy-qwen7b-hard-gta1-7b (39943 samples in split train) Summary Generated on: 2025-09-18 06:29:16 UTC Split: train Column strategy: intersection Samples after merge: 102974 Usage from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-103k-4MP-jedi-ui-vision-gta1-data.image100K<n<1M1 likes543 downloads1y agoHugging Face16uitnlp /vihsdgated Dataset Card for Dataset Name ViHSD - Vietnamese Hate Speech Detection Dataset Dataset Details Dataset Description The dataset contains about 33K annotated comments from social networks. Each has one of three labels: HATE, OFFENSIVE, CLEAN Uses Use directly from the Hugging face dataset loader Direct Use from datasets import load_dataset train = load_dataset("sonlam1102/vihsd", split="train") dev = load_dataset("sonlam1102/vihsd"… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/vihsd.texttext-classification10K<n<100K8 likes454 downloads2y agoHugging Face17uiyunkim-hub /pubmed-abstract Dataset Summary A daily-updated dataset of PubMed abstracts, collected via PubMed’s API and published on Hugging Face Datasets.Each snapshot is versioned by date (e.g., 2025-03-28) so users can track historical changes or use a consistent snapshot for reproducibility. Updated daily Each version tagged by date Abstract-only dataset (no full text) Dataset Structure Column Type Description pmid string Unique PubMed identifier abstract string Abstract text… See the full description on the dataset page: https://huggingface.co/datasets/uiyunkim-hub/pubmed-abstract.text10M<n<100M12 likes441 downloads1y agoHugging Face18Hugnd-UIT /ML-Based-Malicious-Package-Detection DySec A Machine Learning-Based Dynamic Analysis For Detecting Malicious Packages In PyPI Ecosystem Overview Malicious python packages make software supply chains vulnerable by exploiting trust in open-source repositories like PyPI.Lack of real-time behavioral monitoring renders metadata inspection and static code analysis inadequate against advanced attack strategies such as typosquatting, covert remote access activation, and dynamic payload generation. To… See the full description on the dataset page: https://huggingface.co/datasets/Hugnd-UIT/ML-Based-Malicious-Package-Detection.text0 likes440 downloads7d agoHugging Face19CoffeeZongzi /UIISThis dataset is proposed by the ICCV 2023 paper "WaterMask: Instance Segmentation for Underwater Imagery", specific parameters about the dataset can be viewed in the paper The Underwater Image Instance Segmentation (UIIS) dataset contains 4,628 images with pixel-level annotations in seven categories used for the underwater instance segmentation task. The dataset is organized in MS COCO format and the annotation files and images for training and testing are in UDW files. Updatae:… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeZongzi/UIIS.imageimage-segmentation1K<n<10K0 likes420 downloads18d agoHugging Face20isaiahbjork /web-ui-grounding-jsontext100K<n<1M1 likes417 downloads2y agoHugging Face21teleren /ui-navigation-corpus User Interface (Navigation) Corpus Overview This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them. Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent. Dataset Structure The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/teleren/ui-navigation-corpus.imageimage-segmentation1M<n<10M15 likes325 downloads2y agoHugging Face22mrtoy /mobile-ui-design Dataset: Mobile UI Design Detection Introduction This dataset is designed for object detection tasks with a focus on detecting elements in mobile UI designs. The targeted objects include text, images, and groups. The dataset contains images and object detection boxes, including class labels and location information. Dataset Content Load the dataset and take a look at an example: >>> from datasets import load_dataset >>>> ds =… See the full description on the dataset page: https://huggingface.co/datasets/mrtoy/mobile-ui-design.imageobject-detection1K<n<10K88 likes309 downloads3y agoHugging Face23uisp /pali-tripitaka-thai-script-siamrath-version Multi-File CSV Dataset คำอธิบาย พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์ 01/010001.csv: เล่ม 1 หน้า 1 01/010002.csv: เล่ม 1 หน้า 2 ... 02/020001.csv: เล่ม 2 หน้า 1 ... คำอธิบายของแต่ละเล่ม เล่ม ๑: วินย. มหาวิภงฺโค (๑) เล่ม ๒: วินย. มหาวิภงฺโค (๒) เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค เล่ม ๔: วินย. มหาวคฺโค (๑) เล่ม ๕: วินย. มหาวคฺโค (๒) เล่ม ๖: วินย. จุลฺลวคฺโค (๑) เล่ม ๗: วินย. จุลฺลวคฺโค (๒) เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.tabular100K<n<1M3 likes303 downloads2y agoHugging Face24RLAIF /pretext-ui-harbor-runs-v0 Pretext UI Harbor Runs Harbor task-generation and solve-run corpus for the @chenglou/pretext UI task family. The dataset contains flat training indexes plus raw redacted Harbor artifacts. Contents data/train/attempts.jsonl: one row per candidate attempt, with model bucket, reward, Gemini score, prompt, and raw artifact pointers. data/train/tasks.jsonl: materialized task identity/hash index. data/train/sft_conversations.jsonl: OpenAI-style user/assistant conversation rows… See the full description on the dataset page: https://huggingface.co/datasets/RLAIF/pretext-ui-harbor-runs-v0.text10K<n<100K0 likes295 downloads5mo agoHugging Face25uitnlp /ViANLI Dataset Card for “ViANLI” Dataset Summary ViANLI (Vietnamese Adversarial Natural Language Inference) is the first adversarial benchmark dataset for Vietnamese NLI, designed to evaluate model robustness against complex linguistic phenomena. The dataset was constructed using a human-and-machine-in-the-loop approach with multi-round adversarial generation and dual human–machine verification. ViANLI contains over 10,000 high-quality premise–hypothesis pairs across 13 diverse… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/ViANLI.texttext-classification10K<n<100K3 likes280 downloads11mo agoHugging Face26bentrevett /instruct_uie_nertext1M<n<10M0 likes279 downloads1y agoHugging Face27macpaw-research /UiPad UiPad - UI Parsing and Accessibility Dataset 📌 Dataset status: stable release. UiPad was built for the IASA Champ 2024 Challenge and is a complete, fixed research artifact. No further updates are planned. Curated by: MacPaw Way Ltd. Language(s): Mostly EN, UA License: MIT Overview UiPad is a dataset created for the IASA Champ 2024 Challenge, focusing on the accessibility and interface understanding of MacOS applications. With growing interest in AI-driven user interface… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/UiPad.imagequestion-answering1K<n<10K16 likes278 downloads2mo agoHugging Face28mlfoundations-cua-dev /ui-vision-grounding-4MPimage1K<n<10K0 likes252 downloads1y agoHugging Face29MiniLLM /uinst uinst dataset Original version: Unnatural-Instructions This dataset is used to evaluate MiniLLM text10K<n<100K1 likes248 downloads2y agoHugging Face30ura-hcmut /UIT-ViHSD UIT-ViHSD Dataset This is a copy instance of the original dataset provided by UIT. Please visit https://nlp.uit.edu.vn/datasets to obtain a usage permission before using this dataset. texttext-classification10K<n<100K1 likes202 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.