datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies.
Magicoder-Evol-Instruct-110KA decontaminated version of evol-codealpaca-v1. Decontamination is done in the same way as StarCoder (bigcode decontamination process).
UIISThis dataset is proposed by the ICCV 2023 paper "WaterMask: Instance Segmentation for Underwater Imagery", specific parameters about the dataset can be viewed in the paper
The Underwater Image Instance Segmentation (UIIS) dataset contains 4,628 images with pixel-level annotations in seven categories used for the underwater instance segmentation task. The dataset is organized in MS COCO format and the annotation files and images for training and testing are in UDW files.
Updatae:… See the full description on the dataset page: https://huggingface.co/datasets/LiamLian0727/UIIS.BLEnD
BLEnD
This is the official repository of BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages (Submitted to NeurIPS 2024 Datasets and Benchmarks Track).
24/12/05: Updated translation errors25/05/02: Updated multiple choice questions file (v1.1)26/09/15: Added new data collected for SemEval-2026 Task 7, covering 17 additional language-culture pairs (semeval-annotations, semeval-questions, and semeval split of multiple-choice-questions)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/BLEnD.MedVidBench
MedVidBench: A Benchmark for Medical Video Understanding
Introduced in the paper: MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding (CVPR 2026).
📄 Paper: arxiv.org/abs/2512.06581
🌐 Project Page: uii-ai.github.io/MedGRPO
💻 Code: UII-AI/MedGRPO-Code
🤗 Model: UII-AI/uAI-NEXUS-MedVLM-1.0a-7B-RL
🎮 Demo: UII-AI/MedGRPO-Demo
📊 Leaderboard: UII-AI/MedVidBench-Leaderboard
Dataset Description
MedVidBench is a test benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/UII-AI/MedVidBench.wave-uiLICENSE
pali-commentary-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑)
เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒)
เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓)
เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑)
เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒)
เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.tripitaka-siamrath
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาไทยฉบับสยามรัฏฐ จำนวน 45 เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม 1 (754 หน้า): พระวินัยปิฎก เล่ม ๑ มหาวิภังค์ ปฐมภาค
เล่ม 2 (717 หน้า): พระวินัยปิฎก เล่ม ๒ มหาวิภังค์ ทุติภาค
เล่ม 3 (328 หน้า): พระวินัยปิฎก เล่ม ๓ ภิกขุณี วิภังค์
เล่ม 4 (304 หน้า): พระวินัยปิฎก เล่ม ๔ มหาวรรคภาค ๑
เล่ม 5 (278… See the full description on the dataset page: https://huggingface.co/datasets/uisp/tripitaka-siamrath.UI-Grounding-Benchmarks
UI-Grounding-Benchmarks
This is a collection of UI grounding benchmarks:
ScreenSpot
ScreenSpot-V2
ScreenSpot-Pro
OS-World-G
UI-Vision
Thanks for their great work!
This benchmark collection is used in the paper:
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
🖼️ Project Page: https://showlab.github.io/FocusUI/
🏠 Github Repo: https://github.com/showlab/FocusUI
📝 Paper: https://arxiv.org/pdf/2601.03928
Model Zoo
Model
Backbone
🤗… See the full description on the dataset page: https://huggingface.co/datasets/yyyang/UI-Grounding-Benchmarks.UIT-ViQuAD2.0
Vietnamese Question Answering Dataset
Dataset Card for UIT-ViQuAD2.0
Dataset Summary
The HF version for Vietnamese QA dataset created by Nguyen et al. (2020) and released in the shared task.
The original UIT-ViQuAD contains over 23,000 QA pairs based on 174 Vietnamese Wikipedia articles. UIT-ViQuAD2.0 adds over 12K unanswerable questions for the same passage.
The dataset has been processed to remove a few duplicated questions and answers.
Version 2.0 contains… See the full description on the dataset page: https://huggingface.co/datasets/taidng/UIT-ViQuAD2.0.vietnamese_students_feedbackStudents’ feedback is a vital resource for the interdisciplinary research involving the combining of two different
research fields between sentiment analysis and education.
Vietnamese Students’ Feedback Corpus (UIT-VSFC) is the resource consists of over 16,000 sentences which are
human-annotated with two different tasks: sentiment-based and topic-based classifications.
To assess the quality of our corpus, we measure the annotator agreements and classification evaluation on the
UIT-VSFC corpus. As a result, we obtained the inter-annotator agreement of sentiments and topics with more than over
91% and 71% respectively. In addition, we built the baseline model with the Maximum Entropy classifier and achieved
approximately 88% of the sentiment F1-score and over 84% of the topic F1-score.wave-ui-25k
WaveUI-25k
This dataset contains 25k examples of labeled UI elements. It is a subset of a collection of ~80k preprocessed examples assembled from the following sources:
WebUI
RoboFlow
GroundUI-18K
These datasets were preprocessed to have matching schemas and to filter out unwanted examples, such as duplicated, overlapping and low-quality datapoints. We also filtered out many text elements which were not in the main scope of this work.
The WaveUI-25k dataset includes the original… See the full description on the dataset page: https://huggingface.co/datasets/agentsea/wave-ui-25k.Real-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.ui-navigation-corpus
User Interface (Navigation) Corpus
Overview
This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them.
Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent.
Dataset Structure
The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/ijlewis/ui-navigation-corpus.easyr1-103k-4MP-jedi-ui-vision-gta1-data
easyr1-103k-4MP-jedi-ui-vision-gta1-data
Merged dataset composed of the following sources:
datasets/easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP (63031 samples in split train)
datasets/easyr1-grounding-gta1-4MP-easy-qwen7b-hard-gta1-7b (39943 samples in split train)
Summary
Generated on: 2025-09-18 06:29:16 UTC
Split: train
Column strategy: intersection
Samples after merge: 102974
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-103k-4MP-jedi-ui-vision-gta1-data.vihsd
Dataset Card for Dataset Name
ViHSD - Vietnamese Hate Speech Detection Dataset
Dataset Details
Dataset Description
The dataset contains about 33K annotated comments from social networks. Each has one of three labels: HATE, OFFENSIVE, CLEAN
Uses
Use directly from the Hugging face dataset loader
Direct Use
from datasets import load_dataset
train = load_dataset("sonlam1102/vihsd", split="train")
dev = load_dataset("sonlam1102/vihsd"… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/vihsd.pubmed-abstract
Dataset Summary
A daily-updated dataset of PubMed abstracts, collected via PubMed’s API and published on Hugging Face Datasets.Each snapshot is versioned by date (e.g., 2025-03-28) so users can track historical changes or use a consistent snapshot for reproducibility.
Updated daily
Each version tagged by date
Abstract-only dataset (no full text)
Dataset Structure
Column
Type
Description
pmid
string
Unique PubMed identifier
abstract
string
Abstract text… See the full description on the dataset page: https://huggingface.co/datasets/uiyunkim-hub/pubmed-abstract.ML-Based-Malicious-Package-Detection
DySec
A Machine Learning-Based Dynamic Analysis For Detecting Malicious Packages In PyPI Ecosystem
Overview
Malicious python packages make software supply chains vulnerable by exploiting trust in open-source repositories like PyPI.Lack of real-time behavioral monitoring renders metadata inspection and static code analysis inadequate against advanced attack strategies such as typosquatting, covert remote access activation, and dynamic payload generation.
To… See the full description on the dataset page: https://huggingface.co/datasets/Hugnd-UIT/ML-Based-Malicious-Package-Detection.UIISThis dataset is proposed by the ICCV 2023 paper "WaterMask: Instance Segmentation for Underwater Imagery", specific parameters about the dataset can be viewed in the paper
The Underwater Image Instance Segmentation (UIIS) dataset contains 4,628 images with pixel-level annotations in seven categories used for the underwater instance segmentation task. The dataset is organized in MS COCO format and the annotation files and images for training and testing are in UDW files.
Updatae:… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeZongzi/UIIS.web-ui-grounding-jsonui-navigation-corpus
User Interface (Navigation) Corpus
Overview
This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them.
Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent.
Dataset Structure
The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/teleren/ui-navigation-corpus.mobile-ui-design
Dataset: Mobile UI Design Detection
Introduction
This dataset is designed for object detection tasks with a focus on detecting elements in mobile UI designs. The targeted objects include text, images, and groups. The dataset contains images and object detection boxes, including class labels and location information.
Dataset Content
Load the dataset and take a look at an example:
>>> from datasets import load_dataset
>>>> ds =… See the full description on the dataset page: https://huggingface.co/datasets/mrtoy/mobile-ui-design.pali-tripitaka-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
...
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินย. มหาวิภงฺโค (๑)
เล่ม ๒: วินย. มหาวิภงฺโค (๒)
เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค
เล่ม ๔: วินย. มหาวคฺโค (๑)
เล่ม ๕: วินย. มหาวคฺโค (๒)
เล่ม ๖: วินย. จุลฺลวคฺโค (๑)
เล่ม ๗: วินย. จุลฺลวคฺโค (๒)
เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.pretext-ui-harbor-runs-v0
Pretext UI Harbor Runs
Harbor task-generation and solve-run corpus for the @chenglou/pretext UI task family. The dataset contains flat training indexes plus raw redacted Harbor artifacts.
Contents
data/train/attempts.jsonl: one row per candidate attempt, with model bucket, reward, Gemini score, prompt, and raw artifact pointers.
data/train/tasks.jsonl: materialized task identity/hash index.
data/train/sft_conversations.jsonl: OpenAI-style user/assistant conversation rows… See the full description on the dataset page: https://huggingface.co/datasets/RLAIF/pretext-ui-harbor-runs-v0.ViANLI
Dataset Card for “ViANLI”
Dataset Summary
ViANLI (Vietnamese Adversarial Natural Language Inference) is the first adversarial benchmark dataset for Vietnamese NLI, designed to evaluate model robustness against complex linguistic phenomena.
The dataset was constructed using a human-and-machine-in-the-loop approach with multi-round adversarial generation and dual human–machine verification.
ViANLI contains over 10,000 high-quality premise–hypothesis pairs across 13 diverse… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/ViANLI.instruct_uie_nerUiPad
UiPad - UI Parsing and Accessibility Dataset
📌 Dataset status: stable release. UiPad was built for the IASA Champ 2024 Challenge and is a complete, fixed research artifact. No further updates are planned.
Curated by: MacPaw Way Ltd.
Language(s): Mostly EN, UA
License: MIT
Overview
UiPad is a dataset created for the IASA Champ 2024 Challenge, focusing on the accessibility and interface understanding of MacOS applications. With growing interest in AI-driven user interface… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/UiPad.ui-vision-grounding-4MPuinst
uinst dataset
Original version: Unnatural-Instructions
This dataset is used to evaluate MiniLLM
UIT-ViHSD
UIT-ViHSD Dataset
This is a copy instance of the original dataset provided by UIT. Please visit https://nlp.uit.edu.vn/datasets to obtain a usage permission before using this dataset.
