datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies.
Magicoder-Evol-Instruct-110KA decontaminated version of evol-codealpaca-v1. Decontamination is done in the same way as StarCoder (bigcode decontamination process).
BLEnD
BLEnD
This is the official repository of BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages (Submitted to NeurIPS 2024 Datasets and Benchmarks Track).
24/12/05: Updated translation errors25/05/02: Updated multiple choice questions file (v1.1)26/09/15: Added new data collected for SemEval-2026 Task 7, covering 17 additional language-culture pairs (semeval-annotations, semeval-questions, and semeval split of multiple-choice-questions)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/BLEnD.MedVidBench
MedVidBench: A Benchmark for Medical Video Understanding
Introduced in the paper: MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding (CVPR 2026).
📄 Paper: arxiv.org/abs/2512.06581
🌐 Project Page: uii-ai.github.io/MedGRPO
💻 Code: UII-AI/MedGRPO-Code
🤗 Model: UII-AI/uAI-NEXUS-MedVLM-1.0a-7B-RL
🎮 Demo: UII-AI/MedGRPO-Demo
📊 Leaderboard: UII-AI/MedVidBench-Leaderboard
Dataset Description
MedVidBench is a test benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/UII-AI/MedVidBench.Real-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.ui-navigation-corpus
User Interface (Navigation) Corpus
Overview
This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them.
Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent.
Dataset Structure
The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/ijlewis/ui-navigation-corpus.ui-navigation-corpus
User Interface (Navigation) Corpus
Overview
This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them.
Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent.
Dataset Structure
The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/teleren/ui-navigation-corpus.ViANLI
Dataset Card for “ViANLI”
Dataset Summary
ViANLI (Vietnamese Adversarial Natural Language Inference) is the first adversarial benchmark dataset for Vietnamese NLI, designed to evaluate model robustness against complex linguistic phenomena.
The dataset was constructed using a human-and-machine-in-the-loop approach with multi-round adversarial generation and dual human–machine verification.
ViANLI contains over 10,000 high-quality premise–hypothesis pairs across 13 diverse… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/ViANLI.UIT-VSMEC
Introduction
Emotion recognition is a higher approach or special case of sentiment analysis. In this task, the result is not produced in terms of either polarity: positive or negative or in the form of rating (from 1 to 5) but of a more detailed level of sentiment analysis in which the result are depicted in more expressions like sadness, enjoyment, anger, disgust, fear and surprise. Emotion recognition plays a critical role in measuring brand value of a product by recognizing… See the full description on the dataset page: https://huggingface.co/datasets/tridm/UIT-VSMEC.PopMCQ
🎯 PopMCQ
Does your model pick the famous answer or the correct one?
📌 Overview
PopMCQ renders the same question six ways. The question and the correct answer never change — only how popular the three distractors are. That makes option popularity an independent variable, so an accuracy swing across S1–S6 is attributable to popularity rather than to question difficulty.
The swings are large. Under the hardest setting (S2), models pick a popular-but-wrong… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/PopMCQ.UIT-VSFCUI_S1_dataset
Introduction
This repository contains the dataset example for UI-S1-7B, presented in UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning.
Project page: https://github.com/X-PLUG/MobileAgent/tree/main/UI-S1
ui-instruct-4k
UI Instruct 4K
A instruction-completion dataset for finetuning language models to specialize in generating Next.js / ShadCN UI components using React, TypeScript, and Tailwind CSS.
Dataset Summary
This dataset was created with the primary goal of finetuning Qwen 3.5 4B to become a specialist at outputting production-ready Next.js and ShadCN-based UI components. Each example consists of a natural language prompt describing a UI component or layout, paired with a clean… See the full description on the dataset page: https://huggingface.co/datasets/iamdyeus/ui-instruct-4k.OBLIQ-IR-Data
OBLIQ-IR-Data
The training data behind DataScience-UIBK/OBLIQ-IR-3B,
plus the retrieval runs and evaluation outputs for every result in
OBLIQ-IR: Training a Dense Retriever for Oblique Queries (EMNLP 2026).
Oblique retrieval is the setting where relevance is decided by a latent attribute — an implicit stance,
an abstract proof strategy, an authorial style, a lossy recollection of a rhetorical exchange — that has
little or no surface expression in the document.
🤖 Model:… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/OBLIQ-IR-Data.UIT-VSMEC
Introduction
Emotion recognition is a higher approach or special case of sentiment analysis. In this task, the result is not produced in terms of either polarity: positive or negative or in the form of rating (from 1 to 5) but of a more detailed level of sentiment analysis in which the result are depicted in more expressions like sadness, enjoyment, anger, disgust, fear and surprise. Emotion recognition plays a critical role in measuring brand value of a product by recognizing… See the full description on the dataset page: https://huggingface.co/datasets/KoaLee/UIT-VSMEC.agent-ui-efficiency-scores
Agent UI Efficiency Scores
Flat lab-synthetic bakeoff table for the public question: which agent UI is cheapest for a given lab task?
Author
Akash Premkumar (akashnaren)
License
Apache-2.0
Hub files
train.jsonl (63), test.jsonl (14), optional scores.jsonl (77 full)
Related
agent-ui-sft, agent-ui-human, ui-mode-router, agent-ui-mode-pairs
Scope
Rows are original lab fiction for a public agent-UI research question. Identifiers are invented.… See the full description on the dataset page: https://huggingface.co/datasets/akashnaren/agent-ui-efficiency-scores.uipc-assets
UIPC Assets
Simulation assets for libuipc -- a Modern C++20 Library of Unified Incremental Potential Contact.
This dataset collects demos, tests, and examples from IPC-related papers and provides scripts to load them as UIPC scenes via pyuipc.
Structure
All assets live under assets/. Each asset is a named subfolder:
assets/
cube_ground/ # asset name
scene.py # script: build_scene(uipc.Scene) -> adds geometry/objects
cube.msh # mesh data… See the full description on the dataset page: https://huggingface.co/datasets/MuGdxy/uipc-assets.UI-Genie-RM-517kThis repository contains the Reward dataset from the paper UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based
Mobile GUI Agents.
Github: https://github.com/Euphoria16/UI-Genie
bird-platinum
BIRD-Platinum 2.5k v1
Original 2,463-example BIRD-Platinum training-candidate dataset from the ReViSQL repository.
The JSON records contain question_id, db_id, question, evidence, SQL, and grading_method.
This upload preserves the source file unchanged.
ui-distill-html-648
ui-distill-html-648
648 single-file HTML UI components, generated by
Ornith-1.0-35B on a single
RTX 3060 12GB, paired with the build request that produced each one.
Built to fine-tune a 3B model into writing UI
(DogukanUrker/ui-distill-3b), but
it stands on its own — distill your own student from it.
The interesting part
The instructions in this dataset are not the prompts that generated the HTML.
The teacher was driven by a 678-character system prompt: use the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/ui-distill-html-648.ui-weekly-benefit-and-duration-by-state
Unemployment insurance: maximum weekly benefit and weeks of duration by US state
Canonical, always-current version: https://referencesource.org/ui-weekly-benefit-and-duration-by-state/
Machine-readable: https://referencesource.org/ui-weekly-benefit-and-duration-by-state/data.json — this mirror is a point-in-time copy.
Last verified: 2026-10-06
Stale after: 2027-02-14 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 6
For… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ui-weekly-benefit-and-duration-by-state.dreamstruct-human-UI-classes-1kstate-ui-taxable-wage-bases
State unemployment insurance taxable wage bases and new employer tax rates by US state
Canonical, always-current version: https://referencesource.org/state-ui-taxable-wage-bases/
Machine-readable: https://referencesource.org/state-ui-taxable-wage-bases/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-17
Stale after: 2027-08-17 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 14
For each US state… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/state-ui-taxable-wage-bases.agent-ui-human
Agent UI Human
Human-written interface preferences for agents. Companion to akashnaren/agent-ui-sft (synthetic multi-turn tool traces). This set is the opposite shape: one person-shaped request per row, a preferred UI, and a short rationale.
Author
Akash Premkumar (akashnaren)
License
Apache-2.0
Files
train.jsonl (50), test.jsonl (10), agent_ui_human.csv (60 full)
Mirror
Kaggle: akashpnaren/agent-ui-human
How it was made
Public question:… See the full description on the dataset page: https://huggingface.co/datasets/akashnaren/agent-ui-human.ui-design-audit-dataset
UI Design Audit Screenshot Benchmark v2.1
A reproducible synthetic benchmark of 3,000 mobile and web UI screenshots labeled across 12 design-risk categories.
Splits
train: 2,400
validation: 300
test: 300
Labels
small_touch_targets
low_contrast
action_overload
navigation_overload
form_friction
content_density
responsive_risk
modal_overuse
deep_scrolling
weak_hierarchy
interaction_overload
mobile_web_mismatch
Data creation
Every… See the full description on the dataset page: https://huggingface.co/datasets/newazhala/ui-design-audit-dataset.CoALM-ITagent-ui-mode-pairs
Agent UI Mode Pairs
Pairwise preference set for the same research question: which UI should an agent pick?
Each row is one prompt with a preferred and a rejected ui_mode, plus a short rationale. Built from lab-authored text using the same labeling rules as akashnaren/agent-ui-human.
Author
Akash Premkumar (akashnaren)
License
Apache-2.0
Hub files
train.jsonl (65), test.jsonl (16), optional pairs.jsonl (81 full)
Related
agent-ui-human, agent-ui-sft… See the full description on the dataset page: https://huggingface.co/datasets/akashnaren/agent-ui-mode-pairs.Hunter-Alpha-UIGEN-T3-Agent-SFT
Hunter Alpha UIGEN T3
All of the prompts for this dataset were sourced from Tesslate/UIGEN-T3-Dataset-Extended-Reasoning, and the rest were generated.
Unfortunately the model was taken down from openrouter and revealed as xiaomi/mimo-v2-pro in the middle of the dataset creation, and only 2.6k entries were finished when this happened.
Each prompt was given to Hunter-Alpha (The stealth model recently revealed to be xiaomi/mimo-v2-pro) with the follow tools and system prompt:… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/Hunter-Alpha-UIGEN-T3-Agent-SFT.arcwise-plat
Arcwise-Plat
Corrected Arcwise-Plat Text-to-SQL evaluation datasets.
Configurations
full: questions, evidence, and gold SQL may be corrected.
sql_only: only gold SQL is corrected; questions and evidence remain unchanged.
Each configuration contains 498 test examples with unique question_id values.
Schema
Field
Type
Required
Description
question_id
string
yes
Question identifier, unique within each configuration
question
string
yes… See the full description on the dataset page: https://huggingface.co/datasets/uiuc-kang-lab/arcwise-plat.dpo_uid_conj
dpo_uid_conj
Historical AIME response data generated with Qwen3-1.7B for the UID comparison pipeline.
Train: 413 rows. Dev: 53 rows.
Both chosen and rejected responses are verifier-correct, complete, untruncated, and UID-eligible.
The response with the highest uid_product is chosen; the response with the lowest uid_product is rejected.
All four DPO variants use the same prompts. Each prompt must have at least two distinct responses and a score gap of at least 0.05 for all three… See the full description on the dataset page: https://huggingface.co/datasets/talzoomanzoo/dpo_uid_conj.
