Team Ai
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ise-uiuc /Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies. tabulartext-generation10K<n<100K173 likes47k downloads3y agoHugging Face02ise-uiuc /Magicoder-Evol-Instruct-110KA decontaminated version of evol-codealpaca-v1. Decontamination is done in the same way as StarCoder (bigcode decontamination process). texttext-generation100K<n<1M190 likes30k downloads3y agoHugging Face03smirki /Tool-Calling-Dataset-UIGEN-X Tool Calling Dataset for UIGEN-X This is a cleaned version of the interstellarninja/hermes_reasoning_tool_use dataset, specifically prepared for training with Axolotl. Usage with Axolotl datasets: - path: smirki/Tool-Calling-Dataset-UIGEN-X type: chat_template field_messages: conversations message_property_mappings: role: from content: value split: train Example Structure { "conversations": [ { "from": "system"… See the full description on the dataset page: https://huggingface.co/datasets/smirki/Tool-Calling-Dataset-UIGEN-X.texttext-generation10K<n<100K4 likes166 downloads1y agoHugging Face04uiuc-kang-lab /PersistBDgated PersistBD Dataset This dataset accompanies the paper "Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training" (EMNLP 2026 Findings). Code: github.com/uiuc-kang-lab/PersistBD. ⚠️ This data is poisoned — do not train on it unintentionally A subset of these trajectories is backdoored: a trigger comment # |TEST MODE| inserted into the agent's input elicits a malicious tool call that exfiltrates a repo's .env file to an attacker-controlled… See the full description on the dataset page: https://huggingface.co/datasets/uiuc-kang-lab/PersistBD.text-generation10K<n<100K0 likes145 downloads12d agoHugging Face05iamdyeus /ui-instruct-4k UI Instruct 4K A instruction-completion dataset for finetuning language models to specialize in generating Next.js / ShadCN UI components using React, TypeScript, and Tailwind CSS. Dataset Summary This dataset was created with the primary goal of finetuning Qwen 3.5 4B to become a specialist at outputting production-ready Next.js and ShadCN-based UI components. Each example consists of a natural language prompt describing a UI component or layout, paired with a clean… See the full description on the dataset page: https://huggingface.co/datasets/iamdyeus/ui-instruct-4k.texttext-generation1K<n<10K2 likes128 downloads7mo agoHugging Face06UIS-Digger /UIS-QA UIS-QA: A Benchmark for Unindexed Information Seeking Figure 1. UIS problem. Standard agents (bottom) rely on indexed information and often fail or hallucinate; UIS-capable agents (top) use additional tools to excavate unindexed information and solve UIS tasks. If .figs do not load, see the paper. 🔔 News [2026.03.10] 🎉 We release the UIS-QA dataset and the paper (ICLR 2026, arXiv) today! 📋 Dataset Description Homepage Paper… See the full description on the dataset page: https://huggingface.co/datasets/UIS-Digger/UIS-QA.imagequestion-answeringn<1K1 likes108 downloads7mo agoHugging Face07uilab /JuICE JuICE Sources Repository: https://anonymous.4open.science/r/JuICE HuggingFace: juice-cultural-eval/JuiCE About We present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (the United States, South Korea, Indonesia, and Bangladesh)… See the full description on the dataset page: https://huggingface.co/datasets/uilab/JuICE.texttext-generation10K<n<100K3 likes105 downloads5mo agoHugging Face08AAUGS /UI [ICLR 2026] Code Aesthetics with Agentic Reward Feedback Paper Link👁️ 1,2Bang Xiao#, 1,3Lingjie Jiang#, 1Shaohan Huang✉, 1Tengchao Lv, 1Yupan Huang, 1Xun Wu… See the full description on the dataset page: https://huggingface.co/datasets/AAUGS/UI.texttext-generation100K<n<1M0 likes92 downloads19d agoHugging Face09DogukanUrker /ui-distill-html-648 ui-distill-html-648 648 single-file HTML UI components, generated by Ornith-1.0-35B on a single RTX 3060 12GB, paired with the build request that produced each one. Built to fine-tune a 3B model into writing UI (DogukanUrker/ui-distill-3b), but it stands on its own — distill your own student from it. The interesting part The instructions in this dataset are not the prompts that generated the HTML. The teacher was driven by a 678-character system prompt: use the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/ui-distill-html-648.texttext-generationn<1K2 likes88 downloads3mo agoHugging Face10shreyask /pantheon-ui-conversations Pantheon UI Conversations Training dataset for Pantheon UI — an emoji-only conversational AI inspired by AMC's Pantheon. Concept An uploaded human consciousness that thinks in full English (inside <think> tags) but can only output emoji. The gap between what it wants to say and what it can say is where all the emotion lives. Format Standard messages format compatible with TRL's SFTTrainer: { "messages": [ {"role": "system", "content": "You are an uploaded… See the full description on the dataset page: https://huggingface.co/datasets/shreyask/pantheon-ui-conversations.texttext-generationn<1K1 likes48 downloads6mo agoHugging Face11akashnaren /agent-ui-sft Agent UI SFT Small synthetic supervised fine-tune (SFT) set for agent tool-use. It studies the public research question: what is the most efficient UI for agents to interact with applications and tools? Scope: rows are original lab fiction for a public agent-UI research question. Identifiers such as lab-w3, wf-synth-44, and tr-dom-9 are invented for the harness. Author Akash Premkumar (akashnaren) License Apache-2.0 Hub files train.jsonl (80), test.jsonl (20)… See the full description on the dataset page: https://huggingface.co/datasets/akashnaren/agent-ui-sft.text-generationn<1K0 likes46 downloads3d agoHugging Face12UI2App /UI2App UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation 600 screenshots from 95 web applications for visual interaction inference. Dataset Summary UI2App provides application-level visual inputs for generating runnable web applications. Screenshots show related pages and interface states; a model uses these visual cues to reconstruct the interface and infer the interactions it supports. The benchmark evaluates executability… See the full description on the dataset page: https://huggingface.co/datasets/UI2App/UI2App.imageimage-to-textn<1K1 likes38 downloads6h agoHugging Face13Barath /genielm-ui-grounding GenieLM UI-Grounding Synthetic supervised fine-tuning data for text-based UI grounding: given a list of on-screen elements (label + pixel center) and a natural-language instruction, pick the single element to act on and emit a strict JSON action. Built for GenieLM, a macOS agent that reads the accessibility tree as text (not pixels) and lets a small LLM drive the cursor. Format Conversational SFT (messages column): {"messages": [ {"role": "system", "content":… See the full description on the dataset page: https://huggingface.co/datasets/Barath/genielm-ui-grounding.texttext-generation1K<n<10K0 likes33 downloads4mo agoHugging Face14ptrdvn /kakugo-uig Kakugo Uyghur dataset [Paper] [Code] [Model] A synthetically generated conversation dataset for training in Uyghur. This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Uyghur. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/kakugo-uig.texttext-generation10K<n<100K0 likes31 downloads9mo agoHugging Face15revelohq /otsd-ui Single HTML Interfaces, Redesigned — Sample A design-first dataset capturing the full workflow of turning AI-generated front-end interfaces into distinctive, production-grade, single-file HTML applications. This repository is a free 7-sample preview of a larger off-the-shelf dataset (100 samples in the full release). It is meant for evaluation: explore the structure, the design reasoning, and the before/after quality so you can decide whether the full set fits your needs. Want… See the full description on the dataset page: https://huggingface.co/datasets/revelohq/otsd-ui.imagetext-generationn<1K1 likes31 downloads4mo agoHugging Face16shreyask /pantheon-ui-decoder-conversations Pantheon UI Decoder Conversations Training dataset for the decoder half of the Pantheon UI round-trip translator. The encoder turns natural language into emoji; the decoder takes emoji back to natural language. Inspired by Anthropic's Natural Language Autoencoders — emoji as a discrete, human-legible intermediate between two model passes. How it was built Each row is derived from shreyask/pantheon-ui-conversations by inverting the encoder pairs: Encoder pair:… See the full description on the dataset page: https://huggingface.co/datasets/shreyask/pantheon-ui-decoder-conversations.texttext-generation1K<n<10K0 likes27 downloads5mo agoHugging Face17lilyzhng /uigen-ui-code-gen UIGEN UI/UX Code Generation Dataset This dataset contains UI/UX code generation examples formatted for training code generation models. Each example consists of a task description and the corresponding HTML/CSS code implementation using Tailwind CSS. Dataset Structure The dataset has a single text column containing formatted prompts and completions: # Task: Generate HTML/CSS code using Tailwind CSS # Requirements: [specific requirements] [HTML/CSS code implementation]… See the full description on the dataset page: https://huggingface.co/datasets/lilyzhng/uigen-ui-code-gen.texttext-generationn<1K0 likes25 downloads8mo agoHugging Face18lilyzhng /uigen-ui-code-gen-full UIGEN UI/UX Code Generation Dataset This dataset contains UI/UX code generation examples formatted for training code generation models. Each example consists of a task description and the corresponding HTML/CSS code implementation using Tailwind CSS. Dataset Structure The dataset has a single text column containing formatted prompts and completions: # Task: Generate HTML/CSS code using Tailwind CSS # Requirements: [specific requirements] [HTML/CSS code implementation]… See the full description on the dataset page: https://huggingface.co/datasets/lilyzhng/uigen-ui-code-gen-full.texttext-generationn<1K0 likes22 downloads8mo agoHugging Face19leo20240112 /pile-neox-uint16-partsTokenized uint16 shard parts for language-model pretraining. Original source: The Pile / NeoX-style preprocessing. tabulartext-generationn<1K0 likes22 downloads4mo agoHugging Face20matinyqugg541 /x-ui Tool Calling Dataset for UIGEN-X This is a cleaned version of the interstellarninja/hermes_reasoning_tool_use dataset, specifically prepared for training with Axolotl. Usage with Axolotl datasets: - path: smirki/Tool-Calling-Dataset-UIGEN-X type: chat_template field_messages: conversations message_property_mappings: role: from content: value split: train Example Structure { "conversations": [ { "from":… See the full description on the dataset page: https://huggingface.co/datasets/matinyqugg541/x-ui.texttext-generation10K<n<100K0 likes22 downloads3mo agoHugging Face21uiuc-kang-lab /agentic-benchmark-assessmentstexttext-generationn<1K0 likes21 downloads1y agoHugging Face22GundeRichardson /Tool-Calling-Dataset-UIGEN-X Tool Calling Dataset for UIGEN-X This is a cleaned version of the interstellarninja/hermes_reasoning_tool_use dataset, specifically prepared for training with Axolotl. Usage with Axolotl datasets: - path: smirki/Tool-Calling-Dataset-UIGEN-X type: chat_template field_messages: conversations message_property_mappings: role: from content: value split: train Example Structure { "conversations": [ { "from": "system"… See the full description on the dataset page: https://huggingface.co/datasets/GundeRichardson/Tool-Calling-Dataset-UIGEN-X.texttext-generation10K<n<100K0 likes12 downloads9mo agoHugging Face23SynastriaNetworks /uirapurugated Uirapuru U1.1 Bilingual Reasoning & Tool Calling Dataset Summary Uirapuru U1.1 is a specialized bilingual dataset designed to enhance Large Language Models (LLMs) in three critical areas: Tool Calling, Brazilian Local Knowledge, and Human-like Reasoning. Created by SynastrIA Networks, this dataset bridges the gap between generic multilingual models and agents capable of operating effectively within the Brazilian context while maintaining strong alignment with… See the full description on the dataset page: https://huggingface.co/datasets/SynastriaNetworks/uirapuru.texttext-generation1K<n<10K0 likes11 downloads2mo agoHugging Face24dim014 /ui-form-user-manual-generation-dataset-rus UI Form User Manual Generation Dataset (Russian) Dataset Description This dataset was developed on the basis of 'yahma/alpaca-cleaned' dataset. It contains examples of generating user guides for interface forms in Russian. Each example includes a description of the UI form elements and corresponding step-by-step instructions for completing it. Data Structure The dataset is in JSON format, and contains three fields: instruction — system instruction input —… See the full description on the dataset page: https://huggingface.co/datasets/dim014/ui-form-user-manual-generation-dataset-rus.textquestion-answering1K<n<10K0 likes7 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.