akashnaren/agent-ui-sft
Agent UI SFT Small synthetic supervised fine-tune (SFT) set for agent tool-use. It studies the public research question: what is the most efficient UI for agents to interact with applications and tools? Scope: rows are original lab fiction for a public agent-UI research question. Identifiers such as lab-w3, wf-synth-44, and tr-dom-9 are invented for the harness. Author Akash Premkumar (akashnaren) License Apache-2.0 Hub files train.jsonl (80), test.jsonl (20)… See the full description on the dataset page: https://huggingface.co/datasets/akashnaren/agent-ui-sft.
Agent UI SFT
Small synthetic supervised fine-tune (SFT) set for agent tool-use. It studies the public research question: what is the most efficient UI for agents to interact with applications and tools?
Scope: rows are original lab fiction for a public agent-UI research question. Identifiers such as lab-w3, wf-synth-44, and tr-dom-9 are invented for the harness.
Schema
Each JSONL line is one object:
Split balance (checkable on Hub)
Tools by ui_mode (as published)
Observed message-list lengths on the published 100 rows (load the JSONL to reproduce): CLI and structured API averages ~5 messages; DOM and form averages ~7–8 (more hops for the same class of lab job).
How to load
from datasets import load_dataset
ds = load_dataset("akashnaren/agent-ui-sft")
print(ds)
print(ds["train"][0]["id"], ds["train"][0]["ui_mode"])Local files (after huggingface-cli download or cloning the dataset repo):
from datasets import load_dataset
ds = load_dataset("json", data_files={
"train": "train.jsonl",
"test": "test.jsonl",
})Pandas / Polars:
import pandas as pd
train = pd.read_json("train.jsonl", lines=True)
print(train["ui_mode"].value_counts())Example row (abbreviated)
From published train.jsonl, id aui-cli-001 (ui_mode=cli, task_type=ui_efficiency_score):
- user: score lab trace
tr-ui-104for tokens / turns / recoveries (CLI vs form bakeoff). - tool result (fiction):
{"trace_id":"tr-ui-104","tokens":1840,"turns":6,"recoveries":1,"ui_mode":"form","task":"restart_worker"} - final assistant: notes that form path is expensive vs one-shot CLI for that lab job.
Numbers inside tool results are staged lab fiction, not measurements from a live cluster.
Intended use
- Smoke LoRA / SFT on a small open model with OpenAI-style
messages+tools. - Teach routing among
cli,structured_api,dom_click, andform. - Lab exercises: Temporal-like workflow ops, diagnostics-style reads, DOM selector recovery, form validation.
Not intended as: a general tool-use corpus, a production eval suite, or a rehost of Moonshot/Kimi weights or proprietary traces.
Limitations
- 100 rows total — enough for a smoke LoRA, not a general agent corpus.
- Synthetic English only.
- Workflow / metrics / token counts in tool payloads are staged, not live-cluster measurements.
- No dedicated safety / prompt-injection curriculum beyond ordinary lab refusals to invent PII or webhooks.
Links
- Collection: https://huggingface.co/collections/akashnaren/agent-ui-lab-6a9a8e06fec692165b0b3c07
- Human preference companion: https://huggingface.co/datasets/akashnaren/agent-ui-human
- Flat efficiency bakeoff table: https://huggingface.co/datasets/akashnaren/agent-ui-efficiency-scores
- Pairwise UI preferences: https://huggingface.co/datasets/akashnaren/agent-ui-mode-pairs
- Sklearn router trained on this set: https://huggingface.co/akashnaren/ui-mode-router
- Static demo Space: https://huggingface.co/spaces/akashnaren/agent-ui-router
- Kaggle mirror: https://www.kaggle.com/datasets/akashpnaren/agent-ui-sft
- Personal site: https://akashnaren.github.io/
- ORCID: https://orcid.org/0009-0001-8877-9527
- Cursor: https://cursor.com/@akashpn
- Fleet / bot page: https://akashnaren.github.io/bot/
