Team Ai
Datasetpublic

akashnaren/agent-ui-sft

Agent UI SFT Small synthetic supervised fine-tune (SFT) set for agent tool-use. It studies the public research question: what is the most efficient UI for agents to interact with applications and tools? Scope: rows are original lab fiction for a public agent-UI research question. Identifiers such as lab-w3, wf-synth-44, and tr-dom-9 are invented for the harness. Author Akash Premkumar (akashnaren) License Apache-2.0 Hub files train.jsonl (80), test.jsonl (20)… See the full description on the dataset page: https://huggingface.co/datasets/akashnaren/agent-ui-sft.

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
0likes46downloads
Dataset Card

Agent UI SFT

Small synthetic supervised fine-tune (SFT) set for agent tool-use. It studies the public research question: what is the most efficient UI for agents to interact with applications and tools?

Scope: rows are original lab fiction for a public agent-UI research question. Identifiers such as lab-w3, wf-synth-44, and tr-dom-9 are invented for the harness.

AuthorAkash Premkumar (akashnaren)
LicenseApache-2.0
Hub filestrain.jsonl (80), test.jsonl (20)
MirrorKaggle: akashpnaren/agent-ui-sft
Relatedagent-ui-human, agent-ui-efficiency-scores, agent-ui-mode-pairs, ui-mode-router, Space

Schema

Each JSONL line is one object:

fieldtypemeaning
idstringStable id, unique across splits (e.g. aui-cli-001, aui-dom-014)
messageslist[object]OpenAI-style chat: system, user, then assistant (often with tool_calls), tool results, final assistant answer
toolslist[object]JSON-schema tool definitions available for that example
ui_modestring enumcli \structured_api \dom_click \form
task_typestringShort label (e.g. workflow_signal, telemetry_diagnose, form_fill, ui_efficiency_score)

Split balance (checkable on Hub)

splitrowsper `ui_mode`
train8020 × each of 4 modes
test205 × each of 4 modes

Tools by ui_mode (as published)

`ui_mode`tools in traces
clirun_cli, score_ui_trace
structured_apicall_json_api, workflow_op
dom_clicklist_interactive, click_node, read_region
formget_form_schema, patch_form, submit_form

Observed message-list lengths on the published 100 rows (load the JSONL to reproduce): CLI and structured API averages ~5 messages; DOM and form averages ~7–8 (more hops for the same class of lab job).

How to load

python
from datasets import load_dataset

ds = load_dataset("akashnaren/agent-ui-sft")
print(ds)
print(ds["train"][0]["id"], ds["train"][0]["ui_mode"])

Local files (after huggingface-cli download or cloning the dataset repo):

python
from datasets import load_dataset

ds = load_dataset("json", data_files={
    "train": "train.jsonl",
    "test": "test.jsonl",
})

Pandas / Polars:

python
import pandas as pd
train = pd.read_json("train.jsonl", lines=True)
print(train["ui_mode"].value_counts())

Example row (abbreviated)

From published train.jsonl, id aui-cli-001 (ui_mode=cli, task_type=ui_efficiency_score):

  • —user: score lab trace tr-ui-104 for tokens / turns / recoveries (CLI vs form bakeoff).
  • —tool result (fiction): {"trace_id":"tr-ui-104","tokens":1840,"turns":6,"recoveries":1,"ui_mode":"form","task":"restart_worker"}
  • —final assistant: notes that form path is expensive vs one-shot CLI for that lab job.

Numbers inside tool results are staged lab fiction, not measurements from a live cluster.

Intended use

  • —Smoke LoRA / SFT on a small open model with OpenAI-style messages + tools.
  • —Teach routing among cli, structured_api, dom_click, and form.
  • —Lab exercises: Temporal-like workflow ops, diagnostics-style reads, DOM selector recovery, form validation.

Not intended as: a general tool-use corpus, a production eval suite, or a rehost of Moonshot/Kimi weights or proprietary traces.

Limitations

  • —100 rows total — enough for a smoke LoRA, not a general agent corpus.
  • —Synthetic English only.
  • —Workflow / metrics / token counts in tool payloads are staged, not live-cluster measurements.
  • —No dedicated safety / prompt-injection curriculum beyond ordinary lab refusals to invent PII or webhooks.

Links

  • —Collection: https://huggingface.co/collections/akashnaren/agent-ui-lab-6a9a8e06fec692165b0b3c07
  • —Human preference companion: https://huggingface.co/datasets/akashnaren/agent-ui-human
  • —Flat efficiency bakeoff table: https://huggingface.co/datasets/akashnaren/agent-ui-efficiency-scores
  • —Pairwise UI preferences: https://huggingface.co/datasets/akashnaren/agent-ui-mode-pairs
  • —Sklearn router trained on this set: https://huggingface.co/akashnaren/ui-mode-router
  • —Static demo Space: https://huggingface.co/spaces/akashnaren/agent-ui-router
  • —Kaggle mirror: https://www.kaggle.com/datasets/akashpnaren/agent-ui-sft
  • —Personal site: https://akashnaren.github.io/
  • —ORCID: https://orcid.org/0009-0001-8877-9527
  • —Cursor: https://cursor.com/@akashpn
  • —Fleet / bot page: https://akashnaren.github.io/bot/