lerobot/annotation-studio
LeRobot Annotation Studio
Enter a LeRobot v3 dataset, choose annotations, review an estimate, then sign in with Hugging Face to run a CPU Job using the selected vision model through Inference Providers. Jobs and inference are billed to the signed-in user's personal account. The Space itself uses free CPU hardware and has no shared HF access token.
The default tuned Qwen3.8 preset uses the action-phase prompt developed on ABC, timestamped contact sheets at 2 fps, 1,200 frames per request, and up to 80 subtasks per window. It retries failed segmentation with the original 30-second-window preset. This is an experimental prompt preset, not finetuned weights; its small held-out benchmark did not establish a quality improvement.
Models
The default is Qwen3.8 27B / Novita, with up to 32 parallel requests. Standard choices also include Qwen3.6 27B / OVHcloud and Gemma 4 31B / Novita. The selector offers nine curated VLMs from the families listed in this card. This is a practical selection, not a universal quality ranking. All nine available routes passed a live two-image JSON smoke test on 2026-09-25. Only Qwen3.8 27B has a 32-request concurrency test and the recent 33-episode all-annotation run; the new routes use a conservative eight requests, while Flash Next retains its tested two-request limit. Provider/account limits vary.
model_catalog.py owns the allowed model IDs, explicit provider routes, provider-specific generation parameters, checked pricing and context limits. Quotes and jobs pin these together. Model changes invalidate the browser quote; launch checks the selected provider is live and rejects changed catalog prices. The worker uses the pinned generation parameters, including non-Qwen models. There is no silent provider failover, so the reviewed data recipient and rates stay explicit. New model routes require a smoke test and a reviewed rate card.
Rates shown are USD per million tokens. Cached input rates are accounted for; MiniMax M3 switches to its higher pricing tier for prompts of 524,288 tokens or more. The budget reserves the worst tier and full provider context before each request, then reconciles actual reported usage. Small budgets reduce parallelism; a target too small for one conservative reservation is rejected before launch. The action-phase prompt was developed on Flash Next; quality results do not transfer automatically to other models.
Outputs
- Annotated dataset copy: a new standalone LeRobot dataset containing only successfully annotated episodes, with annotations in the standard LeRobot Parquet columns. No separate annotation JSONL files are uploaded. Every RGB camera video is trimmed to those episodes and re-encoded as H.264 (CRF 18); shared source video files are never copied wholesale. Source numeric data, tools and other annotation styles within copied episodes are retained. Indices, tasks, episode metadata and statistics are rebuilt.
annotation_studio/episode_mapping.jsonmaps new episode IDs back to the source. Original license files and the source README are retained. Failed, stopped and unselected episodes are not included. Each completed copy is published atomically with its metadata so a partial run remains loadable. Depth video is currently unsupported. Storage charges are separate from the run estimate and spending target.
Setup is split into four focused steps: Dataset, Annotations, Model, and Review. Examples and provider details expand on demand. Prompt edits and the current draft are retained when moving between steps; changing priced settings requires a new estimate.
The user chooses Private (default) or Public at review. The launch confirmation shows the output visibility. Visibility is pinned in the job manifest and applied only to the new output repository. Older manifests without a visibility field remain private.
Subtasks are generated as a dependency of plans, memory and interjections; the UI makes this explicit. Unchecked output styles are filtered from the saved annotations. All annotation types use one selected RGB video camera per run; views are not combined. Camera names load automatically from the dataset's meta/info.json before estimating, including private datasets after login. Depth and non-video features are excluded. Visual questions have a configurable cadence.
Human demo videos
A separate Video generation category on the Annotations step, next to the text annotations. For every subtask (or up to 1–8 per episode), the job turns the robot view into a first-person human demonstration, following LeRobot's --human_video annotation module:
- the selected vision model reads the robot segment and writes the human task and an image-edit prompt;
Qwen/Qwen-Image-Edit-2509removes the robot from the segment's first full-resolution frame and adds two resting human hands;- the vision model writes a step-by-step video prompt from the edited frame;
MiniMaxAI/MiniMax-H3animates it at 480P (5 s).
Both media models run on fal through HF Inference Providers with the signed-in user's token. Selecting videos also selects Subtasks, since there is one video per subtask.
The estimate uses the selected episodes' total seconds and count, calibrated on the first video run (3 ReBot episodes, 159.7 s, 33 subtasks): 1 subtask per episode + 1 per 5.3 s, up to 1.5× that in the upper bound, with the per-episode cap applied. Each video is priced at fal's list prices: the edit at $0.03 per output megapixel (1024×1024) and the 5 s 480p video at $0.05/s, about $0.28 in total, plus about 1,750 input and 550 output VLM tokens. Parallel generations (1–32, default 8) limits simultaneous media requests across the whole job. The review warns when the expected videos at that parallelism would outlast the job timeout (about 3 min per video).
Videos start only after an episode's annotations pass validation. Each request holds a conservative reservation (model_catalog.HUMAN_VIDEO) from the spending target while in flight and is reconciled to its list price when it returns. A failed video is recorded and the others continue. Videos stop at the spending target, or before the runtime limit (the job timeout when none is set), without paying for an edit whose video could not finish; the episode's annotations are kept. The launch checks that both fal mappings are live.
The output holds annotation_studio/human_videos/episode_XXXXXX/segment_XXX{.mp4, _first_frame.png,_robot_frame.png} (named by source episode) and index.json, which lists each segment's subtask, timestamps, prompts, status and materialized episode_index. These are synthetic media, not language columns. The Jobs page previews completed videos next to the robot frame, served through the Space so private outputs stay private.
Prompts and examples
Every annotation card has an Edit prompt editor with the actual model instructions, required placeholders, and a reset-to-default action. Interjections have separate opening-acknowledgement and mid-task templates. Edits are validated, kept in the browser draft, and pinned in the output run manifest. Subtask edits also apply to internal dependencies and shorter-window recovery. Other unchecked annotation types retain edits in the draft without using them in the run.
Default plans are deterministic remaining-step lists. A custom plan prompt opts into model generation at each boundary; the cost estimate includes those extra calls. Resetting it restores the deterministic default. Longer custom prompts also increase estimated input usage. Instructions can change segment counts and output lengths, so costs remain estimates.
The main screen shows two illustrative examples per annotation type. Subtask examples include timed reach/grasp/move/place phases and a continuous wiping segment, making the default granularity visible before launching a job.
Cost and privacy
Estimates are ranges, based on metadata duration, sampled frames, approximate tokens and the provider's published rates. CPU pricing is retrieved from HF. Before each request, the worker reserves a conservative token cost; successful calls reconcile with usage and uncertain failures retain their reservation. The selected total target reserves the entire job timeout's CPU cost first. This is an application spending guard, not an HF billing guarantee: provider price changes, billing differences and storage charges remain outside it. Set account billing limits on Hugging Face as an additional control.
OAuth tokens stay on the server and are passed only as HF Job secrets. They are not stored in repositories, browser storage, browser outputs or job commands. Private metadata uses temporary per-request caches. Dataset content is sent to the selected inference provider only after the user starts a job.
Reproducibility and operations
Jobs pin the Space code and source dataset revisions. pipeline-source.tar.gz contains the LeRobot source used for this deployment; pipeline-version.json records its Git commit and archive checksum. Each output includes its manifest, per-episode progress and metered inference usage. Completed shards are uploaded incrementally. A budget stop or error produces an explicitly partial result. The worker never executes code from an input dataset.
The deployment helper packages tracked LeRobot source, checks for tracked changes, tests the package checksum and uploads this directory. The existing lerobot/annotate manual-labeling Space is independent.
Local development
cd spaces/annotation-studio/frontend && bun install && bun run build
# From spaces/annotation-studio:
uv run --with-requirements requirements.txt uvicorn server:app --port 7860
uv run --no-sync pytest spaces/annotation-studio/test_studio.py
uv run --no-sync python spaces/annotation-studio/deploy.pyThe Docker Space serves a React frontend and a FastAPI backend from one origin. The interface adapts visualizer design tokens and annotation colors; attribution is in THIRD_PARTY.md. OAuth uses Authorization Code with PKCE and HttpOnly session cookies. Tokens live in server memory; restarting the Space signs users out. A refreshable dashboard lists each user's jobs, progress and cancellation.
Hugging Face login is available on the deployed Space. Local mode is preview only: it cannot launch jobs using a developer's cached credentials.
The Jobs dashboard refreshes every 15 seconds and shows usage-derived cost: provider-reported input/output/cache tokens at the pinned rates, plus HF-reported CPU runtime rounded to minutes. Unknown request charges remain separately reserved; this is not an invoice. Worker usage checkpoints are saved every ten seconds. The billing link opens the user's final HF charges.
Finished annotated copies offer Show annotations, opening the first materialized episode with ?tab=annotations in the visualizer. Private results require signing in there. Older JSONL-only exports retain their Results download link.
