mimo_v2
Datasets
All datasets matching “mimo_v2”MiMo-V2.6-RL-oss
Agentic RL Environments
RL training environments for LLM agents.
Domain
Task Family
Verifier
Code
Software engineering
Executable tests
Cyber
Vulnerability reproduction
Rule checks
General
Knowledge work
Rubric-based judging
Visual
Web development
Visual grading
Music
Symbolic music composition
Rule checks
Docker images: https://hub.docker.com/r/xiaomimimo/mimo-v2.6-rl-oss
Training code: https://github.com/XiaomiMiMo/verl
MiMo-V2.6-RL-harbor-code
MiMo-V2.6-RL Code (Harbor)
Fix a real issue in a real repository. 2,698 Harbor tasks from the Code domain of Xiaomi's
MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained
on, converted so every one runs as a standard Harbor task.
The agent gets an issue and the repository at its base commit. After it finishes, the files the hidden test patch touches are reset, the patch is applied, and the task's own test command decides the reward: 1 when it passes, 0 when it doesn't.… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/MiMo-V2.6-RL-harbor-code.MiMo-V2.6-RL-oss
Agentic RL Environments
RL training environments for LLM agents.
Domain
Task Family
Verifier
Code
Software engineering
Executable tests
Cyber
Vulnerability reproduction
Rule checks
General
Knowledge work
Rubric-based judging
Visual
Web development
Visual grading
Music
Symbolic music composition
Rule checks
Docker images: https://hub.docker.com/r/xiaomimimo/mimo-v2.6-rl-oss
Training code: https://github.com/XiaomiMiMo/verl
MiMo-V2.6-RL-harbor-terminal
MiMo-V2.6-RL Terminal (Harbor)
Terminal-Bench style tasks. 64 Harbor tasks from the Terminal domain of Xiaomi's
MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained
on, converted so every one runs as a standard Harbor task.
Self-contained command-line tasks in the Terminal-Bench format, graded by each task's own pytest suite and anti-hack guard.
Tasks
64
Graded by
pytest + anti-hack guard (deterministic)
Reference step limit
500
Source… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/MiMo-V2.6-RL-harbor-terminal.MiMo-V2.6-RL-harbor-general
MiMo-V2.6-RL General (Harbor)
Work in a simulated company through its MCP systems. 925 Harbor tasks from the General domain of Xiaomi's
MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained
on, converted so every one runs as a standard Harbor task.
Each task is a workplace with 5 to 15 business systems (finance, legal, HR, operations, ...) served over MCP, a workspace of documents, and a brief. The agent works through the systems as an unprivileged user; the task's own… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/MiMo-V2.6-RL-harbor-general.MiMo-V2.6-RL-harbor-music
MiMo-V2.6-RL Music (Harbor)
Compose a piece in ABC notation. 1,000 Harbor tasks from the Music domain of Xiaomi's
MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained
on, converted so every one runs as a standard Harbor task.
The agent writes one piece of music, in ABC notation, to the brief. Xiaomi's scorer turns it into MIDI and measures 18 human-likeness features; a validity gate (notation errors, bar lengths) zeroes invalid pieces.
Tasks
1,000
Graded… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/MiMo-V2.6-RL-harbor-music.
