Team Ai
20 results

mimo_v2

XiaomiMiMo /MiMo-V2.6-RL-oss Agentic RL Environments RL training environments for LLM agents. Domain Task Family Verifier Code Software engineering Executable tests Cyber Vulnerability reproduction Rule checks General Knowledge work Rubric-based judging Visual Web development Visual grading Music Symbolic music composition Rule checks Docker images: https://hub.docker.com/r/xiaomimimo/mimo-v2.6-rl-oss Training code: https://github.com/XiaomiMiMo/verl document1K<n<10K905 likes111k downloads15d agoHugging FaceFineEnvs /MiMo-V2.6-RL-harbor-code MiMo-V2.6-RL Code (Harbor) Fix a real issue in a real repository. 2,698 Harbor tasks from the Code domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on, converted so every one runs as a standard Harbor task. The agent gets an issue and the repository at its base commit. After it finishes, the files the hidden test patch touches are reset, the patch is applied, and the task's own test command decides the reward: 1 when it passes, 0 when it doesn't.… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/MiMo-V2.6-RL-harbor-code.other1K<n<10K4 likes14k downloads2d agoHugging Facekunato /MiMo-V2.6-RL-oss Agentic RL Environments RL training environments for LLM agents. Domain Task Family Verifier Code Software engineering Executable tests Cyber Vulnerability reproduction Rule checks General Knowledge work Rubric-based judging Visual Web development Visual grading Music Symbolic music composition Rule checks Docker images: https://hub.docker.com/r/xiaomimimo/mimo-v2.6-rl-oss Training code: https://github.com/XiaomiMiMo/verl document0 likes11k downloads10d agoHugging FaceFineEnvs /MiMo-V2.6-RL-harbor-terminal MiMo-V2.6-RL Terminal (Harbor) Terminal-Bench style tasks. 64 Harbor tasks from the Terminal domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on, converted so every one runs as a standard Harbor task. Self-contained command-line tasks in the Terminal-Bench format, graded by each task's own pytest suite and anti-hack guard. Tasks 64 Graded by pytest + anti-hack guard (deterministic) Reference step limit 500 Source… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/MiMo-V2.6-RL-harbor-terminal.othern<1K0 likes7.9k downloads2d agoHugging FaceFineEnvs /MiMo-V2.6-RL-harbor-general MiMo-V2.6-RL General (Harbor) Work in a simulated company through its MCP systems. 925 Harbor tasks from the General domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on, converted so every one runs as a standard Harbor task. Each task is a workplace with 5 to 15 business systems (finance, legal, HR, operations, ...) served over MCP, a workspace of documents, and a brief. The agent works through the systems as an unprivileged user; the task's own… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/MiMo-V2.6-RL-harbor-general.othern<1K0 likes5.7k downloads2d agoHugging FaceFineEnvs /MiMo-V2.6-RL-harbor-music MiMo-V2.6-RL Music (Harbor) Compose a piece in ABC notation. 1,000 Harbor tasks from the Music domain of Xiaomi's MiMo-V2.6-RL-oss, the RL environments MiMo-V2.6 was trained on, converted so every one runs as a standard Harbor task. The agent writes one piece of music, in ABC notation, to the brief. Xiaomi's scorer turns it into MIDI and measures 18 human-likeness features; a validity gate (notation errors, bar lengths) zeroes invalid pieces. Tasks 1,000 Graded… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/MiMo-V2.6-RL-harbor-music.other1K<n<10K0 likes5.2k downloads2d agoHugging Face