Team Ai
Apppublic

openenv/repl

sourceHugging Faceupdated 6mo agoView on Hugging Face
1likes
README.md426 linesDownload Raw Back to root
1---2title: REPL Environment Server3emoji: 🎮4colorFrom: blue5colorTo: green6sdk: docker7pinned: false8app_port: 80009base_path: /web10tags:11  - openenv-0.2.312  - openenv13---14 15## Hugging Face Space Deployment16 17This Space is built from OpenEnv environment `repl_env`.18 19- Space URL: `https://huggingface.co/spaces/openenv/repl`20- OpenEnv pinned ref: `0.2.3`21- Hub tag: `openenv`22 23### Connecting from Code24 25```python26from envs.repl_env import Env27 28env = Env(base_url="https://huggingface.co/spaces/openenv/repl")29```30 31# REPL Environment for OpenEnv32 33`repl_env` is an OpenEnv-native Python REPL environment for Recursive Language Model style execution. It now follows the current OpenEnv client/server conventions:34 35- `REPLEnv` is the remote async `EnvClient`36- `.sync()` is the sync wrapper for remote usage37- `LocalREPLEnv` is the explicit in-process helper38- `LocalRLMRunner` is the higher-level orchestration loop for local recursive RLM runs39 40The architecture is intentionally split the same way the official `rlm` and DSPy implementations split things:41 42- the environment executes code and exposes tools43- the runner owns the iterative prompting loop44- recursive behavior lives in backend/controller modules, not in the executor45 46## Overview47 48Inside the REPL, the model can:49 50- inspect `context`51- execute Python code across multiple turns with persistent state52- call `llm_query(...)` and `llm_query_batched(...)`53- call `rlm_query(...)` and `rlm_query_batched(...)` for recursive child runs when configured54- finish with `FINAL(...)`, `FINAL_VAR(...)`, or `answer = {"content": ..., "ready": True}`55 56## Current Architecture57 58Main modules:59 60- [`client.py`](client.py): remote async OpenEnv client61- [`local.py`](local.py): explicit in-process local env helper62- [`runner.py`](runner.py): local RLM orchestration loop63- [`recursive_backends.py`](recursive_backends.py): direct and recursive backend implementations64- [`recursive_controller.py`](recursive_controller.py): server-side backend/broker composition65- [`rubrics.py`](rubrics.py): reward rubrics (OpenEnv RFC 004)66- [`server/repl_environment.py`](server/repl_environment.py): server-side execution environment67- [`server/app.py`](server/app.py): OpenEnv HTTP server app and env factory68 69## What Works Today70 71- Standard remote OpenEnv usage through `REPLEnv`72- Local in-process execution through `LocalREPLEnv`73- Local recursive RLM runs through `LocalRLMRunner`74- Server-backed recursive calls through the current controller/broker path75- Explicit recursion controls:76  - `max_depth`77  - `max_children_total`78  - `max_children_per_batch`79  - `per_child_timeout_s`80  - `result_truncation_limit`81- Lightweight child trace metadata on local runner results82- Rubric-based rewards (OpenEnv RFC 004):83  - `ExactMatchRubric`: binary outcome reward against ground truth84  - `FuzzyMatchRubric`: partial credit for containment matches85  - `CustomMetricRubric`: user-provided `metric(expected, predicted) -> float`86  - `CodeExecutionRubric`: per-step process reward for code errors87  - `REPLRubric`: composite rubric combining outcome + process88  - Ground truth injectable at reset via `expected_answer`89 90## Rewards91 92Rewards follow the OpenEnv Rubric system (RFC 004). The environment uses93`REPLRubric` by default, which combines:94 95- **Outcome reward** (on terminal steps): compares `final_answer` against96  `expected_answer` if provided. Returns 1.0 for match, 0.0 otherwise.97- **Process reward** (on non-terminal steps): returns -0.05 for code98  execution errors, 0.0 for successful steps.99- **Failure reward**: returns -0.1 when max iterations exhausted without an answer.100 101For RL training (GRPO, etc.), pass `expected_answer` at reset time:102 103```python104with LocalREPLEnv() as env:105    env.reset(106        context="...",107        task_prompt="...",108        expected_answer="42",  # ground truth for rubric scoring109    )110    result = env.execute("print(FINAL(42))")111    print(result.reward)  # 1.0 (correct)112```113 114Custom rubrics can be injected at construction:115 116```python117from repl_env import LocalREPLEnv, CustomMetricRubric, REPLRubric118 119def my_metric(expected, predicted):120    return 1.0 if expected.strip() == predicted.strip() else 0.0121 122env = LocalREPLEnv(rubric=REPLRubric(outcome=CustomMetricRubric(my_metric)))123```124 125## Quick Start126 127### Remote Server Usage128 129Async:130 131```python132import asyncio133from repl_env import REPLEnv134 135 136async def main():137    async with REPLEnv(base_url="http://127.0.0.1:8000") as env:138        result = await env.reset(139            context="alpha beta gamma",140            task_prompt="Count the words",141        )142        result = await env.execute("count = len(context.split())")143        result = await env.execute("print(FINAL(count))")144        print(result.done)145 146 147asyncio.run(main())148```149 150Sync:151 152```python153from repl_env import REPLEnv154 155with REPLEnv(base_url="http://127.0.0.1:8000").sync() as env:156    result = env.reset(157        context="alpha beta gamma",158        task_prompt="Count the words",159    )160    result = env.execute("count = len(context.split())")161    result = env.execute("print(FINAL(count))")162    print(result.observation.result.stdout)163```164 165### Local Environment Usage166 167```python168from repl_env import LocalREPLEnv169 170with LocalREPLEnv() as env:171    result = env.reset(172        context="The quick brown fox jumps over the lazy dog",173        task_prompt="Count the words",174    )175    result = env.execute("count = len(context.split())")176    result = env.execute("print(FINAL(count))")177    print(env.state().final_answer)178```179 180### Local Recursive RLM Usage181 182`LocalRLMRunner` takes any `chat_fn(messages, model=None) -> str`. It works183with HF Inference API, vLLM, SGLang, Ollama, or any OpenAI-compatible server.184 185With HF Inference API:186 187```python188from huggingface_hub import InferenceClient189from repl_env import LocalRLMRunner, RLM_SYSTEM_PROMPT190 191client = InferenceClient(model="Qwen/Qwen3.5-9B", timeout=300)192 193def chat_fn(messages, model=None):194    response = client.chat.completions.create(195        model=model or "Qwen/Qwen3.5-9B",196        messages=messages,197        max_tokens=2048,198        temperature=0.6,199        extra_body={"chat_template_kwargs": {"enable_thinking": False}},200    )201    return response.choices[0].message.content202 203runner = LocalRLMRunner(chat_fn, max_iterations=30, max_depth=2)204result = runner.run("The answer is 42", "What number is mentioned?")205print(result.final_answer)206```207 208With a local vLLM server:209 210```python211from openai import OpenAI212from repl_env import LocalRLMRunner213 214client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")215 216def chat_fn(messages, model=None):217    response = client.chat.completions.create(218        model=model or "Qwen/Qwen3.5-9B",219        messages=messages,220        max_tokens=2048,221        temperature=0.6,222    )223    return response.choices[0].message.content224 225runner = LocalRLMRunner(chat_fn, max_iterations=30, max_depth=2)226result = runner.run(context, task)227```228 229### Using Different Models for Outer and Inner Loops230 231The outer loop (code generation) can use a large model while inner232`llm_query`/`rlm_query` calls use a smaller, faster model. Pass a233custom `backend_factory` to the runner:234 235```python236from openai import OpenAI237from huggingface_hub import InferenceClient238from repl_env import LocalRLMRunner239from repl_env.recursive_backends import BackendLimits, LocalChildRLMBackend240 241# Outer loop: large local model via vLLM242vllm = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")243 244def outer_chat(messages, model=None):245    r = vllm.chat.completions.create(246        model="Qwen/Qwen3-32B", messages=messages, max_tokens=2048,247    )248    return r.choices[0].message.content249 250# Inner calls (llm_query/rlm_query): smaller HF-hosted model251hf = InferenceClient(model="Qwen/Qwen3.5-9B")252 253def inner_chat(messages, model=None):254    r = hf.chat.completions.create(255        model=model or "Qwen/Qwen3.5-9B", messages=messages, max_tokens=2048,256        extra_body={"chat_template_kwargs": {"enable_thinking": False}},257    )258    return r.choices[0].message.content259 260def my_backend_factory(llm_chat_fn, **kwargs):261    return LocalChildRLMBackend(262        inner_chat,  # inner calls use the smaller model263        runner_factory=LocalRLMRunner,264        system_prompt=kwargs["system_prompt"],265        max_iterations=kwargs["max_iterations"],266        env_max_iterations_multiplier=kwargs["env_max_iterations_multiplier"],267        depth=kwargs["depth"],268        limits=BackendLimits(max_depth=2),269    )270 271runner = LocalRLMRunner(272    outer_chat,                        # outer loop: large model273    backend_factory=my_backend_factory, # inner calls: small model274    max_iterations=30,275    max_depth=2,276)277result = runner.run(context, task)278```279 280## Server281 282Run the local server:283 284```bash285PYTHONPATH=src:envs uvicorn envs.repl_env.server.app:app --host 127.0.0.1 --port 8000286```287 288The server uses a proper OpenEnv environment factory in [`server/app.py`](server/app.py).289 290## API Surface291 292### Remote Client293 294```python295class REPLEnv(EnvClient[REPLAction, REPLObservation, REPLState]):296    async def reset(...)297    async def execute(code: str)298    async def submit_final_answer(answer: str)299    async def state()300```301 302Use `.sync()` for synchronous code.303 304### Local Helpers305 306```python307class LocalREPLEnv:308    def reset(...)309    def execute(code: str)310    def state()311```312 313```python314class LocalRLMRunner:315    def run(context: str, task_prompt: str, *, model: str | None = None) -> RLMRunResult316```317 318### Actions and Observations319 320`REPLAction`321 322```python323code: str = ""324is_final: bool = False325final_answer: str | None = None326```327 328`REPLObservation`329 330```python331result: CodeBlockResult332context_preview: str | None333context_length: int334available_variables: list[str]335iteration: int336max_iterations: int337done: bool338reward: float | None339metadata: dict340```341 342## Injected REPL Helpers343 344When configured, the REPL namespace exposes:345 346- `llm_query(prompt, model=None)`347- `llm_query_batched(prompts, model=None)`348- `rlm_query(prompt, model=None)`349- `rlm_query_batched(prompts, model=None)`350- `FINAL(value)`351- `FINAL_VAR(name)`352- `SHOW_VARS()`353 354Notes:355 356- `rlm_query` is the recursive child-run surface.357- At max recursion depth, recursion falls back to direct LM calls rather than spawning more children.358- Lifecycle callbacks follow the official `rlm` pattern:359  - `on_subcall_start(depth, model, prompt_preview)`360  - `on_subcall_complete(depth, model, duration, error_or_none)`361 362## Finalization Patterns363 364### `FINAL(...)`365 366```python367result = env.execute("answer = 42")368result = env.execute("print(FINAL(answer))")369```370 371### `FINAL_VAR(...)`372 373```python374result = env.execute("my_answer = '42'")375result = env.execute('print(FINAL_VAR("my_answer"))')376```377 378### `answer` dict379 380```python381result = env.execute("answer['content'] = '42'")382result = env.execute("answer['ready'] = True")383```384 385## Prompt Utilities386 387[`prompts.py`](prompts.py) contains the current message-building and parsing helpers used by the examples and runner.388 389Important exports:390 391- `RLM_SYSTEM_PROMPT`392- `RLM_SYSTEM_PROMPT_QWEN`393- `QueryMetadata`394- `build_rlm_system_prompt(...)`395- `build_user_prompt(...)`396- `extract_code_blocks(...)`397- `format_observations(...)`398 399These prompts were updated to reflect the actual helper surface the environment provides, rather than documenting tools that do not exist.400 401## Examples402 403- [`examples/repl_with_llm.py`](../../examples/repl_with_llm.py)404- [`examples/repl_oolong_simple.py`](../../examples/repl_oolong_simple.py)405 406Default hosted model in the examples is currently `Qwen/Qwen3.5-9B`, but real hosted inference still depends on provider availability and token access.407 408## Environment Variables409 410Server-side configuration in [`server/app.py`](server/app.py):411 412- `LLM_MODEL`413- `HF_TOKEN`414- `REPL_MAX_ITERATIONS`415- `REPL_MAX_OUTPUT_LENGTH`416- `REPL_CONTEXT_PREVIEW_LENGTH`417- `REPL_RLM_MAX_DEPTH`418- `REPL_RLM_MAX_ITERATIONS`419 420## References421 422- [RLM Paper (arXiv:2512.24601)](https://huggingface.co/papers/2512.24601)423- [RLM Implementation](https://github.com/alexzhang13/rlm)424- [Alex Zhang's RLM Blog](https://alexzhang13.github.io/blog/2025/rlm/)425- [Prime Intellect RLM Blog](https://www.primeintellect.ai/blog/rlm)426