pdevyani/dev-assistant-env
Dev Assistant OpenEnv Environment
๐ Overview
This project implements an OpenEnv-compatible environment that simulates real-world developer tasks. The environment allows an AI agent to perform tasks such as fixing syntax errors, cleaning code, and reviewing code.
It is designed for evaluating AI systems in practical software engineering workflows.
๐ฏ Motivation
Developers spend significant time debugging, formatting, and reviewing code. This environment replicates these real-world tasks so that AI agents can be evaluated on meaningful and useful work instead of toy problems.
๐ง OpenEnv Interface
The environment follows the OpenEnv specification:
reset()โ Returns initial observationstep(action)โ Returns(observation, reward, done, info)state()โ Returns current state
๐ Observation Space
The agent receives a task as input in text form.
Example: Fix syntax error: print('Hello'
โก Action Space
The agent must return a solution as text.
Example: print('Hello')
๐ฏ Reward Function
1.0โ Correct answer0.0โ Incorrect answer- Can be extended to give partial rewards
The reward is deterministic and reproducible.
๐งช Tasks
๐ข Easy Task
Fix simple syntax errors Example: Missing parentheses
๐ก Medium Task
Clean and format messy code Example: Improve readability and indentation
๐ด Hard Task
Perform code review Example: Detect bugs and suggest improvements
๐ Grading
Each task includes a programmatic grader:
- Returns score from
0.0to1.0 - Fully correct โ
1.0 - Incorrect โ
0.0
๐ Project Structure
openenv-dev-assistant/ โ โโโ env.py โโโ models.py โโโ tasks.py โโโ grader.py โโโ run_baseline.py โโโ requirements.txt โโโ openenv.yaml โโโ Dockerfile โโโ README.md
โ๏ธ Deployment (Hugging Face Spaces)
- Go to Hugging Face Spaces
- Create new Space
- Select Docker as SDK
- Upload all project files
- Add tag:
openenv
โ Requirements Checklist
- [x] Real-world task simulation
- [x] OpenEnv interface implemented
- [x] At least 3 tasks (easy, medium, hard)
- [x] Programmatic grading
- [x] Reward function
- [x] Docker support
- [x] Documentation
๐ Future Improvements
- Integrate OpenAI API for real agent evaluation
- Add more complex coding tasks
- Improve reward function with partial scoring
- Expand dataset of tasks
๐ Author
Team Neon
