thepikachu/architecture-env
ArchitectureEnv
ArchitectureEnv is an OpenEnv environment for incremental software architecture design. An agent designs real-world systems step by step — choosing the right technologies, wiring components together, and submitting only when the architecture is complete. Reward is deterministic, in [0.0, 1.0], with no LLM judge involved.
Built for the Meta PyTorch × OpenEnv Hackathon Grand Finale, April 2026.
Links
Quick start in Google Colab
Open a Colab notebook with a T4 GPU runtime and run:
# 1) Clone the repo — pulls all env files, agent logic, and inference scripts
!git clone https://huggingface.co/spaces/thepikachu/architecture-env
%cd architecture-env
# 2) Install dependencies
!pip install -q openenv-core
!pip install -q "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"
# 3) Run the agentic inference demo
!python agentic_inference.pyNo file uploads needed. The inference script pulls the fine-tuned adapter from `thepikachu/architecture-sft-model` automatically.

How to run locally
1) Install dependencies
pip install openenv-core
pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"2) Run agentic inference
python agentic_inference.py
# Or point explicitly at the HF model repo:
MODEL_REPO_ID=thepikachu/architecture-sft-model python agentic_inference.py3) Run the training notebook
Open `notebooks/ArchitectureEnv_SFT_HF_Deploy_Notebook.ipynb` and run all cells. The notebook covers SFT dataset generation, Qwen2.5-3B fine-tuning with Unsloth, GRPO exploration, evaluation, and HF upload.
What the environment does
The agent receives a task like "Design a real-time chat system" and issues one action per step:
add api_server
add rabbitmq
connect api_server rabbitmq
add redis
connect worker redis
submitSix tasks across three difficulty levels:
Reward = component coverage + connection coverage + family-fit bonus + bonus items − coherence penalty.
The agentic inference system
On top of the fine-tuned SFT model we built a Planner → LLM → Critic → Negotiator loop that runs every step:
- PlannerAgent — deterministic priority ordering: required items → required connections → bonus items
- LLM (Proposer) — the fine-tuned Qwen2.5-3B model generates the next action, guided by the planner hint
- CriticAgent — validates the proposal against live environment state before execution
- NegotiatorAgent — repairs rejected proposals using the planner as ground truth, logging every decision
This is what closes the gap between a raw SFT model (~0.76 avg) and the final demo results.
Results
SFT checkpoint (direct inference, no agentic layer)
Run via trained_inference_sft.py immediately after training.
SFT + GRPO checkpoint (direct inference, no agentic layer)
Run via trained_inference_sft_grpo.py. SFT already saturated the task — GRPO added no benefit and hurt ml_platform.
ml_platform dropped because the GRPO model submitted before adding the required s3 storage component and wiring all connections — SFT had already saturated these tasks, so GRPO found a shortcut that bypassed the hard steps.Agentic demo (Planner → LLM → Critic → Negotiator)
Run via agentic_inference.py on 2026-04-25, total time 3 minutes 14 seconds.
Note on ride_sharing: The Negotiator got stuck in aadd stripe/payment_gateway already presentloop after the required components were complete, burning the step budget. The env already had reward=0.97 at step 16 before the loop started — this is a known Negotiator edge case, not a model failure.
Reward curves (agentic demo)
BASELINE
chat_system [ ▁▂▃▄▄▅▅▆▆▇▇] final=0.930
ecommerce_platform [ ▁▃▃▄▅▅▆▆▆▆▇] final=0.900
youtube_platform [ ▁▂▂▃▃▄▄▅▅▅▆▆▇▇] final=0.930
ride_sharing [ ▁▂▃▃▄▄▅▅▅▆▆▇▇] final=0.930
ml_platform [ ▁▂▃▃▄▄▅▅▅▆▆▇▇] final=0.930
url_shortener [▁▂▃▄▅▆▇▇] final=0.930
IMPROVED
chat_system [ ▁▂▃▄▄▅▅▆▆▇▇▇▇██] final=1.000
ecommerce_platform [ ▁▃▃▄▅▅▆▆▆▆▇▇▇▇] final=0.990
youtube_platform [ ▁▂▂▃▃▄▄▅▅▅▆▆▇▇▇▇█] final=1.000
ride_sharing [ ▁▂▃▃▄▄▅▅▅▆▆▇▇▇▇▇▇▇▇…] final=0.890
ml_platform [ ▁▂▃▃▄▄▅▅▅▆▆▇▇▇▇] final=0.990
url_shortener [▁▂▂▃▄▅▆▇▇▇▇█] final=1.000Training evidence
SFT training loss

150 steps (10 epochs, 120 examples, batch size 8, Qwen2.5-3B-Instruct 4-bit QLoRA). Loss dropped from 3.565 at step 10 to a plateau of ~0.006 by step 70, with a final logged train_loss of 0.208 (averaged across all steps including the steep early descent).
Training runtime: 793 seconds (~13 min) on Tesla T4.
Reward curve — SFT → GRPO

50 GRPO steps on top of the SFT checkpoint (120 prompts, batch size 4).
Pre-training eval avg reward: 0.522 → Post-training eval avg reward: 0.538. The step-level reward peaks at 0.987 by step 50. The post-eval uses a single-shot (non-agentic) greedy decode on the 6 tasks , it showed near-zero net gain vs SFT on most tasks, with ml_platform actually regressing. We shipped the SFT checkpoint instead. Check out Blog.md for the full reasoning.
Training logs
All training outputs are captured as cell outputs in `notebooks/ArchitectureEnv_SFT_HF_Deploy_Notebook.ipynb`.
Sample from agentic inference log (Cell 33)
[BASELINE] chat_system → score=0.930 steps=12
[IMPROVED] chat_system → score=1.000 steps=16 bonus=['auth','notification_service','observability','presence_service']
[BASELINE] youtube_platform → score=0.930 steps=15
[IMPROVED] youtube_platform → score=1.000 steps=18 bonus=['auth','observability','recommendation_worker']
BASELINE vs IMPROVED
chat_system baseline=0.930 improved=1.000 delta=+0.070 ++
ecommerce_platform baseline=0.900 improved=0.990 delta=+0.090 +++
youtube_platform baseline=0.930 improved=1.000 delta=+0.070 ++
ride_sharing baseline=0.930 improved=0.890 delta=-0.040
ml_platform baseline=0.930 improved=0.990 delta=+0.060 ++
url_shortener baseline=0.930 improved=1.000 delta=+0.070 ++
Average reward gain: +0.053
Total time: 0:03:14Repo layout
architecture-env/
├── archive/
│ └── inference/
│ ├── inference.py
│ ├── trained_inference_sft.py
│ └── trained_inference_sft_grpo.py
├── components/
│ └── planner_critic.py # Planner, Critic, Negotiator agents
├── config/
│ └── component_catalog.py # component catalogue + family definitions
├── datasets/
│ └── sft_dataset.jsonl # generated SFT training dataset
├── notebooks/
│ └── ArchitectureEnv_SFT_HF_Deploy_Notebook.ipynb # end-to-end training notebook
├── plots/
│ ├── loss_curve.png # SFT training loss
│ └── reward_curve.png # SFT → GRPO reward progression
├── server/
│ ├── __init__.py
│ ├── app.py # OpenEnv FastAPI entry point
│ └── architecture_env_environment.py # core env logic + reward function
├── training/
│ └── __init__.py
├── agentic_inference.py # main demo — planner+critic+LLM loop
├── client.py # env client helper
├── components.json # component registry
├── Dockerfile
├── encyclopedia_rules.py # family→component mappings + system prompt
├── models.py # ArchitectureAction / ArchitectureObservation
├── openenv.yaml # OpenEnv manifest
├── pyproject.toml
├── requirements.txt
└── README.mdStack
- Environment: OpenEnv + FastAPI + Pydantic
- Training: Unsloth (4-bit QLoRA) + HuggingFace TRL (SFT + GRPOTrainer)
- Base model: Qwen2.5-3B-Instruct
- Inference: Planner-Critic-Negotiator multi-agent loop
- Deployment: HuggingFace Spaces (Docker)
