Team Ai
Apppublic

thepikachu/architecture-env

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes
App README

ArchitectureEnv

ArchitectureEnv is an OpenEnv environment for incremental software architecture design. An agent designs real-world systems step by step — choosing the right technologies, wiring components together, and submitting only when the architecture is complete. Reward is deterministic, in [0.0, 1.0], with no LLM judge involved.

Built for the Meta PyTorch × OpenEnv Hackathon Grand Finale, April 2026.


Links

🤗 HF Space (live env)https://huggingface.co/spaces/thepikachu/architecture-env
🧠 Trained modelhttps://huggingface.co/thepikachu/architecture-sft-model
📓 Training notebookhttps://huggingface.co/spaces/thepikachu/architecture-env/blob/main/notebooks/ArchitectureEnvSFTHFDeployNotebook.ipynb
📝 Blog writeupBlog.md

Quick start in Google Colab

Open a Colab notebook with a T4 GPU runtime and run:

python
# 1) Clone the repo — pulls all env files, agent logic, and inference scripts
!git clone https://huggingface.co/spaces/thepikachu/architecture-env
%cd architecture-env

# 2) Install dependencies
!pip install -q openenv-core
!pip install -q "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"

# 3) Run the agentic inference demo
!python agentic_inference.py

No file uploads needed. The inference script pulls the fine-tuned adapter from `thepikachu/architecture-sft-model` automatically.

![Open In Colab](https://colab.research.google.com/#fileId=https%3A//huggingface.co/spaces/thepikachu/architecture-env/blob/main/notebooks/ArchitectureEnvSFTHFDeployNotebook.ipynb)


How to run locally

1) Install dependencies

bash
pip install openenv-core
pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"

2) Run agentic inference

bash
python agentic_inference.py

# Or point explicitly at the HF model repo:
MODEL_REPO_ID=thepikachu/architecture-sft-model python agentic_inference.py

3) Run the training notebook

Open `notebooks/ArchitectureEnv_SFT_HF_Deploy_Notebook.ipynb` and run all cells. The notebook covers SFT dataset generation, Qwen2.5-3B fine-tuning with Unsloth, GRPO exploration, evaluation, and HF upload.


What the environment does

The agent receives a task like "Design a real-time chat system" and issues one action per step:

add api_server
add rabbitmq
connect api_server rabbitmq
add redis
connect worker redis
submit

Six tasks across three difficulty levels:

TaskDifficultyKey challenge
url_shortenerEasyLoad balancer, cache, DB, abuse prevention
chat_systemMediumWebSocket fanout, presence, async broker
ecommerce_platformMediumSearch, payments, async workers
youtube_platformHardTranscoding pipeline, CDN, recommendations
ride_sharingHardGeospatial indexing, real-time matching
ml_platformHardFeature store, model registry, inference server

Reward = component coverage + connection coverage + family-fit bonus + bonus items − coherence penalty.


The agentic inference system

On top of the fine-tuned SFT model we built a Planner → LLM → Critic → Negotiator loop that runs every step:

  • —PlannerAgent — deterministic priority ordering: required items → required connections → bonus items
  • —LLM (Proposer) — the fine-tuned Qwen2.5-3B model generates the next action, guided by the planner hint
  • —CriticAgent — validates the proposal against live environment state before execution
  • —NegotiatorAgent — repairs rejected proposals using the planner as ground truth, logging every decision

This is what closes the gap between a raw SFT model (~0.76 avg) and the final demo results.


Results

SFT checkpoint (direct inference, no agentic layer)

Run via trained_inference_sft.py immediately after training.

TaskScoreSteps
url_shortener1.00011
chat_system1.00016
ecommerce_platform0.99015
youtube_platform1.00018
ride_sharing1.00018
ml_platform0.99017
Average0.997—

SFT + GRPO checkpoint (direct inference, no agentic layer)

Run via trained_inference_sft_grpo.py. SFT already saturated the task — GRPO added no benefit and hurt ml_platform.

TaskScoreSteps
url_shortener1.00011
chat_system1.00016
ecommerce_platform0.99015
youtube_platform1.00018
ride_sharing1.00017
ml_platform0.58010
Average0.928—
ml_platform dropped because the GRPO model submitted before adding the required s3 storage component and wiring all connections — SFT had already saturated these tasks, so GRPO found a shortcut that bypassed the hard steps.

Agentic demo (Planner → LLM → Critic → Negotiator)

Run via agentic_inference.py on 2026-04-25, total time 3 minutes 14 seconds.

TaskBaseline (raw SFT)Improved (agentic)Δ
chat_system0.9301.000+0.070
ecommerce_platform0.9000.990+0.090
youtube_platform0.9301.000+0.070
ride_sharing0.9300.890−0.040
ml_platform0.9300.990+0.060
url_shortener0.9301.000+0.070
Average0.9250.978+0.053
Note on ride_sharing: The Negotiator got stuck in a add stripe / payment_gateway already present loop after the required components were complete, burning the step budget. The env already had reward=0.97 at step 16 before the loop started — this is a known Negotiator edge case, not a model failure.

Reward curves (agentic demo)

BASELINE
  chat_system            [ ▁▂▃▄▄▅▅▆▆▇▇] final=0.930
  ecommerce_platform     [ ▁▃▃▄▅▅▆▆▆▆▇] final=0.900
  youtube_platform       [ ▁▂▂▃▃▄▄▅▅▅▆▆▇▇] final=0.930
  ride_sharing           [ ▁▂▃▃▄▄▅▅▅▆▆▇▇] final=0.930
  ml_platform            [ ▁▂▃▃▄▄▅▅▅▆▆▇▇] final=0.930
  url_shortener          [▁▂▃▄▅▆▇▇] final=0.930

IMPROVED
  chat_system            [ ▁▂▃▄▄▅▅▆▆▇▇▇▇██] final=1.000
  ecommerce_platform     [ ▁▃▃▄▅▅▆▆▆▆▇▇▇▇] final=0.990
  youtube_platform       [ ▁▂▂▃▃▄▄▅▅▅▆▆▇▇▇▇█] final=1.000
  ride_sharing           [ ▁▂▃▃▄▄▅▅▅▆▆▇▇▇▇▇▇▇▇…] final=0.890
  ml_platform            [ ▁▂▃▃▄▄▅▅▅▆▆▇▇▇▇] final=0.990
  url_shortener          [▁▂▂▃▄▅▆▇▇▇▇█] final=1.000

Training evidence

SFT training loss

SFT Loss Curve

150 steps (10 epochs, 120 examples, batch size 8, Qwen2.5-3B-Instruct 4-bit QLoRA). Loss dropped from 3.565 at step 10 to a plateau of ~0.006 by step 70, with a final logged train_loss of 0.208 (averaged across all steps including the steep early descent).

StepLossLREpoch
103.5651.60e-40.33
201.7821.95e-40.67
300.5551.88e-41.00
400.1261.81e-41.33
500.0431.74e-41.67
600.0151.67e-42.00
700.0091.60e-42.33
800.0071.53e-42.67
90–150~0.006decaying3–10

Training runtime: 793 seconds (~13 min) on Tesla T4.

Reward curve — SFT → GRPO

Reward Curve

50 GRPO steps on top of the SFT checkpoint (120 prompts, batch size 4).

StepMean RewardKL
50.8981.852
100.8941.798
150.9191.863
200.9251.625
250.9501.751
300.9621.413
350.9201.893
400.8232.130
450.8822.061
500.9871.506

Pre-training eval avg reward: 0.522 → Post-training eval avg reward: 0.538. The step-level reward peaks at 0.987 by step 50. The post-eval uses a single-shot (non-agentic) greedy decode on the 6 tasks , it showed near-zero net gain vs SFT on most tasks, with ml_platform actually regressing. We shipped the SFT checkpoint instead. Check out Blog.md for the full reasoning.


Training logs

SourceDescription
Notebook Cell 8 outputSFT training — 150 steps, loss per step (see table above)
Notebook Cell 14 outputSFT inference — per-task step logs, final scores
Notebook Cell 16 outputGRPO training — 50 steps, reward/KL per step (see table above)
Notebook Cell 18 outputGRPO inference — per-task step logs, final scores
Notebook Cell 33 outputAgentic inference demo — baseline + improved, all 6 tasks
datasets/sft_dataset.jsonl120 SFT training examples
plots/loss_curve.pngSFT loss curve (generated by notebook Cell 36)
plots/reward_curve.pngSFT → GRPO reward progression (generated by notebook Cell 38)

All training outputs are captured as cell outputs in `notebooks/ArchitectureEnv_SFT_HF_Deploy_Notebook.ipynb`.

Sample from agentic inference log (Cell 33)

[BASELINE] chat_system      → score=0.930  steps=12
[IMPROVED] chat_system      → score=1.000  steps=16  bonus=['auth','notification_service','observability','presence_service']

[BASELINE] youtube_platform → score=0.930  steps=15
[IMPROVED] youtube_platform → score=1.000  steps=18  bonus=['auth','observability','recommendation_worker']

BASELINE vs IMPROVED
  chat_system        baseline=0.930  improved=1.000  delta=+0.070  ++
  ecommerce_platform baseline=0.900  improved=0.990  delta=+0.090  +++
  youtube_platform   baseline=0.930  improved=1.000  delta=+0.070  ++
  ride_sharing       baseline=0.930  improved=0.890  delta=-0.040
  ml_platform        baseline=0.930  improved=0.990  delta=+0.060  ++
  url_shortener      baseline=0.930  improved=1.000  delta=+0.070  ++

  Average reward gain: +0.053
  Total time: 0:03:14

Repo layout

architecture-env/
├── archive/
│   └── inference/
│       ├── inference.py
│       ├── trained_inference_sft.py
│       └── trained_inference_sft_grpo.py
├── components/
│   └── planner_critic.py          # Planner, Critic, Negotiator agents
├── config/
│   └── component_catalog.py       # component catalogue + family definitions
├── datasets/
│   └── sft_dataset.jsonl          # generated SFT training dataset
├── notebooks/
│   └── ArchitectureEnv_SFT_HF_Deploy_Notebook.ipynb   # end-to-end training notebook
├── plots/
│   ├── loss_curve.png             # SFT training loss
│   └── reward_curve.png           # SFT → GRPO reward progression
├── server/
│   ├── __init__.py
│   ├── app.py                     # OpenEnv FastAPI entry point
│   └── architecture_env_environment.py   # core env logic + reward function
├── training/
│   └── __init__.py
├── agentic_inference.py           # main demo — planner+critic+LLM loop
├── client.py                      # env client helper
├── components.json                # component registry
├── Dockerfile
├── encyclopedia_rules.py          # family→component mappings + system prompt
├── models.py                      # ArchitectureAction / ArchitectureObservation
├── openenv.yaml                   # OpenEnv manifest
├── pyproject.toml
├── requirements.txt
└── README.md

Stack

  • —Environment: OpenEnv + FastAPI + Pydantic
  • —Training: Unsloth (4-bit QLoRA) + HuggingFace TRL (SFT + GRPOTrainer)
  • —Base model: Qwen2.5-3B-Instruct
  • —Inference: Planner-Critic-Negotiator multi-agent loop
  • —Deployment: HuggingFace Spaces (Docker)