ZhaoYue800/coding-agent-harness
0
Coding Agent Harness
A minimal but complete Coding Agent Harness built from scratch in Python, with a deep focus on governance (guardrails / HITL state machine / scope fence).
Install
pip install coding-agent-harnessConfigure API Key
First-time setup — store your DeepSeek API key securely in the OS keyring:
harness config set-key
# prompts for key (hidden input), stores in keyringCheck if configured (does not reveal the key):
harness config show-keyRemove the key:
harness config clear-keyAlternatively, set DEEPSEEK_API_KEY environment variable via a .env file (plaintext — see Security below).
Run
harness # start REPL
harness --mock # mock LLM mode (no API needed, for testing)Security
- API key is stored in the OS keyring (Windows Credential Manager / macOS Keychain / Linux Secret Service) — never in plaintext config.
.envfallback is plaintext; process environment is visible to other processes on shared hosts. Keyring is recommended.- Key is never logged, never committed to Git.
.envis in.gitignore. - Guardrail blocks dangerous commands (
rm -rf /,dd,format) and requires approval for risky ones (rm,sudo). - Scope fence restricts file operations to configured directories.
Directory Structure
coding-agent-harness/
├── src/harness/
│ ├── __init__.py # Package metadata
│ ├── __main__.py # python -m harness entry
│ ├── models.py # Dataclasses: Message, Action, ToolResult, GuardrailResult, Signal
│ ├── llm.py # LLM abstract base, MockLLM, DeepSeekLLM
│ ├── tools.py # ToolRegistry + 5 preset tools (read/write/bash/glob/grep)
│ ├── guardrail.py # Governance core: 3 layers (detection/HITL/scope fence) [DEEP FOCUS]
│ ├── feedback.py # Result parser: ToolResult → Signal
│ ├── memory.py # JSON-backed cross-session memory
│ ├── config.py # YAML config loader with defaults
│ ├── credentials.py # Keyring + env credential manager
│ ├── parser.py # Parse LLM output into Action objects
│ ├── agent_loop.py # Main loop: context → LLM → parse → guardrail → execute → feedback
│ ├── cli.py # CLI REPL + config subcommands
│ └── web.py # Flask web UI
├── templates/
│ └── index.html # Chat UI page
├── tests/ # 70 unit tests (all MockLLM-driven, no network)
├── demo/
│ └── run_demo.py # §A.6 mechanism demonstration (3 scenarios)
├── pyproject.toml # Package metadata, deps, entry point
├── .gitlab-ci.yml # CI config with unit-test job
├── SPEC.md # Design document
├── PLAN.md # Implementation plan (13 tasks)
├── SPEC_PROCESS.md # Brainstorming process + cold-start verification
├── AGENT_LOG.md # Agent work log
└── REFLECTION.md # Reflection reportDeployment
The WebUI is deployed to Hugging Face Spaces:
- URL: https://ZhaoYue800-coding-agent-harness.hf.space
- Platform: Hugging Face Spaces (Docker SDK, free tier)
- Build: Docker image from
Dockerfile - Start command:
gunicorn harness.web:app --bind 0.0.0.0:7860 - CI/CD: GitHub Actions runs tests on every push; HF Spaces auto-rebuilds on push to main
- Key configuration on target machine: Set
DEEPSEEK_API_KEYas an environment variable in HF Space Settings → Variables. Default runs in mock mode (HARNESS_MOCK=1).
Limitations
- Python 3.10+ required
- Single-process, no parallelism
- LLM action protocol uses
<action>XML tags - Max 20 tool calls per user turn
Development
git clone <repo>
cd coding-agent-harness
pip install -e ".[dev]"
pytest -v
python demo/run_demo.pyLicense
MIT
