broadfield/openjev-cpu-gradio-kit
⚖️ OpenJev on CPU — a zero-GPU decision API Space Runs the official openjev/openjev-GGUF OpenJev-Q4_K_M.gguf (16.5 GB, ~4.8 bits, text-only) with llama-cpp-python on CPU only (n_gpu_layers=0) and serves the exact same /v1/systemone decision protocol as the hosted openjev-server. One forward pass per question — no generation, no chain-of-thought. Same prompt layout, same temperature (T=0.85), same noul calibration (noul_t=1.829074), same confidence formulas. Choice / score /… See the full description on the dataset page: https://huggingface.co/datasets/broadfield/openjev-cpu-gradio-kit.
⚖️ OpenJev on CPU — a zero-GPU decision API Space
Runs the official **openjev/openjev-GGUF** OpenJev-Q4_K_M.gguf (16.5 GB, ~4.8 bits, text-only) with llama-cpp-python on CPU only (n_gpu_layers=0) and serves the exact same `/v1/systemone` decision protocol as the hosted openjev-server.
- One forward pass per question — no generation, no chain-of-thought.
- Same prompt layout, same temperature (
T=0.85), samenoulcalibration (noul_t=1.829074), same confidence formulas. - Choice / score / noul answer shapes are byte-for-byte the openjev-server contract.
⚠️ Critical: the free tier cannot host this
The smallest official OpenJev GGUF is 16.5 GB (OpenJev-Q4_K_M.gguf). Hugging Face free Spaces have two CPU-free options:
So a true zero-GPU Space that actually runs the model needs either:
- A PRO account → create the Gradio Space on CPU basic. It will likely OOM on the 16.5 GB model though — use CPU upgrade ($0.03/hr) for real use, or switch
core_engine.pyto the `Q5_K_M` / `Q6_K` / `Q8_0` files at your own RAM budget. - A free account → you cannot create a Gradio/Docker compute Space at all (HF blocks it: "hosting Gradio and Docker Spaces on free cpu-basic requires a PRO subscription"). Option: host this app on your own CPU machine / VPS (docker-compose or bare
uvicorn), or run the **Docker recipe** on a CPU box.
This repo is the complete, runnable artifact for either path — it just cannot itself spawn a compute Space on a free account.
Run it locally (zero GPU, any 32 GB RAM machine)
pip install -r requirements.txt
python app.py
# → Gradio UI at http://localhost:7860Run it as a HuggingFace Space (needs PRO for CPU tiers)
- Create a Gradio Space.
- Upload
app.py,core_engine.py,readout_core.py,requirements.txt. - In Settings → Hardware: choose CPU upgrade ($0.03/hr, 8 vCPU, 32 GB RAM) — or CPU basic if you only smoke-test (won't fit 16.5 GB, but the app boots and shows the loader).
- Save → HF builds, downloads the GGUF from
openjev/openjev-GGUF, loads it, and serves.
Zero-GPU API (same as openjev-server /v1/systemone)
curl -s localhost:7860/v1/systemone -H 'Content-Type: application/json' -d '{
"state": {"page": "checkout", "elements": [{"tag": "button", "text": "Place order", "index": 1}]},
"questions": {
"q1": {
"type": "choice",
"instructions": "Which action completes the purchase?",
"criteria": {"click_place_order": "click Place order", "go_back": "return to cart"}
}
}
}'Response (unchanged from openjev-server):
{
"id": "oj-<uuid>",
"model": "openjev/openjev-GGUF/OpenJev-Q4_K_M.gguf",
"answers": {
"q1": {
"type": "choice",
"choice": "click_place_order",
"probabilities": {"click_place_order": 0.97, "go_back": 0.03},
"confidence": 0.94
}
},
"usage": {"input_tokens": 123, "output_tokens": 0}
}Endpoints
(The optional FastAPI passthrough in `app.py` is enabled with `OPENJEV_HTTP=1`; by default the Gradio UI + the raw-API tab serve the same contract without an extra server.)
Files
Model
All from openjev/openjev-GGUF.
