Team Ai
Datasetpublic

broadfield/openjev-cpu-gradio-kit

⚖️ OpenJev on CPU — a zero-GPU decision API Space Runs the official openjev/openjev-GGUF OpenJev-Q4_K_M.gguf (16.5 GB, ~4.8 bits, text-only) with llama-cpp-python on CPU only (n_gpu_layers=0) and serves the exact same /v1/systemone decision protocol as the hosted openjev-server. One forward pass per question — no generation, no chain-of-thought. Same prompt layout, same temperature (T=0.85), same noul calibration (noul_t=1.829074), same confidence formulas. Choice / score /… See the full description on the dataset page: https://huggingface.co/datasets/broadfield/openjev-cpu-gradio-kit.

sourceHugging Faceupdated 9d agoView on Hugging Face
0likes45downloads
Dataset Card

⚖️ OpenJev on CPU — a zero-GPU decision API Space

Runs the official **openjev/openjev-GGUF** OpenJev-Q4_K_M.gguf (16.5 GB, ~4.8 bits, text-only) with llama-cpp-python on CPU only (n_gpu_layers=0) and serves the exact same `/v1/systemone` decision protocol as the hosted openjev-server.

  • —One forward pass per question — no generation, no chain-of-thought.
  • —Same prompt layout, same temperature (T=0.85), same noul calibration (noul_t=1.829074), same confidence formulas.
  • —Choice / score / noul answer shapes are byte-for-byte the openjev-server contract.

⚠️ Critical: the free tier cannot host this

The smallest official OpenJev GGUF is 16.5 GB (OpenJev-Q4_K_M.gguf). Hugging Face free Spaces have two CPU-free options:

TierCPURAMDiskPriceCan it host?
Static––50 GBFree for everyone❌ static pages only — cannot run Python
CPU basic (Gradio)2 vCPU16 GB50 GBFree, but requires a paid (PRO) plan to create⚠️ 16.5 GB model + runtime does not fit 16 GB RAM
CPU upgrade (Gradio)8 vCPU32 GB50 GB$0.03/hr✅ fits — zero GPU, still CPU

So a true zero-GPU Space that actually runs the model needs either:

  1. 1.A PRO account → create the Gradio Space on CPU basic. It will likely OOM on the 16.5 GB model though — use CPU upgrade ($0.03/hr) for real use, or switch core_engine.py to the `Q5_K_M` / `Q6_K` / `Q8_0` files at your own RAM budget.
  2. 2.A free account → you cannot create a Gradio/Docker compute Space at all (HF blocks it: "hosting Gradio and Docker Spaces on free cpu-basic requires a PRO subscription"). Option: host this app on your own CPU machine / VPS (docker-compose or bare uvicorn), or run the **Docker recipe** on a CPU box.

This repo is the complete, runnable artifact for either path — it just cannot itself spawn a compute Space on a free account.


Run it locally (zero GPU, any 32 GB RAM machine)

bash
pip install -r requirements.txt
python app.py
# → Gradio UI at http://localhost:7860

Run it as a HuggingFace Space (needs PRO for CPU tiers)

  1. 1.Create a Gradio Space.
  2. 2.Upload app.py, core_engine.py, readout_core.py, requirements.txt.
  3. 3.In Settings → Hardware: choose CPU upgrade ($0.03/hr, 8 vCPU, 32 GB RAM) — or CPU basic if you only smoke-test (won't fit 16.5 GB, but the app boots and shows the loader).
  4. 4.Save → HF builds, downloads the GGUF from openjev/openjev-GGUF, loads it, and serves.

Zero-GPU API (same as openjev-server /v1/systemone)

bash
curl -s localhost:7860/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": {"page": "checkout", "elements": [{"tag": "button", "text": "Place order", "index": 1}]},
  "questions": {
    "q1": {
      "type": "choice",
      "instructions": "Which action completes the purchase?",
      "criteria": {"click_place_order": "click Place order", "go_back": "return to cart"}
    }
  }
}'

Response (unchanged from openjev-server):

json
{
  "id": "oj-<uuid>",
  "model": "openjev/openjev-GGUF/OpenJev-Q4_K_M.gguf",
  "answers": {
    "q1": {
      "type": "choice",
      "choice": "click_place_order",
      "probabilities": {"click_place_order": 0.97, "go_back": 0.03},
      "confidence": 0.94
    }
  },
  "usage": {"input_tokens": 123, "output_tokens": 0}
}

Endpoints

RouteDescription
GET /healthzliveness
GET /readyzmodel loaded?
GET /v1/versionbackend / profile / flags (same as openjev-server)
POST /v1/systemonethe one-pass decision API (choice/score/noul)
POST /v1/prewarmprefill warm-up

(The optional FastAPI passthrough in `app.py` is enabled with `OPENJEV_HTTP=1`; by default the Gradio UI + the raw-API tab serve the same contract without an extra server.)

Files

FilePurpose
app.pyGradio UI (Chat + Raw API tabs), the /v1/systemone handler, optional FastAPI passthrough
core_engine.pyllama-cpp-python CPU engine, letter-token table, top-K logprob readout
readout_core.pyfaithful port of openjev-server prompt/readout math (choice/score/noul, calibration, confidence)
requirements.txtgradio + llama-cpp-python (CPU wheel, zero GPU deps)

Model

FileBitsSizeNotes
OpenJev-Q4_K_M.gguf~4.816.5 GBrecommended — the default this repo loads
OpenJev-Q5_K_M.gguf~5.719.2 GB24 GB cards with short context
OpenJev-Q6_K.gguf~6.622.1 GB32 GB+
OpenJev-Q8_0.gguf8.528.6 GBclosest to the 16-bit weights

All from openjev/openjev-GGUF.