Team Ai
Apppublic

mstrasser/Jeff-adapters-demo

sourceHugging Faceapache-2.0updated 5d agoView on Hugging Face
0likes
App README
Superseded: this Space is from an earlier Jeff release and is no longer updated. The current Jeff release is Jeff v1.3: mstrasser/jeff-base.

Jeff v1.2 with adapters

Jeff is a small decision model (a fine-tune of Qwen3.5-0.8B). You give it a situation and one or more questions with named options; it returns a calibrated probability for every option in one forward pass. It does not write text.

This demo loads the Jeff v1.2 base model once, with small LoRA adapters beside it (each tens of MB). Pick an adapter; the editor fills with an example request for it. Change the text or the options and press Run. You see:

  • —a bar chart of the probabilities for the first question (yes and no for a yes/no question, every level for a score question, the 15 most likely options for a choice);
  • —the time of the model call itself, in milliseconds (parsing and drawing are not included);
  • —the full answer, in the same shape the Jeff server returns.

"Answer twice" asks each question a second time with the options in reverse order and averages the two answers. It evens out a small model's lean towards options by position, at twice the cost.

The time shown covers the model call only, on a shared GPU (ZeroGPU): it leaves out the wait for a GPU, and the first call after a pause can be slower. It is not a benchmark; the model cards give measured latencies.

Code, server and clients: github.com/firelex/jeff.

Settings

Environment variables (Space settings, "Variables"):

VariableDefaultMeaning
JEFF_BASE_REPOmstrasser/Jeff-Qwen3.5-0.8Bthe base model repository
JEFF_BASE_REVISIONv1.2its revision; the adapters only load on the exact base they were trained on
JEFF_ADAPTER_REVISIONmainrevision of every adapter repository
JEFF_ADAPTERS_LOADevery adapter in adapters.jsoncomma-separated adapter names to load, for example spam,guard,tools
JEFF_DEVICECUDA when present, else CPUcuda, mps or cpu

The Space stops at start-up with a message naming the adapter if an adapter repository cannot be downloaded. To run with fewer adapters, set JEFF_ADAPTERS_LOAD.

Run it on your own machine

bash
pip install -r requirements.txt
python app.py                                   # downloads the base and the adapters from Hugging Face
JEFF_ADAPTERS_LOAD=spam,guard python app.py     # only two adapters
# local folders instead of downloads: a base checkpoint folder, and a folder with one sub-folder per adapter
JEFF_BASE_DIR=Jeff-Qwen3.5-0.8B JEFF_ADAPTERS_DIR=adapters python app.py

Without a GPU it works on the CPU, but slowly (seconds per request); for fast local serving use jeff-serve from the repository, which has an Apple-silicon (MLX) backend.