Team Ai
Modelpublic

mstrasser/jeff-base

sourceHugging Faceapache-2.0updated 7h agoView on Hugging Face
2likes72downloads
Model Card

jeff-base

Jeff v1.3: the base model for Jeff's LoRA adapters. Jeff is a small decision model, a fine-tune of Qwen3.5-0.8B. You describe a situation (the state) and list the options in plain words; Jeff returns a calibrated probability for each option from one forward pass, with no generated text to parse. This repository is versioned by tag: use revision `v1.3`.

Catalogue, results and docs: [jeffhub.ai](https://jeffhub.ai).

Jeff v1.3 is a change in direction. Zero-shot on everything is no longer the goal: the base is always meant to be used with an adapter. The base is the foundation the adapters are trained on, and an adapter is where Jeff becomes good at a task. If you want zero-shot use without an adapter, use Jeff v1.2.

The base: live-last

The v1.3 base is trained on the same data as v1.2, with one change: the live-last prompt layout. The fixed part of the prompt (instructions and options) comes first and the changing input comes last, so prefix caching works and repeated decisions over the same options get faster.

The price: the model now reads all the options before it sees the input, and a 0.8B model is much worse at going back over a long option list than at reading the options with the input already in mind. On the general panel (4,599 questions) the two bases are level: 78.6% against 78.8%, calibration error 0.028 against 0.024. On unfamiliar tasks with long option lists, the base on its own falls apart:

Task (no adapter)Optionsv1.2 basev1.3 base
support-intents7–6485.1%24.2%
legal-clauses10066.0%7.4%
tools4–13657.8%30.2%
triage2–3667.1%47.6%

With the right adapter, almost all of it comes back: v1.3 + adapter is within −0.7 to +0.3 points of v1.2 + adapter on every task except legal-clauses (83.6% against 85.7%). Once the adapter has learned the options, putting them first costs almost nothing, and the fixed part of the prompt can be cached. We decided the speed-up was worth it.

What each adapter adds

Each adapter is scored on its full held-out test set, which it never trained on, three ways on the same rows: the untrained Qwen3.5-0.8B that Jeff is built from, the Jeff v1.3 base alone, and the v1.3 base with the adapter. Across the 13 measured adapters, mean accuracy goes from 30.4% untrained to 38.6% for the base alone and 92.8% with the adapter.

AdapterTest rowsQwen3.5-0.8B untrainedJeff base v1.3 aloneJeff base v1.3 + adapter
aml5,12036.0% · 0.02240.5% · 0.10395.0% · 0.012
emotion5,40812.6% · 0.04525.3% · 0.08060.5% · 0.018
ground4,16028.9% · 0.06150.9% · 0.07996.6% · 0.007
guard6,55243.8% · 0.06449.4% · 0.15898.2% · 0.004
legal-clauses9,89512.5% · 0.0947.4% · 0.03983.6% · 0.011
nav3,30012.6% · 0.03813.6% · 0.15697.3% · 0.006
sanctions4,90934.4% · 0.01068.3% · 0.080100.0% · 0.001
soc4,92922.1% · 0.01133.8% · 0.03294.1% · 0.014
spam3,89758.1% · 0.04769.8% · 0.04698.1% · 0.006
support-intents5,57733.9% · 0.16424.2% · 0.09596.3% · 0.003
tools5,15717.9% · 0.06330.2% · 0.01697.2% · 0.007
trading-desk5,00038.6% · 0.01541.1% · 0.11598.1% · 0.010
triage7,25644.1% · 0.09847.6% · 0.05991.5% · 0.015
codenot measured yet
code-routernot measured yet

Each cell: accuracy · calibration error (ECE, 15 bins, after each model's own fitted temperature; lower is better, 0 is perfect). Calibration error measures how far the stated confidence is from the real hit rate: 0.01 means the stated confidence is, on average, about 1 percentage point away from how often those answers are right.

Jeff-Code. The code and code-router adapters make two decisions for Qwen3.8-27B in the Jeff-Code coding agent (github.com/firelex/jeff-code). With Jeff's thinking threshold at 0.6 (step threshold 0.40), Jeff-Code matches Qwen3.8-27B's pass rate: 62.4% against 62.8% (paired difference −0.2 points, 95% interval −2.6 to +2.1) over 1,242 paired tasks from six benchmarks, run side by side, and is 47% faster (32% less time) per task on average. Results per benchmark.

v1.3 + adapter against v1.2 + adapter

Each pair is measured on exactly the same test rows. Where an adapter's v1.3 test set changed, the comparison uses its copy of the v1.2 test (named in brackets).

Adapter test setTest rowsv1.2 + adapterv1.3 + adapterChange (points)
emotion5,40860.6%60.5%−0.1
ground4,16097.0%96.6%−0.4
guard6,55298.4%98.2%−0.2
legal-clauses9,89585.7%83.6%−2.1
nav (heldout_synthetic-v12)3,30097.0%97.3%+0.3
spam (test-v12)3,60398.4%98.1%−0.3
support-intents5,57796.8%96.3%−0.5
tools5,15797.9%97.2%−0.7
triage7,25691.8%91.5%−0.3

All results and their sources: jeffhub.ai/results, and in one file, jeffhub-v1.3.json.

The adapters

Adapters are not merged into the base. You load one base and all the adapters you need, and pick one per request. Each adapter records the exact base it was trained on, and the server refuses an adapter trained on a different one, so a v1.2 adapter does not load on v1.3. All nine v1.2 adapters are retrained on v1.3, with about 10% of the base model's own training data mixed in as a precaution (its effect has not been measured).

In a request, an adapter is named by its short name ("model": "soc", "model": "code"): on a Jeff server, names starting with "jeff" mean the base model.

How to use it

Adapter serving is part of Jeff's server, on the main branch of firelex/jeff.

bash
git clone https://github.com/firelex/jeff && cd jeff
uv sync --no-default-groups --extra lora                # CPU; add --extra cuda (NVIDIA GPU) or --extra mac (Apple silicon)
uv run --no-default-groups hf download mstrasser/jeff-base --revision v1.3 --local-dir checkpoints/jeff-base
uv run --no-default-groups hf download mstrasser/jeff-adapter-support-intents --revision v1.3 \
  --local-dir adapters/support-intents
JEFF_CHECKPOINT=checkpoints/jeff-base JEFF_ADAPTERS=adapters/ PORT=8765 \
  uv run --no-default-groups jeff-serve                 # on a Mac, add JEFF_BACKEND=mlx

Each adapter lives in its own folder inside one adapters folder; the folder's name is the name you use in requests. Name the adapter as the model:

bash
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
  "model": "support-intents",
  "state": {"service": "Customer support chat of an online shop",
            "message": "I sent the jacket back two weeks ago and still have not seen the money."},
  "questions": {
    "intent": {"type": "choice",
               "instructions": "What does the customer want? Choose the request that best matches what the customer is asking for in their message.",
               "criteria": {"track_refund": "Check the status of a refund they are expecting.",
                            "get_refund": "Get their money back for a purchase.",
                            "track_order": "Find out where their order is or its current status."}}
  }
}'

Guides: getting started, request format, serving adapters.

llama.cpp. v1.3 also ships as GGUF, in Q80 and Q4K_M: mstrasser/jeff-base-gguf, plus one small LoRA GGUF per adapter. See Running Jeff with llama.cpp.

What stays fixed for the life of v1.3

  • —The request format: state, questions and instructions, with the changing state field last.
  • —The option rules: named options, keys never bare numbers.
  • —The answer format: a probability for every option.

A data set written to the data guidelines trains on v1.3 as it is.

Files

FileWhat it is
model.safetensorsThe weights (sha256 d324dd6c9bb61b30af65564135b33f6892c30a9b2bd22667b2e09b9c8118cf77; every v1.3 adapter checks it)
readout.safetensorsJeff's readout over the answer codes
decision_config.jsonAnswer codes and their token ids, temperature and prompt layout (live-last)
config.json, tokenizer.json, tokenizer_config.json, chat_template.jinja, processor_config.jsonThe Qwen3.5 configuration and tokenizer

Training

Built onQwen/Qwen3.5-0.8B
DataThe v1.2 base training data, unchanged
Change from v1.2The live-last prompt layout
Training codeThe git_commit recorded in decision_config.json is the training machine's copy and was not published. It builds exactly the same prompt as main of firelex/jeff (from commit 6d0d7da) for a text state and for an object with at least one field; the format is in docs/v1.3-request-format.md
Checkpointll-v12-final (run ll-v12, step 1113)

Not measured yet

Nothing on JeffHub is estimated. These are still to come for v1.3:

  • —code and code-router: test scores (untrained, v1.3 base, v1.3 + adapter);
  • —jeff-serve speed and GPU memory with the v1.3 base and adapters, with prompt reuse;
  • —calibration charts, commonest confusions and accuracy per answer for each v1.3 adapter;
  • —example responses recorded from the v1.3 adapters;
  • —Jeff-Code: the detailed result files behind the maintainer-supplied numbers;
  • —Jeff's memory on a Mac.

Limitations

  • —Use it with an adapter. Without one, the v1.3 base is weak on long, unfamiliar option lists (table above). For zero-shot use, use v1.2.
  • —Small models don't reason. Expect fast, calibrated choices between the options you describe, not multi-step reasoning, and no generated text.
  • —English and text only.

Licence and data

Weights: Apache-2.0, as for v1.2. The model is a fine-tune of Qwen3.5-0.8B by the Qwen team (Alibaba Cloud).

Qwen3.5-0.8B notice: these weights were modified from Qwen3.5-0.8B by the Jeff project. Qwen3.5-0.8B is Copyright 2026 Alibaba Cloud and licensed under the Apache License, Version 2.0; a copy of that licence is in `LICENSE`.

Training data: the same data as the v1.2 base, which mixes public data sets under various licences, some of them share-alike (CC BY-SA), with code-built and synthetic questions. The training data is not released. The v1.2 data sources and their licences are listed in docs/data-sources.md in the Jeff repository.

To confirm: JeffHub does not yet list the base model's data sources and their licences for v1.3; the list above is the v1.2 one, which the v1.3 base reuses unchanged.

Each adapter's own data sources and licences, including non-commercial restrictions (sanctions and soc are CC BY-NC 4.0), are on its model card and its JeffHub page.

Links

Jeff is an independent project. It uses the same request format as Jev but is not affiliated with or endorsed by TypeSafe, the makers of Jev.