Team Ai
Modelpublic

IFM/K2-Type-0.9B

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
24likes1.2kdownloads
Model Card

K2-Type-0.9B

A 0.9B decision model in the style of TypeSafe's Jev (K2-Horizon-0.9B backbone; 1.08B parameters stored, including the base model's unused language-model head). You send one state (text or JSON) and any number of typed questions; it returns a probability for every option of every question from one forward pass. It never generates text.

Question typeYou giveYou get
noula statement, optional definitions of false / trueP(true)
choice1-255 named options with optional descriptionsthe best option, its confidence, all probabilities
score2-255 ordered levelsexpected level, probability per level

Questions share the state but cannot see each other (block-causal attention mask), so adding a question never changes another's answer. Built on IFM/K2-Horizon-0.9B.

Results

JevBench public set (231 items, jevbench commit 26eb72d, typesafe adapter against this repo's server, one H200, no network):

TierCorrect
standard (72, original.jsonl)66
easy (48, easy.jsonl)47
hard (111, hard.jsonl)63
Total176 / 231 = 0.762

Brier 0.328, ECE 0.065; latency p50 27 ms, p95 60 ms per decision; mean 590 input tokens per decision. For reference, public accuracy on the JevBench v1.4 board: Gemma 4 E2B + LoRA (system-one-open) 0.732, decider-2b 0.710, kev 0.6B 0.667, kev 4B 0.662, Qwen3.5-4B entries 0.74-0.82, Jev 1.13.0 0.866. The official JevBench score also uses sealed items; every listed system scores well below its public accuracy there.

Snake: the same weights play Snake from a text board (one choice + four yes/no questions per move): 66 food per game on 12x12 on average (max 101).

Quickstart

bash
pip install -U huggingface_hub
hf download IFM/K2-Type-0.9B --local-dir K2-Type-0.9B && cd K2-Type-0.9B
pip install -r requirements.txt
python -m jev.serve --run . --port 8000          # from this repo's root; needs one CUDA GPU

The first request after start-up takes a few seconds (CUDA warm-up); later requests take about 20-60 ms.

bash
curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": {"subject": "Charged twice", "body": "You billed my card twice for March. Refund one or I cancel."},
  "questions": {
    "queue":  {"type": "choice", "instructions": "Which queue handles this?",
               "criteria": {"billing": "Payments and refunds", "technical": "Bugs and login", "general": "Anything else"}},
    "angry":  {"type": "noul", "instructions": "The customer sounds angry."},
    "urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["Low", "Normal", "High", "Critical"]}
  }}'

The request and response follow TypeSafe's /v1/systemone wire format, so clients written for Jev or Kev work unchanged. GET /health reports the model name and temperature.

This repository contains the weights and the minimal code needed to serve them (jev/: input encoding, the pointer head, and the HTTP server). Training code and data are not released at this time.

Requires transformers >= 5.17 (remote code); tested on torch 2.8. Do not use the backbone for text generation: its weights were trained for the decision head, and the language-model head is left from the base model. Use the jev/ server.

How it works

  • —Input layout: <state> ... | <q> question <opt> option </opt> ... <decide> | <q> ..., using five reserved tokens of the base tokenizer (decision_config.json).
  • —A pointer head (pointer_head.safetensors) scores each option's </opt> hidden state against the question's <decide> hidden state; a softmax at temperature 1.478 gives the probabilities.
  • —Training, in short: full fine-tune on ~354k decision records (public classification/NLI/QA sets, Kev's decision data, game positions with exact or search labels, and synthetic decision items written and blind-verified by a large model), targets mixed with soft labels from a 7B teacher decision model; then 50 iterations of PPO on Snake; then a temperature refit on held-out calibration data. Training data were checked against every Kev evaluation suite and the 231 JevBench public items: no exact or containment overlap.

Limits

  • —One pass, no reasoning: multi-step arithmetic, date differences and very long documents (> 8192 tokens) are weak.
  • —Calibration is fitted on Kev's calibration suite; on other distributions probabilities can be over- or under-confident (JevBench public ECE 0.065).
  • —Answers are bounded by the options you give; it cannot say "none of these" unless you offer that option.