Team Ai
Modelpublic

IFM/K2-Type-0.9B

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
24likes1.2kdownloads
README.md100 linesDownload Raw Back to root
1---2license: apache-2.03base_model: IFM/K2-Horizon-0.9B4language:5- en6tags:7- decision-model8- jev9- classification10- calibration11- pointer-head12library_name: transformers13---14 15# K2-Type-0.9B16 17A 0.9B decision model in the style of TypeSafe's Jev (K2-Horizon-0.9B backbone; 1.08B parameters stored, including the base model's unused language-model head). You send one **state** (text or JSON) and any number of typed18**questions**; it returns a probability for every option of every question from **one forward pass**. It never19generates text.20 21| Question type | You give | You get |22| --- | --- | --- |23| `noul` | a statement, optional definitions of `false` / `true` | P(true) |24| `choice` | 1-255 named options with optional descriptions | the best option, its confidence, all probabilities |25| `score` | 2-255 ordered levels | expected level, probability per level |26 27Questions share the state but cannot see each other (block-causal attention mask), so adding a question never changes28another's answer. Built on [IFM/K2-Horizon-0.9B](https://huggingface.co/IFM/K2-Horizon-0.9B).29 30## Results31 32**JevBench public set** (231 items, [jevbench](https://github.com/fstandhartinger/jevbench) commit 26eb72d,33`typesafe` adapter against this repo's server, one H200, no network):34 35| Tier | Correct |36| --- | --- |37| standard (72, `original.jsonl`) | 66 |38| easy (48, `easy.jsonl`) | 47 |39| hard (111, `hard.jsonl`) | 63 |40| **Total** | **176 / 231 = 0.762** |41 42Brier 0.328, ECE 0.065; latency p50 27 ms, p95 60 ms per decision; mean 590 input tokens per decision.43For reference, public accuracy on the JevBench v1.4 board: Gemma 4 E2B + LoRA (system-one-open) 0.732, decider-2b440.710, kev 0.6B 0.667, kev 4B 0.662, Qwen3.5-4B entries 0.74-0.82, Jev 1.13.0 0.866. The official JevBench score also45uses sealed items; every listed system scores well below its public accuracy there.46 47 48**Snake**: the same weights play Snake from a text board (one choice + four yes/no questions per move):4966 food per game on 12x12 on average (max 101).50 51## Quickstart52 53```bash54pip install -U huggingface_hub55hf download IFM/K2-Type-0.9B --local-dir K2-Type-0.9B && cd K2-Type-0.9B56pip install -r requirements.txt57python -m jev.serve --run . --port 8000          # from this repo's root; needs one CUDA GPU58```59 60The first request after start-up takes a few seconds (CUDA warm-up); later requests take about 20-60 ms.61 62```bash63curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{64  "state": {"subject": "Charged twice", "body": "You billed my card twice for March. Refund one or I cancel."},65  "questions": {66    "queue":  {"type": "choice", "instructions": "Which queue handles this?",67               "criteria": {"billing": "Payments and refunds", "technical": "Bugs and login", "general": "Anything else"}},68    "angry":  {"type": "noul", "instructions": "The customer sounds angry."},69    "urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["Low", "Normal", "High", "Critical"]}70  }}'71```72 73The request and response follow TypeSafe's `/v1/systemone` wire format, so clients written for Jev or Kev work74unchanged. `GET /health` reports the model name and temperature.75 76This repository contains the weights and the minimal code needed to serve them (`jev/`: input encoding, the pointer77head, and the HTTP server). Training code and data are not released at this time.78 79Requires transformers >= 5.17 (remote code); tested on torch 2.8. Do not use the backbone for text generation: its weights were trained80for the decision head, and the language-model head is left from the base model. Use the `jev/` server.81 82## How it works83 84- Input layout: `<state> ... | <q> question <opt> option </opt> ... <decide> | <q> ...`, using five reserved tokens of85  the base tokenizer (`decision_config.json`).86- A pointer head (`pointer_head.safetensors`) scores each option's `</opt>` hidden state against the question's87  `<decide>` hidden state; a softmax at temperature 1.478 gives the probabilities.88- Training, in short: full fine-tune on ~354k decision records (public classification/NLI/QA sets, Kev's decision89  data, game positions with exact or search labels, and synthetic decision items written and blind-verified by a large90  model), targets mixed with soft labels from a 7B teacher decision model; then 50 iterations of PPO on Snake; then a91  temperature refit on held-out calibration data. Training data were checked against every Kev evaluation suite and92  the 231 JevBench public items: no exact or containment overlap.93 94## Limits95 96- One pass, no reasoning: multi-step arithmetic, date differences and very long documents (> 8192 tokens) are weak.97- Calibration is fitted on Kev's calibration suite; on other distributions probabilities can be over- or98  under-confident (JevBench public ECE 0.065).99- Answers are bounded by the options you give; it cannot say "none of these" unless you offer that option.100