kortexa-ai/shingi-27b
Shingi 27B — 審議
Shingi 27B is a local decision model. You give it context (text, and optionally images), a question and the possible answers. It returns a choice, calibrated probabilities or an ordinal score. It reads the logit of every candidate answer directly and generates no text.
The model is a single 7.2 GB ternary GGUF file (PQ2_0). It is Bonsai 2 27B with retrained block scales: the ternary weights are unchanged, only their scales were trained. It needs no adapter.
Run it
On Linux with an NVIDIA GPU and the CUDA toolkit, or on a Mac with Apple Silicon:
curl -fsSL https://raw.githubusercontent.com/kortexa-ai/shingi-27b/main/run.sh | bashThis builds the pinned Prism runtime, downloads this model into your Hugging Face cache and serves the API on http://127.0.0.1:8765. Use --port to change the port. The code, API examples and options are in kortexa-ai/shingi-27b.
Requirements
Speed
Median latency per decision, one request at a time, loading excluded:
With image input on the RTX 4090, one image and one question take about 0.66 s (median) and the model uses about 9.2 GB.
Questions in one request share a single pass over the images and the state, so extra questions are cheap: on an RTX PRO 6000, eight questions about one image take about 1.8 s instead of 8.7 s.
Scores
External suites, canonical choice order, 16K context:
These suites also informed the choice of training data, so treat them as development results rather than a clean held-out test.
Images
Shingi reads images through the Bonsai 2 27B vision projector (mmproj.gguf, included here). Send up to 8 images per request: /v1/decisions follows SGLang's decision endpoint, and /v1/systemone takes an optional images list.
Zero-shot:
The model was not trained on images; these come from the base model's vision with Shingi's decision training on top.
Training data
Public datasets, used under their stated licenses: Banking77, CLINC150, MMLU, HelpSteer2, HelpSteer, Measuring Hate Speech, Civil Comments, GoEmotions, LEDGAR (LexGLUE), MASSIVE, CommonsenseQA, WinoGrande, GSM8K, HelpSteer3, TabFact, TAT-QA, CaseHOLD and Unfair-ToS (LexGLUE), TimeQA, Mind2Web, MedMCQA and MedQA. Several came through the jev-bench repackaging.
About a third of the training tokens are model-generated: a private synthetic dataset of decision tasks, plus the Apache-2.0 teacher data from Mapika/decider.
Limitations
- English only.
- Much slower on a Mac than on an NVIDIA GPU: on an M4 Pro, a short decision takes about 1.2 s and a 1,100-token one about 12.6 s. MLX was measured alongside Metal and was about 10% slower on the M4, so the Mac build uses Metal.
License
Apache-2.0. Derived from Bonsai 2 27B by Prism ML (Apache-2.0), which descends from Qwen3.8-27B; see NOTICE. The training datasets keep their own licenses.
