Team Ai
Modelpublic

kortexa-ai/shingi-27b

sourceHugging Faceapache-2.0updated 2h agoView on Hugging Face
3likes949downloads
Model Card

Shingi 27B — 審議

[image]

Shingi 27B is a local decision model. You give it context (text, and optionally images), a question and the possible answers. It returns a choice, calibrated probabilities or an ordinal score. It reads the logit of every candidate answer directly and generates no text.

The model is a single 7.2 GB ternary GGUF file (PQ2_0). It is Bonsai 2 27B with retrained block scales: the ternary weights are unchanged, only their scales were trained. It needs no adapter.

Run it

On Linux with an NVIDIA GPU and the CUDA toolkit, or on a Mac with Apple Silicon:

bash
curl -fsSL https://raw.githubusercontent.com/kortexa-ai/shingi-27b/main/run.sh | bash

This builds the pinned Prism runtime, downloads this model into your Hugging Face cache and serves the API on http://127.0.0.1:8765. Use --port to change the port. The code, API examples and options are in kortexa-ai/shingi-27b.

Requirements

ResourceNeed
GPUNVIDIA with at least 20 GB: RTX 4090, RTX PRO 6000 or DGX Spark (tested). The model uses about 8 GiB at the 16K context, about 9 GiB with image input.
System RAMabout 1 GB for the server process, plus page cache for the 7.2 GB file
Disk7.2 GB for the weights, plus the runtime build
MacApple Silicon with 24 GB or more of unified memory (16 GB minimum), using Metal
SoftwareLinux (x86-64 or aarch64) with the CUDA toolkit, or macOS with the Xcode command line tools; CMake and a C++17 compiler

Speed

Median latency per decision, one request at a time, loading excluded:

RequestRTX PRO 6000RTX 4090DGX Spark
Short yes/no (~120 tokens)66 ms85 ms148 ms
Short choice, 3 options73 ms95 ms181 ms
Choice, 10 options (~1,300 tokens)246 ms306 ms802 ms
Long choice, 3 options (~3,000 tokens)881 ms1,045 ms2,984 ms
GPU memory in use8.3 GB8.2 GB7.9 GB

With image input on the RTX 4090, one image and one question take about 0.66 s (median) and the model uses about 9.2 GB.

Questions in one request share a single pass over the images and the state, so extra questions are cheap: on an RTX PRO 6000, eight questions about one image take about 1.8 s instead of 8.7 s.

Scores

External suites, canonical choice order, 16K context:

SuiteItemsAccuracy
JevBench public23185.7%
JevBench hard11172.1%
DecisionBench58675.3%
DecisionBench hard29365.2%
This/That7,30567.9%

These suites also informed the choice of training data, so treat them as development results rather than a clean held-out test.

Images

Shingi reads images through the Bonsai 2 27B vision projector (mmproj.gguf, included here). Send up to 8 images per request: /v1/decisions follows SGLang's decision endpoint, and /v1/systemone takes an optional images list.

Zero-shot:

SuiteItemsAccuracy
VSR (spatial yes/no)30082.0%
A-OKVQA (4-way)30089.0%
VQAv2 yes/no30089.0%

The model was not trained on images; these come from the base model's vision with Shingi's decision training on top.

Training data

Public datasets, used under their stated licenses: Banking77, CLINC150, MMLU, HelpSteer2, HelpSteer, Measuring Hate Speech, Civil Comments, GoEmotions, LEDGAR (LexGLUE), MASSIVE, CommonsenseQA, WinoGrande, GSM8K, HelpSteer3, TabFact, TAT-QA, CaseHOLD and Unfair-ToS (LexGLUE), TimeQA, Mind2Web, MedMCQA and MedQA. Several came through the jev-bench repackaging.

About a third of the training tokens are model-generated: a private synthetic dataset of decision tasks, plus the Apache-2.0 teacher data from Mapika/decider.

Limitations

  • —English only.
  • —Much slower on a Mac than on an NVIDIA GPU: on an M4 Pro, a short decision takes about 1.2 s and a 1,100-token one about 12.6 s. MLX was measured alongside Metal and was about 10% slower on the M4, so the Mac build uses Metal.

License

Apache-2.0. Derived from Bonsai 2 27B by Prism ML (Apache-2.0), which descends from Qwen3.8-27B; see NOTICE. The training datasets keep their own licenses.