Team Ai
Modelpublic

wxsys/qwen-servitor

sourceHugging Faceapache-2.0updated 2d agoView on Hugging Face
1likes1.1kdownloads
Model Card

πŸ’€ Qwen-Servitor

<p align="center"> <strong>Decapitated 0.8B Hybrid Gated DeltaNet Decision Gate & OASIS SARIF v2.1.0 Arbiter</strong> </p>

<p align="center"> <a href="https://huggingface.co/wxsys/qwen-servitor"><img src="https://img.shields.io/badge/πŸ€—%20HuggingFace-wxsys%2Fqwen--servitor-ffd21e.svg" alt="Hugging Face Model"></a> <img src="https://img.shields.io/badge/License-Apache2.0-red.svg" alt="License"> <img src="https://img.shields.io/badge/Architecture-HybridGatedDeltaNet-darkred.svg" alt="Architecture"> <img src="https://img.shields.io/badge/ContextWindow-262kTokens-blueviolet.svg" alt="Context Window"> <img src="https://img.shields.io/badge/GenerativeHead-Excised(Non--Autoregressive)-black.svg" alt="Non-Autoregressive"> <img src="https://img.shields.io/badge/OperationalMode-PureLogits-critical.svg" alt="Pure Logits"> <img src="https://img.shields.io/badge/Latency-11.5ms(C%2B%2B%2FAVX2)-success.svg" alt="Sub-15ms Latency"> </p>


Overview

Qwen-Servitor is a headless neural decision engine designed for automated pull request review, pre-commit gating, and static code security triage. It operates as the fast, lightweight component in the Servitor family, complementing the larger `wxsys/spark-servitor` (1.7B, Sliding Window Attention, 1M context).

Where Spark-Servitor provides deep multi-file receptive capacity for massive monorepos, Qwen-Servitor is optimized for speed and minimal memory footprint. Built upon the hybrid backbone of Qwen3.5-0.8B (18 Gated DeltaNet layers and 6 Full Attention layers), the 155M-parameter generative projection layer (lm_head) is excised. Hidden states route directly into a multi-task neural judgment cortex that outputs structured verdicts, calibrated risk scores, pathology tags, and exploit flows in a single forward evaluation pass.


Technical Architecture

text
                       [ RAW DIFF / 262K TOKEN CONTEXT ]
                                      β”‚
                                      β–Ό
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β”‚               QWEN 3.5 HYBRID NEURAL SPINE                  β”‚
       β”‚   β€’ 18x Gated DeltaNet Layers (O(1) Recurrent State)        β”‚
       β”‚   β€’ 6x Gated Full Attention Layers (Global Context)         β”‚
       β”‚   β€’ HySparse2 Cross-Layer KV Sharing (33% GEMM Reduction)   β”‚
       β”‚   β€’ AST-Landmark Eviction (Caps attention cache at 4k tokensβ”‚
       β”‚   β€’ 155M-parameter Generative Head: EXCISED / REMOVED       β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                      β”‚
                         [ Final Hidden State (h_T) ]
                                      β”‚
                                      β–Ό
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β”‚          LATENT RECURRENT PONDERING (k = 1..4)              β”‚
       β”‚   β€’ Internal recurrence loop without token emission         β”‚
       β”‚   β€’ Early halting at confidence threshold p_halt >= 0.95    β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                      β”‚
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
             β–Ό                        β–Ό                        β–Ό
     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
     β”‚ VERDICT HEAD  β”‚        β”‚   RISK HEAD   β”‚        β”‚ PATHOLOGY HEADβ”‚
     β”‚ 3-Way Output  β”‚        β”‚  Sigmoid(1)   β”‚        β”‚ 8-Class Multi β”‚
     β”‚ β€’ APPROVE     β”‚        β”‚  Continuous   β”‚        β”‚ β€’ SECURITY    β”‚
     β”‚ β€’ QUARANTINE  β”‚        β”‚  Risk Score   β”‚        β”‚ β€’ DEADLOCK    β”‚
     β”‚ β€’ REJECT      β”‚        β”‚  (0.00-1.00)  β”‚        β”‚ β€’ PERF_COLLAP β”‚
     β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β”‚
             β–Ό (If QUARANTINE: Formal Neuro-Symbolic Hand-Off)
     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
     β”‚ Z3 SMT Solver β”‚ ──> [ Final Verified Gate ]
     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Specifications

ParameterSpecification
Base ArchitectureDecapitated Qwen/Qwen3.5-0.8B-Base
Hidden Dimension ($d_{\text{model}}$)1024
Attention Layers24 total (18 Gated DeltaNet Linear Layers + 6 Full Attention Layers)
Linear Attention MechanismGated DeltaNet with associative state recurrence ($O(1)$ inference memory)
Full Attention LayersInterleaved at 4-layer intervals with HySparse2 Cross-Layer KV Sharing
Context WindowUp to 262,144 tokens (AST-Landmark Eviction caps attention cache at 4,096 tokens)
Decision MechanismHierarchical Multi-Task Cortex with Latent Recurrent Pondering ($k = 1..4$)
Execution Latency11.52 ms median on standard CPU (AVX2 / C++ binary)
Active Memory~450 MB RAM (AWQ INT4) / ~500 MB (W8A8)

Capabilities & Enhancements (v2.5)

1. Single-Pass 3-Way Verdict Gate

Diffs route directly into three deterministic states:

  • β€”`APPROVE`: Patch contains verified changes with no detected pathology (risk $\le 0.20$).
  • β€”`QUARANTINE`: Borderline invariant changes, formal boundary shifts, or ambiguous concurrency locks routed to SMT verification.
  • β€”`REJECT`: Confirmed vulnerability or regression (risk $\ge 0.80$).

2. Parameter-Aware Anti-False-Alarm Taint Tracking

Inspects abstract syntax trees to distinguish between unsafe string concatenation and safe parameterized query execution. Prepared statements, argument tuples (?, %s, $1), and sanitized call wrappers are recognized as valid sanitizers across Python, Go, and TypeScript, eliminating false alarms on database calls.

3. Scope-Aware Indentation Normalization

Diff hunks often contain partial indentations or dangling block clauses (except:, finally:, case, unclosed function blocks). The canonicalizer normalizes relative indentation and wraps unclosed control clauses prior to AST parsing, preventing syntax parser failures on partial diff hunks.

4. Dynamic Auto-Remediation with Closed Self-Verifying Loop

When a defect is identified, the remediation engine synthesizes targeted, git-applicable patch diffs (parameterized queries, argument lists, deferred mutex releases). Each proposed patch undergoes verification through the evaluation engine; only patches that achieve APPROVE with risk $\le 0.35$ are accepted.

5. Outlier-Preserved INT8 Embedding Table Compaction

Quantizes the large token embedding table ($248,320 \times 1024$) using symmetric per-token INT8 scales while preserving top norm outliers in floating-point precision. This cuts embedding storage footprint by ~49.6% while maintaining directional cosine similarity $\ge 0.999$ and 0% vocabulary loss.

6. OASIS SARIF v2.1.0 Exploit Flow Reconstruction

For any detected vulnerability, Qwen-Servitor reconstructs the exploit data-flow path:

text
[1. SOURCE: Untrusted Input] ──> [2. PROPAGATION: Variable Flow] ──> [3. SANITIZER: Status] ──> [4. SINK: Vulnerable Call]

Emits standard SARIF codeFlows and threadFlows compatible with GitHub Advanced Security and VS Code SARIF Viewer.


Available Model Formats

FormatPath in RepositorySize on DiskActive RAMRuntime Environment
AWQ INT4 Nativeint4/qwen-servitor-awq_int4.servitor422.7 MB~450 MBStandalone C++/Rust binary (11.52 ms latency)
W8A8 Nativeint4/qwen-servitor-w8a8.servitor751.5 MB~500 MBHigh-precision native C-ABI execution
GGUF Q8_0gguf/qwen-servitor-q8_0.gguf774.0 MB~850 MBllama.cpp / Local CPU inference
GGUF BF16gguf/qwen-servitor-bf16.gguf1.45 GB~1.6 GBFull precision GGUF runner
Native BF16bf16/model.safetensors1.45 GB~1.6 GBPyTorch GPU / Server pipelines
Cortex Headcortex_head.pt74.0 MB~100 MBDecapitated CAMQP multi-task head

Benchmark Scorecard

Evaluated across the test suite on standard CPU (AVX2 / FMA instruction set, DDR3-1600 memory):

BenchmarkValueTargetStatus
Median Native Inference Latency11.52 ms< 15.0 msMet
Throughput (Single Core)72.7 ops/sec> 50 ops/secMet
Unit & Integration Test Suite181 / 181 Passed100%Met
Prepared Statement False Alarms0 / 1000%Met
Diff Hunk Indentation Parse Errors0 / 1000%Met
Embedding Compaction Cosine Sim0.9997$\ge 0.9900$Met

Quickstart

Python API

python
from qwen_servitor import Servitor

# Initialize engine (loads weights from local path or Hugging Face Hub)
servitor = Servitor.summon("wxsys/qwen-servitor")

git_patch = """
--- a/auth/session.py
+++ b/auth/session.py
@@ -12,4 +12,3 @@ def verify_token(token: str) -> bool:
-    if not validate_hmac(token):
-        raise SecurityException("Invalid HMAC signature")
+    return True  # TODO: temporary debug bypass
"""

verdict = servitor.judge(git_patch)

print(f"Status: {verdict.status}")        # "REJECTED"
print(f"Risk: {verdict.risk:.4f}")         # 0.9984
print(f"Pathologies: {verdict.pathology}") # ["SECURITY_VULN"]
print(f"Latency: {verdict.latency_ms:.2f} ms")

Dynamic Auto-Remediation

python
from qwen_servitor.remediation import AutoRemediator

remediator = AutoRemediator(servitor=servitor)
remediations = remediator.remediate(git_patch, sarif_filepath="auth/session.py")

for candidate in remediations:
    print(f"Strategy: {candidate.strategy}")
    print(f"Verified Clean: {candidate.verified}")
    print(f"Verified Risk: {candidate.verification_verdict.risk:.4f}")
    print("Synthesized Patch:")
    print(candidate.patch_diff)

CLI and Pre-Commit Hook

bash
# Install git pre-commit hook targeting sub-15ms execution
servitor hook install --strict --veto-threshold 0.80

# Evaluate a specific unified diff and emit SARIF v2.1.0
servitor judge patch.diff --format sarif > report.sarif

Related Repositories

  • β€”**wxsys/spark-servitor**: Scaled 1.7B variant with Sliding Window Attention (SWA 512) and 1M token context for deep monorepo analysis.

License

  • β€”Base model weights are adapted from Alibaba Cloud's Qwen3.5 series under the Apache 2.0 License.
  • β€”Servitor architecture, neural cortex, and native runtime are licensed under the Apache 2.0 License.