Team Ai
Modelpublic

developerabu/qefro-decision-jef-c1

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
0likes17downloads
Model Card

Qefro Decision-Jef C1 (PyTorch & ONNX FP32)

Qefro Decision-Jef C1 is a high-performance, enterprise-grade decision and tool-routing transformer fine-tuned on the Qefro multi-intent marketplace distribution. It serves as the authoritative reranking decider in a two-stage hierarchical agent router:

User Request
     │
     ▼
multilingual-e5-base
     │ (Top-3 Functional Candidates)
     ▼
Decision-Jef C1 + "unknown" Fallback (K=4 Options)
     │
     ▼
Final Chosen Capability / Fallback

This repository provides both the PyTorch checkpoint (`c1_best.pt`) and the production-certified ONNX FP32 model (`c1_decision_jef.onnx` + `c1_decision_jef.onnx.data`) with 100% exact semantic and behavioral parity.


Key Performance Highlights

  • —Frozen Production Benchmark (525 Cases):
  • —Overall Accuracy: 92.38% (485 / 525 cases)
  • —Macro F1 Score: 0.9161
  • —Micro F1 Score: 0.9238
  • —Candidate Pool Recall (E5 Top-3 + unknown): 94.48% (496 / 525 cases)
  • —Semantic Parity (PyTorch vs ONNX):
  • —Prediction Mismatches: 0 / 525 cases (100.00% agreement)
  • —Max Logit Difference: 1.64e-04 (mean: 1.97e-05)
  • —Max Probability Difference: 1.06e-05 (mean: 1.26e-07)
  • —Order Stability: Identical flip rates across PyTorch, ONNX CPU, and ONNX CUDA.
  • —Inference Latency & Throughput:
  • —Tesla T4 GPU (CUDAExecutionProvider): 8.94 ms (p50), 10.81 ms (p99) | 111.4 req/s throughput
  • —Host CPU (CPUExecutionProvider): 152.00 ms (p50) | 6.2 req/s throughput
  • —Memory Footprint: 1.67 GB VRAM (GPU), 7.7 GB RSS (CPU)

Per-Capability Performance Breakdown (525 Frozen Cases)

Evaluated across all 14 enterprise capabilities in the Qefro platform:

CapabilityDescriptionSupportPrecisionRecallF1 ScoreAccuracy
orders.listList recent or historical customer orders470.90200.97870.938897.9%
orders.statusCheck live status/tracking of an order441.00000.90910.952490.9%
orders.cancelRequest cancellation of an order430.85110.93020.888993.0%
orders.refundRequest refund for an order or item411.00000.92680.962092.7%
catalog.searchSearch product catalog for items/stock420.97370.88100.925088.1%
customers.searchLook up customer profiles and CRM data380.94440.89470.918989.5%
invoices.getRetrieve invoices, tax documents, or receipts381.00000.94740.973094.7%
messages.sendSend individual messages or customer notifications350.89190.94290.916794.3%
campaigns.sendTrigger bulk marketing/promotional campaigns350.96880.88570.925488.6%
records.deleteDelete database entries or customer records351.00001.00001.0000100.0%
sales.summaryAggregate analytics, GMV, and revenue metrics350.97221.00000.9859100.0%
tickets.createCreate customer support / escalation tickets380.88371.00000.9383100.0%
ambiguousMark underspecified or conflicting requests250.93330.56000.700056.0%
unknownFallback for out-of-domain requests290.68290.96550.800096.6%
Macro Average—5250.92880.92090.9161—
Micro / Overall—5250.92380.92380.923892.38%

Architectural Details: Segment Isolation & Sliding Attention

Decision-Jef employs a ModernBERT backbone with learned decision (q_proj) and option (o_proj) projection heads. The ONNX model natively computes two 4D boolean attention masks:

  1. 1.Segment-Isolated Full Attention: State tokens (Segment 0) are mathematically prevented from cross-attending to Question/Option tokens (Segment 1+). This preserves state contextualization independently of candidate ordering.
  2. 2.Position-ID Sliding Attention: Sliding window layers ($w = 128$, $ ext{half} = 64$) calculate distance using restarted position_ids rather than absolute token indices, ensuring single-question geometric consistency.

Quickstart: Python Inference with ONNX Runtime

1. Install Dependencies

bash
pip install onnxruntime-gpu transformers numpy

2. Run Inference

python
import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer

# Load tokenizer and ONNX session
model_id = "developerabu/qefro-decision-jef-c1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
session = ort.InferenceSession("c1_decision_jef.onnx", providers=["CUDAExecutionProvider", "CPUExecutionProvider"])

# Example input query and candidates
user_query = "Please cancel order ORD-99214 immediately"
candidates = ["orders.cancel", "orders.status", "orders.refund", "unknown"]
descriptions = {
    "orders.cancel": "Request cancellation of an order",
    "orders.status": "Check live status/tracking of an order",
    "orders.refund": "Request refund for an order or item",
    "unknown": "Fallback when request is out-of-domain or unsupported"
}

# Format question block
# Segment 0: <s> user query </s>
# Segment 1: <unused2> instructions <unused0> opt0 ... <unused1>
bos_id = tokenizer.bos_token_id or tokenizer.cls_token_id
eos_id = tokenizer.eos_token_id or tokenizer.sep_token_id
opt_id = tokenizer.convert_tokens_to_ids("<unused0>")
dec_id = tokenizer.convert_tokens_to_ids("<unused1>")
q_id = tokenizer.convert_tokens_to_ids("<unused2>")

state_ids = [bos_id] + tokenizer(user_query, add_special_tokens=False)["input_ids"] + [eos_id]
seg_ids = [0] * len(state_ids)
pos_ids = list(range(len(state_ids)))

# Question block
q_base = len(state_ids)
q_text = "choice question: Which Qefro capability should handle this request?"
q_tokens = [q_id] + tokenizer(q_text, add_special_tokens=False)["input_ids"]

opt_positions = []
for cand in candidates:
    opt_positions.append(len(state_ids) + len(q_tokens))
    q_tokens.append(opt_id)
    q_tokens.extend(tokenizer(descriptions[cand], add_special_tokens=False)["input_ids"][:48])

dec_position = len(state_ids) + len(q_tokens)
q_tokens.append(dec_id)

all_ids = state_ids + q_tokens
all_segs = seg_ids + [1] * len(q_tokens)
all_poss = pos_ids + list(range(q_base, q_base + len(q_tokens)))

# Prepare ONNX tensors
ort_inputs = {
    "input_ids":      np.array([all_ids], dtype=np.int64),
    "attention_mask": np.ones((1, len(all_ids)), dtype=np.int64),
    "segment_ids":    np.array([all_segs], dtype=np.int64),
    "position_ids":   np.array([all_poss], dtype=np.int64),
    "batch_idx":      np.array([0], dtype=np.int64),
    "dec_pos":        np.array([dec_position], dtype=np.int64),
    "opt_pos":        np.array([opt_positions], dtype=np.int64),
    "opt_mask":       np.ones((1, len(candidates)), dtype=bool),
}

logits, probabilities = session.run(["logits", "probabilities"], ort_inputs)
best_idx = int(np.argmax(probabilities[0]))
print(f"Selected Capability: {candidates[best_idx]} (Confidence: {probabilities[0][best_idx]:.4f})")

Benchmark Artifacts & Reproducibility

  • —c1_best.pt: PyTorch weights checkpoint (Epoch 3, Validation Accuracy: 85.19%, Test Accuracy: 88.89%).
  • —c1_decision_jef.onnx: Standalone FP32 ONNX computational graph.
  • —c1_decision_jef.onnx.data: External model weights (1.23 GB).
  • —evaluation_results_525.json: Full per-case evaluation results across all 525 frozen benchmark instances.
  • —capabilities.json: Canonical dictionary of Qefro capability criteria.

Citation & License

Developed and released by Qefro AI for enterprise agent routing. Licensed under the Apache 2.0 License.