Team Ai
Modelpublic

sameer-saraf-quant-ai/slm-workflow-planner-7b-v3

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
0likes
Model Card

SLM Workflow Planner 7B v3 — Fork-Suppression Alignment (Best Overall)

Model Description

LoRA adapter for Qwen/Qwen2.5-7B-Instruct fine-tuned as a workflow execution planner. This is the v3 model — the best-performing checkpoint across all training phases, trained in three stages:

  1. 1.Stage A: Base policy training on 554K samples from 89 diverse workflow graphs (iter 800)
  2. 2.Stage B: Contrastive alignment on 20K curated samples with clean decision boundaries (iter 100)
  3. 3.Stage C: Fork-suppression alignment on 4.6K targeted samples to fix FORK over-triggering (iter 200)

The model makes real-time decisions about workflow transitions by analyzing state signals, eligible nodes, and topology information.

Decision Types

DecisionDescription
NEXTProceed to the next sequential step
RETRYRetry the current step (within budget)
FORKLaunch parallel execution branches
JOINSynchronize parallel branches
METAEscalate — anomaly detected, human intervention needed

Performance (76-scenario evaluation suite)

Category**v3 SLM**v2 SLMGPT-4.1GPT-4o-miniBase SLM
NEXT15/22 (68%)8/22 (36%)6/22 (27%)2/22 (9%)16/22 (73%)
RETRY12/12 (100%)7/12 (58%)11/12 (92%)12/12 (100%)3/12 (25%)
FORK12/14 (86%)14/15 (93%)14/14 (100%)14/14 (100%)1/14 (7%)
JOIN6/15 (40%)10/15 (67%)10/15 (67%)12/15 (80%)0/15 (0%)
META0/13 (0%)3/12 (25%)0/13 (0%)0/13 (0%)8/13 (62%)
TOTAL45/76 (59.2%)42/76 (55.3%)41/76 (53.9%)40/76 (52.6%)28/76 (36.8%)

Key Results

  • —🏆 Best overall accuracy: 59.2% — outperforms all previous versions and GPT-4.1
  • —🔥 RETRY: 100% — perfect retry handling (was 58% in v2)
  • —🔥 FORK: 86% — strong parallel execution decisions with correct suppression
  • —🔥 NEXT: 68% — massive improvement over v2 (36%) without collapse to NEXT
  • —⚡ Balanced policy — the only checkpoint that achieves strong NEXT + RETRY + FORK simultaneously
  • —⚡ 4x faster inference than base model, runs locally on Apple Silicon

Architecture Evolution

VersionStrategyTotalNEXTRETRYFORKJOINMETA
v1 (base)800-iter policy training36.8%73%25%7%0%62%
v2+ contrastive alignment55.3%36%58%93%67%25%
v3+ fork suppression59.2%68%100%86%40%0%

v3 fixes v2's FORK over-triggering problem. v2 had learned "forkable → FORK" blindly. v3 correctly learns "forkable AND conditions favorable → FORK, otherwise NEXT".

Training Details

Three-Stage Training

Stage A: Base Policy (iter 800)

  • —Dataset: 554K instruction pairs from 89 workflow graphs
  • —8 structural families (linear, retry, fork-join, escalation, etc.)
  • —Balanced decision distribution: NEXT 36%, JOIN 27%, META 13%, FORK 12%, RETRY 12%

Stage B: Contrastive Alignment (iter 100)

  • —Dataset: 20K curated samples with clean decision boundaries
  • —Contrastive pairs: FORK positives + hard negatives, JOIN positives + hard negatives
  • —Proportional representation across all decision types

Stage C: Fork Suppression (iter 200)

  • —Dataset: 4,600 targeted samples
  • —Focus: "forkable but blocked → NEXT" hard negatives
  • —Teaches: resource pressure, parallel depth, uncertainty block FORK
  • —Stabilizers: RETRY and NEXT anchors to prevent forgetting

LoRA Configuration

ParameterValue
Rank16
Alpha (scale)32 (2.0x)
Dropout0.02
Target layersLast 28 of 32
Target modulesqproj, kproj, vproj, oproj

Training Configuration

ParameterValue
FrameworkMLX (Apple Silicon native)
HardwareApple M4 Pro, 48GB unified memory
Stage A iters800
Stage B iters100
Stage C iters200
Batch size4
Learning rate2e-5 (fork-suppression stage)
Sequence length512
Prompt maskingYes (loss only on assistant tokens)

Usage

With MLX (Apple Silicon)

python
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load(
    "Qwen/Qwen2.5-7B-Instruct",
    adapter_path="ssaraf1/slm-workflow-planner-7b-v3"
)

messages = [
    {"role": "system", "content": "You are a workflow planner. Given the current workflow state, eligible nodes, and topology information, classify the decision type. Respond with exactly one of: NEXT, RETRY, FORK, JOIN, META"},
    {"role": "user", "content": "Current node: VERIFY_POLICY (SYSTEM)\nOutcome: success\n\nState:\n  goal_progress=0.35\n  parallel_active=0\n  resource_pressure=0.1\n  retry_count=0\n\nEligible nodes:\n  1. FRAUD_SCREENING (SYSTEM) → produces: fraud_score\n  2. DAMAGE_ASSESSMENT (AGENT) → produces: damage_report\n\nForkable sets: [{FRAUD_SCREENING, DAMAGE_ASSESSMENT}]\nJoin-ready: []\n\nWhat decision type?"}
]

prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
sampler = make_sampler(temp=0.0)
response = generate(model, tokenizer, prompt=prompt, max_tokens=10, sampler=sampler)
print(response)  # Expected: FORK (low pressure, independent actors)

What Makes v3 Special

Fork Suppression — Correct Policy Boundaries

v2 over-triggered FORK whenever forkable_sets was present. v3 learned the correct policy:

ScenarioTopologyStatev2 Decisionv3 Decision
Low pressure + independentForkableGo parallelFORK ✅FORK ✅
High resource pressureForkableDon't parallelizeFORK ❌NEXT ✅
Already in parallelForkableToo deepFORK ❌NEXT ✅
High uncertaintyForkableRiskyFORK ❌NEXT ✅
First retry failureNot forkableRetry availableNEXT ❌RETRY ✅

Remaining Challenges (v4 targets)

  • —JOIN: 40% — model struggles with join synchronization
  • —META: 0% — anomaly detection not yet learned
  • —These require a unified alignment approach (not sequential patching)

Files

  • —adapters.safetensors — LoRA adapter weights (Stage A + B + C)
  • —adapter_config.json — LoRA configuration for MLX

Citation

Part of the Agentic Factory project — building autonomous workflow orchestration with SLM-powered planning on Apple Silicon.