Team Ai
Modelpublic

boying07/CPU-Based-AI-Guardrail

sourceHugging Faceupdated 3mo agoView on Hugging Face
2likes
Model Card

HomayShield ๐Ÿ”’

CPU-Based AI Guardrail for Turkish & English Security Filtering

HomayShield is a lightweight CPU-based AI guardrail designed to detect malicious, adversarial, and suspicious prompts targeting AI systems.

Unlike LLM-based guardrails, HomayShield is optimized for CPU-only inference, making it practical for organizations operating in resource-constrained or on-prem environments.


Overview

HomayShield provides AI security filtering for:

  • โ€”LLM applications
  • โ€”Chatbots
  • โ€”AI agents
  • โ€”RAG systems
  • โ€”Internal AI assistants
  • โ€”Enterprise AI pipelines

Supported languages:

  • โ€”Turkish ๐Ÿ‡น๐Ÿ‡ท
  • โ€”English ๐Ÿ‡ฌ๐Ÿ‡ง
  • โ€”Mixed Turkish-English prompts

Key Features

  • โ€”โœ… CPU-friendly inference
  • โ€”โœ… Shared encoder architecture
  • โ€”โœ… Low-latency detection
  • โ€”โœ… No GPU required in production
  • โ€”โœ… Semantic attack detection
  • โ€”โœ… Classifier-based attack detection
  • โ€”โœ… Hybrid decision engine

Architecture

HomayShield uses a shared encoder design:

Screenshot 2026-06-26 at 14.30.30

Detection Strategy

HomayShield combines two detection mechanisms.

1. Semantic Detection

Incoming prompt embeddings are compared against known attack embeddings.

Detects:

  • โ€”Prompt injection
  • โ€”Jailbreak attacks
  • โ€”Instruction override
  • โ€”Adversarial prompts
  • โ€”Semantic attack variants

2. Classifier Detection

Classifier predicts attack probability from embeddings.

Detects:

  • โ€”Known attack patterns
  • โ€”Learned malicious behaviors
  • โ€”Structured attack prompts

Inference Modes

OR Logic

Attack if either semantic or classifier score exceeds threshold.

Best for:

  • โ€”Security-first environments
  • โ€”Low false negatives

Weighted Fusion

Weighted combination of semantic + classifier scores.

Best for:

  • โ€”Balanced detection
  • โ€”Tunable sensitivity

Single Signal

Use only:

  • โ€”Semantic detection or
  • โ€”Classifier detection

Best for:

  • โ€”Benchmarking
  • โ€”Lightweight deployments

Training

Training consists of two stages.

Stage 1 โ€” Encoder Training

Loss: CosineEmbeddingLoss

Goal:

  • โ€”Cluster similar attacks
  • โ€”Separate benign and malicious prompts

Stage 2 โ€” Classifier Training

Loss: BCEWithLogitsLoss

Outputs:

  • โ€”Encoder weights
  • โ€”Classifier weights
  • โ€”Attack embedding bank

Training Data

HomayShield was trained using a multilingual dataset containing:

  • โ€”Benign prompts
  • โ€”Adversarial prompts
  • โ€”Turkish prompts
  • โ€”English prompts
  • โ€”Mixed-language prompts

Attack categories include:

  • โ€”Prompt injection
  • โ€”Jailbreak
  • โ€”Instruction override
  • โ€”Prompt leakage
  • โ€”Data exfiltration
  • โ€”Tool abuse
  • โ€”Code injection

Files

This repository contains:

  • โ€”homayshield_encoder.pt
  • โ€”homayshield_classifier.pt
  • โ€”homayshield_attack_bank.npy

Usage

Example:

Folder Structure

text
HomayShield/
โ”‚
โ”œโ”€โ”€ datasets/
โ”‚   โ”œโ”€โ”€ token_level_adversarial_tr_v2.jsonl
โ”‚   โ”œโ”€โ”€ token_level_adversarial_en_v2.jsonl
โ”‚   โ””โ”€โ”€ final_classifier_merged_all.jsonl
โ”‚
โ”œโ”€โ”€ output/
โ”‚   โ””โ”€โ”€ Homayv6/
โ”‚       โ”œโ”€โ”€ homayshield_encoder.pt
โ”‚       โ”œโ”€โ”€ homayshield_classifier.pt
โ”‚       โ””โ”€โ”€ homayshield_attack_bank.npy
โ”‚
โ”œโ”€โ”€ training2.py
โ”œโ”€โ”€ inference3.py

Training Command

bash
python training2.py \
  --train \
  ./datasets/token_level_adversarial_tr_v2.jsonl \
  ./datasets/token_level_adversarial_en_v2.jsonl \
  ./datasets/final_classifier_merged_all.jsonl \
  --output-dir ./output/Homayv6

Output Files After Training

Training generates:

text
output/Homayv6/
โ”œโ”€โ”€ homayshield_encoder.pt
โ”œโ”€โ”€ homayshield_classifier.pt
โ””โ”€โ”€ homayshield_attack_bank.npy

Inference Command

bash
python inference.py

Inference loads:

  • โ€”homayshield_encoder.pt
  • โ€”homayshield_classifier.pt
  • โ€”homayshield_attack_bank.npy

from:

text
./output/Homayv6/

Inference modes:

  • โ€”OR
  • โ€”Fusion
  • โ€”Semantic Only
  • โ€”Classifier Only

Limitations

HomayShield is not intended to replace advanced LLM-based guardrails.

Compared to LLM guardrails:

Advantages:

  • โ€”Lower infrastructure cost
  • โ€”Faster CPU inference
  • โ€”Easier deployment

Tradeoffs:

  • โ€”Lower reasoning capability
  • โ€”Less contextual understanding
  • โ€”Reduced zero-day detection

Intended Use

Recommended for:

  • โ€”Enterprise AI security
  • โ€”SOC environments
  • โ€”On-prem AI systems
  • โ€”Air-gapped deployments
  • โ€”CPU-only environments

Example Usage

Screenshot 2026-06-26 at 10.36.34


Final Verdict (Attack Detection)

ThresholdAttack RecallPrecision
0.57100%78.2%
0.5880.2%~100%
0.5938.6%100%

Your guardrail is highly effective for attack detection, especially due to the semantic layer. Attack Detection Rating: Semantic Layer: 9.5/10 Classifier Layer: 7.5/10 Overall Attack Detection: 9/10

Philosophy

AI security should not be limited to organizations with GPU infrastructure.

Even lightweight CPU-based guardrails can provide meaningful protection for real-world AI systems.

ChatGPT Image Jun 26, 2026 at 12_02_58 AM(2)