Team Ai
Modelpublic

model-rampage/BareTorch-500M-Base

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes222downloads
Model Card

๐Ÿป๐Ÿ”ฅ BareTorch-500M-Base

BareTorch-500M Base is a foundational sub-quadratic language model built under a pure GEMM-compliant, kernel-free paradigm. The model combines CS-LRAD (Chunk-Segmented Low-Rank Associative Delta Engine) recurrent layers with standard Transformer multi-head self-attention in a 3:1 interleaved hybrid topology ($3\times\text{CS-LRAD} \to 1\times\text{Transformer}$).

By structuring sub-quadratic state updates into block-parallel chunk segments ($C=32$), BareTorch bypasses the compilation and hardware lock-in of custom CUDA or Triton kernels, running with $O(N)$ execution and memory efficiency natively across NVIDIA CUDA, Apple Silicon MLX, WebGPU, and TPUs.


๐Ÿ“ Model Architecture Specifications

  • โ€”Parameters: ~500M (498.2M active parameters)
  • โ€”Hidden Dimension ($d_{model}$): 1,152
  • โ€”Total Layers: 24 (Interleaved 18x CS-LRAD + 6x Transformer)
  • โ€”Attention Heads: 16 ($d_{head} = 72$)
  • โ€”CS-LRAD Subspace Rank ($r$): 8
  • โ€”Chunk Size ($C$): 32 tokens
  • โ€”Tokenizer: HuggingFaceTB/SmolLM2-360M (Vocab size: 49,152)
  • โ€”Context Window: Up to 32,768 tokens

๐Ÿ‹๏ธ Pre-Training Runway Specs & Loss Convergence

  • โ€”Training Runway: 100 Billion Tokens ($190,735$ optimization steps)
  • โ€”Hardware Cluster: $4\times$ NVIDIA H100 SXM (80GB VRAM)
  • โ€”Global Batch Size: $524,288$ tokens/step ($256$ sequences of length $2048$)
  • โ€”Optimizer & LR: AdamW ($ ext{LR}_{peak} = 6 \times 10^{-4}$, weight decay $0.1$, cosine decay scheduler with $2,000$ warmup steps)

Pre-Training Evaluation Loss Milestones

Optimization StepTokens ProcessedTrain LossEval Loss
Step 1,000~0.52 Billion5.2876--
Step 10,000~5.24 Billion2.76322.6943
Step 50,000~26.21 Billion2.52752.4836
Step 100,000~52.43 Billion2.43242.3965
Step 150,000~78.64 Billion2.33552.3054
Step 190,735 (Final)100.0 Billion2.29012.2690

๐Ÿ“Š Zero-Shot Downstream Benchmarks

Task / BenchmarkMetricScore
HellaSwagAcc (Norm)43.69%
ARC EasyAcc (Norm)53.58%
ARC ChallengeAcc (Norm)28.92%
WinoGrandeAccuracy51.30%
MMLU (Overall 57-Subject Average)Accuracy24.70%
โ”œโ”€ MMLU STEMAccuracy23.53%
โ”œโ”€ MMLU HumanitiesAccuracy24.87%
โ”œโ”€ MMLU Social SciencesAccuracy24.86%
โ””โ”€ MMLU OtherAccuracy25.46%

โšก Long-Context Inference Hardware Scaling (32,768 Context)

BareTorch replaces context-dependent KV-caches with constant-sized $O(1)$ recurrent state updates, eliminating memory bus bottlenecks and out-of-memory crashes on long-context workloads.

1. Discrete CUDA GPU (NVIDIA RTX 4090 - 24GB)

Baseline PairContextPrefill LatencyLocal GPU DecodePeak VRAMAdvantage
Qwen3 0.6B Match32,768314.89 ms vs 4,343.26 ms164.49 tok/s vs 13.15 tok/s1.66 GB vs 8.58 GB12.51x Faster (-80.6% VRAM)
SmolLM2 1.7B Match32,7681,134.20 ms vs 4,292.54 ms98.92 tok/s vs 15.79 tok/s4.39 GB vs 15.57 GB6.26x Faster (-71.8% VRAM)
Llama 3.2 1B Match32,768598.36 ms vs 2,960.03 ms169.24 tok/s vs 24.44 tok/s3.02 GB vs 4.67 GB6.92x Faster (-35.4% VRAM)
Gemma 2 2B Match32,7681,049.72 ms (Baseline: ๐Ÿ’ฅ OOM)101.51 tok/s (Baseline: ๐Ÿ’ฅ OOM)5.96 GB (Baseline: ๐Ÿ’ฅ OOM)Prevents OOM Crashes

2. Apple Silicon Unified Memory (M1 MacBook Pro 16GB - Native MLX)

Baseline PairContextPrefill LatencyLocal GPU DecodePeak VRAMAdvantage
SmolLM2 1.7B Match32,76829.56 s vs 59.03 s29.69 tok/s vs 0.66 tok/s3.78 GB vs 9.97 GB44.98x Faster (-62.1% VRAM)
Llama 3.2 1B Match32,76818.69 s vs 41.06 s28.42 tok/s vs 3.29 tok/s2.81 GB vs 3.71 GB8.64x Faster (-24.4% VRAM)
Qwen3 0.6B Match32,7687.06 s (Baseline: ๐Ÿ’ฅ OOM)79.55 tok/s (Baseline: ๐Ÿ’ฅ OOM)1.37 GB (Baseline: ๐Ÿ’ฅ OOM)Prevents OOM Crashes

๐Ÿ’ป Usage & Code Example

python
import torch
from transformers import AutoTokenizer
from baretorch.integration.configuration_baretorch import BareTorchConfig
from baretorch.integration.modeling_baretorch import BareTorchForCausalLM

model_id = "model-rampage/BareTorch-500M-Base"
tokenizer = AutoTokenizer.from_pretrained("HuggingFaceTB/SmolLM2-360M")
model = BareTorchForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16).cuda()

prompt = "The key innovation of pure GEMM sub-quadratic architectures is"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

with torch.no_grad():
    outputs = model.generate(**inputs, max_new_tokens=100)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

๐Ÿ“œ Citation

bibtex
@article{kovacevic2026baretorch,
  title={BareTorch: Challenging State-of-The-Art Sequence Mixing Topologies via Kernel-Free, Pure GEMM-Compliant Architectures},
  author={Kovacevic Buvinic, Martin Ignacio},
  journal={BareTorch Framework Laboratory Technical Report},
  year={2026}
}