model-rampage/BareTorch-500M-Base
๐ป๐ฅ BareTorch-500M-Base
BareTorch-500M Base is a foundational sub-quadratic language model built under a pure GEMM-compliant, kernel-free paradigm. The model combines CS-LRAD (Chunk-Segmented Low-Rank Associative Delta Engine) recurrent layers with standard Transformer multi-head self-attention in a 3:1 interleaved hybrid topology ($3\times\text{CS-LRAD} \to 1\times\text{Transformer}$).
By structuring sub-quadratic state updates into block-parallel chunk segments ($C=32$), BareTorch bypasses the compilation and hardware lock-in of custom CUDA or Triton kernels, running with $O(N)$ execution and memory efficiency natively across NVIDIA CUDA, Apple Silicon MLX, WebGPU, and TPUs.
๐ Model Architecture Specifications
- Parameters: ~500M (498.2M active parameters)
- Hidden Dimension ($d_{model}$): 1,152
- Total Layers: 24 (Interleaved 18x CS-LRAD + 6x Transformer)
- Attention Heads: 16 ($d_{head} = 72$)
- CS-LRAD Subspace Rank ($r$): 8
- Chunk Size ($C$): 32 tokens
- Tokenizer:
HuggingFaceTB/SmolLM2-360M(Vocab size: 49,152) - Context Window: Up to 32,768 tokens
๐๏ธ Pre-Training Runway Specs & Loss Convergence
- Training Runway: 100 Billion Tokens ($190,735$ optimization steps)
- Hardware Cluster: $4\times$ NVIDIA H100 SXM (80GB VRAM)
- Global Batch Size: $524,288$ tokens/step ($256$ sequences of length $2048$)
- Optimizer & LR: AdamW ($ ext{LR}_{peak} = 6 \times 10^{-4}$, weight decay $0.1$, cosine decay scheduler with $2,000$ warmup steps)
Pre-Training Evaluation Loss Milestones
๐ Zero-Shot Downstream Benchmarks
โก Long-Context Inference Hardware Scaling (32,768 Context)
BareTorch replaces context-dependent KV-caches with constant-sized $O(1)$ recurrent state updates, eliminating memory bus bottlenecks and out-of-memory crashes on long-context workloads.
1. Discrete CUDA GPU (NVIDIA RTX 4090 - 24GB)
2. Apple Silicon Unified Memory (M1 MacBook Pro 16GB - Native MLX)
๐ป Usage & Code Example
import torch
from transformers import AutoTokenizer
from baretorch.integration.configuration_baretorch import BareTorchConfig
from baretorch.integration.modeling_baretorch import BareTorchForCausalLM
model_id = "model-rampage/BareTorch-500M-Base"
tokenizer = AutoTokenizer.from_pretrained("HuggingFaceTB/SmolLM2-360M")
model = BareTorchForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16).cuda()
prompt = "The key innovation of pure GEMM sub-quadratic architectures is"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))๐ Citation
@article{kovacevic2026baretorch,
title={BareTorch: Challenging State-of-The-Art Sequence Mixing Topologies via Kernel-Free, Pure GEMM-Compliant Architectures},
author={Kovacevic Buvinic, Martin Ignacio},
journal={BareTorch Framework Laboratory Technical Report},
year={2026}
}