aabbdev/RWKV7-1.5B-SMI-20260822
RWKV7-1.5B-SMI-20260822
This is an RWKV-7 model fine-tuned for the State Model Interface (SMI). It preserves the parent architecture and appends exactly ten structural tokens.
Provenance
Training mixture
The locked corpus artifact contains 93,235,868 assistant target tokens across 134,295 rows. This full-sft stage selected buckets short, medium: 74,229,330 target tokens across 132,586 rows.
The values below are copied from smi_corpus_manifest.json; they are not estimates.
SMI usage and protocol
The tokenizer assigns these atomic, append-only IDs: <|ctrl|>=65536, <|sys|>=65537, <|dev|>=65538, <|caps|>=65539, <|usr|>=65540, <|obs|>=65541, <|think|>=65542, <|out|>=65543, <|act|>=65544, <|eot|>=65545. Compile trusted message structure to token IDs with an SMI-compatible compiler; do not interpolate untrusted payload text into structural markers. Runtime turns end with <|eot|> (ID 65545). Generation stops on either ID 0 or ID 65545.
The preserved chat_template.jinja, smi_token_ids.json, and tokenizer artifacts are the training-time protocol contract. Consumers should hash-pin this repository and use trust_remote_code=True for the bundled model implementation.
Training configuration
Evaluation
Values are copied from the closed-schema smi_evaluation.json v2. Main cases SHA-256: aca1b98413377a3bffa6fed28d024777e34195abfb3ef9e11abeca08433739a7. Multi-turn cases SHA-256: d5a407b61e700805ab1a58eb7cd830bf1f5b395c1355416f580316e85033d9ce.
Top-1 parity: 16 / 16.
Loading
Install the supported runtime first:
python -m pip install "transformers>=5.3,<6" "huggingface-hub>=1.5,<2"import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, PreTrainedConfig
model_id = "aabbdev/RWKV7-1.5B-SMI-20260822"
tokenizer = AutoTokenizer.from_pretrained(model_id, config=PreTrainedConfig())
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype=torch.bfloat16,
)OpenAI-compatible serving
python -m pip install -r inference/requirements.txt
python inference/serve.py --host 127.0.0.1 --port 8000The tokenizer response template maps SMI thinking, output, and actions to reasoning_content, content, and OpenAI tool_calls. Tool observations are sent back as standard role="tool" messages with the returned tool_call_id. Continuous batching is intentionally rejected because RWKV uses recurrent state, not a paged KV cache. The launcher requires transformers[serving]>=5.15,<6; direct model loading remains compatible with Transformers 5.3+.
Known limitations
- SMI structural-token discipline is a serialization boundary, not a complete security sandbox or a guarantee that generated tool calls are safe to execute.
- Fine-tuning and the reported benchmark do not establish broad factuality, safety, multilingual quality, or production suitability.
- Recurrent-cache rollback for assisted/speculative decoding is unsupported.
- The optional optimized runtime has hardware-, dtype-, and shape-specific limits and falls back to eager PyTorch outside validated boundaries.
- No evaluation values are inferred: when
smi_evaluation.jsonis absent, this card makes no quantitative training-final or benchmark claim.
License and notices
The derived weight-license identifier is reported as apache-2.0 from release metadata; other means that this publisher makes no specific weight-license claim. The generated remote code and inference bundle are distributed under Apache-2.0; see LICENSE and NOTICE.
