Rebine/Qwen3.5-Embedding-0.8B-Memory
Qwen3.5-Embedding-0.8B-Memory
An embedding model fine-tuned from Qwen/Qwen3.5-0.8B-Base for agent memory retrieval and related domain-specific semantic retrieval.
This is the in-domain specialist sibling of `Rebine/Qwen3.5-Embedding-0.8B`. Same base model, same architecture, same encoding protocol; different training data and recipe. The two are separate, coexisting releases — this one is trained on a re-built memory-retrieval dataset and is stronger on the target domain, while the earlier release generalizes better out of domain.
Model Overview
- Model Type: Text Embedding
- Supported Languages: Chinese, English
- Number of Parameters: 752M (752,393,024)
- Context Length: 2048 tokens
- Embedding Dimension: 1024 (supports MRL dimensions 128, 256, 512, 768, 1024)
Model Details
Intended Use
Chinese/English semantic retrieval for agent memory, personal knowledge bases, conversation archives, technical documentation, and related domain-specific search, including short keyword queries and natural questions.
Encoding Protocol
Queries:
Instruct: Given a query, retrieve relevant passages that answer the query
Query: {query}Memories/passages: no instruction prefix.
Use last-token pooling (the last non-pad position of attention_mask) and L2 normalization. If using Matryoshka truncation, truncate to 128, 256, 512, or 768 dimensions before the final L2 normalization.
Use right padding: under left padding the "last non-pad token" is not the true end of the text.
This is not a sentence-transformers package with a built-in pooling wrapper; implement the pooling and query prefix as described above.
import torch
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer
model_id = "Rebine/Qwen3.5-Embedding-0.8B-Memory"
INSTRUCTION = "Given a query, retrieve relevant passages that answer the query"
tok = AutoTokenizer.from_pretrained(model_id)
tok.padding_side = "right"
model = AutoModel.from_pretrained(model_id, dtype=torch.bfloat16, device_map="cuda").eval()
@torch.no_grad()
def encode(texts, is_query=False):
if is_query:
texts = [f"Instruct: {INSTRUCTION}\nQuery: {t}" for t in texts]
enc = tok(texts, padding=True, truncation=True, max_length=2048,
return_tensors="pt").to(model.device)
hidden = model(**enc, use_cache=False).last_hidden_state
idx = enc["attention_mask"].sum(dim=1) - 1
vec = hidden[torch.arange(hidden.size(0)), idx][:, :1024]
return F.normalize(vec.float(), dim=-1)Evaluation
Full-corpus retrieval, dense cosine, K = 1/3/6/10 (K = 6 matches OpenClaw's maxResults = 6). Values are Recall@6 unless stated otherwise. Every row uses the same tokenizer protocol, query instruction, pooling, normalization, corpus and query set.
vs. Qwen/Qwen3-Embedding-0.6B
vs. Rebine/Qwen3.5-Embedding-0.8B (previous recipe)
Matryoshka dimensions (Recall@6)
Headline metrics at 1024 dims
Known Limitations
- In-domain specialist. On out-of-domain benchmarks (LoCoMo en/zh) it scores below both the previous 0.8B release and the general-purpose 0.6B baseline. For cross-domain/general retrieval use `Rebine/Qwen3.5-Embedding-0.8B`.
- No refusal behaviour. Queries with no answer in the corpus still return results; "no answer" handling must be implemented by a score threshold upstream.
- Retrieval quality is coupled to chunking: training passages were cut at ~800 tokens. Changing the chunk size degrades results.
- Inputs longer than 2048 tokens are truncated on the right; with last-token pooling a truncated input does not pool over its true ending.
Quantizations
GGUF versions: `Rebine/Qwen3.5-Embedding-0.8B-Memory-GGUF` — F16, Q80, and **Q4K_M (recommended for deployment)**.
Base Model and License
Derived from Qwen/Qwen3.5-0.8B-Base and distributed under the base model's Apache-2.0 license. Please review the base model license and usage requirements.
