Team Ai
Modelpublic

Rebine/Qwen3.5-Embedding-0.8B-Memory

sourceHugging Faceapache-2.0updated 7d agoView on Hugging Face
0likes31downloads
Model Card

Qwen3.5-Embedding-0.8B-Memory

An embedding model fine-tuned from Qwen/Qwen3.5-0.8B-Base for agent memory retrieval and related domain-specific semantic retrieval.

This is the in-domain specialist sibling of `Rebine/Qwen3.5-Embedding-0.8B`. Same base model, same architecture, same encoding protocol; different training data and recipe. The two are separate, coexisting releases — this one is trained on a re-built memory-retrieval dataset and is stronger on the target domain, while the earlier release generalizes better out of domain.

Model Overview

  • —Model Type: Text Embedding
  • —Supported Languages: Chinese, English
  • —Number of Parameters: 752M (752,393,024)
  • —Context Length: 2048 tokens
  • —Embedding Dimension: 1024 (supports MRL dimensions 128, 256, 512, 768, 1024)

Model Details

PropertyValue
Base model`Qwen/Qwen3.5-0.8B-Base`
ArchitectureQwen3.5 hybrid text model
Layers24 (18 linear-attention + 6 full-attention)
Hidden size1024
Intermediate size3584
Vocabulary size248,320
Embedding output1024 dimensions
PoolingLast-token pooling
NormalizationL2 normalization (after Matryoshka truncation)
SimilarityCosine similarity or normalized dot product
Training precisionBF16
MRL dimensions128, 256, 512, 768, 1024
LicenseApache-2.0

Intended Use

Chinese/English semantic retrieval for agent memory, personal knowledge bases, conversation archives, technical documentation, and related domain-specific search, including short keyword queries and natural questions.

Encoding Protocol

Queries:

text
Instruct: Given a query, retrieve relevant passages that answer the query
Query: {query}

Memories/passages: no instruction prefix.

Use last-token pooling (the last non-pad position of attention_mask) and L2 normalization. If using Matryoshka truncation, truncate to 128, 256, 512, or 768 dimensions before the final L2 normalization.

Use right padding: under left padding the "last non-pad token" is not the true end of the text.

This is not a sentence-transformers package with a built-in pooling wrapper; implement the pooling and query prefix as described above.

python
import torch
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer

model_id = "Rebine/Qwen3.5-Embedding-0.8B-Memory"
INSTRUCTION = "Given a query, retrieve relevant passages that answer the query"

tok = AutoTokenizer.from_pretrained(model_id)
tok.padding_side = "right"
model = AutoModel.from_pretrained(model_id, dtype=torch.bfloat16, device_map="cuda").eval()

@torch.no_grad()
def encode(texts, is_query=False):
    if is_query:
        texts = [f"Instruct: {INSTRUCTION}\nQuery: {t}" for t in texts]
    enc = tok(texts, padding=True, truncation=True, max_length=2048,
              return_tensors="pt").to(model.device)
    hidden = model(**enc, use_cache=False).last_hidden_state
    idx = enc["attention_mask"].sum(dim=1) - 1
    vec = hidden[torch.arange(hidden.size(0)), idx][:, :1024]
    return F.normalize(vec.float(), dim=-1)

Evaluation

Full-corpus retrieval, dense cosine, K = 1/3/6/10 (K = 6 matches OpenClaw's maxResults = 6). Values are Recall@6 unless stated otherwise. Every row uses the same tokenizer protocol, query instruction, pooling, normalization, corpus and query set.

vs. Qwen/Qwen3-Embedding-0.6B

SetDomainCasesMemory R@6Qwen3-Embedding-0.6B R@6
validin-domain (synthetic queries)91383.46%70.43%
goldenin-domain (real queries)10078.00%69.00%
locomoout-of-domain, English1,97767.53%74.41%
locomo_zhout-of-domain, Chinese1,97762.32%72.18%

vs. Rebine/Qwen3.5-Embedding-0.8B (previous recipe)

SetMemory R@6Rebine R@6
valid83.46%80.28%
golden (real queries)78.00%74.00%
locomo67.53%75.77%
locomo_zh62.32%72.53%

Matryoshka dimensions (Recall@6)

Dimvalidgolden
12866.16%65.00%
25675.03%70.00%
51280.94%71.00%
76883.35%77.00%
102483.46%78.00%

Headline metrics at 1024 dims

SetR@1R@3R@6R@10MRR@6
valid58.49%75.36%83.46%86.86%67.59%
golden49.00%69.00%78.00%82.00%59.80%

Known Limitations

  • —In-domain specialist. On out-of-domain benchmarks (LoCoMo en/zh) it scores below both the previous 0.8B release and the general-purpose 0.6B baseline. For cross-domain/general retrieval use `Rebine/Qwen3.5-Embedding-0.8B`.
  • —No refusal behaviour. Queries with no answer in the corpus still return results; "no answer" handling must be implemented by a score threshold upstream.
  • —Retrieval quality is coupled to chunking: training passages were cut at ~800 tokens. Changing the chunk size degrades results.
  • —Inputs longer than 2048 tokens are truncated on the right; with last-token pooling a truncated input does not pool over its true ending.

Quantizations

GGUF versions: `Rebine/Qwen3.5-Embedding-0.8B-Memory-GGUF` — F16, Q80, and **Q4K_M (recommended for deployment)**.

Base Model and License

Derived from Qwen/Qwen3.5-0.8B-Base and distributed under the base model's Apache-2.0 license. Please review the base model license and usage requirements.