Team Ai
Modelpublic

Longlong418/knowme-memory-gate-model-grpo

sourceHugging Faceupdated 7d agoView on Hugging Face
0likes369downloads
Model Card

KnowMe Memory Gate

面向个人 Agent 长期记忆检索的轻量级本地模型。

本模型基于 Qwen3-1.7B,经过 LoRA SFT + GRPO 后训练,用于完成两个任务:

  1. 1.判断当前用户输入是否需要检索长期记忆;
  2. 2.在需要检索时,生成适合 SQLite FTS5 / BM25 的高信号检索 query。

对应项目:

  • —KnowMe: https://github.com/Longlong418/KnowMe
  • —Post-training: https://github.com/Longlong418/KnowMe-memorygatemodelposttrain

模型信息

项目内容
Base ModelQwen3-1.7B
TrainingLoRA SFT + GRPO
TaskMemory Retrieval Gate + Query Generation
OutputJSON
Retrieval BackendSQLite FTS5 / BM25
DeploymentOllama / llama.cpp / Transformers

模型输出格式:

json
{
  "retrieve": true,
  "query": "新服装品牌 首次购物 体验",
  "reason": "需要查看相关历史信息"
}

字段说明:

  • —retrieve:是否需要读取长期记忆;
  • —query:真正发送给检索器的关键词;
  • —reason:给用户展示的简短原因。

评测结果

Standard Test

ModelMacro F1 ↑Precision ↑Recall ↑JSON Valid ↑Format Valid ↑
Qwen3-1.7B (Base)0.49600.52630.84511.00000.8134
Qwen3-1.7B + SFT0.95770.94520.97181.00000.9930
Qwen3-1.7B + SFT + GRPO0.96130.94560.97891.00001.0000
DeepSeek V4.1 Flash0.77460.77860.76761.00000.9824
MiMo-V2.6-Flash0.76760.76760.76760.99650.9648

Distractor-Augmented Retrieval

在候选记忆池中加入额外干扰记忆,测试 query 的检索和排序能力。

ModelHit Rate ↑Hit@1 ↑Hit@3 ↑Hit@5 ↑MRR ↑Conditional MRR ↑
Qwen3-1.7B (Base)0.43660.33800.43660.43660.38260.5906
Qwen3-1.7B + SFT0.83100.66200.82390.83100.73300.7597
Qwen3-1.7B + SFT + GRPO0.88730.67610.85920.88730.75820.7746
DeepSeek V4.1 Flash0.66900.48590.66900.66900.56690.7594
MiMo-V2.6-Flash0.62680.51410.62680.62680.56340.7692

Checkpoint 仅使用 dev set 选择,test set 只用于最终结果汇报。


输入格式

模型训练时使用单个 user message,不使用独立的 system role。

推荐输入:

text
You are a retrieval gate for a personal assistant's long-term memory.
Given the user's current message, decide whether answering well requires the user's stored long-term memory.
Long-term memory may contain facts, preferences, past events, prior conversations, decisions, plans, ongoing work, or personal constraints.

Reply with ONLY this JSON, nothing else:
{"retrieve": true/false, "query": "<2-6 high-signal search keywords if true, else empty>", "reason": "<一句中文,给用户看,不超过15字>"}

Rules:
- General knowledge, math, coding, small talk, translation, rewriting, and other self-contained requests usually do not need memory.
- Retrieve when important information needed to answer is missing from the current message but may exist in long-term memory.
- Retrieve when relevant stored preferences, constraints, previous decisions, ongoing projects, or personal history would materially improve the answer or prevent a conflicting answer.
- Do NOT retrieve merely because the message mentions the user, their life, another person, a project, or the past.
- If the current message already provides all personal information needed to answer well, do NOT retrieve.
- When retrieve=true, query must contain 2-6 concise, high-signal keywords derived ONLY from the current message. Do not invent hidden facts or answers.
- For Chinese queries, separate important search terms with ASCII spaces.
- When retrieve=false, query must be an empty string.
- Never return an array.

User message: {当前用户输入}

Chat Template

部署时需要保持与训练一致的 Qwen3 chat template。

训练时实际输入结构为:

text
<|im_start|>user
{完整 Gate Prompt}<|im_end|>
<|im_start|>assistant
<think>

</think>

对于 GGUF / Ollama,推荐使用:

text
{{- range .Messages }}
<|im_start|>{{ .Role }}
{{ .Content }}<|im_end|>
{{- end }}
<|im_start|>assistant
<think>

</think>
不建议使用裸 {{ .Prompt }} 模板,否则推理分布会与训练阶段不一致。

Ollama

如果使用 GGUF 版本,在 GGUF 文件同目录创建 Modelfile:

`text
FROM ./knowme-memory-gate-model-grpo-f16.gguf

PARAMETER temperature 0
PARAMETER num_predict 128
PARAMETER stop "<|im_end|>"

TEMPLATE """{{- range .Messages }}
<|im_start|>{{ .Role }}
{{ .Content }}<|im_end|>
{{- end }}
<|im_start|>assistant
<think>

</think>

"""

创建模型:

bash
ollama create knowme-memory-gate -f Modelfile

运行:

bash
ollama run knowme-memory-gate

OpenAI-compatible API

Ollama 默认提供本地 API:

text
http://127.0.0.1:11434/v1

示例:

python
from openai import OpenAI

client = OpenAI(
    base_url="http://127.0.0.1:11434/v1",
    api_key="ollama",
)

response = client.chat.completions.create(
    model="knowme-memory-gate",
    messages=[
        {
            "role": "user",
            "content": gate_prompt,
        }
    ],
    temperature=0,
    max_tokens=128,
)

print(response.choices[0].message.content)

其中 gate_prompt 应为上文完整 Gate Prompt 与当前用户消息拼接后的文本。


与 KnowMe 的关系

本模型是 KnowMe 长期记忆模块中的 Memory Gate。

调用链:

text
User Message
    ↓
Memory Gate
    ↓
retrieve=false ──→ Skip memory retrieval
    ↓
retrieve=true
    ↓
Generate Query
    ↓
SQLite FTS5 / BM25
    ↓
Retrieve Facts / Episodes
    ↓
Agent Context

模型只负责:

text
是否需要检索
      +
生成什么 Query

真正的记忆存储和检索由 KnowMe 完成。


数据

训练与评测数据由公开 memory benchmark 构造,包括:

  • —LoCoMo
  • —LongMemEval
  • —PersonaMem-v2
  • —RHELM

训练仓库与真实用户数据完全隔离,不读取 KnowMe 的 .knowme/state.db,也不会复制真实用户记忆用于训练或评测。


说明

  • —本模型仍保留 Qwen3-1.7B 的基础语言能力,但推荐只作为 Memory Gate 使用。
  • —部署时建议保持训练阶段的 Prompt、Chat Template 和 thinking=false 设置。
  • —GGUF F16 适合高精度本地测试;低显存设备可进一步量化为 Q4_K_M。