Team Ai
Modelpublic

LiquidAI/LFM2.5-Encoder-350M-Diffusion

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
26likes1.3kdownloads
README.md121 linesDownload Raw Back to root
1---2language:3- en4tags:5- liquid6- lfm27- lfm2.58- bidirectional9- masked-lm10- encoder11- diffusion-language-model12- masked-diffusion13- mdlm14- instruction-tuned15library_name: transformers16license: other17license_name: lfm1.018license_link: LICENSE19pipeline_tag: text-generation20base_model:21    - LiquidAI/LFM2.5-Encoder-350M22---23 24<div align="center">25  <img26    src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png"27    alt="Liquid AI"28    style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;"29  />30  <div style="display: flex; justify-content: center; gap: 0.5em; margin-bottom: 1em;">31    <a href="https://playground.liquid.ai/"><strong>Try LFM</strong></a> •32    <a href="https://docs.liquid.ai/lfm/getting-started/welcome"><strong>Docs</strong></a> •33    <a href="https://leap.liquid.ai/"><strong>LEAP</strong></a> •34    <a href="https://discord.com/invite/liquid-ai"><strong>Discord</strong></a>35  </div>36</div>37 38# LFM2.5-Encoder-350M-Diffusion39 40A full fine-tune of [LFM2.5-Encoder-350M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M) as a masked-diffusion instruction model that generates text by iteratively unmasking tokens instead of decoding left to right.41 42The model was SFT-trained on [`mlabonne/open-perfectblend`](https://huggingface.co/datasets/mlabonne/open-perfectblend), a dataset of roughly 1.39M conversations, for 3 epochs.43 44Masked diffusion is a natural extension of masked-language modeling: the model starts from masked answer tokens, repeatedly predicts all masked positions, fills the most confident tokens, and continues until the answer is complete.45 46Find more details about our encoders in our [blog post](https://www.liquid.ai/blog/lfm2-5-encoders).47 48> [!NOTE]49> 💻 **Demos**: Try this fine-tuned model running in a CPU-only Hugging Face space:50> **[Masked-diffusion text generation](https://huggingface.co/spaces/LiquidAI/masked-diffusion)** — run the encoder as a chatbot that generates text by iteratively unmasking instead of left to right.51 52## Usage53 54Install the required packages:55 56```bash57pip install torch transformers58```59 60Run masked-diffusion text generation:61 62```python63import torch64from transformers import AutoModelForMaskedLM, AutoTokenizer65 66model_id = "LiquidAI/LFM2.5-Encoder-350M-Diffusion"67 68tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)69model = AutoModelForMaskedLM.from_pretrained(model_id, trust_remote_code=True).eval()70 71messages = [{"role": "user", "content": "Give one short tip for writing clearer code."}]72prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)73inputs = tokenizer(prompt, return_tensors="pt")74 75num_new_tokens = 1276mask_id = tokenizer.mask_token_id77input_ids = torch.cat(78    [inputs.input_ids, torch.full((1, num_new_tokens), mask_id, dtype=torch.long)],79    dim=1,80)81attention_mask = torch.ones_like(input_ids)82 83with torch.no_grad():84    for _ in range(num_new_tokens):85        mask_positions = (input_ids[0] == mask_id).nonzero(as_tuple=True)[0]86        if len(mask_positions) == 0:87            break88 89        logits = model(input_ids=input_ids, attention_mask=attention_mask).logits[0, mask_positions]90        logits[:, len(tokenizer):] = -torch.inf91        for token_id in tokenizer.all_special_ids:92            if token_id != tokenizer.eos_token_id:93                logits[:, token_id] = -torch.inf94 95        probs = logits.softmax(dim=-1)96        confidence, token_ids = probs.max(dim=-1)97        best = confidence.argmax()98        input_ids[0, mask_positions[best]] = token_ids[best]99 100generated = input_ids[0, inputs.input_ids.shape[1]:]101text = tokenizer.decode(generated, skip_special_tokens=True).split("[/Answer]")[0]102print(text.strip())103```104 105## 📬 Contact106 107- Got questions or want to connect? [Join our Discord community](https://discord.com/invite/liquid-ai)108- If you are interested in custom solutions with edge deployment, please contact [our sales team](https://www.liquid.ai/contact).109 110## Citation111 112```bibtex113@article{liquidAI2026Encoders,114  author = {Liquid AI},115  title = {LFM2.5-Encoders: Fast at Long Context, Even on CPU},116  journal = {Liquid AI Blog},117  year = {2026},118  note = {www.liquid.ai/blog/lfm2-5-encoders},119}120```121