datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moq-lfm25-230m-resultslfm2.5-2.6b-atlas
lfm2.5-2.6b-atlas
A brain atlas for LiquidAI/LFM2.5-2.6B, a 30-layer hybrid language model with 22 short-convolution blocks and eight grouped-query attention blocks. The atlas maps activation statistics, topic preferences, authentic/corporate prompt contrasts, weight spectra, and candidate intervention directions across the model.
The early contrast is nearly absent at layer 0 and becomes much stronger around the first attention block. Later, a correlated ML-oriented group… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/lfm2.5-2.6b-atlas.lfm25-fr-corpus-v1
LFM2.5-Audio FR/EN — assistant vocal avec appel d'outils
14360 dialogues bilingues et 68.5 h de parole synthétique, pour entraîner un
modèle parole à parole qui décide d'appeler un outil, lit son résultat et répond à voix haute,
en français comme en anglais.
Deux vues du même corpus :
Les briques audio (ci-dessous) : un clip par tour, écoutable dans le visualiseur, avec son
texte, sa voix et son taux d'erreur de ré-écoute.
Les dialogues (dataset/*.jsonl) : la conversation… See the full description on the dataset page: https://huggingface.co/datasets/Rcarvalo/lfm25-fr-corpus-v1.LFM2-Terminal-SFT-Processedfineweb-edu-10B-lfm25-chunked-512
FineWeb-Edu 10BT — LFM2.5 chunked (512 tokens)
Pre-tokenized, packed chunks of HuggingFaceFW/fineweb-edu sample-10BT, tokenized with the LFM2.5 tokenizer and cut into fixed 512-token sequences.
What's in it
Field
Value
Source
HuggingFaceFW/fineweb-edu sample-10BT (train)
Tokenizer
LiquidAI/LFM2.5-1.2B-Base
Chunk size
512 tokens
Total chunks
19,300,635
Total tokens
9,881,925,120
Columns
input_ids (int32, length 512)
Format
HF Arrow (mmap-able… See the full description on the dataset page: https://huggingface.co/datasets/dme5245/fineweb-edu-10B-lfm25-chunked-512.lfm25-hermes-tool-trace-analysis
LFM2.5 Hermes Tool Trace Analysis
This repository documents a local Mac fine-tuning run for LiquidAI/LFM2.5-8B-A1B aimed at Hermes-style agent tool use.
It contains compact, reproducible artifacts only: configs, manifests, eval JSON, logs, and the focused tool-call repair dataset. Large model checkpoints are released separately under sjakek/LFM-2.5-8B-1B-hermes-ft.
Current Release State
The first public GGUF export from the tool-router repair was withdrawn because… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/lfm25-hermes-tool-trace-analysis.LFM2.5-8B-A1B-KO-CPT-DATA
LFM2.5-8B-A1B Korean CPT Data
Prepared Korean continued-pretraining data for LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL.
Files
data/ko_cpt_mix_full_lfmstyle_20260627.jsonl: prepared full CPT corpus with one JSON object per line and a text field
metadata/ko_cpt_mix_full_lfmstyle_20260627.stats.json: corpus statistics
metadata/ko_cpt_mix_full_lfmstyle_20260627.stats.json.full_report.json: per-source preprocessing report
metadata/ko_cpt_sources_full_20260627.json:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-DATA.lfm2-bilingual-fr-partialfdm-40ch-fresh-lfm2to-tool-call-datasets-LFM2.5-pythonic
to-tool-call-datasets → LFM2.5 Pythonic tool-call format
A derivative of zhangdw/to-tool-call-datasets (apache-2.0)
normalized for fine-tuning Liquid AI LFM2 / LFM2.5 models, whose native tool-call format is Pythonic:
<|im_start|>assistant
<|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|>
Nine public tool-calling corpora (APIGen-MT, ButtonInstruct, Glaive v2, GraphSyn, LoopTool, τ-bench train, ToolACE, When2Call, xLAM-60k)… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/to-tool-call-datasets-LFM2.5-pythonic.LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-Raw
LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-Raw
Stage2 raw LFM chat JSONL shards: Korean domain, behavior, SWE/coding, reasoning, finance, legal, Text2SQL.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-SFT-Stage2-Diverse-KoSWE-Reasoning-LFMChat-Raw.lfm2-bilingual-pilot-125hantidoom-mix-v1.0-LFM2.5-1.2B-Base
Antidoom Mix v1.0 – FTPO Pairs for LFM2.5-1.2B-Base
[!Note]
📓 Tutorial: This is the dataset used in Antidoom Training with TRL
Pre-extracted FTPO (Final Token Preference Optimization) training pairs for eliminating doom loops in LiquidAI/LFM2.5-1.2B-Base.
Derived from LiquidAI/antidoom-mix-v1.0 by generating completions with LFM2.5-1.2B-Base.
Schema
Each row contains:
Field
Type
Description
full_prompt
string
The formatted prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/LiquidAI/antidoom-mix-v1.0-LFM2.5-1.2B-Base.LFM2.5-KO-CPT-Full-LFMStyle-Shards-20260627
LFM2.5-KO-CPT-Full-LFMStyle-Shards-20260627
Source-separated LFM-style CPT shards: Korean Wiki, finance, legal raw/tasks/RAG/bar answers, and terminal ToolBench.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-CPT-Full-LFMStyle-Shards-20260627.med-lfm2.5-1.2b-autocomplete-assets
assets
genshiai-daichi/med-lfm2.5-1.2b-autocomplete(gated)のカードで使うデモGIFを、ゲート外で配信するための補助リポジトリです。
lfm2-tool-aware-dataset-v1
LFM2-Tool-Aware Dataset (v1)
Synthetic speech dataset for fine-tuning LFM2.5-Audio-class audio LLMs to be tool-aware: read a tool list from the system prompt, acknowledge briefly when a user query matches a listed tool, refuse politely when no tool covers the query, and otherwise behave as a normal conversational model.
Used to train matbee/lfm2.5-audio-tool-aware-v1 (96.6% accuracy on the held-out eval split).
What this teaches a model
Voice-assistant systems often… See the full description on the dataset page: https://huggingface.co/datasets/matbee/lfm2-tool-aware-dataset-v1.lfm2-tool-aware-dataset-v2
LFM2-Tool-Aware Dataset (v2)
Synthetic speech dataset for fine-tuning LFM2.5-Audio-class audio LLMs to handle both turns of a tool-augmented voice flow: acknowledge briefly on turn 1, then narrate the dispatcher's result on turn 2 after the coordinator injects it via set_context().
Used to train matbee/lfm2.5-audio-tool-aware-v2 (~97% accuracy on the eval split, including the new turn-2 narration class).
What's new in v2
The v1 dataset taught the model to ack-and-stop… See the full description on the dataset page: https://huggingface.co/datasets/matbee/lfm2-tool-aware-dataset-v2.LFM2.5-1.2B-Translate-DatasetLFM2-Terminal-SFT-TokenizedLFM25-Terminal-ToolBench-Full-Tokenized
LFM2.5 Terminal ToolBench Full Tokenized Dataset
LFM2.5-8B-A1B train-ready token IDs for the Terminal + ToolBench full SFT run.
Contents
lfm25_8b_a1b_terminal_full_toolbench_full_train_ready_v1: 197373 rows, 17.67 GiB, features: input_ids, seq_lengths, labels
Notes
This dataset stores token IDs and labels, not raw conversations.
It was used by the LFM2.5-8B-A1B Terminal ToolBench full SFT config.
Features: input_ids, seq_lengths, labels.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM25-Terminal-ToolBench-Full-Tokenized.LFM2.5-350M-IT-dataset-MedTech-WorkshopNemotron-RL-IF-Gym-LFM2.5-prompts
Nemotron-RL instruction-following & abstention Gym tasks → LFM2.5 prompt format
Verifiers are NVIDIA NeMo-Gym rule checkers, not IFEval. Each row keeps the full verifier metadata in extra (verifier regex/string-match specs, schema_str/schema_type, exp_cal_state, answer) so the checks can be re-implemented or run through NeMo Gym's resources servers. qa_abstention rows carry a per-row license of CC BY-SA 4.0 from the source.
A derivative of [five nvidia/Nemotron-RL-* NeMo-Gym… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/Nemotron-RL-IF-Gym-LFM2.5-prompts.LFM2.5-8B-A1B-Terminal-ECHO-RLVR-Rolloutsantidoom-mix-v1.0-LFM2.5-1.2B-Base
Antidoom Mix v1.0 – FTPO Pairs for LFM2.5-1.2B-Base
[!Note]
📓 Tutorial: This is the dataset used in Antidoom Training with TRL
Pre-extracted FTPO (Final Token Preference Optimization) training pairs for eliminating doom loops in LiquidAI/LFM2.5-1.2B-Base.
Derived from LiquidAI/antidoom-mix-v1.0 by generating completions with LFM2.5-1.2B-Base.
Schema
Each row contains:
Field
Type
Description
full_prompt
string
The formatted prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/kepom/antidoom-mix-v1.0-LFM2.5-1.2B-Base.lfm25-datasetLFM2.5-Fable-Router-Curriculum
Autonoma Fable router curriculum — reproducibility package
This public repository documents and reproduces the router-only curriculum for an experimental LFM2.5-2.6B Fable hybrid. It intentionally does not redistribute records derived from a source whose license is unknown.
Included
fable-router-curriculum.v2.json: immutable training/data policy.
build_fable_router_curriculum.py: deterministic local builder.
REPRODUCIBILITY_MANIFEST.json: pinned sources… See the full description on the dataset page: https://huggingface.co/datasets/Swordnael/LFM2.5-Fable-Router-Curriculum.LFM2.5-KO-SFT-Stage1-Legal-Terminal-LFMChat-8K
LFM2.5-KO-SFT-Stage1-Legal-Terminal-LFMChat-8K
Stage1 8k Korean legal/terminal/tool-use prepared SFT arrays.
This dataset is part of the LFM2.5-8B-A1B-KO-SFT / Agentic SFT workflow.
Main SFT model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-SFT
CPT base model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-CPT-FULL
Agentic follow-up model: https://huggingface.co/LLM-OS-Models/LFM2.5-8B-A1B-KO-Agentic-SFT
SFT GitHub: https://github.com/gyunggyung/LFM25-KO-SFT… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/LFM2.5-KO-SFT-Stage1-Legal-Terminal-LFMChat-8K.LFM2.5-1.2B-Thinking-0006-Dataset
KoarAI/LFM2.5-1.2B-Thinking-0006-Dataset
Next-generation, curated multi-teacher High-Density Reasoning, Multilingual CoT & Adversarial Calibration Dataset specifically engineered for full fine-tuning of LiquidAI LFM2.5-1.2B (Base/Instruct) into an elite on-device reasoning intelligence.
🌟 Key Highlights & Architectural Innovations
Zero-Shortcutting (Strict CoT Verification):
All trivial/1-line thinking templates (<think> Understand intent </think>) have been… See the full description on the dataset page: https://huggingface.co/datasets/KoarAI/LFM2.5-1.2B-Thinking-0006-Dataset.AI-Bicycle-LFM2-VL-450M
🧠 Multistage Japanese Dialogue & VQA Dataset
本データセットは、複数の公開データセットを統合・日本語化し、
自然な日本語対話および視覚応答(Vision Reaction)データを生成したものです。
出典を明示することで、商用利用が可能です。
ライセンスを尊重したうえで翻訳・加工を行い、3つのステージで構成されています。
注意: Vision Reaction とは、画像に対して自然な応答をすることです。
例えば、海の画像を入力したとき、一般的なVQAタスクでは「青い空と青い海が広がっています。....」のような説明をします。
Vision Reactionでは、「お、きれいな海だな!」というように、自然なリアクションをします。
📘 データ構成
Stage 1:en_multiturn.jsonl
元データ:allenai/soda
ライセンス:Creative Commons Attribution 4.0 International (CC BY 4.0)… See the full description on the dataset page: https://huggingface.co/datasets/HayatoHongo/AI-Bicycle-LFM2-VL-450M.IF-multi-constraints-upto5-LFM2.5-prompts
IF_multi_constraints_upto5 → LFM2.5 prompt format (for RLVR / rejection sampling / DPO)
A derivative of allenai/IF_multi_constraints_upto5 (odc-by)
normalized for fine-tuning Liquid AI LFM2 / LFM2.5 models, whose native tool-call format is Pythonic:
<|im_start|>assistant
<|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|>
Prompt-only rows (prompt_only = true): Tulu-SFT instructions with up to 5 verifiable constraints from IFEval… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/IF-multi-constraints-upto5-LFM2.5-prompts.
