datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lfm2.5-2.6b-atlas
lfm2.5-2.6b-atlas
A brain atlas for LiquidAI/LFM2.5-2.6B, a 30-layer hybrid language model with 22 short-convolution blocks and eight grouped-query attention blocks. The atlas maps activation statistics, topic preferences, authentic/corporate prompt contrasts, weight spectra, and candidate intervention directions across the model.
The early contrast is nearly absent at layer 0 and becomes much stronger around the first attention block. Later, a correlated ML-oriented group… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/lfm2.5-2.6b-atlas.to-tool-call-datasets-LFM2.5-pythonic
to-tool-call-datasets → LFM2.5 Pythonic tool-call format
A derivative of zhangdw/to-tool-call-datasets (apache-2.0)
normalized for fine-tuning Liquid AI LFM2 / LFM2.5 models, whose native tool-call format is Pythonic:
<|im_start|>assistant
<|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|>
Nine public tool-calling corpora (APIGen-MT, ButtonInstruct, Glaive v2, GraphSyn, LoopTool, τ-bench train, ToolACE, When2Call, xLAM-60k)… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/to-tool-call-datasets-LFM2.5-pythonic.Nemotron-RL-IF-Gym-LFM2.5-prompts
Nemotron-RL instruction-following & abstention Gym tasks → LFM2.5 prompt format
Verifiers are NVIDIA NeMo-Gym rule checkers, not IFEval. Each row keeps the full verifier metadata in extra (verifier regex/string-match specs, schema_str/schema_type, exp_cal_state, answer) so the checks can be re-implemented or run through NeMo Gym's resources servers. qa_abstention rows carry a per-row license of CC BY-SA 4.0 from the source.
A derivative of [five nvidia/Nemotron-RL-* NeMo-Gym… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/Nemotron-RL-IF-Gym-LFM2.5-prompts.Nemotron-Agentic-v1-LFM2.5-pythonic
Nemotron-Agentic-v1 → LFM2.5 Pythonic tool-call format
A derivative of nvidia/Nemotron-Agentic-v1 (cc-by-4.0)
normalized for fine-tuning Liquid AI LFM2 / LFM2.5 models, whose native tool-call format is Pythonic:
<|im_start|>assistant
<|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|>
Multi-turn conversational tool-use trajectories (interactive_agent: goal decomposition with persona-seeded users; tool_calling: general function… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/Nemotron-Agentic-v1-LFM2.5-pythonic.IF-RLVR-IFEval-prompts-LFM2.5
IF-RLVR prompts with IFEval verifier spec → LFM2.5 prompt format
A derivative of [nvidia/Nemotron-RL-instruction_following + allenai/RLVR-IFeval (via sungyub/if-verl-unified)](https://huggingface.co/datasets/nvidia/Nemotron-RL-instruction_following + allenai/RLVR-IFeval (via sungyub/if-verl-unified)) (odc-by)
normalized for fine-tuning Liquid AI LFM2 / LFM2.5 models, whose native tool-call format is Pythonic:
<|im_start|>assistant
<|tool_call_start|>[get_weather(location='Paris… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/IF-RLVR-IFEval-prompts-LFM2.5.Nemotron-SFT-Agentic-v2-LFM2.5-pythonic
Nemotron-SFT-Agentic-v2 → LFM2.5 Pythonic tool-call format
A derivative of nvidia/Nemotron-SFT-Agentic-v2
(CC-BY-4.0) normalized for supervised fine-tuning of Liquid AI LFM2 / LFM2.5 models, whose
native tool-call format is Pythonic:
<|im_start|>assistant
<|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|>
Every row was rendered through the official LiquidAI/LFM2.5-VL-3B chat template (identical
to the LFM2.5 text models'… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/Nemotron-SFT-Agentic-v2-LFM2.5-pythonic.IF-multi-constraints-upto5-LFM2.5-prompts
IF_multi_constraints_upto5 → LFM2.5 prompt format (for RLVR / rejection sampling / DPO)
A derivative of allenai/IF_multi_constraints_upto5 (odc-by)
normalized for fine-tuning Liquid AI LFM2 / LFM2.5 models, whose native tool-call format is Pythonic:
<|im_start|>assistant
<|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|>
Prompt-only rows (prompt_only = true): Tulu-SFT instructions with up to 5 verifiable constraints from IFEval… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/IF-multi-constraints-upto5-LFM2.5-prompts.IF-multi-constraints-upto5-SFT-LFM2.5
IF_multi_constraints_upto5_SFT → LFM2.5 chat format
A derivative of UniLu/IF_multi_constraints_upto5_SFT (odc-by)
normalized for fine-tuning Liquid AI LFM2 / LFM2.5 models, whose native tool-call format is Pythonic:
<|im_start|>assistant
<|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|>
SFT-ready precise-instruction-following pairs: the allenai IF-RLVR prompts answered by Gemma-4-31B-it and filtered with the official IFBench… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/IF-multi-constraints-upto5-SFT-LFM2.5.Nemotron-SFT-Agentic-v2-LFM2.5-pythonic-dryrun
[DRY RUN — 2,000 rows/split] Nemotron-SFT-Agentic-v2 → LFM2.5 Pythonic tool-call format
A derivative of nvidia/Nemotron-SFT-Agentic-v2
(CC-BY-4.0) normalized for supervised fine-tuning of Liquid AI LFM2 / LFM2.5 models, whose
native tool-call format is Pythonic:
<|im_start|>assistant
<|tool_call_start|>[get_weather(location='Paris, France', unit='celsius')]<|tool_call_end|><|im_end|>
Every row was rendered through the official LiquidAI/LFM2.5-VL-3B chat template (identical
to… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/Nemotron-SFT-Agentic-v2-LFM2.5-pythonic-dryrun.lfm2.5-calibration-pack
Calibration Pack v0 — 5,000,000 Tokens Dataset
Mô hình mục tiêu: LFM2.5-2.6BQuy mô: 5,000,000 tokens (13,975 sequences)Ngày tạo: 2026-09-12 16:25:56
1. Cơ cấu Phân vùng (Partitions)
Phân vùng
Tệp Parquet
Mẫu (Seqs)
Số Tokens
Dung lượng
Short Language / General
partitions/01_short_language_1.5m.parquet
2,159
1,500,000
5.05 MB
Reasoning & Arithmetic
partitions/02_reasoning_1.0m.parquet
8,433
1,000,000
2.14 MB
Targeted Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/Neeze/lfm2.5-calibration-pack.scriber-lfm2.5-350m-polishing-de-training-v1
Scriber LFM2.5 German STT post-processing data
This repository contains the exact 2,000 German source/target pairs used to
train the final Scriber LFM2.5 350M local post-processing model.
The matching model is
Buttermilk03/scriber-lfm2.5-350m-polishing-de-qad-v1.
The complete production recipe and the lessons that determined it are in
TRAINING.md; machine-readable settings are in
training_recipe.json.
Data
Each JSONL row contains:
source: flat German… See the full description on the dataset page: https://huggingface.co/datasets/Buttermilk03/scriber-lfm2.5-350m-polishing-de-training-v1.HardGen-LFM2.5-pythonic
HardGen (FunReason-MT) → LFM2.5 Pythonic tool-call format
⚠️ Evaluation contamination notice. The source was generated by sampling in the Berkeley Function-Calling Leaderboard (BFCL) multi-turn environment (Gorilla file system, trading bot, etc.). Do not train on it if you report BFCL numbers; use it for analysis, as a hard held-out set, or with full awareness of the overlap.
A derivative of Bingguang/HardGen (apache-2.0)
normalized for fine-tuning Liquid AI LFM2 / LFM2.5… See the full description on the dataset page: https://huggingface.co/datasets/hungbenjamin402/HardGen-LFM2.5-pythonic.lfm2-gqa-only-tts-benchmark_metricshomedepot-lfm25-scored
