datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mix60k-math-code-sft
mix60k: 50/50 OpenMathInstruct-2 + OpenCodeInstruct SFT mixture
This dataset is from the paper Register Tokens for Bounded-State Reasoning in Diffusion Language Models.
The exact 60,000-trace SFT mixture used to train the LLaDA-8B-Base and Dream-7B-Base
main triad in the dLLM Registers project.
Composition: ~30K math traces from OpenMathInstruct-2 and ~30K code traces from OpenCodeInstruct.
License: governed by the upstream licenses of OpenMathInstruct-2 and OpenCodeInstruct.… See the full description on the dataset page: https://huggingface.co/datasets/albertge/mix60k-math-code-sft.agentic-code-sft-mix-v1
Agentic Code SFT Mix v1
Local derived SFT mixture for code-agent/tool-use training.
This is not a single upstream dataset. It is a filtered local mixture built from:
nvidia/OpenCodeInstruct, split train
nvidia/Nemotron-SFT-OpenCode-v1, splits general, bash_only_tool, bash_only_tool_skills, question_tool, agent_skills, agent_skills_question_tool
nvidia/Nemotron-SFT-SWE-v2, split agentless
nvidia/Nemotron-SFT-SWE-v2, file data/swe.jsonl
The output schema is JSONL with messages… See the full description on the dataset page: https://huggingface.co/datasets/synquid/agentic-code-sft-mix-v1.
