datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mix60k-math-code-sft
mix60k: 50/50 OpenMathInstruct-2 + OpenCodeInstruct SFT mixture
This dataset is from the paper Register Tokens for Bounded-State Reasoning in Diffusion Language Models.
The exact 60,000-trace SFT mixture used to train the LLaDA-8B-Base and Dream-7B-Base
main triad in the dLLM Registers project.
Composition: ~30K math traces from OpenMathInstruct-2 and ~30K code traces from OpenCodeInstruct.
License: governed by the upstream licenses of OpenMathInstruct-2 and OpenCodeInstruct.… See the full description on the dataset page: https://huggingface.co/datasets/albertge/mix60k-math-code-sft.Constrained_Indic_CodemixingUnity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam
Unity Code and GPT-Generated GDD Pairs Dataset
This dataset contains paired samples of Unity game mechanic scripts and their corresponding GPT-4 generated Game Design Documents (GDDs). It is intended for training and benchmarking LLMs in game code generation from design specifications.
Format
Each entry is stored as a .jsonl file with:
"input": GPT-4 generated GDD describing a specific game and its mechanics
"output": Unity C# scripts implementing the described mechanic… See the full description on the dataset page: https://huggingface.co/datasets/AmnaHassan/Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam.kanitakorn-deepseek-v47-v46-if-code-lcb-bridge-mix
Kanitakorn v47 v46 + IF/Code/LCB Bridge Mix
Conservative later-stage bridge mix for non-Thai-family <=14B candidates.
Delta from v46:
reuses all v46 rows unchanged, including inherited identity-attribution rows
adds 197 de-duplicated local Codex-generated rows targeting Thai instruction following, code-output formatting, and short math/code verification
excludes higher-risk dataset/train LiveCodeBench and public benchmark rows
keeps single-model training only; no BoN… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v47-v46-if-code-lcb-bridge-mix.agentic-code-sft-mix-v1
Agentic Code SFT Mix v1
Local derived SFT mixture for code-agent/tool-use training.
This is not a single upstream dataset. It is a filtered local mixture built from:
nvidia/OpenCodeInstruct, split train
nvidia/Nemotron-SFT-OpenCode-v1, splits general, bash_only_tool, bash_only_tool_skills, question_tool, agent_skills, agent_skills_question_tool
nvidia/Nemotron-SFT-SWE-v2, split agentless
nvidia/Nemotron-SFT-SWE-v2, file data/swe.jsonl
The output schema is JSONL with messages… See the full description on the dataset page: https://huggingface.co/datasets/synquid/agentic-code-sft-mix-v1.hindi-english-code-mixedThis dataset was compiled from various open sources online including some asr datasets and some percenteage of data generated using prompt engineering on generative llms. Some sources used are listed down below:
https://github.com/l3cube-pune/code-mixed-nlp?tab=readme-ov-file
https://github.com/piyushmakhija5/hinglishNorm
https://github.com/ishan00/translation-for-code-switching-acl/tree/master
ko-en-code-mixing-sts
Korean–English Code-Mixing STS Dataset
This dataset contains 1,500 Korean–English code-mixed pairs derived from KLUE-STS. We keep the original sentence_a and apply insertion-only code-mixing to sentence_b using an LLM (Gemini 2.5 Flash), recording where and how code-mixing occurred.
Interactive Dashboard
🌐 Explore the dataset interactively: https://watchstep.github.io/ko-en-cm/
The dashboard provides:
Interactive data exploration and filtering
Sample visualization with… See the full description on the dataset page: https://huggingface.co/datasets/watchstep/ko-en-code-mixing-sts.code_rlvr_mixture_sftMixtureVitae-pseudo_code_of_thoughtThis is a WIP dataset to augment https://huggingface.co/datasets/nampdn-ai/tiny-codes to create a psuedo-code of thought for reasoning about common sense situations.
There are bugs(!) in the translation process which we are in the process of fixing. Use at your discretion.
code_rlvr_mixture_dponative_script_codemixed
