Team Ai
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01albertge /mix60k-math-code-sft mix60k: 50/50 OpenMathInstruct-2 + OpenCodeInstruct SFT mixture This dataset is from the paper Register Tokens for Bounded-State Reasoning in Diffusion Language Models. The exact 60,000-trace SFT mixture used to train the LLaDA-8B-Base and Dream-7B-Base main triad in the dLLM Registers project. Composition: ~30K math traces from OpenMathInstruct-2 and ~30K code traces from OpenCodeInstruct. License: governed by the upstream licenses of OpenMathInstruct-2 and OpenCodeInstruct.… See the full description on the dataset page: https://huggingface.co/datasets/albertge/mix60k-math-code-sft.texttext-generation10K<n<100K0 likes163 downloads23d agoHugging Face02Anvesh-Lankala /Constrained_Indic_Codemixingtext1K<n<10K0 likes124 downloads2mo agoHugging Face03AmnaHassan /Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam Unity Code and GPT-Generated GDD Pairs Dataset This dataset contains paired samples of Unity game mechanic scripts and their corresponding GPT-4 generated Game Design Documents (GDDs). It is intended for training and benchmarking LLMs in game code generation from design specifications. Format Each entry is stored as a .jsonl file with: "input": GPT-4 generated GDD describing a specific game and its mechanics "output": Unity C# scripts implementing the described mechanic… See the full description on the dataset page: https://huggingface.co/datasets/AmnaHassan/Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam.textn<1K4 likes45 downloads1y agoHugging Face04Jnx03 /kanitakorn-deepseek-v47-v46-if-code-lcb-bridge-mix Kanitakorn v47 v46 + IF/Code/LCB Bridge Mix Conservative later-stage bridge mix for non-Thai-family <=14B candidates. Delta from v46: reuses all v46 rows unchanged, including inherited identity-attribution rows adds 197 de-duplicated local Codex-generated rows targeting Thai instruction following, code-output formatting, and short math/code verification excludes higher-risk dataset/train LiveCodeBench and public benchmark rows keeps single-model training only; no BoN… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v47-v46-if-code-lcb-bridge-mix.text1K<n<10K0 likes38 downloads4mo agoHugging Face05synquid /agentic-code-sft-mix-v1 Agentic Code SFT Mix v1 Local derived SFT mixture for code-agent/tool-use training. This is not a single upstream dataset. It is a filtered local mixture built from: nvidia/OpenCodeInstruct, split train nvidia/Nemotron-SFT-OpenCode-v1, splits general, bash_only_tool, bash_only_tool_skills, question_tool, agent_skills, agent_skills_question_tool nvidia/Nemotron-SFT-SWE-v2, split agentless nvidia/Nemotron-SFT-SWE-v2, file data/swe.jsonl The output schema is JSONL with messages… See the full description on the dataset page: https://huggingface.co/datasets/synquid/agentic-code-sft-mix-v1.texttext-generation10K<n<100K0 likes33 downloads4mo agoHugging Face06ar5entum /hindi-english-code-mixedThis dataset was compiled from various open sources online including some asr datasets and some percenteage of data generated using prompt engineering on generative llms. Some sources used are listed down below: https://github.com/l3cube-pune/code-mixed-nlp?tab=readme-ov-file https://github.com/piyushmakhija5/hinglishNorm https://github.com/ishan00/translation-for-code-switching-acl/tree/master text100K<n<1M0 likes31 downloads2y agoHugging Face07watchstep /ko-en-code-mixing-sts Korean–English Code-Mixing STS Dataset This dataset contains 1,500 Korean–English code-mixed pairs derived from KLUE-STS. We keep the original sentence_a and apply insertion-only code-mixing to sentence_b using an LLM (Gemini 2.5 Flash), recording where and how code-mixing occurred. Interactive Dashboard 🌐 Explore the dataset interactively: https://watchstep.github.io/ko-en-cm/ The dashboard provides: Interactive data exploration and filtering Sample visualization with… See the full description on the dataset page: https://huggingface.co/datasets/watchstep/ko-en-code-mixing-sts.tabularsentence-similarity1K<n<10K0 likes16 downloads1y agoHugging Face08hamishivi /code_rlvr_mixture_sfttabular10K<n<100K0 likes8 downloads1y agoHugging Face09ontocord /MixtureVitae-pseudo_code_of_thoughtgatedThis is a WIP dataset to augment https://huggingface.co/datasets/nampdn-ai/tiny-codes to create a psuedo-code of thought for reasoning about common sense situations. There are bugs(!) in the translation process which we are in the process of fixing. Use at your discretion. text100K<n<1M0 likes6 downloads1y agoHugging Face10hamishivi /code_rlvr_mixture_dpotabular10K<n<100K0 likes6 downloads1y agoHugging Face11cs23s036 /native_script_codemixedgatedtext1M<n<10M0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.