Team Ai
Datasetpublic

Pinkstack/LuauDev-instructions-SFT-preview

LuauDev-SFT-PREVIEW THIS IS A PREVIEW VARIANT OF LUAUDEV. non preview: Pinkstack/LuauDev-instructions-SFT-full This is an SFT dataset meant for training Luau(Roblox's coding language) oriented large language models. Once the full version would be out it would be the biggest Luau instruction-style dataset ever released. These are the models which were used for data generation: (no specific order) DiffusionGemma 26B A4B Deepseek v4 Flash 0731 Nemotron 3 Ultra 550B A55B dots3 note… See the full description on the dataset page: https://huggingface.co/datasets/Pinkstack/LuauDev-instructions-SFT-preview.

sourceHugging Facemitupdated 14d agoView on Hugging Face
2likes160downloads
Dataset Card

LuauDev-SFT-PREVIEW

THIS IS A PREVIEW VARIANT OF LUAUDEV. non preview: Pinkstack/LuauDev-instructions-SFT-full

This is an SFT dataset meant for training Luau(Roblox's coding language) oriented large language models.

Once the full version would be out it would be the biggest Luau instruction-style dataset ever released.

These are the models which were used for data generation:

(no specific order)

  • —DiffusionGemma 26B A4B
  • —Deepseek v4 Flash 0731
  • —Nemotron 3 Ultra 550B A55B
  • —dots3 note prev
  • —GPT OSS 120b
  • —Muse Glimmer 30B
  • —GPT OSS 20b
  • —Ling 3.0 flash

And more...

each row has exactly 5 assistant and user turns. Making it to date, the biggest Luau post training dataset ever made publicly.

The dataset is organized like this:

(columns)

conversation```

Same exact format as other sharegpt datasets.

Compared to luaucoder instructions v3 dataset:

- Higher quality data overall
- Better system prompt
- Real debugging samples⁕
- Significantly improved user agent

⁕ The user agent did detect when the assistant generated mistakes, and did tell the assistant to fix those mistakes. (eg; bugs)