VIRUS374/ai-developer-dataset
AI Developer Dataset A large-scale instruction-tuning dataset for fine-tuning an open-weight LLM into a universal AI developer assistant. The model trained on this dataset should be especially good at: PROGRAMMING + WEB DEVELOPMENT + UI/UX + ANIMATIONS + BOTS + SCRIPTS + AUTOMATION + BACKEND + API + DATABASES + DEBUGGING + LINUX + DEPLOYMENT + AI DEVELOPMENT + SECURITY. ๐ Statistics Total examples: 1,124,699 File size: 2.13 GB Format: JSONL (conversational)โฆ See the full description on the dataset page: https://huggingface.co/datasets/VIRUS374/ai-developer-dataset.
AI Developer Dataset
A large-scale instruction-tuning dataset for fine-tuning an open-weight LLM into a universal AI developer assistant.
The model trained on this dataset should be especially good at: PROGRAMMING + WEB DEVELOPMENT + UI/UX + ANIMATIONS + BOTS + SCRIPTS + AUTOMATION + BACKEND + API + DATABASES + DEBUGGING + LINUX + DEPLOYMENT + AI DEVELOPMENT + SECURITY.
๐ Statistics
- Total examples: 1,124,699
- File size: 2.13 GB
- Format: JSONL (conversational)
- Average assistant length: ~1,222 characters
- Schema:
{"messages":[{"role":"system|user|assistant","content":"..."}, ...], "metadata":{"category","subcategory","complexity","source"}}
๐ Sources
This dataset is a curated merge of high-quality open-source code instruction datasets from HuggingFace, plus handcrafted examples in Russian covering UI/UX, security, and other specialized topics.
๐ Categories
๐ฏ Complexity Distribution
๐ Schema
Each line is a single JSON object:
{
"messages": [
{"role": "system", "content": "Category: programming\nSubcategory: code_alpaca\n..."},
{"role": "user", "content": "Write a Python function to..."},
{"role": "assistant", "content": "Here is the implementation..."}
],
"metadata": {
"category": "programming",
"subcategory": "code_alpaca",
"complexity": "intermediate",
"source": "hf:HuggingFaceH4/CodeAlpaca_20K"
}
}๐ ๏ธ Quality Control
Every example passes:
- JSON-serializable (validated on 100% of examples)
- Non-empty user and assistant content
- Assistant length between 200 and 25,000 characters
- Per-source deduplication by user-content fingerprint (SHA-256, first 16 hex chars)
- Schema validation (messages + metadata required)
๐ Usage
from datasets import load_dataset
# Load the full dataset
ds = load_dataset("VIRUS374/ai-developer-dataset", data_files="dataset.jsonl", split="train")
print(ds)
# Dataset({
# features: ["messages", "metadata"],
# num_rows: 1,124,699
# })
# Or use streaming for memory-efficient iteration
ds = load_dataset("VIRUS374/ai-developer-dataset", data_files="dataset.jsonl", split="train", streaming=True)
for example in ds:
messages = example["messages"]
# ... feed to tokenizer with chat template๐๏ธ Fine-tuning Recipe (QLoRA)
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
from trl import SFTTrainer
from datasets import load_dataset
import torch
MODEL = "meta-llama/Llama-3.1-8B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(MODEL)
tokenizer.pad_token = tokenizer.eos_token
model = AutoModelForCausalLM.from_pretrained(
MODEL,
torch_dtype=torch.bfloat16,
load_in_4bit=True,
device_map="auto",
)
model = prepare_model_for_kbit_training(model)
lora_config = LoraConfig(
r=16, lora_alpha=32,
target_modules=["q_proj", "v_proj", "k_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
lora_dropout=0.05, bias="none", task_type="CAUSAL_LM",
)
model = get_peft_model(model, lora_config)
ds = load_dataset("VIRUS374/ai-developer-dataset", data_files="dataset.jsonl", split="train")
def format_example(example):
text = tokenizer.apply_chat_template(
example["messages"], tokenize=False, add_generation_prompt=False,
)
return {"text": text}
ds = ds.map(format_example, remove_columns=ds.column_names)
trainer = SFTTrainer(
model=model,
train_dataset=ds,
args=TrainingArguments(
output_dir="./output",
num_train_epochs=3,
per_device_train_batch_size=2,
gradient_accumulation_steps=8,
warmup_steps=50,
learning_rate=2e-4,
bf16=True,
logging_steps=10,
save_steps=500,
),
max_seq_length=4096,
)
trainer.train()๐ License & Attribution
This dataset aggregates multiple open-source datasets. Each retains its original license:
- HuggingFaceH4/CodeAlpaca_20K โ Apache 2.0
- HuggingFaceH4/Code-Feedback โ Apache 2.0
- nickrosh/Evol-Instruct-Code-80k-v1 โ Apache 2.0
- WizardLMTeam/WizardLM_evol_instruct_70k โ Apache 2.0 (Non-commercial)
- TokenBender/code_instructions_122k_alpaca_style โ Apache 2.0
- iamtarun/code_instructions_120k_alpaca โ Apache 2.0
- m-a-p/Code-Feedback โ CC BY-NC 4.0
- Locutusque/code-feedback-sharegpt โ Apache 2.0
- ajibawa-2023/Code-74k-ShareGPT โ Apache 2.0
- ajibawa-2023/Python-Code-23k-ShareGPT โ Apache 2.0
- b-mc2/sql-create-context โ CC BY 4.0
- glaiveai/glaive-code-assistant โ Apache 2.0
- glaiveai/glaive-code-assistant-v2 โ Apache 2.0
- mwitiderrick/glaive-code-assistant โ Apache 2.0
- Data-Agora/magpie_code_gpt4o_mini_50000 โ Apache 2.0
- Handcrafted examples (VIRUS374) โ MIT
Important: If you use this dataset for commercial purposes, please review the WizardLM and m-a-p licenses โ they may have non-commercial restrictions.
๐ Acknowledgements
Thanks to all the original dataset creators:
- HuggingFace H4 team
- WizardLM team
- TokenBender
- iamtarun
- m-a-p
- Locutusque
- ajibawa-2023
- b-mc2
- glaiveai
- mwitiderrick
- Data-Agora
๐ Citation
@dataset{virus374_ai_developer_2024,
title = {AI Developer Dataset},
author = {VIRUS374},
year = {2024},
url = {https://huggingface.co/datasets/VIRUS374/ai-developer-dataset},
note = {Aggregated from open-source code instruction datasets + handcrafted examples}
}