Team Ai
Datasetpublic

VIRUS374/ai-developer-dataset

AI Developer Dataset A large-scale instruction-tuning dataset for fine-tuning an open-weight LLM into a universal AI developer assistant. The model trained on this dataset should be especially good at: PROGRAMMING + WEB DEVELOPMENT + UI/UX + ANIMATIONS + BOTS + SCRIPTS + AUTOMATION + BACKEND + API + DATABASES + DEBUGGING + LINUX + DEPLOYMENT + AI DEVELOPMENT + SECURITY. ๐Ÿ“Š Statistics Total examples: 1,124,699 File size: 2.13 GB Format: JSONL (conversational)โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/VIRUS374/ai-developer-dataset.

sourceHugging Facemitupdated 5d agoView on Hugging Face
0likes52downloads
Dataset Card

AI Developer Dataset

A large-scale instruction-tuning dataset for fine-tuning an open-weight LLM into a universal AI developer assistant.

The model trained on this dataset should be especially good at: PROGRAMMING + WEB DEVELOPMENT + UI/UX + ANIMATIONS + BOTS + SCRIPTS + AUTOMATION + BACKEND + API + DATABASES + DEBUGGING + LINUX + DEPLOYMENT + AI DEVELOPMENT + SECURITY.

๐Ÿ“Š Statistics

  • โ€”Total examples: 1,124,699
  • โ€”File size: 2.13 GB
  • โ€”Format: JSONL (conversational)
  • โ€”Average assistant length: ~1,222 characters
  • โ€”Schema: {"messages":[{"role":"system|user|assistant","content":"..."}, ...], "metadata":{"category","subcategory","complexity","source"}}

๐Ÿ“š Sources

This dataset is a curated merge of high-quality open-source code instruction datasets from HuggingFace, plus handcrafted examples in Russian covering UI/UX, security, and other specialized topics.

SourceExamples
hf:glaiveai/glaive-code-assistant-v2213,670
hf:glaiveai/glaive-code-assistant135,195
hf:mwitiderrick/glaive-code-assistant104,336
hf:b-mc2/sql-create-context78,573
hf:TokenBender/code_instructions_122k_alpaca_style78,433
hf:iamtarun/code_instructions_120k_alpaca78,433
hf:ajibawa-2023/Code-74k-ShareGPT68,790
hf:Locutusque/code-feedback-sharegpt62,866
hf:m-a-p/Code-Feedback62,866
hf:HuggingFaceH4/Code-Feedback61,923
hf:WizardLMTeam/WizardLM_evol_instruct_70k60,411
hf:nickrosh/Evol-Instruct-Code-80k-v158,265
hf:Data-Agora/magpie_code_gpt4o_mini_5000036,818
hf:ajibawa-2023/Python-Code-23k-ShareGPT17,737
hf:HuggingFaceH4/CodeAlpaca_20K6,307
handcrafted76

๐Ÿ“ Categories

CategoryExamples%
programming1,046,06093.0%
backendapidb78,5797.0%
uiuxdesign100.0%
web_development90.0%
bots_automation80.0%
security80.0%
ai_ml70.0%
debugging60.0%
linuxserverdeploy60.0%
websearchdocs60.0%

๐ŸŽฏ Complexity Distribution

LevelExamples%
intermediate686,51961.0%
advanced438,16639.0%
beginner120.0%
expert20.0%

๐Ÿ“‹ Schema

Each line is a single JSON object:

json
{
  "messages": [
    {"role": "system", "content": "Category: programming\nSubcategory: code_alpaca\n..."},
    {"role": "user", "content": "Write a Python function to..."},
    {"role": "assistant", "content": "Here is the implementation..."}
  ],
  "metadata": {
    "category": "programming",
    "subcategory": "code_alpaca",
    "complexity": "intermediate",
    "source": "hf:HuggingFaceH4/CodeAlpaca_20K"
  }
}

๐Ÿ› ๏ธ Quality Control

Every example passes:

  • โ€”JSON-serializable (validated on 100% of examples)
  • โ€”Non-empty user and assistant content
  • โ€”Assistant length between 200 and 25,000 characters
  • โ€”Per-source deduplication by user-content fingerprint (SHA-256, first 16 hex chars)
  • โ€”Schema validation (messages + metadata required)

๐Ÿš€ Usage

python
from datasets import load_dataset

# Load the full dataset
ds = load_dataset("VIRUS374/ai-developer-dataset", data_files="dataset.jsonl", split="train")
print(ds)
# Dataset({
#     features: ["messages", "metadata"],
#     num_rows: 1,124,699
# })

# Or use streaming for memory-efficient iteration
ds = load_dataset("VIRUS374/ai-developer-dataset", data_files="dataset.jsonl", split="train", streaming=True)
for example in ds:
    messages = example["messages"]
    # ... feed to tokenizer with chat template

๐Ÿ‹๏ธ Fine-tuning Recipe (QLoRA)

python
from transformers import AutoModelForCausalLM, AutoTokenizer, TrainingArguments
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
from trl import SFTTrainer
from datasets import load_dataset
import torch

MODEL = "meta-llama/Llama-3.1-8B-Instruct"

tokenizer = AutoTokenizer.from_pretrained(MODEL)
tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(
    MODEL,
    torch_dtype=torch.bfloat16,
    load_in_4bit=True,
    device_map="auto",
)
model = prepare_model_for_kbit_training(model)

lora_config = LoraConfig(
    r=16, lora_alpha=32,
    target_modules=["q_proj", "v_proj", "k_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
    lora_dropout=0.05, bias="none", task_type="CAUSAL_LM",
)
model = get_peft_model(model, lora_config)

ds = load_dataset("VIRUS374/ai-developer-dataset", data_files="dataset.jsonl", split="train")

def format_example(example):
    text = tokenizer.apply_chat_template(
        example["messages"], tokenize=False, add_generation_prompt=False,
    )
    return {"text": text}

ds = ds.map(format_example, remove_columns=ds.column_names)

trainer = SFTTrainer(
    model=model,
    train_dataset=ds,
    args=TrainingArguments(
        output_dir="./output",
        num_train_epochs=3,
        per_device_train_batch_size=2,
        gradient_accumulation_steps=8,
        warmup_steps=50,
        learning_rate=2e-4,
        bf16=True,
        logging_steps=10,
        save_steps=500,
    ),
    max_seq_length=4096,
)
trainer.train()

๐Ÿ“œ License & Attribution

This dataset aggregates multiple open-source datasets. Each retains its original license:

  • โ€”HuggingFaceH4/CodeAlpaca_20K โ€” Apache 2.0
  • โ€”HuggingFaceH4/Code-Feedback โ€” Apache 2.0
  • โ€”nickrosh/Evol-Instruct-Code-80k-v1 โ€” Apache 2.0
  • โ€”WizardLMTeam/WizardLM_evol_instruct_70k โ€” Apache 2.0 (Non-commercial)
  • โ€”TokenBender/code_instructions_122k_alpaca_style โ€” Apache 2.0
  • โ€”iamtarun/code_instructions_120k_alpaca โ€” Apache 2.0
  • โ€”m-a-p/Code-Feedback โ€” CC BY-NC 4.0
  • โ€”Locutusque/code-feedback-sharegpt โ€” Apache 2.0
  • โ€”ajibawa-2023/Code-74k-ShareGPT โ€” Apache 2.0
  • โ€”ajibawa-2023/Python-Code-23k-ShareGPT โ€” Apache 2.0
  • โ€”b-mc2/sql-create-context โ€” CC BY 4.0
  • โ€”glaiveai/glaive-code-assistant โ€” Apache 2.0
  • โ€”glaiveai/glaive-code-assistant-v2 โ€” Apache 2.0
  • โ€”mwitiderrick/glaive-code-assistant โ€” Apache 2.0
  • โ€”Data-Agora/magpie_code_gpt4o_mini_50000 โ€” Apache 2.0
  • โ€”Handcrafted examples (VIRUS374) โ€” MIT

Important: If you use this dataset for commercial purposes, please review the WizardLM and m-a-p licenses โ€” they may have non-commercial restrictions.

๐Ÿ™ Acknowledgements

Thanks to all the original dataset creators:

  • โ€”HuggingFace H4 team
  • โ€”WizardLM team
  • โ€”TokenBender
  • โ€”iamtarun
  • โ€”m-a-p
  • โ€”Locutusque
  • โ€”ajibawa-2023
  • โ€”b-mc2
  • โ€”glaiveai
  • โ€”mwitiderrick
  • โ€”Data-Agora

๐Ÿ“ Citation

bibtex
@dataset{virus374_ai_developer_2024,
  title  = {AI Developer Dataset},
  author = {VIRUS374},
  year   = {2024},
  url    = {https://huggingface.co/datasets/VIRUS374/ai-developer-dataset},
  note   = {Aggregated from open-source code instruction datasets + handcrafted examples}
}