Team Ai
Modelpublic

x0root/qwen2-7b-orca-math-lora

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes32downloads
Model Card

qwen2-7b-orca-math-lora

A LoRA fine-tune of Qwen2-7B-Instruct trained with supervised fine-tuning on a curated blend of mathematical reasoning and general instruction-following data. Training was performed using Unsloth for memory-efficient adaptation on a single GPU.


Model Details

PropertyValue
Base modelQwen/Qwen2-7B-Instruct
Model familyQwen2
Parameter count7B
Fine-tuning methodLoRA (PEFT)
Quantization (training)4-bit NormalFloat (bitsandbytes)
Chat templateChatML
Context length2048 tokens
LanguageEnglish
LicenseApache 2.0

Training Details

LoRA Configuration

HyperparameterValue
Rank (r)8
Alpha8
Dropout0
Biasnone
RSLoRATrue
Target modulesqproj, kproj, vproj, oproj, gateproj, upproj, down_proj
Gradient checkpointingUnsloth (memory-optimized)

Training Hyperparameters

HyperparameterValue
TrainerSFTTrainer (TRL)
Max steps300
Per-device batch size1
Gradient accumulation steps8
Effective batch size8
Learning rate2e-4
LR schedulerLinear
OptimizerAdamW (8-bit)
Weight decay0.01
Precisionbf16 (fp16 fallback if bf16 unavailable)
PackingFalse
Training objectiveResponses only (instruction tokens masked)
Seed3407

Training Data

The model was trained on a concatenated and shuffled mixture of three datasets (seed 3407):

DatasetSplitSamples
openai/gsm8ktrain (full)~7,473
microsoft/orca-math-word-problems-200ktrain4,000
HuggingFaceH4/ultrachat_200ktrain_sft2,000

All examples were formatted using the ChatML conversation template before training. The loss was computed on assistant responses only; user turns and system prompts were excluded from the gradient.


Intended Use

This model is suited for tasks involving:

  • —Grade-school and competition-level math word problems
  • —Step-by-step arithmetic and algebraic reasoning
  • —General instruction following and question answering in English

It is not intended for safety-critical applications, factual knowledge retrieval, or domains outside its training distribution.


Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "x0root/qwen2-7b-orca-math-lora"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")

messages = [
    {"role": "user", "content": "A train travels 300 km in 4 hours. What is its average speed?"}
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

For faster inference with the original 4-bit quantized weights, load via Unsloth:

python
from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="x0root/qwen2-7b-orca-math-lora",
    max_seq_length=2048,
    load_in_4bit=True,
)
FastLanguageModel.for_inference(model)

Training Framework

ComponentVersion / Note
UnslothLatest at training time
TRL<= 0.24.0
Transformers<= 5.5.0
Datasets< 4.4.0
AccelerateLatest at training time
PEFTLatest at training time
bitsandbytesLatest at training time
HardwareSingle CUDA GPU