ruby2022/NYCU-IAlI-DL2026-LLM3-GRPO
020
NYCU IAI-DL 2026 LLM3 GRPO โ Qwen3-14B LoRA Adapter
This repository contains:
- LoRA adapter fine-tuned from
Qwen/Qwen3-14Bvia GRPO reinforcement learning - Training data (
train-reasoning-v2.csv) used for GRPO training
๐ Results
๐ Competition
Kaggle: nycu-i-al-i-dl-2026-llm-3-grpo
Task: Fine-tune a Chinese-released LLM using GRPO so that it correctly answers Traditional Chinese single-choice questions (A/B/C/D).
๐ง LoRA Configuration
๐ Files
๐ Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
import torch
base_model_name = "Qwen/Qwen3-14B"
adapter_name = "ruby2022/NYCU-IAlI-DL2026-LLM3-GRPO"
tokenizer = AutoTokenizer.from_pretrained(adapter_name)
model = AutoModelForCausalLM.from_pretrained(
base_model_name,
dtype=torch.bfloat16,
device_map="auto"
)
model = PeftModel.from_pretrained(model, adapter_name)
model = model.merge_and_unload()๐๏ธ Training Details
- Algorithm: GRPO (Group Relative Policy Optimization)
- Key innovation: Option Shuffling augmentation to prevent reward collapse
- Reward functions:
correctness_reward(+2.0/โ1.0) +format_reward(+0.5) - Inference: Majority Voting (N=8, temperature=0.6)
- Thinking mode: Qwen3 native
<think>...</think>enabled throughout
See the GitHub repository for full source code, training scripts, and logs.
โ๏ธ Hardware
- GPU: NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM)
- Training: Single GPU, ~500 steps GRPO
