Team Ai
Modelpublic

ruby2022/NYCU-IAlI-DL2026-LLM3-GRPO

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes20downloads
Model Card

NYCU IAI-DL 2026 LLM3 GRPO โ€” Qwen3-14B LoRA Adapter

This repository contains:

  • โ€”LoRA adapter fine-tuned from Qwen/Qwen3-14B via GRPO reinforcement learning
  • โ€”Training data (train-reasoning-v2.csv) used for GRPO training

๐Ÿ“Š Results

SubmissionInferencePublicPrivate
Qwen3-14B + GRPOMajority Vote N=80.6940.698
Qwen3-14B + GRPOGreedy (N=1)0.6870.685

๐Ÿ† Competition

Kaggle: nycu-i-al-i-dl-2026-llm-3-grpo

Task: Fine-tune a Chinese-released LLM using GRPO so that it correctly answers Traditional Chinese single-choice questions (A/B/C/D).

๐Ÿ”ง LoRA Configuration

ParameterValue
Base modelQwen/Qwen3-14B
Rank (r)16
Alpha16
Dropout0.0
Target modulesqproj, kproj, vproj, oproj, gateproj, upproj, down_proj
Trainable params~1.5%
Training methodGRPO (Group Relative Policy Optimization)

๐Ÿ“ Files

FileDescription
adapter_config.jsonLoRA adapter configuration
adapter_model.safetensorsLoRA adapter weights (~245 MB)
tokenizer.jsonTokenizer
tokenizer_config.jsonTokenizer configuration
chat_template.jinjaQwen3 chat template with thinking enabled
train-reasoning-v2.csvTraining dataset (3,199 samples)
kaggle_test_set_792.csvTest set (792 questions, no labels)

๐Ÿš€ Usage

python
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
import torch

base_model_name = "Qwen/Qwen3-14B"
adapter_name = "ruby2022/NYCU-IAlI-DL2026-LLM3-GRPO"

tokenizer = AutoTokenizer.from_pretrained(adapter_name)
model = AutoModelForCausalLM.from_pretrained(
    base_model_name,
    dtype=torch.bfloat16,
    device_map="auto"
)
model = PeftModel.from_pretrained(model, adapter_name)
model = model.merge_and_unload()

๐Ÿ—๏ธ Training Details

  • โ€”Algorithm: GRPO (Group Relative Policy Optimization)
  • โ€”Key innovation: Option Shuffling augmentation to prevent reward collapse
  • โ€”Reward functions: correctness_reward (+2.0/โˆ’1.0) + format_reward (+0.5)
  • โ€”Inference: Majority Voting (N=8, temperature=0.6)
  • โ€”Thinking mode: Qwen3 native <think>...</think> enabled throughout

See the GitHub repository for full source code, training scripts, and logs.

โš™๏ธ Hardware

  • โ€”GPU: NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM)
  • โ€”Training: Single GPU, ~500 steps GRPO