Team Ai
Modelpublic

GetSoloTech/Gemma3-Code-Reasoning-4B-GGUF

sourceHugging Faceupdated 1y agoView on Hugging Face
6likes402downloads
Model Card

Gemma3-Code-Reasoning-4B-GGUF

This repository contains GGUF (GGML Universal Format) quantized versions of the GetSoloTech/Gemma3-Code-Reasoning-4B model, optimized for local inference with various quantization levels to balance performance and resource usage.

๐ŸŽฏ Model Overview

This is a LoRA-finetuned version of gemma-3-4b-it specifically optimized for competitive programming and code reasoning tasks. The model has been trained on the high-quality Code-Reasoning dataset to enhance its capabilities in solving complex programming problems with detailed reasoning.

๐Ÿš€ Key Features

  • โ€”Enhanced Code Reasoning: Specifically trained on competitive programming problems
  • โ€”Thinking Capabilities: Inherits the advanced reasoning capabilities from the base model
  • โ€”High-Quality Solutions: Trained on solutions with โ‰ฅ85% test case pass rates
  • โ€”Structured Output: Optimized for generating well-reasoned programming solutions
  • โ€”Efficient Training: Uses LoRA adapters for efficient parameter updates
  • โ€”Multiple Quantization Levels: Available in various GGUF formats for different hardware capabilities

๐Ÿ“ Available GGUF Models

Model FileSizeQuantizationUse Case
Gemma3-Code-Reasoning-4B.f16.gguf7.77 GBFP16Highest quality, requires more VRAM
Gemma3-Code-Reasoning-4B.Q8_0.gguf4.13 GBQ8_0High quality, good balance
Gemma3-Code-Reasoning-4B.Q6_K.gguf3.19 GBQ6_KGood quality, moderate VRAM usage
Gemma3-Code-Reasoning-4B.Q5_K_M.gguf2.83 GBQ5KMBalanced quality and size
Gemma3-Code-Reasoning-4B.Q4_K_M.gguf2.49 GBQ4KMGood compression, reasonable quality
Gemma3-Code-Reasoning-4B.Q3_K_M.gguf2.1 GBQ3KMSmaller size, moderate quality
Gemma3-Code-Reasoning-4B.Q2_K.gguf1.73 GBQ2_KSmallest size, basic quality
Gemma3-Code-Reasoning-4B.IQ4_XS.gguf2.28 GBIQ4_XSIntel optimized, good quality

๐Ÿ”ง Usage

Using with llama.cpp

bash
# Download a GGUF model file
wget https://huggingface.co/GetSoloTech/Gemma3-Code-Reasoning-4B-GGUF/resolve/main/Gemma3-Code-Reasoning-4B.Q4_K_M.gguf

# Run inference with llama.cpp
./llama.cpp/main -m Gemma3-Code-Reasoning-4B.Q4_K_M.gguf -n 4096 --repeat_penalty 1.1 -p "You are an expert competitive programmer. Solve this problem: [YOUR_PROBLEM_HERE]"

Using with Python (llama-cpp-python)

python
from llama_cpp import Llama

# Load the model
llm = Llama(
    model_path="./Gemma3-Code-Reasoning-4B.Q4_K_M.gguf",
    n_ctx=4096,
    n_threads=4
)

# Prepare the prompt
prompt = """You are an expert competitive programmer. Read the problem and produce a correct, efficient solution. Include reasoning if helpful.

Problem: [YOUR_PROGRAMMING_PROBLEM_HERE]

Solution:"""

# Generate response
output = llm(
    prompt,
    max_tokens=4096,
    temperature=1.0,
    top_p=0.95,
    top_k=64,
    repeat_penalty=1.1
)

print(output['choices'][0]['text'])

๐ŸŽ›๏ธ Recommended Settings

  • โ€”Temperature: 1.0
  • โ€”Top-p: 0.95
  • โ€”Top-k: 64
  • โ€”Max New Tokens: 4096 (adjust based on problem complexity)
  • โ€”Repeat Penalty: 1.1

๐Ÿ’ป Hardware Requirements

QuantizationMinimum VRAMRecommended VRAMCPU RAM
FP168 GB12 GB16 GB
Q8_05 GB8 GB12 GB
Q6_K4 GB6 GB10 GB
Q5KM3 GB5 GB8 GB
Q4KM3 GB4 GB6 GB
Q3KM2 GB3 GB4 GB
Q2_K2 GB2 GB3 GB
IQ4_XS3 GB4 GB6 GB

๐Ÿ“ˆ Performance Expectations

This finetuned model is expected to show improved performance on:

  • โ€”Competitive Programming Problems: Better understanding of problem constraints and requirements
  • โ€”Code Generation: More accurate and efficient solutions
  • โ€”Reasoning Quality: Enhanced step-by-step reasoning for complex problems
  • โ€”Solution Completeness: More comprehensive solutions with proper edge case handling

๐Ÿ”— Related Resources

๐Ÿค Contributing

This model was created using the Unsloth framework and the Code-Reasoning dataset. For questions about:

๐Ÿ™ Acknowledgments

  • โ€”Gemma Team for the excellent base model
  • โ€”Unsloth Team for the efficient training framework
  • โ€”NVIDIA Research for the original OpenCodeReasoning-2 dataset
  • โ€”llama.cpp community for the GGUF format and tools

๐Ÿ“ž Contact

For questions about this GGUF converted model, please open an issue in the repository.


Note: This model is specifically optimized for competitive programming and code reasoning tasks. Choose the appropriate quantization level based on your hardware capabilities and quality requirements.