Team Ai
Modelpublic

Felipe97/llama-cpp-compiled

sourceHugging Faceupdated 21d agoView on Hugging Face
0likes1.2kdownloads
README.md20 linesDownload Raw Back to training
1# llama.cpp/examples/training2 3This directory contains examples related to language model training using llama.cpp/GGML.4So far finetuning is technically functional (for FP32 models and limited hardware setups) but the code is very much WIP.5Finetuning of Stories 260K and LLaMA 3.2 1b seems to work with 24 GB of memory.6**For CPU training, compile llama.cpp without any additional backends such as CUDA.**7**For CUDA training, use the maximum number of GPU layers.**8 9Flash attention is disabled during training because `FLASH_ATTN_EXT` has no backward pass.10 11Proof of concept:12 13``` sh14export model_name=llama_3.2-1b && export quantization=f3215./build/bin/llama-finetune --file wikitext-2-raw/wiki.test.raw -ngl 999 --model models/${model_name}-${quantization}.gguf -c 512 -b 512 -ub 51216./build/bin/llama-perplexity --file wikitext-2-raw/wiki.test.raw -ngl 999 --model finetuned-model.gguf17```18 19The perplexity value of the finetuned model should be lower after training on the test set for 2 epochs.20