Team Ai
Apppublic

KBaba7/llama.cpp

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes
README.md51 linesDownload Raw Back to run
1# llama.cpp/example/run2 3The purpose of this example is to demonstrate a minimal usage of llama.cpp for running models.4 5```bash6llama-run granite3-moe7```8 9```bash10Description:11  Runs a llm12 13Usage:14  llama-run [options] model [prompt]15 16Options:17  -c, --context-size <value>18      Context size (default: 2048)19  -n, -ngl, --ngl <value>20      Number of GPU layers (default: 0)21  --temp <value>22      Temperature (default: 0.8)23  -v, --verbose, --log-verbose24      Set verbosity level to infinity (i.e. log all messages, useful for debugging)25  -h, --help26      Show help message27 28Commands:29  model30      Model is a string with an optional prefix of31      huggingface:// (hf://), ollama://, https:// or file://.32      If no protocol is specified and a file exists in the specified33      path, file:// is assumed, otherwise if a file does not exist in34      the specified path, ollama:// is assumed. Models that are being35      pulled are downloaded with .partial extension while being36      downloaded and then renamed as the file without the .partial37      extension when complete.38 39Examples:40  llama-run llama341  llama-run ollama://granite-code42  llama-run ollama://smollm:135m43  llama-run hf://QuantFactory/SmolLM-135M-GGUF/SmolLM-135M.Q2_K.gguf44  llama-run huggingface://bartowski/SmolLM-1.7B-Instruct-v0.2-GGUF/SmolLM-1.7B-Instruct-v0.2-IQ3_M.gguf45  llama-run https://example.com/some-file1.gguf46  llama-run some-file2.gguf47  llama-run file://some-file3.gguf48  llama-run --ngl 999 some-file4.gguf49  llama-run --ngl 999 some-file5.gguf Hello World50```51