KBaba7/llama.cpp
0
1# llama.cpp/example/tts2This example demonstrates the Text To Speech feature. It uses a3[model](https://www.outeai.com/blog/outetts-0.2-500m) from4[outeai](https://www.outeai.com/).5 6## Quickstart7If you have built llama.cpp with `-DLLAMA_CURL=ON` you can simply run the8following command and the required models will be downloaded automatically:9```console10$ build/bin/llama-tts --tts-oute-default -p "Hello world" && aplay output.wav11```12For details about the models and how to convert them to the required format13see the following sections.14 15### Model conversion16Checkout or download the model that contains the LLM model:17```console18$ pushd models19$ git clone --branch main --single-branch --depth 1 https://huggingface.co/OuteAI/OuteTTS-0.2-500M20$ cd OuteTTS-0.2-500M && git lfs install && git lfs pull21$ popd22```23Convert the model to .gguf format:24```console25(venv) python convert_hf_to_gguf.py models/OuteTTS-0.2-500M \26 --outfile models/outetts-0.2-0.5B-f16.gguf --outtype f1627```28The generated model will be `models/outetts-0.2-0.5B-f16.gguf`.29 30We can optionally quantize this to Q8_0 using the following command:31```console32$ build/bin/llama-quantize models/outetts-0.2-0.5B-f16.gguf \33 models/outetts-0.2-0.5B-q8_0.gguf q8_034```35The quantized model will be `models/outetts-0.2-0.5B-q8_0.gguf`.36 37Next we do something simlar for the audio decoder. First download or checkout38the model for the voice decoder:39```console40$ pushd models41$ git clone --branch main --single-branch --depth 1 https://huggingface.co/novateur/WavTokenizer-large-speech-75token42$ cd WavTokenizer-large-speech-75token && git lfs install && git lfs pull43$ popd44```45This model file is PyTorch checkpoint (.ckpt) and we first need to convert it to46huggingface format:47```console48(venv) python examples/tts/convert_pt_to_hf.py \49 models/WavTokenizer-large-speech-75token/wavtokenizer_large_speech_320_24k.ckpt50...51Model has been successfully converted and saved to models/WavTokenizer-large-speech-75token/model.safetensors52Metadata has been saved to models/WavTokenizer-large-speech-75token/index.json53Config has been saved to models/WavTokenizer-large-speech-75tokenconfig.json54```55Then we can convert the huggingface format to gguf:56```console57(venv) python convert_hf_to_gguf.py models/WavTokenizer-large-speech-75token \58 --outfile models/wavtokenizer-large-75-f16.gguf --outtype f1659...60INFO:hf-to-gguf:Model successfully exported to models/wavtokenizer-large-75-f16.gguf61```62 63### Running the example64 65With both of the models generated, the LLM model and the voice decoder model,66we can run the example:67```console68$ build/bin/llama-tts -m ./models/outetts-0.2-0.5B-q8_0.gguf \69 -mv ./models/wavtokenizer-large-75-f16.gguf \70 -p "Hello world"71...72main: audio written to file 'output.wav'73```74The output.wav file will contain the audio of the prompt. This can be heard75by playing the file with a media player. On Linux the following command will76play the audio:77```console78$ aplay output.wav79```80 81### Running the example with llama-server82Running this example with `llama-server` is also possible and requires two83server instances to be started. One will serve the LLM model and the other84will serve the voice decoder model.85 86The LLM model server can be started with the following command:87```console88$ ./build/bin/llama-server -m ./models/outetts-0.2-0.5B-q8_0.gguf --port 802089```90 91And the voice decoder model server can be started using:92```console93./build/bin/llama-server -m ./models/wavtokenizer-large-75-f16.gguf --port 8021 --embeddings --pooling none94```95 96Then we can run [tts-outetts.py](tts-outetts.py) to generate the audio.97 98First create a virtual environment for python and install the required99dependencies (this in only required to be done once):100```console101$ python3 -m venv venv102$ source venv/bin/activate103(venv) pip install requests numpy104```105 106And then run the python script using:107```conole108(venv) python ./examples/tts/tts-outetts.py http://localhost:8020 http://localhost:8021 "Hello world"109spectrogram generated: n_codes: 90, n_embd: 1282110converting to audio ...111audio generated: 28800 samples112audio written to file "output.wav"113```114And to play the audio we can again use aplay or any other media player:115```console116$ aplay output.wav117```118 