Brunobkr/llama.cpp_AlgMor24_github
ΩFFFΣLLIa • llama.cpp • AlgMor24 ██████╗ ███████╗███████╗███████╗██╗ ██╗ ██╗ █████╗ ██╔═══██╗██╔════╝██╔════╝██╔════╝██║ ██║ ██║██╔══██╗ ██║ ██║█████╗ █████╗ █████╗ ██║ ██║ ██║███████║ ██║ ██║██╔══╝ ██╔══╝ ██╔══╝ ██║ ██║ ██║██╔══██║ ╚██████╔╝██║ ██║ ███████╗███████╗███████╗██║██║ ██║ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝╚══════╝╚══════╝╚═╝╚═╝ ╚═╝ High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.
03.1k
1# libmtmd dev guide2 3## History4 5Please refer to [multimodal.md](../../docs/multimodal.md) for a broader context.6 7In short:8- `libmtmd` started as a wrapper around `libllava` / `clip.cpp`9- Various components that used to be in `clip.cpp` are moved progressively to mtmd. For example, preprocessor is now part of mtmd10 11## Terminologies12 13- mtmd: **M**ul**T**i**M**o**D**al14- bitmap: representing a raw input data, for example: RGB image, PCM audio15- tiles / slices: for llava-uhd-style models, the preprocessor breaks a large input into smaller square images called tiles or slices16- chunk: a mtmd_input_chunk represents a preprocessed input that can then be passed through `mtmd_encode()`17 18## Pipeline19 20A typical pipeline of the core libmtmd is as follows:21- A bitmap (RGB image or PCM audio) is created22- Bitmap and the text prompt is provided to `mtmd_tokenize()` that breaks the input into chunks23 - The tokenizer function first expands a "lazy" bitmap if it finds one. Typically, this is used by video, so that one media token corresponds to one input bitmap24 - For models that support "fused" temporal frames like Qwen-VL, the tokenizer tries to merge pair of consecutive frames into one batch25 - The preprocessor will then be called, which produces a list of chunks26 - Depending on the model itself, special tokens will be injected to separate image chunks (i.e. llava-uhd-style models)27- Multiple bitmaps may be batched together to form a larger `mtmd_batch()`28- Single image or batch is encoded, via `mtmd_encode()` or `mtmd_batch_encode()`29- Get the output embeddings30 31## Helper32 33We provide a set of helper functions via `mtmd_helper` to make using libmtmd easier. The helper provides:34- Image, audio and video file decoding (for example, decode raw JPEG into RGB bitmap)35- Manage `llama_batch` and calls to `llama_decode`36 37## Audio generation support38 39Audio generation is added to mtmd in PR [#26254](https://github.com/ggml-org/llama.cpp/pull/26254)40 41Currently, we support the 3-stage pipeline below which should cover most TTS models:42- Stage 1: Backbone / Semantic Stage: Backbone model accepts text prompt and reference voice as input43- Stage 2: Acoustic Detail Generator: A model takes the hidden state from backbone and generate audio details (usually as audio codes or mel-spectrogram)44- Stage 3: Waveform Reconstruction: Convert the semantic and acoustic data from previous stages to the final waveform45 46For example, Qwen3-TTS:47- Reference voice is encoded using ECAPA-TDNN speaker encoder (`speaker_encoder`)48- Text prompt and reference voice are processed via a backbone (`talker.model`)49- A model converts sampled semantic token and hidden state from stage 2 into a list of 15 acoustic codes (`talker.code_predictor`)50- 16 generated codes are converted into waveform (`code2wav`)51 52### API design constraints53 54Due to wide variety of audio generation pipelines, the `mtmd_gen_audio` system is designed to be flexible and reusable by new models.55 56`mtmd_gen_audio` is split into 2 main API:57- Core API `mtmd.h`: handles main inference. Important: the API surface must be stateless; caller must handle state management and audio frame accumulation.58- Helper API `mtmd-helper.h`: provides a model-agnostic stateful API. Usage example can be found in the `tools/tts` directory.59 60### Checklist for porting new audio generation models to mtmd61 621. Establish a list of reusable and missing components from the current mtmd implementation.632. For GGUF conversion:64 - Backbone model should be converted to a normal text model (loadable via `libllama`)65 - If model used hard-coded embedding row ID, append them to token embeddings and assign token name for them (see `qwen3tts.py`)66 - If model have a specific output logits head for audio codes (usually semantic code), keep the head as-is and pad the logits at inference time (see `src/models/qwen3vl.cpp`)67 - Sidecar models (code2wav, bigvgan, etc) must live inside the mmproj GGUF (but can be in different `clip_context` if necessary)68 - Note: it should use `ggml_build_forward_select` to select graphs if multiple graphs living in the same context69 - Reuse existing GGUF metadata key name and tensor name whenever possible; think twice before adding extensive changes to GGUF writer. For example, Qwen3-TTS hard-code part of the hparams to `clip.cpp` as they won't likely to change.70 - For tensor naming:71 - Prefixed with `a.*` for tensors used by speaker encoder pipeline72 - Prefixed with `a.gen.*` for generation stages (code / mel-spectrogram / PCM generation)733. Make sure most of the changes happen inside `mtmd-helper-gen.cpp`. A good PR looks like this:74 - 10-20% changes is to add new backbone (text) model and conversion75 - 60% changes inside `mtmd-helper-gen.cpp`76 - 10% changes inside `libmtmd` and `clip.cpp` systems77 - The rest downstream code (CLI, server) should have no changes at all784. Update usage documentation in `tools/tts/README.md`79 80IMPORTANT: If your model needs changes that don't fit the existing infrastructure, **open an issue first for discussion**.81 82No-go checklist (these will get the PR rejected and require discussion before proceeding):83- Violating the API design constraints stated above84- Adding a new model-specific binary: the API and binary surface must stay model-agnostic85 