ik_llama_opt/src
Samuel Oliveira Alves 09a88c9ae5
Add MTP decoding support for GLM-4.x MoE (#1270)
* wip: port MTP architecture

Ports the Multi-Token Prediction (MTP) architecture to the older `llama.cpp` codebase used by `ikllama`.

Changes include:
- Updating `llama_batch` to support `mtp_params`.
- Modifying `llama_decode_internal` (and `encode`) to handle MTP operations (Warmup, Update, Draft).
- Adding public APIs for MTP state management (`llama_set_draft_input_hidden_state`).
- Adapting the embedding extraction logic to skip MTP update passes.

* Refactors `server_slot` to support generic speculative decoding (MTP or Draft Model).

* core: enable hybrid outputs (logits + embeddings) for MTP support

* fix(mtp): correct KV-cache slot finding for updates

* fix(mtp): persist hidden states to prevent context corruption during drafting

* refactor(mtp): clean unused code

* fix(mtp): update server to new functions name

* fix(mtp): fix graph and save hidden state

* mtp: refactor integration, context params and kv cache search

* mtp: fix hidden state extraction and speculative acceptance flow

* server: fix MTP warmup for long prompts and reset token buffer

* llama: refactor MTP operation state to context parameters

* server: fix n_past calculation in MTP acceptance

* llama: fix mtp enable flags

* speculative: refactor MTP to use common_speculative interface

* context: remove unused signatures

* clip: fix deprecated enum-enum conversion warning

* common: fix format string crash in help message

* context: fix mtp activation logic
2026-02-22 18:14:39 +01:00
..
CMakeLists.txt Factor out delta net (#1286) 2026-02-18 17:16:17 +01:00
llama-arch.cpp Qwen3.5-MoE: fix regenerating message error (#1295) 2026-02-21 18:24:12 +01:00
llama-arch.h Qwen3.5-MoE: fix regenerating message error (#1295) 2026-02-21 18:24:12 +01:00
llama-build-context.cpp Add MTP decoding support for GLM-4.x MoE (#1270) 2026-02-22 18:14:39 +01:00
llama-build-context.h Add MTP decoding support for GLM-4.x MoE (#1270) 2026-02-22 18:14:39 +01:00
llama-context.h Add MTP decoding support for GLM-4.x MoE (#1270) 2026-02-22 18:14:39 +01:00
llama-cparams.h Add MTP decoding support for GLM-4.x MoE (#1270) 2026-02-22 18:14:39 +01:00
llama-delta-net.cpp Qwen3.5-MoE support (#1288) 2026-02-21 08:33:06 +01:00
llama-delta-net.h Factor out delta net (#1286) 2026-02-18 17:16:17 +01:00
llama-grammar.cpp llama : add token matching support to llama-grammar (#1220) 2026-02-03 07:57:17 +02:00
llama-grammar.h llama : add token matching support to llama-grammar (#1220) 2026-02-03 07:57:17 +02:00
llama-hparams.cpp Add MTP decoding support for GLM-4.x MoE (#1270) 2026-02-22 18:14:39 +01:00
llama-hparams.h WIP: Qwen3Next (#1266) 2026-02-16 06:50:28 +01:00
llama-impl.h server: stop processing the prompt when client disconnects (#1134) 2026-01-13 07:56:59 +02:00
llama-load-tensors.cpp Add MTP decoding support for GLM-4.x MoE (#1270) 2026-02-22 18:14:39 +01:00
llama-mmap.cpp Enable CUDA graphs for MoE models + GPT-OSS support (#689) 2025-08-15 09:18:07 +03:00
llama-mmap.h Enable CUDA graphs for MoE models + GPT-OSS support (#689) 2025-08-15 09:18:07 +03:00
llama-model-loader.cpp Be able to read uint32_t and bool arrays from GGUFs (#1252) 2026-02-07 19:20:15 +02:00
llama-model-loader.h Merge ffn_up and ffn_gate experts tensors (#1137) 2026-01-12 18:30:53 +02:00
llama-model.cpp Qwen3.5-MoE: fix regenerating message error (#1295) 2026-02-21 18:24:12 +01:00
llama-model.h Add MTP decoding support for GLM-4.x MoE (#1270) 2026-02-22 18:14:39 +01:00
llama-quantize.cpp Fix very low bpw missing imatrix check (#1284) 2026-02-19 08:15:26 +01:00
llama-sampling.cpp Fix adaptive p sampler bug with string ban (#1287) 2026-02-20 07:11:36 +01:00
llama-sampling.h Fix adaptive p sampler bug with string ban (#1287) 2026-02-20 07:11:36 +01:00
llama-vocab.cpp Qwen3.5-MoE support (#1288) 2026-02-21 08:33:06 +01:00
llama-vocab.h Qwen3.5-MoE support (#1288) 2026-02-21 08:33:06 +01:00
llama.cpp Add MTP decoding support for GLM-4.x MoE (#1270) 2026-02-22 18:14:39 +01:00
unicode-data.cpp Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
unicode-data.h Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
unicode.cpp Server: refactor and rename functions (#1151) 2026-01-18 08:16:57 +02:00
unicode.h Enable CUDA graphs for MoE models + GPT-OSS support (#689) 2025-08-15 09:18:07 +03:00