ik_llama_opt/docs
Samuel Oliveira Alves f4f4b3ff26
Allow dual speculative decoding (#1789)
* wip: test logic to use multiple specs

* feat: introduce composite speculative decoding stages

* handle MTP context and draft invalidation

* fix: allow gemma mtp for speculative stages

* fix: normalize spec stage keys

* refactor: remove enable_mtp flag and improve speculative stage handling

* fix: update cached text tokens handling for stage chains

* feat: implement sync for external MTP after non-MTP accept
2026-05-15 10:10:40 +03:00
..
backend Merge mainline - Aug 12 2024 (#17) 2024-08-12 15:14:32 +02:00
development Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
android.md Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
autoparser.md Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
build.md Update repository clone instructions in build.md (#1753) 2026-05-07 12:57:06 +03:00
docker.md Update Docker documentation with important notice 2026-03-15 12:35:04 +01:00
function-calling.md common : introduce composable PEG parser combinators for chat parsing and new jinja template engine (#1369) 2026-03-09 11:03:33 +01:00
install.md Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
llguidance.md Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
parameters.md Allow dual speculative decoding (#1789) 2026-05-15 10:10:40 +03:00
speculative.md spec : add self speculative decoding, ngram and refactor (#1261) 2026-02-13 19:04:55 +01:00