ik_llama_opt/docs
Thireus ☠ 6d78a87c4c
perplexity: signal-driven hot-swap mode for persistent per-tensor PPL/KLD benchmarking (extends #1989) (#2131)
* perplexity: add signal-driven hot-swap mode for persistent KLD/PPL benchmarking

llama-perplexity can now stay resident and be driven through control/status files (reload/compute/exit): it reloads only the tensors that changed on disk and recomputes PPL/KLD without ever reloading the full model. File-based signalling works on Windows, macOS and Linux. The reload returning-to-original path now refreshes tensor data from disk instead of
reattaching stale weights.

* Not Cygwin specific
2026-07-14 12:56:03 +03:00
..
backend Merge mainline - Aug 12 2024 (#17) 2024-08-12 15:14:32 +02:00
development perplexity: signal-driven hot-swap mode for persistent per-tensor PPL/KLD benchmarking (extends #1989) (#2131) 2026-07-14 12:56:03 +03:00
android.md Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
autoparser.md common: handle Laguna chat delimiters (#1943) 2026-06-10 07:46:19 +02:00
build.md Update repository clone instructions in build.md (#1753) 2026-05-07 12:57:06 +03:00
docker.md Update Docker documentation with important notice 2026-03-15 12:35:04 +01:00
function-calling.md common : introduce composable PEG parser combinators for chat parsing and new jinja template engine (#1369) 2026-03-09 11:03:33 +01:00
install.md Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
llguidance.md Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
parameters.md model: add openPangu-2.0-Flash (92B-A6B) with MLA-latent cache, DSA/SWA, mHC, and multi-head MTP (#2065) 2026-07-11 12:29:20 +03:00
speculative.md model: add openPangu-2.0-Flash (92B-A6B) with MLA-latent cache, DSA/SWA, mHC, and multi-head MTP (#2065) 2026-07-11 12:29:20 +03:00