* perplexity: add signal-driven hot-swap mode for persistent KLD/PPL benchmarking llama-perplexity can now stay resident and be driven through control/status files (reload/compute/exit): it reloads only the tensors that changed on disk and recomputes PPL/KLD without ever reloading the full model. File-based signalling works on Windows, macOS and Linux. The reload returning-to-original path now refreshes tensor data from disk instead of reattaching stale weights. * Not Cygwin specific |
||
|---|---|---|
| .. | ||
| llama-star | ||
| HOWTO-add-model.md | ||
| debugging-tests.md | ||
| on-demand-tensor-reload.md | ||
| on-demand-tensor-reload.mmd | ||
| parsing.md | ||
| token_generation_performance_tips.md | ||