* perplexity: add signal-driven hot-swap mode for persistent KLD/PPL benchmarking llama-perplexity can now stay resident and be driven through control/status files (reload/compute/exit): it reloads only the tensors that changed on disk and recomputes PPL/KLD without ever reloading the full model. File-based signalling works on Windows, macOS and Linux. The reload returning-to-original path now refreshes tensor data from disk instead of reattaching stale weights. * Not Cygwin specific |
||
|---|---|---|
| .. | ||
| backend | ||
| development | ||
| android.md | ||
| autoparser.md | ||
| build.md | ||
| docker.md | ||
| function-calling.md | ||
| install.md | ||
| llguidance.md | ||
| parameters.md | ||
| speculative.md | ||