* Add --prefetch-experts to stream mmap'd MoE experts into page cache * Drop fds, fault experts in with MADV_POPULATE_READ instead of pread * Remove stale note about pread workers * Move MoE prefetch behind ggml_backend_prefetch_* wrappers * Cleanup stale comments * Add --prefetch-experts-threads, drop GGML_MOE_PREFETCH_THREADS env var |
||
|---|---|---|
| .. | ||
| llama.h | ||