ik_llama_opt/ggml
Kawrakow 15dddc60b3
Qwen3.8-Flash-Next: faster TG on CUDA (#2373)
* Qwen3.8-Flash-Next: faster TG on CUDA

* Only one thread should write to the destination
2026-08-28 18:11:17 +02:00
..
cmake Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
include Quantization fudge factors (#2361) 2026-08-27 17:35:04 +02:00
src Qwen3.8-Flash-Next: faster TG on CUDA (#2373) 2026-08-28 18:11:17 +02:00
.gitignore Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
CMakeLists.txt Chunked experts (CPU) (#2202) 2026-07-30 13:16:02 +03:00