ik_llama_opt/ggml
Kawrakow 76c1942716
Allow Q8_0 cache in the CUDA DSA implementation (#2276)
* CUDA indexer topk: this is better for PP

* Don't overstep

* Cleanup

* Allow Q8_0 cache in the CUDA DSA implementation
2026-08-08 17:18:36 +03:00
..
cmake Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
include DS4: faster long-context TG (#2201) 2026-07-30 13:13:42 +03:00
src Allow Q8_0 cache in the CUDA DSA implementation (#2276) 2026-08-08 17:18:36 +03:00
.gitignore Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
CMakeLists.txt Chunked experts (CPU) (#2202) 2026-07-30 13:16:02 +03:00