ik_llama_opt/ggml
Kawrakow 17d2db910f Use cuBLAS for large batches and quants with block size 16 (#559)
Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2025-06-27 17:43:51 +02:00
..
cmake Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
include Fix non rpc build error (#506) 2025-06-08 17:27:00 +03:00
src Use cuBLAS for large batches and quants with block size 16 (#559) 2025-06-27 17:43:51 +02:00
.gitignore Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
CMakeLists.txt Better strategy for GPU offload (#520) 2025-06-12 19:25:11 +03:00