ik_llama_opt/ggml
Kawrakow defa6945b3 CUDA: fuse copies to K and V cache (#921)
* Fuse copies to K- and V-cache on CUDA

* Adapt to latest main

---------

Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2025-11-08 18:13:58 +02:00
..
cmake Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
include CUDA: set compute parameters via command line arguments (#910) 2025-11-07 07:11:23 +02:00
src CUDA: fuse copies to K and V cache (#921) 2025-11-08 18:13:58 +02:00
.gitignore Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
CMakeLists.txt Disable CUDA fusion by default for now (#903) 2025-11-05 10:58:12 +02:00