* WIP: indexer_topk on CUDA * Forgot these * WIP * WIP * This seems to work * Minor * Fix bug. Fix suggested by @sayap using GLM-5.2 * GLM-DSA: much better PP long context performance (CUDA) * DSA: Better way to build the attention mask |
||
|---|---|---|
| .. | ||
| ggml-alloc.h | ||
| ggml-backend.h | ||
| ggml-cann.h | ||
| ggml-cpp.h | ||
| ggml-cuda.h | ||
| ggml-metal.h | ||
| ggml-rpc.h | ||
| ggml-sycl.h | ||
| ggml-vulkan.h | ||
| ggml.h | ||