* Adding ds4_comp op with CPU implementation * ds4_comp on CUDA * ds4_comp: ratio = 4 specialization Surprisingly small performance gain * Also handle HCA via ds4_comp But much smaller gain, if any. * Delete commented out stuff * Remove the [(size_t) il] noise * Minor * Fix quantized cache |
||
|---|---|---|
| .. | ||
| ggml-alloc.h | ||
| ggml-backend.h | ||
| ggml-cann.h | ||
| ggml-cpp.h | ||
| ggml-cuda.h | ||
| ggml-metal.h | ||
| ggml-rpc.h | ||
| ggml-sycl.h | ||
| ggml-vulkan.h | ||
| ggml.h | ||