* CUDA indexer topk: this is better for PP * Don't overstep * Cleanup * Allow Q8_0 cache in the CUDA DSA implementation * DS4: do not cast caches to f32 * Fix massive inefficiency in CUDA Q->f32/f16 and f32/f16->Q copies * Re-enable -ictk | --indexer-cache-type-k |
||
|---|---|---|
| .. | ||
| cmake | ||
| include | ||
| src | ||
| .gitignore | ||
| CMakeLists.txt | ||