* Adding ds4_comp op with CPU implementation * ds4_comp on CUDA * ds4_comp: ratio = 4 specialization Surprisingly small performance gain * Also handle HCA via ds4_comp But much smaller gain, if any. * Delete commented out stuff * Remove the [(size_t) il] noise * Minor * Fix quantized cache |
||
|---|---|---|
| .. | ||
| cmake | ||
| include | ||
| src | ||
| .gitignore | ||
| CMakeLists.txt | ||