On P100 (GP100, sm_60) the fp16 vec kernel used for decode (batch<=8) accumulates the online-softmax denominator and the P*V product in fp16, flipping ~3-4% of decode top-1 tokens vs an all-fp32 reference (llama.cpp#25593). Decode is memory-bandwidth-bound on P100, so routing sm_60 decode to the in-tree vec_f32 kernel is free (tg128 ~identical). Gated on cc == CC_PASCAL && Q->ne[1] <= 8 (decode only) inside the !fp16_mma_available block, so the prefill tile_f16 path, the D=256 prefill vec path, and fast_fp16_available() are untouched, and the is_pascal_mla_absorbed_decode early-return (MLA) is unaffected. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| cmake | ||
| include | ||
| src | ||
| .gitignore | ||
| CMakeLists.txt | ||