* Also take into account KV cache * Take into account attn_wkv_b and mla = 3 compute buffers |
||
|---|---|---|
| .. | ||
| llama.h | ||
* Also take into account KV cache * Take into account attn_wkv_b and mla = 3 compute buffers |
||
|---|---|---|
| .. | ||
| llama.h | ||
Powered by TurnKey Linux.