Kawrakow
1b53a58bf9
Enable split mode graph for Gemma4-12B ( #1922 )
2026-06-05 10:59:22 +02:00
Kawrakow
8960c5ba5e
Add extra nodes when dealing with MLA and amb ( #1899 )
2026-05-29 15:17:24 +03:00
Kawrakow
3bf7e836c2
Allow Hadamard transform for head sizes that are not power of 2 ( #1883 )
...
* Disable K Hadamard transform if K-head size is not a power of 2
* Allow Hadamard transform for head sizes that are not power of 2
* Give more details why Hadamard is not possible
* Arghh
2026-05-27 18:29:32 +03:00
Kawrakow
c5dc847d0a
Fix Gemma4-E4B compute graph ( #1855 )
2026-05-21 12:46:28 +03:00
Kawrakow
a407b9ca3d
Fix Qwen3.6-MoE low MTP acceptance rate ( #1815 )
...
* Fix Qwen3.6-MoE low MTP acceptance rate
* Fix Gemma4 MTP
2026-05-18 07:26:17 +03:00
Samuel Oliveira Alves
0fcffdb64d
feat: map Gemma 4 tensor and support with imatrix ( #1796 )
2026-05-14 09:01:24 +03:00
Kawrakow
86b5d076c5
Gemma4 MTP: avoid casting KV cache to f32 ( #1786 )
2026-05-13 09:11:27 +03:00
Samuel Oliveira Alves
c2b8bca807
Add MTP Support for Gemma 4 ( #1744 )
...
* gemma-mtp: build the arch to load the MTP model
* gemma-mtp: fix mtp kv state
* gemma-mtp: refactor some functions and create gguf
* gemma-mtp: make usable for embeddings models variant
* gemma-mtp: fix qwen mtp load in graph split
* gemma-mtp: refactor tensor creation and adjust output tensor handling
* Gemma 4 MTP: improve tensor handling, and adjust split mode logic
2026-05-10 07:44:20 +03:00
Kawrakow
8befd92ea5
Refactor model compute graphs ( #1651 )
...
* Refactor model compute graphs
* Remove unused function
2026-04-18 17:08:43 +02:00