..
graphs
Fix Minimax M3 crash when -muge merges up/gate experts
2026-06-15 15:00:32 +02:00
CMakeLists.txt
Add preliminary MiniMax-M3 support
2026-06-14 12:23:20 +00:00
llama-arch.cpp
Add preliminary MiniMax-M3 support
2026-06-14 12:23:20 +00:00
llama-arch.h
Add preliminary MiniMax-M3 support
2026-06-14 12:23:20 +00:00
llama-build-context.cpp
Split mode graph for MiniMax-M3
2026-06-15 08:41:34 +00:00
llama-build-context.h
Split mode graph for MiniMax-M3
2026-06-15 08:41:34 +00:00
llama-context.h
Add extra nodes when dealing with MLA and amb ( #1899 )
2026-05-29 15:17:24 +03:00
llama-cparams.h
MLA tensor parallelism under -sm graph (DEEPSEEK2/GLM_DSA/MISTRAL4) ( #1821 )
2026-05-19 08:36:17 +03:00
llama-delta-net.cpp
MTP: faster recurrent state restore ( #1791 )
2026-05-13 11:00:24 +03:00
llama-delta-net.h
MTP: faster recurrent state restore ( #1791 )
2026-05-13 11:00:24 +03:00
llama-expert-io.h
Add --defer-experts flag to defer expert mmap residency on Linux ( #1634 )
2026-04-16 08:54:44 +02:00
llama-grammar.cpp
Bug fixes ( #1940 )
2026-06-10 07:45:49 +02:00
llama-grammar.h
llama : add token matching support to llama-grammar ( #1220 )
2026-02-03 07:57:17 +02:00
llama-hparams.cpp
Add preliminary MiniMax-M3 support
2026-06-14 12:23:20 +00:00
llama-hparams.h
model: add Poolside Laguna XS.2 support ( #1911 )
2026-06-08 18:33:12 +02:00
llama-impl.h
Full graph parallel for Qwen3.5 (dense and MoE) ( #1388 )
2026-03-10 09:08:24 +01:00
llama-load-tensors.cpp
Add preliminary MiniMax-M3 support
2026-06-14 12:23:20 +00:00
llama-mmap.cpp
Add --defer-experts flag to defer expert mmap residency on Linux ( #1634 )
2026-04-16 08:54:44 +02:00
llama-mmap.h
Add --defer-experts flag to defer expert mmap residency on Linux ( #1634 )
2026-04-16 08:54:44 +02:00
llama-model-loader.cpp
fix: use int8_t for GGUF bool array loading instead of platform-dependent bool ( #1648 )
2026-04-17 07:25:07 +02:00
llama-model-loader.h
MTP: option to use re-quantized output tensor for better TG performance ( #1809 )
2026-05-16 14:40:18 +03:00
llama-model.cpp
Add preliminary MiniMax-M3 support
2026-06-14 12:23:20 +00:00
llama-model.h
model: add Poolside Laguna XS.2 support ( #1911 )
2026-06-08 18:33:12 +02:00
llama-quantize.cpp
quantize: add exception for Gemma4 ( #1897 )
2026-05-29 10:54:21 +03:00
llama-quantize.h
Allow using -rtr and -muge together ( #1444 )
2026-03-16 18:26:26 +01:00
llama-sampling.cpp
Bug fixes ( #1940 )
2026-06-10 07:45:49 +02:00
llama-sampling.h
Add adaptive sampling clone and free functions to manage memory ( #1851 )
2026-05-21 08:11:17 +03:00
llama-spec-features.cpp
Support for alternative Gemma4 assistant ( #1937 )
2026-06-09 09:30:12 +02:00
llama-spec-features.h
Move embedding management to speculative ( #1825 )
2026-05-20 17:42:48 +03:00
llama-vocab.cpp
model: add Cohere2-MoE North Mini Code support ( #1945 )
2026-06-10 15:28:27 +02:00
llama-vocab.h
Add Mellum2 architecture support ( #1919 )
2026-06-04 14:28:02 +02:00
llama.cpp
Add preliminary MiniMax-M3 support
2026-06-14 12:23:20 +00:00
unicode-data.cpp
Merge mainline llama.cpp ( #3 )
2024-07-27 07:55:01 +02:00
unicode-data.h
Merge mainline llama.cpp ( #3 )
2024-07-27 07:55:01 +02:00
unicode-script-data.cpp
Add Unicode allowlist ( #1597 )
2026-04-10 18:22:57 +02:00
unicode.cpp
Gemma4 tokenizer fixes ( #1603 )
2026-04-09 15:33:28 +02:00
unicode.h
Add Unicode allowlist ( #1597 )
2026-04-10 18:22:57 +02:00