* initial map to load deepseek 4 arch * wip * wip: match graph build and attn logic for dpv4 * wip: Enhance DeepSeek-V4 architecture with new tensor types and sqrtsoftplus gating function * Update DeepSeek-V4 to support raw key indexing with read/write indices * fix mismatch in attn_raw * Enable FA with CSA/HCA * Fix logit mismatch with FA path * Clean traces and logs for debug * Refactor DSV4 tensor handling for MTP execution and improve raw context management * Refactor DeepSeek4 tensor operations: replace manual weighted sum and post-processing with new helper functions * Share mHC pre-projection and fix packed DSV4 writes * DSV4: add shared top-k selection and improve mask handling * Fix DSV4 c2048 view stride and duplicate loader instantiation * Reuse shared RMS normalization in DSV4 graph * Replace DSV4 indexer rotation with shared Hadamard * Share CSA visibility mask with DSV4 LID * dsv4: document dependency ordering and reset state * Remove DSV4 zero-dependency graph shim * Fix DSV4 packed stream execution * Remove DSV4 l_out backend override * Enable DSV4 quantized K-only cache * Revert "Enable DSV4 quantized K-only cache" This reverts commit 04f9b425321f62ba60e16d1bea2f8de714cfe855. * Fix DSV4 quantized cache accounting * Fail closed on unsupported DSV4 cache lifecycle operations * Various optimizations * llama: fix GGML_METAL=ON build - missing ggml-metal.h include in llama-dflash.cpp (#2134) llama-dflash.cpp calls ggml_backend_is_metal() and ggml_backend_metal_set_n_cb() inside an #ifdef GGML_USE_METAL block but never includes ggml-metal.h, so any Metal-enabled build fails to compile. Add the same guarded include llama.cpp already uses. * New op: ggml_sum_rows_ext (#2132) * Add ggml_sum_rows_ext * openPangu: use ggml_sum_rows_ext also in mhc_post * openPangu: use ggml_sum_rows_ext also in mhc_tail * Minor * Reuse shared inverse RoPE operation for DSV4 * Reuse maintainer CUDA concat implementation * WIP * hc_pre * hc_post * Remove unnecessary mask manipulations * WIP * Take into account swiglu limits * Turn on fused indexer by default * Give names to mat mul results * More named ops * dsv4: do not uselessly copy the KV cache +20% TG at 32k tokens * mask_to_index and make CPU FA work with that * Much better CPU-only, CUDA still not functional * Better CPU TG I'm now at 9.7 t/s for zero context and 6.5 t/s for context of 32k. PP is 120 t/s for short context and 101 t/s at 32k. * Even better CPU TG I'm now at 8.1 t/s for context of 32k tokens. * Turn off DSA on CUDA for now * Fix CUDA DSA * Remove again the unnecessary softmax result buffer * Experiments * Various * More named ops * Forgot to uncomment --------- Co-authored-by: samuel <samueloliveira32df@gmail.com> Co-authored-by: hchengit <95317477+hchengit@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| Apertus-8B-Instruct.jinja | ||
| Apriel-1.6-15b-Thinker-fixed.jinja | ||
| Bielik-11B-v3.0-Instruct.jinja | ||
| ByteDance-Seed-OSS.jinja | ||
| Cohere2MoE.jinja | ||
| CohereForAI-c4ai-command-r-plus-tool_use.jinja | ||
| CohereForAI-c4ai-command-r7b-12-2024-tool_use.jinja | ||
| GLM-4.6.jinja | ||
| GLM-4.7-Flash.jinja | ||
| GigaChat3-10B-A1.8B.jinja | ||
| GigaChat3.1-10B-A1.8B.jinja | ||
| HuggingFaceTB-SmolLM3-3B.jinja | ||
| Kimi-K2-Instruct.jinja | ||
| Kimi-K2-Thinking.jinja | ||
| LFM2-8B-A1B.jinja | ||
| LFM2.5-Instruct.jinja | ||
| MiMo-VL.jinja | ||
| MiroThinker.jinja | ||
| Mistral-Small-3.2-24B-Instruct-2506.jinja | ||
| NVIDIA-Nemotron-3-Nano-30B-A3B-BF16.jinja | ||
| NousResearch-Hermes-2-Pro-Llama-3-8B-tool_use.jinja | ||
| NousResearch-Hermes-3-Llama-3.1-8B-tool_use.jinja | ||
| Qwen-QwQ-32B.jinja | ||
| Qwen-Qwen2.5-7B-Instruct.jinja | ||
| Qwen-Qwen3-0.6B.jinja | ||
| Qwen3-Coder.jinja | ||
| Qwen3.5-4B.jinja | ||
| README.md | ||
| Reka-Edge.jinja | ||
| StepFun3.5-Flash.jinja | ||
| deepseek-ai-DeepSeek-R1-Distill-Llama-8B.jinja | ||
| deepseek-ai-DeepSeek-R1-Distill-Qwen-32B.jinja | ||
| deepseek-ai-DeepSeek-V3.1.jinja | ||
| deepseek-ai-DeepSeek-V3.2.jinja | ||
| deepseek-ai-DeepSeek-V4.jinja | ||
| fireworks-ai-llama-3-firefunction-v2.jinja | ||
| google-gemma-2-2b-it.jinja | ||
| google-gemma-4-31B-it-interleaved.jinja | ||
| google-gemma-4-31B-it.jinja | ||
| ibm-granite-granite-3.3-2B-Instruct.jinja | ||
| llama-cpp-deepseek-r1.jinja | ||
| llama-cpp-rwkv-world.jinja | ||
| meetkai-functionary-medium-v3.1.jinja | ||
| meetkai-functionary-medium-v3.2.jinja | ||
| meta-llama-Llama-3.1-8B-Instruct.jinja | ||
| meta-llama-Llama-3.2-3B-Instruct.jinja | ||
| meta-llama-Llama-3.3-70B-Instruct.jinja | ||
| microsoft-Phi-3.5-mini-instruct.jinja | ||
| mistralai-Ministral-3-14B-Reasoning-2512.jinja | ||
| mistralai-Mistral-Nemo-Instruct-2407.jinja | ||
| moonshotai-Kimi-K2.jinja | ||
| openai-gpt-oss-120b.jinja | ||
| stepfun-ai-Step-3.5-Flash.jinja | ||
| unsloth-Apriel-1.5.jinja | ||
| unsloth-mistral-Devstral-Small-2507.jinja | ||
README.md
These templates can be updated with the following commands:
./scripts/get_chat_template.py CohereForAI/c4ai-command-r-plus tool_use > models/templates/CohereForAI-c4ai-command-r-plus-tool_use.jinja
./scripts/get_chat_template.py CohereForAI/c4ai-command-r7b-12-2024 default > models/templates/CohereForAI-c4ai-command-r7b-12-2024-default.jinja
./scripts/get_chat_template.py CohereForAI/c4ai-command-r7b-12-2024 rag > models/templates/CohereForAI-c4ai-command-r7b-12-2024-rag.jinja
./scripts/get_chat_template.py CohereForAI/c4ai-command-r7b-12-2024 tool_use > models/templates/CohereForAI-c4ai-command-r7b-12-2024-tool_use.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-R1-Distill-Llama-8B > models/templates/deepseek-ai-DeepSeek-R1-Distill-Llama-8B.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-R1-Distill-Qwen-32B > models/templates/deepseek-ai-DeepSeek-R1-Distill-Qwen-32B.jinja
./scripts/get_chat_template.py fireworks-ai/llama-3-firefunction-v2 > models/templates/fireworks-ai-llama-3-firefunction-v2.jinja
./scripts/get_chat_template.py google/gemma-2-2b-it > models/templates/google-gemma-2-2b-it.jinja
./scripts/get_chat_template.py meetkai/functionary-medium-v3.1 > models/templates/meetkai-functionary-medium-v3.1.jinja
./scripts/get_chat_template.py meetkai/functionary-medium-v3.2 > models/templates/meetkai-functionary-medium-v3.2.jinja
./scripts/get_chat_template.py meta-llama/Llama-3.1-8B-Instruct > models/templates/meta-llama-Llama-3.1-8B-Instruct.jinja
./scripts/get_chat_template.py meta-llama/Llama-3.2-3B-Instruct > models/templates/meta-llama-Llama-3.2-3B-Instruct.jinja
./scripts/get_chat_template.py meta-llama/Llama-3.3-70B-Instruct > models/templates/meta-llama-Llama-3.3-70B-Instruct.jinja
./scripts/get_chat_template.py microsoft/Phi-3.5-mini-instruct > models/templates/microsoft-Phi-3.5-mini-instruct.jinja
./scripts/get_chat_template.py mistralai/Mistral-Nemo-Instruct-2407 > models/templates/mistralai-Mistral-Nemo-Instruct-2407.jinja
./scripts/get_chat_template.py NousResearch/Hermes-2-Pro-Llama-3-8B tool_use > models/templates/NousResearch-Hermes-2-Pro-Llama-3-8B-tool_use.jinja
./scripts/get_chat_template.py NousResearch/Hermes-3-Llama-3.1-8B tool_use > models/templates/NousResearch-Hermes-3-Llama-3.1-8B-tool_use.jinja
./scripts/get_chat_template.py Qwen/Qwen2.5-7B-Instruct > models/templates/Qwen-Qwen2.5-7B-Instruct.jinja
./scripts/get_chat_template.py Qwen/QwQ-32B > models/templates/Qwen-QwQ-32B.jinja
./scripts/get_chat_template.py Qwen/Qwen3-0.6B > models/templates/Qwen-Qwen3-0.6B.jinja
./scripts/get_chat_template.py zai-org/GLM-4.5 > models/templates/zai-org-GLM-4.5.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-V3.1 > models/templates/deepseek-ai-DeepSeek-V3.1.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-V4 > models/templates/deepseek-ai-DeepSeek-V4.jinja