ik_llama_opt/models/templates
Yap Sok Ann 6b55d2c750
Fix DSV4 tool calls and reasoning (#2242)
* Fix DSV4 tool calls and reasoning

There are multiple changes. The most important one is the wiring, to
avoid falling back to the autoparser. With autoparser, all arguments
will be forced by the generated grammar to have the `string="true"`
attribute, which then breaks prompt caching, as it would diverge from
what is rendered by the template. Parallel tool calls also doesn't work
when falling back to autoparser.

Other changes:
* Order tool results by tool call order.
* Consume `</think>` instead of `<think></think>` when thinking is
  disabled.
* Use "preserved thinking" mode when any tool is defined, otherwise use
  "interleaved thinking" mode, e.g. for multi-turns chat. Set template
  arg `drop_thinking` to false to force "preserved thinking" mode even
  when no tool is defined.
* Add a message to system prompt when reasoning effort is set to max.

The changes were made by following:
1. The Technical Report: https://arxiv.org/abs/2606.19348
2. Reference implementatin: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py
3. VLLM implementation: https://github.com/vllm-project/vllm/blob/main/vllm/tokenizers/deepseek_v4_encoding.py

For the last bullet point about max reasoning effort, the reference
implementation diverges from the other 2, so we follow the techinical
report and the VLLM implementation, for now. This needs more testing.

* Add back trailing newline

* Update the reasoning effort instruction to follow the reference impl

Using the prompt counting test from @coder543, 0731 does have a special
instruction for "high" and another special instruction for "max".

This will break preview, but assuming most people will use the 0731
release, it should be fine.

[1] https://www.reddit.com/r/DeepSeek/comments/1vdqjwr/openrouter_reasoning_effort_levels_are_broken_for/
2026-08-04 19:28:06 +03:00
..
Apertus-8B-Instruct.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
Apriel-1.6-15b-Thinker-fixed.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
Bielik-11B-v3.0-Instruct.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
ByteDance-Seed-OSS.jinja tests: add Seed-OSS chat template fixture (#2014) 2026-06-23 09:35:28 +02:00
Cohere2MoE.jinja chat: add Cohere2MoE North Code parser (#1968) 2026-06-16 15:27:30 +02:00
CohereForAI-c4ai-command-r-plus-tool_use.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
CohereForAI-c4ai-command-r7b-12-2024-tool_use.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
GLM-4.6.jinja common: Generalized XML-style tool-call parsing with streaming support (#958) 2025-11-18 15:29:58 +01:00
GLM-4.7-Flash.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
GigaChat3-10B-A1.8B.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
GigaChat3.1-10B-A1.8B.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
HuggingFaceTB-SmolLM3-3B.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
Kimi-K2-Instruct.jinja fix kimi-k2 tool call (#996) 2025-11-24 06:51:16 +01:00
Kimi-K2-Thinking.jinja fix kimi-k2 tool call (#996) 2025-11-24 06:51:16 +01:00
LFM2-8B-A1B.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
LFM2.5-Instruct.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
MiMo-VL.jinja common: Generalized XML-style tool-call parsing with streaming support (#958) 2025-11-18 15:29:58 +01:00
MiroThinker.jinja Improve MiroThinker chat template compatibility with the new Jinja template engine (#1404) 2026-03-13 08:11:17 +01:00
Mistral-Small-3.2-24B-Instruct-2506.jinja add jinja template support (#677) 2025-08-09 12:50:30 +00:00
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16.jinja common : introduce composable PEG parser combinators for chat parsing and new jinja template engine (#1369) 2026-03-09 11:03:33 +01:00
NousResearch-Hermes-2-Pro-Llama-3-8B-tool_use.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
NousResearch-Hermes-3-Llama-3.1-8B-tool_use.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
Qwen-QwQ-32B.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
Qwen-Qwen2.5-7B-Instruct.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
Qwen-Qwen3-0.6B.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
Qwen3-Coder.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
Qwen3.5-4B.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
README.md DS4: slowly approaching a meaningful performance (#2165) 2026-07-22 17:18:57 +03:00
Reka-Edge.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
StepFun3.5-Flash.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
deepseek-ai-DeepSeek-R1-Distill-Llama-8B.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
deepseek-ai-DeepSeek-R1-Distill-Qwen-32B.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
deepseek-ai-DeepSeek-V3.1.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
deepseek-ai-DeepSeek-V3.2.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
deepseek-ai-DeepSeek-V4.jinja Fix DSV4 tool calls and reasoning (#2242) 2026-08-04 19:28:06 +03:00
fireworks-ai-llama-3-firefunction-v2.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
google-gemma-2-2b-it.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
google-gemma-4-31B-it-interleaved.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
google-gemma-4-31B-it.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
ibm-granite-granite-3.3-2B-Instruct.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
llama-cpp-deepseek-r1.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
llama-cpp-rwkv-world.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
meetkai-functionary-medium-v3.1.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
meetkai-functionary-medium-v3.2.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
meta-llama-Llama-3.1-8B-Instruct.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
meta-llama-Llama-3.2-3B-Instruct.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
meta-llama-Llama-3.3-70B-Instruct.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
microsoft-Phi-3.5-mini-instruct.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
mistralai-Ministral-3-14B-Reasoning-2512.jinja common : introduce composable PEG parser combinators for chat parsing and new jinja template engine (#1369) 2026-03-09 11:03:33 +01:00
mistralai-Mistral-Nemo-Instruct-2407.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
moonshotai-Kimi-K2.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
openai-gpt-oss-120b.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
stepfun-ai-Step-3.5-Flash.jinja common : introduce composable PEG parser combinators for chat parsing and new jinja template engine (#1369) 2026-03-09 11:03:33 +01:00
unsloth-Apriel-1.5.jinja Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
unsloth-mistral-Devstral-Small-2507.jinja Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00

README.md

These templates can be updated with the following commands:

./scripts/get_chat_template.py CohereForAI/c4ai-command-r-plus tool_use      > models/templates/CohereForAI-c4ai-command-r-plus-tool_use.jinja
./scripts/get_chat_template.py CohereForAI/c4ai-command-r7b-12-2024 default  > models/templates/CohereForAI-c4ai-command-r7b-12-2024-default.jinja
./scripts/get_chat_template.py CohereForAI/c4ai-command-r7b-12-2024 rag      > models/templates/CohereForAI-c4ai-command-r7b-12-2024-rag.jinja
./scripts/get_chat_template.py CohereForAI/c4ai-command-r7b-12-2024 tool_use > models/templates/CohereForAI-c4ai-command-r7b-12-2024-tool_use.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-R1-Distill-Llama-8B      > models/templates/deepseek-ai-DeepSeek-R1-Distill-Llama-8B.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-R1-Distill-Qwen-32B      > models/templates/deepseek-ai-DeepSeek-R1-Distill-Qwen-32B.jinja
./scripts/get_chat_template.py fireworks-ai/llama-3-firefunction-v2          > models/templates/fireworks-ai-llama-3-firefunction-v2.jinja
./scripts/get_chat_template.py google/gemma-2-2b-it                          > models/templates/google-gemma-2-2b-it.jinja
./scripts/get_chat_template.py meetkai/functionary-medium-v3.1               > models/templates/meetkai-functionary-medium-v3.1.jinja
./scripts/get_chat_template.py meetkai/functionary-medium-v3.2               > models/templates/meetkai-functionary-medium-v3.2.jinja
./scripts/get_chat_template.py meta-llama/Llama-3.1-8B-Instruct              > models/templates/meta-llama-Llama-3.1-8B-Instruct.jinja
./scripts/get_chat_template.py meta-llama/Llama-3.2-3B-Instruct              > models/templates/meta-llama-Llama-3.2-3B-Instruct.jinja
./scripts/get_chat_template.py meta-llama/Llama-3.3-70B-Instruct             > models/templates/meta-llama-Llama-3.3-70B-Instruct.jinja
./scripts/get_chat_template.py microsoft/Phi-3.5-mini-instruct               > models/templates/microsoft-Phi-3.5-mini-instruct.jinja
./scripts/get_chat_template.py mistralai/Mistral-Nemo-Instruct-2407          > models/templates/mistralai-Mistral-Nemo-Instruct-2407.jinja
./scripts/get_chat_template.py NousResearch/Hermes-2-Pro-Llama-3-8B tool_use > models/templates/NousResearch-Hermes-2-Pro-Llama-3-8B-tool_use.jinja
./scripts/get_chat_template.py NousResearch/Hermes-3-Llama-3.1-8B tool_use   > models/templates/NousResearch-Hermes-3-Llama-3.1-8B-tool_use.jinja
./scripts/get_chat_template.py Qwen/Qwen2.5-7B-Instruct                      > models/templates/Qwen-Qwen2.5-7B-Instruct.jinja
./scripts/get_chat_template.py Qwen/QwQ-32B                                  > models/templates/Qwen-QwQ-32B.jinja
./scripts/get_chat_template.py Qwen/Qwen3-0.6B                               > models/templates/Qwen-Qwen3-0.6B.jinja
./scripts/get_chat_template.py zai-org/GLM-4.5                               > models/templates/zai-org-GLM-4.5.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-V3.1                     > models/templates/deepseek-ai-DeepSeek-V3.1.jinja
./scripts/get_chat_template.py deepseek-ai/DeepSeek-V4                       > models/templates/deepseek-ai-DeepSeek-V4.jinja