ik_llama_opt/common
Yap Sok Ann 6b55d2c750
Fix DSV4 tool calls and reasoning (#2242)
* Fix DSV4 tool calls and reasoning

There are multiple changes. The most important one is the wiring, to
avoid falling back to the autoparser. With autoparser, all arguments
will be forced by the generated grammar to have the `string="true"`
attribute, which then breaks prompt caching, as it would diverge from
what is rendered by the template. Parallel tool calls also doesn't work
when falling back to autoparser.

Other changes:
* Order tool results by tool call order.
* Consume `</think>` instead of `<think></think>` when thinking is
  disabled.
* Use "preserved thinking" mode when any tool is defined, otherwise use
  "interleaved thinking" mode, e.g. for multi-turns chat. Set template
  arg `drop_thinking` to false to force "preserved thinking" mode even
  when no tool is defined.
* Add a message to system prompt when reasoning effort is set to max.

The changes were made by following:
1. The Technical Report: https://arxiv.org/abs/2606.19348
2. Reference implementatin: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731/blob/main/encoding/encoding_dsv4.py
3. VLLM implementation: https://github.com/vllm-project/vllm/blob/main/vllm/tokenizers/deepseek_v4_encoding.py

For the last bullet point about max reasoning effort, the reference
implementation diverges from the other 2, so we follow the techinical
report and the VLLM implementation, for now. This needs more testing.

* Add back trailing newline

* Update the reasoning effort instruction to follow the reference impl

Using the prompt counting test from @coder543, 0731 does have a special
instruction for "high" and another special instruction for "max".

This will break preview, but assuming most people will use the 0731
release, it should be fine.

[1] https://www.reddit.com/r/DeepSeek/comments/1vdqjwr/openrouter_reasoning_effort_levels_are_broken_for/
2026-08-04 19:28:06 +03:00
..
cmake Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
jinja Chores : Typos fixing round 3 (project wide, ggml dir included, comments and user facing msg only) (#2249) 2026-08-04 07:15:28 +03:00
CMakeLists.txt server: enable mcp proxy (#1904) 2026-06-04 15:43:07 +02:00
base64.hpp llava : expose as a shared library for downstream projects (#3613) 2023-11-07 00:36:23 +03:00
build-info.cpp.in build : link against build info instead of compiling against it (#3879) 2023-11-02 08:50:16 +02:00
chat-auto-parser-generator.cpp common: gate empty-start reasoning extraction (#1955) 2026-06-12 07:16:24 +02:00
chat-auto-parser-helpers.cpp Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
chat-auto-parser-helpers.h Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
chat-auto-parser.h common: handle Laguna chat delimiters (#1943) 2026-06-10 07:46:19 +02:00
chat-diff-analyzer.cpp model: add openPangu-2.0-Flash (92B-A6B) with MLA-latent cache, DSA/SWA, mHC, and multi-head MTP (#2065) 2026-07-11 12:29:20 +03:00
chat-peg-parser.cpp fix: MiniMax-M3 streaming parser when tool calls start before `</mm:think>` (#2085) 2026-07-09 18:12:47 +03:00
chat-peg-parser.h fix: MiniMax-M3 streaming parser when tool calls start before `</mm:think>` (#2085) 2026-07-09 18:12:47 +03:00
chat.cpp Fix DSV4 tool calls and reasoning (#2242) 2026-08-04 19:28:06 +03:00
chat.h fix: MiniMax-M3 streaming parser when tool calls start before `</mm:think>` (#2085) 2026-07-09 18:12:47 +03:00
common.cpp Chores : tidy up more typos project wide (ggml directory excluded), new -ptcall alias (#2237) 2026-08-03 08:01:18 +03:00
common.h Chores : tidy up more typos project wide (ggml directory excluded), new -ptcall alias (#2237) 2026-08-03 08:01:18 +03:00
console.cpp check C++ code with -Wmissing-declarations (#3184) 2023-09-15 15:38:27 -04:00
console.h gguf : new file format with flexible meta data (beta) (#2398) 2023-08-21 23:07:43 +03:00
http.h server: enable mcp proxy (#1904) 2026-06-04 15:43:07 +02:00
json-partial.cpp common : introduce composable PEG parser combinators for chat parsing and new jinja template engine (#1369) 2026-03-09 11:03:33 +01:00
json-partial.h Move minja and nlohmann/json to vendor (#802) 2025-09-27 09:12:35 +02:00
json-schema-to-grammar.cpp Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
json-schema-to-grammar.h common : introduce composable PEG parser combinators for chat parsing and new jinja template engine (#1369) 2026-03-09 11:03:33 +01:00
llguidance.cpp Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
log.cpp Refactor chat and server file (#1062) 2025-12-15 08:27:20 +01:00
log.h Chores : tidy up more typos project wide (ggml directory excluded), new -ptcall alias (#2237) 2026-08-03 08:01:18 +03:00
ngram-cache.cpp spec : add self speculative decoding, ngram and refactor (#1261) 2026-02-13 19:04:55 +01:00
ngram-cache.h spec : add self speculative decoding, ngram and refactor (#1261) 2026-02-13 19:04:55 +01:00
ngram-map.cpp Chores : tidy up more typos project wide (ggml directory excluded), new -ptcall alias (#2237) 2026-08-03 08:01:18 +03:00
ngram-map.h Chores : tidy up more typos project wide (ggml directory excluded), new -ptcall alias (#2237) 2026-08-03 08:01:18 +03:00
ngram-mod.cpp spec : add self speculative decoding, ngram and refactor (#1261) 2026-02-13 19:04:55 +01:00
ngram-mod.h spec : add self speculative decoding, ngram and refactor (#1261) 2026-02-13 19:04:55 +01:00
peg-parser.cpp Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
peg-parser.h Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
reasoning-budget.cpp Remove reasoning budget logs (#1846) 2026-05-20 07:12:02 +03:00
reasoning-budget.h AutoParser: improve reasoning budget and handling of space/newline in tool calls (#1819) 2026-05-19 08:34:19 +03:00
regex-partial.cpp Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
regex-partial.h Tool calls support from mainline (#723) 2025-09-01 08:38:49 +03:00
sampling.cpp sampling: fix out-of-bounds logits read when the vocab has no newline token (#2188) 2026-07-26 10:59:58 +03:00
sampling.h Extend expiring logit bias to other sampling parameters (#1770) 2026-05-23 19:19:12 +03:00
spec-tuner.cpp feat: allow dflash to work with spec auto tune (#2112) 2026-07-12 07:49:03 +03:00
spec-tuner.h feat: allow dflash to work with spec auto tune (#2112) 2026-07-12 07:49:03 +03:00
speculative-dflash-impl.h clean logs 2026-06-14 21:07:57 -03:00
speculative.cpp openpangu: support server context checkpoints and prompt reuse (#2245) 2026-08-03 09:23:01 +03:00
speculative.h Feat speculative benchmark standard (#2208) 2026-07-30 18:38:48 +03:00
suffix-tree.cpp Standardize speculative decoding arguments on the server (#1908) 2026-06-04 15:44:57 +02:00
suffix-tree.h Self-decoding: Adds support for suffix decoding (#1646) 2026-04-18 16:10:10 +02:00
train.cpp Server: refactor and rename functions (#1151) 2026-01-18 08:16:57 +02:00
train.h sync : ggml (backend v2) (#3912) 2023-11-13 14:16:23 +02:00
unicode.cpp Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00
unicode.h Autoparser - complete refactoring of parser architecture (#1376) 2026-04-22 10:04:13 +02:00