* common: fix coding mistakes (typos in identifiers, flags and log strings)
Fix misspelled identifiers and user-facing strings across common, server
and model loading:
- allow_ruless -> allow_rules (misspelled identifier used in the allowlist
CLI parsing and the server slot/context code)
- get_formated_timings/get_formated_generation -> get_formatted_*
- 'termionated' -> 'terminated' in the fit-margin assert message
- 'defaulr' -> 'default' in the YAML dump
- 'overriden' -> 'overridden' in tensor buffer type override logs
- 'becausee' -> 'because' in the output-tensor split log
- 'etected NaNs' -> 'detected NaNs' in the imatrix error message
* common: fix comment typos across src, common, include and examples
Fix misspelled words in code comments:
- llama.h: 'typy' -> 'type', 'transfrom' -> 'transform', 'ecoder' ->
'encoder', 'indicies' -> 'indices', 'Intializes' -> 'Initializes'
- common.h: 'embendings' -> 'embeddings', 'pr' -> 'or' in the
fused-indexer-topk comment
- chat.cpp: 'overridde' -> 'override'
- ngram-map: 'occurences' -> 'occurrences', 'stastistics' -> 'statistics'
- speculative.cpp: 'dont'/'inehit' -> 'don't'/'inherit'
- llama-mmap.cpp: 'dont't' -> 'don't'
- llama-model.h: 'hcurrently andle' -> 'currently handle'
- build_gemma3/4.cpp: 'emdeddings' -> 'embeddings'
- examples: 'quantizuation', 'logprobe', 'throught', 'retrun', 'swich',
'convinient', 'temporally' (-> 'temporary'), 'temproal', 'preceed'
* common: remove duplicate definitions and duplicate help entries
- clip-impl.h: drop the second, identical #define TN_FFN_GATE
- common.cpp: remove the duplicate '-t, --threads N' help entry that was
misplaced in the export-lora section (already listed in the general
section)
- common.cpp: merge the two 'embedding' help groups into a single group
so the embedding options are listed together
- llama.cpp: remove the redundant LLAMA_MAX_LAYERS define (llama-hparams.h
already defines the same value and is included by llama.cpp)
* common: fix remaining typos (accomodate, recommanded, occurences, occassionally)
- accomodate -> accommodate in src/llama.cpp comment
- recommanded -> recommended in quantize.cpp user-facing output
- occurences -> occurrences in test-chat.cpp JSON string
- occassionally -> occasionally in vendor/stb/stb_image_resize2.h comment
Note: tokenizer.ggml.seperator_token_id kept as-is to match GGUF spec
* common: remove duplicate help entries
- remove the duplicate '--reasoning-budget N' help entry that was repeated
in the main section (introduced in e0596bf614 'Autoparser - complete
refactoring of parser architecture (PR 1376)')
- remove the second '--parallel-tool-calls' help entry that advertised the
'-ptc' short flag, which belongs to '--print-token-count' (introduced in
e0596bf614 'Autoparser - complete refactoring of parser architecture
(PR 1376)'); the '-ptc' alias was non-functional for '--parallel-tool-calls'
because the parser only binds it to '--print-token-count'
The canonical help entries are kept:
- '--reasoning-budget N' is listed once
- '--parallel-tool-calls' is listed once (without the conflicting '-ptc' alias)
* common: remove duplicate LOG_ENABLE define
- the '#undef LOG_ENABLE / #define LOG_ENABLE() // dummy stub' pair was
repeated verbatim inside the LOG_DISABLE_LOGS section
- remove the second occurrence (introduced in a2588b53e1 'main : log
file (PR 2748)')
* llama-bench: align MLA and attention-max-batch flags with common tools
llama-bench used '--mla-attn' and '--attn-max-batch' while the common
CLI parsing (common/common.cpp) uses '--mla-use' and
'--attention-max-batch' for the same features. This made the flags
inconsistent across tools.
- update the help text to advertise the canonical names
'--mla-use' and '--attention-max-batch'
- keep the old '--mla-attn' and '--attn-max-batch' spellings working
as aliases so existing scripts are not broken
The divergent names were introduced in 3e536b95b0 'Add optional MLA
(PR 188)'.
* fix typos in comments and user-facing strings
- ngram-map.cpp: 'Do we haven a existing' -> 'Do we have an existing'
(introduced in 1cb7e1bf39 'spec : add self speculative decoding,
ngram and refactor (PR 1261)')
- build_mamba.cpp: 'weigth' -> 'weight' (introduced in 8befd92ea5
'Refactor model compute graphs (PR 1651)')
- gguf-split.cpp: 'one of splits have 0 tensors' -> 'one of the splits
has 0 tensors' (introduced in 75b580db0a 'split: allow
--split-max-size option (PR 6343)')
- gguf-split.cpp: 'merged from %d split' -> 'merged from %d splits'
(introduced in 1b5523dc79 'gguf-split: split and merge gguf per
batch of tensors (PR 6135)')
- convert-llama2c-to-ggml.cpp: missing opening quote in the help line,
'(default %s\\')' -> '(default '%s\\')' (introduced in bb9ebb4394
'Adding support for llama2.c models (PR 2559)')
* harmonize British and American spelling to American English
The codebase uses American English (e.g. --embd-normalize, --color),
but a few strings/comments still used British spellings. Unify them:
- 'normalisation' -> 'normalization' in common.h, common.cpp help text
and code comment, and llama-build-context.cpp comment
- 'colorise' -> 'colorize' in the --color help text (common.cpp)
- 'behaviour' -> 'behavior' in a chat.cpp warning and a llama.cpp comment
- also fix 'openai' -> 'OpenAI' capitalization in the embedding help
text and common.h comment (embedding output format is OpenAI-style)
* common: fix help text formatting inconsistencies
- '-smf16'/'--split-mode-f16' and '-smf32'/'--split-mode-f32' help
entries displayed hardcoded 'true'/'false' as the default value;
show the actual state derived from params.reduce_type instead
- '-no-mmad' help entry had 'fused_mmad?' without a space before the
ternary operator
- '--reasoning-tokens' help continuation lines used tab characters for
indentation while the sibling '--reasoning-format' entry uses spaces;
convert to consistent space indentation
* common: revert smf16/smf32 help text default display change
Revert the '-smf16'/'--split-mode-f16' and '-smf32'/'--split-mode-f32'
help entries back to their original hardcoded 'true'/'false' default
display. The change to derive the default from params.reduce_type was
not desired; the split-mode options are legacy and the hardcoded
defaults reflect their intended meaning.
The other formatting fixes in the same area (fused_mmad ternary
spacing and the reasoning-tokens tab-to-space indentation) are kept.
* llama-bench: fix help text column alignment
The --mla-use and --attention-max-batch help lines introduced by the
flag alignment landed one column off from the sibling entries
((default: at column 51 instead of 50). Adjust the padding so all
help lines align.
* common: fix help text defaults for graph-reduce-type and log-format
Mismatch 1: -grt, --graph-reduce-type help shows default "f32", but actual default (common.h:463) is "f16" and llama.cpp uses GGML_TYPE_F16.
Mismatch 2: --log-format help shows default "json", but actual default (common.h:536 log_json=false) is text.
* common: add -ptcall short flag for --parallel-tool-calls
* typo
* server: spec checkpoints for recurrent models
* fix: save/restore sampler state during speculative checkpoint
When speculative decoding rejects draft tokens and restores the
recurrent state checkpoint, the sampler (RNG, grammar, prev tokens)
must also be restored to maintain consistency. Without this, the
sampler state reflects the rejected draft tokens, leading to
potential divergence.
Uses common_sampler_clone() to snapshot the sampler before the
speculative batch decode, and restores it on rejection.
* server: snapshot recurrent state in tensor
* reset ngram mod state for rejected tokens
* server: refactor checkpoint state logic
* speculative: fix sampler for checkpoints
* recurrent model: implement recurrent kernel checkpoint
* recurrent model: refactor api
* spec: free rbudget before overwriting
* spec : add self speculative decoding and ngram-mod and refactor
common : use common_ prefix for common library function
llama : use LLAMA_TOKEN_NULL
spec : add self speculative decoding (no draft model required) + refactor
spec : add ngram-mod
spec : various improvements ton ngram-map + docs
spec : fix the check-rate logic of ngram-simple
common : add common_speculative_is_compat()
spec : simplify time measurement using common_time_meas
refactor common_sampler_init
refactor common_token_to_piece
refactor and fix cur_p bug
clean up
* spec : remove check rate
* spec: show warnings instead of abort
---------
Co-authored-by: firecoperana <firecoperana>
Co-authored-by: Sascha Rogmann <59577610+srogmann@users.noreply.github.com>