DFlashDraftModel.modify_tensors rewrites flat-named norm.weight and layers.N.* to their model.* form but not embed_tokens.weight, so a draft that carries its own (self-contained) token embeddings with flat naming fails with 'Can not map tensor embed_tokens.weight'. Handle it the same way as norm.weight so self-contained drafts convert.
Co-authored-by: Joel Farthing <262452229+joelfarthing@users.noreply.github.com>
The openai gpt-oss models, and the z-lab gpt-oss DFlash drafts that reuse their
tokenizer, ship the o200k_harmony BPE tokenizer, which shares the GPT-4o
pre-tokenizer regex already handled by LLAMA_VOCAB_PRE_TYPE_GPT4O. Its checksum
was unrecognized, so get_vocab_base_pre() fell through to the "BPE pre-tokenizer
was not recognized" warning and conversion failed.
Register the gpt-oss-20b tokenizer hash against the existing "gpt-4o" pre-type
in convert_hf_to_gguf_update.py and add the matching entry in
get_vocab_base_pre(), following the dual-location pattern already used for the
z-lab Qwen3.5 DFlash drafts.
Co-authored-by: Joel Farthing <262452229+joelfarthing@users.noreply.github.com>
* Support MiMo DFlash draft conversion
* Fix MiMo2 DFlash capture row pruning
* Fix MiMo DFlash draft RoPE and value scale
* Honor partial_rotary_factor in DFlash draft RoPE dim count
The draft set rope.dimension_count to the full head_dim (128), ignoring the
MiMo DFlash draft's partial_rotary_factor=0.5. The correct count is
head_dim*partial_rotary_factor=64; the remaining dims are NoPE. With the full
head_dim the upper half of each head receives position rotation it was never
trained for, which roughly halves draft acceptance on code (~26% -> ~60% once
corrected). RoPE base (5e6) and value scale (0.612) were already correct.
* Filter weight-map shard discovery to files that exist
get_model_part_names_from_weight_map() returned shard names straight from the
index weight_map without checking they exist. A model dir with a stale
model.safetensors.index.json but no safetensors shards would then set
is_safetensors=True and skip the pytorch_model*.bin fallback, failing later when
opening the missing files. Filter to shards present on disk so a stale index
falls through to the other weight formats.
* DFlash: store backbone_rotary_base in dedicated GGUF key
backbone_rotary_base (the target model's RoPE theta used when encoding
context K/V) was written to rope.freq_base, clobbering the draft
model's own rope_theta. For MiMo this swapped 10000 → 5000000 in the
draft attention path.
Fix: write backbone_rotary_base to a dedicated dflash.backbone_rotary_base
GGUF key and read it into hparams.dflash_backbone_rotary_base. In
build_dflash_kv_cache, use target_freq_base (the new hparam when set,
falling back to freq_base) for the context-K RoPE call. The draft model's
own rope.freq_base is now set correctly from rope_theta.
Existing MiMo DFlash GGUFs must be reconverted.
---------
Co-authored-by: Joel Farthing <262452229+joelfarthing@users.noreply.github.com>
* DFlash: bound intra-block draft tokens to the SWA window
The SWA mask builder applied the sliding-window distance check only to the
cross-context section; the intra-block draft-token loop masked causal-only,
so a draft token could attend to earlier block tokens beyond n_swa. Apply
the same window bound ((j - block_k) < swa_window) in both the F16 and F32
paths so it matches the cross-context section.
Behavior-neutral for dense models: the SWA mask tensor is only allocated
when the model has SWA layers (build_dflash.cpp needs_swa_mask gate), so
for dense targets the changed block is unreachable.
* DFlash: enable sliding-window attention for draft models
DFlash drafts can be trained with sliding-window attention for long context,
but the runtime ignored it: the draft loader never read the window keys and the
converter never emitted them, so SWA-trained drafts always ran full-attention.
Enable it end to end and fix the dormant SWA graph path it exposes:
- convert_hf_to_gguf.py (DFlashDraftModel): emit attention.sliding_window + an
all-layers sliding_window_pattern when the source config sets use_sliding_window.
- llama-hparams.cpp (LLM_ARCH_DFLASH_DRAFT): read sliding_window + pattern into
n_swa / swa_layers.
- build_dflash.cpp + llama-dflash.cpp: the SWA mask path had never run; an all-SWA
draft turned the full kq_mask into a dead graph node the scheduler never backs
with a buffer, then the input-set wrote it unconditionally (GGML_ASSERT buf!=NULL).
Create + set each mask only when a layer uses it; derive mask dims from whichever
mask is live. Dense/mixed drafts are byte-identical.
Validated on gemma-4-26B-A4B at long context (cross_ctx 8176 > window 2048): no
crash, no short-context regression, SWA-on recovers long-context draft acceptance.
* DFlash: derive draft SWA pattern from layer_types
The converter emitted an all-layers SWA pattern ([True]*n_layers). The z-lab
DFlash drafts are sliding-window on every layer except a final full-attention
(global) layer, so this ran that global layer as sliding-window and clipped its
long-context view. Read layer_types and emit the matching per-layer pattern
(sliding_attention -> True), falling back to all-SWA only when layer_types is
absent.
---------
Co-authored-by: Joel Farthing <262452229+joelfarthing@users.noreply.github.com>
* Conditionally write moe_shared_expert_intermediate_size
Ling-1T config.json does *not* have `moe_shared_expert_intermediate_size`.
Ling-flash-2.0a *does* have it.
This small patch just makes the gguf_writer conditionally detect as
needed.
* Fix Ling-1T missing moe_shared_expert_intermediate_size
Thanks CISC for the proper patch to include the needed values!
* convert_hf_to_gguf for Kimi-K2-Instruct
Adapt mainline `PR14653` for tokenizer while maintaining proper MLA
tensors. Tested with this workflow using deepseek fp8_cast_bf16.py and
triton-cpu to upcast the fp8 safetensors to bf16 safetensors then used
this convert_hf_to_gguf.
* Add Kimi-K2 chat template
moonshotai/Kimi-K2-Instruct
https://github.com/ikawrakow/ik_llama.cpp/pull/609#issuecomment-3071259454
* kimi-k2 add ass to template to get response
* Add llama.cpp changes for dots1 support
* Add python changes for dots1 support
* Fix to make it convert
* Remove V reshaping, remove BOS by default for dots1 and fix warmup to handle models without BOS
* Minor fix
* Remove commented lines
* Legacy quants conversion schemes in convert_hf_to_gguf.py
This, notably in order to make smaller conversions to generate an iMatrix file.
`Q4_0`,`Q4_1` are here using embeddings, output, attn_k and attn_v in q5_0.
`Q5_0`,`Q5_1` are here using embeddings, output, attn_k and attn_v in q8_0.
Adapted from the following llama.cpp mainline PR : https://github.com/ggml-org/llama.cpp/pull/9022
Original author @chentyjpm
Also, 2 forgotten mentions of FTYPE IQ3_KL in llama.cpp file.
* forgotten IQ5_KS case mention
* lora : fix llama conversion script with ROPE_FREQS
* convert : refactor rope_freqs generation
This should also fix vocab-only conversion for Phi-3.
* convert : adapt MiniCPM3 to separate rope_freqs insertion
MiniCPM3's tokenizer is treated as a SentencePiece tokenizer to avoid
having to run its custom Python code which mixes tokenization
in the same file as tool calls.
gguf-py : add long and short RoPE factors to tensor mappings
Empty, but the key names are used to populate the mappings.
---------
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
Co-authored-by: Francis Couture-Harpin <git@compilade.net>
* conflict resolution
* Changes to make work and add longrope support
* Changes to n_attention_wv rule
* Untested support of 253B
* DeciLMCausalModel now reads rope_theta from config.json properly
* Remove errant Granite mentions
* Better n_attention_vw rule
* Update vocab.py
---------
Co-authored-by: Yee Man Chan <ymchan@gmail.com>
Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
* Deepseek MLA Optimizations
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
* Make MLA optional
* Remove some unnecessary copies in the MLA attention
* Deepseek MLA Optimizations V2 (#195)
* Avoid allocating MHA KV cache when MLA is turned on
* Added missing gguf-py file
* Added final optimizations
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
* Make sure we do have wk_b and wv_b before enabling MLA
---------
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
* Use type_k and type_v to set the types of the MLA caches
They were hard-coded at f16.
On my Ryzen-7950X with native bf16 support I get a fairly
significant PP performance boost with bf16 KV-cache:
PP-4096 = 320 t/s up from 292 t/s with fp16 KV-cache.
* Better gemm strategy when nth > nhead
It gives a ~10% PP performance boost for DeepSeek-Lite with 32 threads
(with or without MLA).
Before this commit, when nth > nhead heads were processed
sequentially with all nth threads participating in each
matrix multiplication. Now we ind the gcd of nhead and
nth and split threads into nth/gcd groups, each group
processing nhead/gcd heads.
---------
Co-authored-by: Saood Karim <saood05@gmail.com>
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
* Merging mainline - WIP
* Merging mainline - WIP
AVX2 and CUDA appear to work.
CUDA performance seems slightly (~1-2%) lower as it is so often
the case with llama.cpp/ggml after some "improvements" have been made.
* Merging mainline - fix Metal
* Remove check
---------
Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>