ik_llama_opt/gguf-py/gguf
Joel Farthing 29a54f4b04
DFlash: support MiMo-V2.5-Pro draft conversion and runtime (#2048)
* Support MiMo DFlash draft conversion

* Fix MiMo2 DFlash capture row pruning

* Fix MiMo DFlash draft RoPE and value scale

* Honor partial_rotary_factor in DFlash draft RoPE dim count

The draft set rope.dimension_count to the full head_dim (128), ignoring the
MiMo DFlash draft's partial_rotary_factor=0.5. The correct count is
head_dim*partial_rotary_factor=64; the remaining dims are NoPE. With the full
head_dim the upper half of each head receives position rotation it was never
trained for, which roughly halves draft acceptance on code (~26% -> ~60% once
corrected). RoPE base (5e6) and value scale (0.612) were already correct.

* Filter weight-map shard discovery to files that exist

get_model_part_names_from_weight_map() returned shard names straight from the
index weight_map without checking they exist. A model dir with a stale
model.safetensors.index.json but no safetensors shards would then set
is_safetensors=True and skip the pytorch_model*.bin fallback, failing later when
opening the missing files. Filter to shards present on disk so a stale index
falls through to the other weight formats.

* DFlash: store backbone_rotary_base in dedicated GGUF key

backbone_rotary_base (the target model's RoPE theta used when encoding
context K/V) was written to rope.freq_base, clobbering the draft
model's own rope_theta. For MiMo this swapped 10000 → 5000000 in the
draft attention path.

Fix: write backbone_rotary_base to a dedicated dflash.backbone_rotary_base
GGUF key and read it into hparams.dflash_backbone_rotary_base. In
build_dflash_kv_cache, use target_freq_base (the new hparam when set,
falling back to freq_base) for the context-K RoPE call. The draft model's
own rope.freq_base is now set correctly from rope_theta.

Existing MiMo DFlash GGUFs must be reconverted.

---------

Co-authored-by: Joel Farthing <262452229+joelfarthing@users.noreply.github.com>
2026-06-29 13:26:29 +02:00
..
__init__.py Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
constants.py DFlash: support MiMo-V2.5-Pro draft conversion and runtime (#2048) 2026-06-29 13:26:29 +02:00
gguf.py gguf-py: Refactor and allow reading/modifying existing GGUF files (#3981) 2023-11-11 08:04:50 +03:00
gguf_reader.py Make gguf-py stuff work with numpy 2.0 (#991) 2025-11-20 10:20:55 +01:00
gguf_writer.py DFlash: support MiMo-V2.5-Pro draft conversion and runtime (#2048) 2026-06-29 13:26:29 +02:00
lazy.py Merge mainline - Aug 12 2024 (#17) 2024-08-12 15:14:32 +02:00
metadata.py Merge mainline - Aug 12 2024 (#17) 2024-08-12 15:14:32 +02:00
py.typed convert : various script cleanups/fixes + merges and special token handling (#2842) 2023-08-30 11:25:50 +03:00
quants.py convert_hf_to_gguf.py : conversion from hf weights to Q6_0 (#483) 2025-06-03 09:30:30 +03:00
tensor_mapping.py DFlash: support MiMo-V2.5-Pro draft conversion and runtime (#2048) 2026-06-29 13:26:29 +02:00
utility.py Merge mainline llama.cpp (#3) 2024-07-27 07:55:01 +02:00
vocab.py model: add Cohere2-MoE North Mini Code support (#1945) 2026-06-10 15:28:27 +02:00