* Support MiMo DFlash draft conversion * Fix MiMo2 DFlash capture row pruning * Fix MiMo DFlash draft RoPE and value scale * Honor partial_rotary_factor in DFlash draft RoPE dim count The draft set rope.dimension_count to the full head_dim (128), ignoring the MiMo DFlash draft's partial_rotary_factor=0.5. The correct count is head_dim*partial_rotary_factor=64; the remaining dims are NoPE. With the full head_dim the upper half of each head receives position rotation it was never trained for, which roughly halves draft acceptance on code (~26% -> ~60% once corrected). RoPE base (5e6) and value scale (0.612) were already correct. * Filter weight-map shard discovery to files that exist get_model_part_names_from_weight_map() returned shard names straight from the index weight_map without checking they exist. A model dir with a stale model.safetensors.index.json but no safetensors shards would then set is_safetensors=True and skip the pytorch_model*.bin fallback, failing later when opening the missing files. Filter to shards present on disk so a stale index falls through to the other weight formats. * DFlash: store backbone_rotary_base in dedicated GGUF key backbone_rotary_base (the target model's RoPE theta used when encoding context K/V) was written to rope.freq_base, clobbering the draft model's own rope_theta. For MiMo this swapped 10000 → 5000000 in the draft attention path. Fix: write backbone_rotary_base to a dedicated dflash.backbone_rotary_base GGUF key and read it into hparams.dflash_backbone_rotary_base. In build_dflash_kv_cache, use target_freq_base (the new hparam when set, falling back to freq_base) for the context-K RoPE call. The draft model's own rope.freq_base is now set correctly from rope_theta. Existing MiMo DFlash GGUFs must be reconverted. --------- Co-authored-by: Joel Farthing <262452229+joelfarthing@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| constants.py | ||
| gguf.py | ||
| gguf_reader.py | ||
| gguf_writer.py | ||
| lazy.py | ||
| metadata.py | ||
| py.typed | ||
| quants.py | ||
| tensor_mapping.py | ||
| utility.py | ||
| vocab.py | ||