* Load standalone Qwen3.5 MTP GGUFs passed with -md
A predictor-only MTP GGUF reports the full block count (n_main +
nextn_predict_layers) but only ships the NextN block, so loading one
with -md failed:
check_tensor_dims: tensor 'blk.0.attn_norm.weight' not found
create_qwen35_tensors() and create_qwen35moe_tensors() create every
main block as required. Detect the predictor-only case the same way
create_step35_tensors() does and mark the absent blocks
TENSOR_SKIP|TENSOR_NOT_REQUIRED.
Qwen3.5 also has to use the common MTP package contract, otherwise the
predictor-only GGUF is never classified as a companion, and the target
is not classified TARGET_ONLY - which is what makes it export the
hidden states the companion consumes.
The remaining two hunks cover cases the above newly reaches: a
predictor-only GGUF passed as -m now loads far enough to abort in the
graph builder, and its empty main blocks reach split_recurrent_tensors()
under -sm graph.
* Qwen3.5 MTP: require q_proj in predictor-only GGUFs, check companion arch
Review follow-up.
A dense NextN block loads q_proj as optional because it can be shared
with the last main block. A predictor-only GGUF has no main blocks, so
one built that way loaded with wq == nullptr and then hung. Require the
tensor in that case so the load fails naming it. eh_proj, attn_q and the
MLP are all optional on that block, so the tail probe stays on enorm,
which is required - the comment there said only eh_proj.
Adding Qwen3.5 to the common MTP package contract also made
common_speculative_has_recognized_mtp_companion() accept any GGUF
classified COMPANION, with no architecture check of the kind the Step
and DeepSeek branches have. Add it, plus the predictor count. Dense and
MoE are separate architectures, so the comparison is on the arch itself.
OpenAI-compatible clients (e.g. pi coding agent) send
max_completion_tokens for the output token cap on custom
openai-completions providers. The server only read n_predict and
max_tokens, so the cap was silently dropped and n_predict fell back
to -1 (unlimited). This allowed runaway generations of 40k+ tokens
on long agent sessions.
Matches upstream llama.cpp behavior where max_completion_tokens is an
alias of n_predict (tools/server/server-schema.cpp).
* Update parameters.md
- Add new parameters
- Update modified parameters
- Add graph parallel new arch
- Fix some words case
* Update README.md
- Update supported models list
- Add the new features
* model: Ling-3.0 (bailingmoe3) runtime support
* model: Ling-3.0-tiny support
---------
Co-authored-by: Joel Farthing <262452229+joelfarthing@users.noreply.github.com>
The draft block was seeded at last_target_pos (the newest committed
feature row = id_last's predecessor), placing the whole block one
position early vs mainline's [id_last @ n_past, ...] convention and
colliding the seed with the newest cross-KV row. Shift the common-side
batch and the graph-side SWA mask base coherently.
* Adding Muse-Glimmer support
* Different rms_eps for post norm ops
* Need attn_post_norm split for Muse-Flimmer
* WIP: split mode graph
* Forgot this file
* Clean it up
* Minor
* Muse-glimmer: Slightly better split mode graph (+2% TG)
build_llama passes nullptr for the rope_factors_in argument, so
rope_freqs.weight never reaches ggml_rope_ext and llama3 rope frequency
scaling is not applied. The LLAMA_SPLIT_MODE_GRAPH branch inside
build_std_attention falls back to model.layers[il].rope_freqs, so only
the standard path is affected.
The DFLASH case was inserted into the NEOX fall-through group in
llama_rope_type(), so the 39 cases above it now return LLAMA_ROPE_TYPE_NORM
instead of LLAMA_ROPE_TYPE_NEOX.
Move DFLASH into the NORM group so it keeps its intended rope type and the
others fall through to NEOX again.
* CUDA indexer topk: this is better for PP
* Don't overstep
* Cleanup
* Allow Q8_0 cache in the CUDA DSA implementation
* DS4: do not cast caches to f32
* Fix massive inefficiency in CUDA Q->f32/f16 and f32/f16->Q copies
* Re-enable -ictk | --indexer-cache-type-k
* CUDA indexer topk: this is better for PP
* Don't overstep
* Cleanup
* Allow Q8_0 cache in the CUDA DSA implementation
* DS4: do not cast caches to f32