* vulkan : use ggml_row_size for types with a per-row scale
Types that declare a row_meta_size store a per-row scale ahead of the row's
blocks, so a row is not ggml_type_size()*ne/ggml_blck_size() bytes. This
under-sized src0 in the four quantized mat-mul paths, and made
ggml_vk_dim01_contiguous() report such a tensor non-contiguous, which in turn
made supports_op reject it. No change for row_meta_size == 0.
* vulkan : add IQ4_KS and IQ4_KT support
A row of these types is one f32 scale followed by the row's blocks, so rows are
not a whole number of blocks apart and the usual block-indexed addressing does
not work. They are read through a uint32_t alias of binding 0 and addressed by
word; types.comp holds the alias, the stride and the decode, so each shader only
expresses its own addressing and no push constant layouts change.
Covers to_fp16, get_rows, mul_mat_vec (incl. MUL_MAT_ID) and scalar + coopmat1
mul_mm. get_rows addresses by row rather than through nb01/02/03, which cannot
express a per-row scale, so supports_op accepts only a contiguous src0 for these
two types. coopmat2 is excluded because coopMatLoadTensorNV addresses through a
uniform grid tensor layout, which cannot describe the row prefix; the two
mat-mat getters return nullptr there and the callers fall back to F16.
* tests : add IQ4_KS/IQ4_KT decode validation
Re-implements in C++ the indexing each of the four shader families uses and
diffs it against ggml's to_float over several row and block counts. CPU only: it
validates the format transcription, not the compiled shaders.
* Load standalone Qwen3.5 MTP GGUFs passed with -md
A predictor-only MTP GGUF reports the full block count (n_main +
nextn_predict_layers) but only ships the NextN block, so loading one
with -md failed:
check_tensor_dims: tensor 'blk.0.attn_norm.weight' not found
create_qwen35_tensors() and create_qwen35moe_tensors() create every
main block as required. Detect the predictor-only case the same way
create_step35_tensors() does and mark the absent blocks
TENSOR_SKIP|TENSOR_NOT_REQUIRED.
Qwen3.5 also has to use the common MTP package contract, otherwise the
predictor-only GGUF is never classified as a companion, and the target
is not classified TARGET_ONLY - which is what makes it export the
hidden states the companion consumes.
The remaining two hunks cover cases the above newly reaches: a
predictor-only GGUF passed as -m now loads far enough to abort in the
graph builder, and its empty main blocks reach split_recurrent_tensors()
under -sm graph.
* Qwen3.5 MTP: require q_proj in predictor-only GGUFs, check companion arch
Review follow-up.
A dense NextN block loads q_proj as optional because it can be shared
with the last main block. A predictor-only GGUF has no main blocks, so
one built that way loaded with wq == nullptr and then hung. Require the
tensor in that case so the load fails naming it. eh_proj, attn_q and the
MLP are all optional on that block, so the tail probe stays on enorm,
which is required - the comment there said only eh_proj.
Adding Qwen3.5 to the common MTP package contract also made
common_speculative_has_recognized_mtp_companion() accept any GGUF
classified COMPANION, with no architecture check of the kind the Step
and DeepSeek branches have. Add it, plus the predictor count. Dense and
MoE are separate architectures, so the comparison is on the arch itself.
OpenAI-compatible clients (e.g. pi coding agent) send
max_completion_tokens for the output token cap on custom
openai-completions providers. The server only read n_predict and
max_tokens, so the cap was silently dropped and n_predict fell back
to -1 (unlimited). This allowed runaway generations of 40k+ tokens
on long agent sessions.
Matches upstream llama.cpp behavior where max_completion_tokens is an
alias of n_predict (tools/server/server-schema.cpp).
* Update parameters.md
- Add new parameters
- Update modified parameters
- Add graph parallel new arch
- Fix some words case
* Update README.md
- Update supported models list
- Add the new features
* model: Ling-3.0 (bailingmoe3) runtime support
* model: Ling-3.0-tiny support
---------
Co-authored-by: Joel Farthing <262452229+joelfarthing@users.noreply.github.com>
The draft block was seeded at last_target_pos (the newest committed
feature row = id_last's predecessor), placing the whole block one
position early vs mainline's [id_last @ n_past, ...] convention and
colliding the seed with the newest cross-KV row. Shift the common-side
batch and the graph-side SWA mask base coherently.
* Adding Muse-Glimmer support
* Different rms_eps for post norm ops
* Need attn_post_norm split for Muse-Flimmer
* WIP: split mode graph
* Forgot this file
* Clean it up
* Minor
* Muse-glimmer: Slightly better split mode graph (+2% TG)
build_llama passes nullptr for the rope_factors_in argument, so
rope_freqs.weight never reaches ggml_rope_ext and llama3 rope frequency
scaling is not applied. The LLAMA_SPLIT_MODE_GRAPH branch inside
build_std_attention falls back to model.layers[il].rope_freqs, so only
the standard path is affected.
The DFLASH case was inserted into the NEOX fall-through group in
llama_rope_type(), so the 39 cases above it now return LLAMA_ROPE_TYPE_NORM
instead of LLAMA_ROPE_TYPE_NEOX.
Move DFLASH into the NORM group so it keeps its intended rope type and the
others fall through to NEOX again.
* CUDA indexer topk: this is better for PP
* Don't overstep
* Cleanup
* Allow Q8_0 cache in the CUDA DSA implementation
* DS4: do not cast caches to f32
* Fix massive inefficiency in CUDA Q->f32/f16 and f32/f16->Q copies
* Re-enable -ictk | --indexer-cache-type-k
* CUDA indexer topk: this is better for PP
* Don't overstep
* Cleanup
* Allow Q8_0 cache in the CUDA DSA implementation
* DS4: do not cast caches to f32