hybrid-llama/turboquant/tools/gguf-split
Marvin 1dd0700988 hybrid-llama: merge ik_llama IQK CPU GEMM into TurboQuant fork
Base: AtomicBot-ai/atomic-llama-cpp-turboquant @ cd5609390. IQK source: ikawrakow/ik_llama.cpp @ fe215a8c (ggml/src/iqk only).

- GGML_IQK_MUL_MAT / GGML_IQK_FLASH_ATTENTION options (default OFF)

- 57 IQK repacked types, blocks, traits; IQK hooks in ggml_compute_forward_mul_mat

- ggml-cpu with IQK ON builds and links; IQK OFF build unaffected

Assisted-by: opencode (Muse Spark)
2026-09-05 18:09:49 -03:00
..
CMakeLists.txt hybrid-llama: merge ik_llama IQK CPU GEMM into TurboQuant fork 2026-09-05 18:09:49 -03:00
README.md hybrid-llama: merge ik_llama IQK CPU GEMM into TurboQuant fork 2026-09-05 18:09:49 -03:00
gguf-split.cpp hybrid-llama: merge ik_llama IQK CPU GEMM into TurboQuant fork 2026-09-05 18:09:49 -03:00
tests.sh hybrid-llama: merge ik_llama IQK CPU GEMM into TurboQuant fork 2026-09-05 18:09:49 -03:00

README.md

GGUF split Example

CLI to split / merge GGUF files.

Command line options:

  • --split: split GGUF to multiple GGUF, default operation.
  • --split-max-size: max size per split in M or G, f.ex. 500M or 2G.
  • --split-max-tensors: maximum tensors in each split: default(128)
  • --merge: merge multiple GGUF to a single GGUF. You only need to specify the name of the first GGUF to merge, the name of the merged GGUF, and the CLI will find the other GGUFs it needs within the same folder.