390 lines
16 KiB
Markdown
390 lines
16 KiB
Markdown
# IQK + TurboQuant Merge: Complete Coding Handoff
|
|
|
|
## Objective
|
|
|
|
Merge ik_llama's IQK CPU GEMM optimization layer into the AtomicBot-ai/atomic-llama-cpp-turboquant fork to get ~5x prefill speedup on Ling 3.0 Flash (bailingmoe3) while retaining TurboQuant's bugfixes (tool calls, token handling, KV cache compression).
|
|
|
|
## Repository URLs
|
|
|
|
- **TurboQuant (target)**: https://github.com/AtomicBot-ai/atomic-llama-cpp-turboquant
|
|
- **ik_llama (source)**: https://github.com/ikawrakow/ik_llama.cpp
|
|
|
|
## Expected Outcome
|
|
|
|
- Prefill: 60 t/s → ~300 t/s (5x improvement)
|
|
- Decode: unchanged (~9-11 t/s on DDR4)
|
|
- All TurboQuant features intact (KV cache compression, bailingmoe3 fixes, MTP/NextN)
|
|
|
|
## Architecture Comparison
|
|
|
|
| Aspect | TurboQuant | ik_llama |
|
|
|---|---|---|
|
|
| Code structure | Modular (separate backend libs) | Monolithic (single ggml.c) |
|
|
| CPU backend file | `ggml/src/ggml-cpu/ggml-cpu.c` (4K lines) | `ggml/src/ggml.c` (32K lines) |
|
|
| Forward dispatch | `void ggml_compute_forward(params, tensor)` | `int ggml_compute_forward(params, tensor, cgraph, node_n)` |
|
|
| Params struct | `params->threadpool` (ggml-cpu-impl.h) | `params->shared` (inline in ggml.c) |
|
|
| Type traits | `type_traits_cpu[]` | `type_traits[]` |
|
|
| IQK present? | No | Yes |
|
|
| KV cache compression | Yes (TURBO types) | No |
|
|
| bailingmoe3 | Yes (bugfixed) | May have issues |
|
|
|
|
## Why They Don't Conflict
|
|
|
|
The two forks modify different layers:
|
|
- **TurboQuant**: Higher-level (KV cache quantization types, architecture-specific fixes, MTP speculative decoding)
|
|
- **IQK**: Lower-level (CPU GEMM primitives in `ggml/src/iqk/`)
|
|
|
|
IQK accelerates `ggml_mul_mat` — the primitive every model calls. TurboQuant's changes are above that layer. The KV cache compression (TURBO types) and Hadamard fast path are in different code regions.
|
|
|
|
## Step 1: Copy IQK Source Files
|
|
|
|
Copy the entire `ggml/src/iqk/` directory from ik_llama into TurboQuant:
|
|
|
|
```bash
|
|
# From the TurboQuant repo root
|
|
git remote add ik_llama https://github.com/ikawrakow/ik_llama.cpp
|
|
git fetch ik_llama
|
|
git checkout ik_llama/master -- ggml/src/iqk/
|
|
```
|
|
|
|
This brings in 36 files (~2.1 MB total):
|
|
|
|
**Core files (always compiled):**
|
|
- `iqk/iqk_config.h` (1.7 KB) — build config
|
|
- `iqk/iqk_common.h` (42 KB) — shared utilities
|
|
- `iqk/iqk_utils.h` (12 KB) — helper macros
|
|
- `iqk/iqk_quantize.cpp` (439 KB) — weight quantization
|
|
- `iqk/iqk_quantize.h` (34 KB) — quantization API
|
|
- `iqk/iqk_cpu_ops.cpp` (43 KB) — SIMD-accelerated CPU ops (top-k, argsort, etc.)
|
|
- `iqk/iqk_cpu_ops.h` (2 KB) — ops API
|
|
|
|
**GEMM files (compiled when GGML_IQK_MUL_MAT=ON):**
|
|
- `iqk/iqk_mul_mat.cpp` (99 KB) — main dispatch + node fusion
|
|
- `iqk/iqk_mul_mat.h` (4.5 KB) — matmul API
|
|
- `iqk/iqk_kda.cpp` (17 KB) — KDA (Gated Delta Net) fused path
|
|
- `iqk/iqk_flash_attn.cpp` (32 KB) — CPU flash attention
|
|
- `iqk/iqk_flash_impl.h` (2 KB) — FA implementation
|
|
- `iqk/iqk_gemm_kquants.cpp` (250 KB) — Q2_K through Q6_K kernels
|
|
- `iqk/iqk_gemm_kquants.h` (0.4 KB)
|
|
- `iqk/iqk_gemm_ktquants.cpp` (126 KB) — Trellis quant kernels
|
|
- `iqk/iqk_gemm_ktquants.h` (0.3 KB)
|
|
- `iqk/iqk_gemm_iquants.cpp` (202 KB) — IQ2_XXS through IQ4_XS kernels
|
|
- `iqk/iqk_gemm_iquants.h` (0.3 KB)
|
|
- `iqk/iqk_gemm_iqk_quants.cpp` (288 KB) — IQ2_K through IQ6_K kernels
|
|
- `iqk/iqk_gemm_iqk_quants.h` (0.3 KB)
|
|
- `iqk/iqk_gemm_legacy_quants.cpp` (172 KB) — Q4_0, Q4_1, Q5_0, Q5_1, Q8_0
|
|
- `iqk/iqk_gemm_legacy_quants.h` (0.5 KB)
|
|
- `iqk/iqk_gemm_floats.cpp` (48 KB) — FP16/FP32/BF16 kernels
|
|
- `iqk/iqk_gemm_floats.h` (0.3 KB)
|
|
- `iqk/iqk_gemm_1bit.cpp` (188 KB) — Bitnet kernels
|
|
- `iqk/iqk_gemm_1bit.h` (0.3 KB)
|
|
|
|
**Flash attention template files:**
|
|
- `iqk/fa/iqk_fa_templates.h` — FA template definitions
|
|
- `iqk/fa/iqk_fa_64_64.cpp` through `iqk_fa_576_512.cpp` (9 files, ~1.6 KB each) — size-specialized FA kernels
|
|
|
|
## Step 2: CMake Integration
|
|
|
|
### 2a. Add cmake option
|
|
|
|
In `ggml/src/CMakeLists.txt` (or create a new section), add:
|
|
|
|
```cmake
|
|
option(GGML_IQK_MUL_MAT "Enable optimized IQK matrix multiplications" OFF)
|
|
option(GGML_IQK_FLASH_ATTENTION "Enable IQK Flash Attention kernels" OFF)
|
|
```
|
|
|
|
### 2b. Add IQK sources to ggml-cpu backend
|
|
|
|
In `ggml/src/ggml-cpu/CMakeLists.txt`, inside the `ggml_add_cpu_backend_variant_impl()` function, before the `target_sources(${GGML_CPU_NAME} PRIVATE ${GGML_CPU_SOURCES})` line, add:
|
|
|
|
```cmake
|
|
# IQK sources — unconditional (quantize + CPU ops)
|
|
set(GGML_CPU_IQK_SOURCES
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_quantize.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_cpu_ops.cpp
|
|
)
|
|
set(GGML_CPU_IQK_HEADERS
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_config.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_common.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_utils.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_quantize.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_cpu_ops.h
|
|
)
|
|
|
|
if(GGML_IQK_MUL_MAT)
|
|
list(APPEND GGML_CPU_IQK_SOURCES
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_mul_mat.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_kda.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_flash_attn.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_floats.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_kquants.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_ktquants.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_iquants.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_iqk_quants.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_1bit.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_legacy_quants.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/fa/iqk_fa_64_64.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/fa/iqk_fa_96_96.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/fa/iqk_fa_128_128.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/fa/iqk_fa_192_128.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/fa/iqk_fa_192_192.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/fa/iqk_fa_256_256.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/fa/iqk_fa_320_256.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/fa/iqk_fa_512_512.cpp
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/fa/iqk_fa_576_512.cpp
|
|
)
|
|
list(APPEND GGML_CPU_IQK_HEADERS
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_mul_mat.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_flash_impl.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/fa/iqk_fa_templates.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_floats.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_kquants.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_ktquants.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_iquants.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_iqk_quants.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_1bit.h
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk/iqk_gemm_legacy_quants.h
|
|
)
|
|
endif()
|
|
|
|
target_sources(${GGML_CPU_NAME} PRIVATE
|
|
${GGML_CPU_SOURCES}
|
|
${GGML_CPU_IQK_SOURCES}
|
|
)
|
|
target_sources(${GGML_CPU_NAME} PRIVATE FILE_SET HEADERS
|
|
${GGML_CPU_HEADERS}
|
|
${GGML_CPU_IQK_HEADERS}
|
|
)
|
|
target_include_directories(${GGML_CPU_NAME} PRIVATE
|
|
${CMAKE_CURRENT_SOURCE_DIR}/../iqk
|
|
)
|
|
```
|
|
|
|
### 2c. Add compile definitions
|
|
|
|
In the same `ggml_add_cpu_backend_variant_impl()` function:
|
|
|
|
```cmake
|
|
if(GGML_IQK_MUL_MAT)
|
|
target_compile_definitions(${GGML_CPU_NAME} PRIVATE GGML_USE_IQK_MULMAT)
|
|
endif()
|
|
if(GGML_IQK_FLASH_ATTENTION)
|
|
target_compile_definitions(${GGML_CPU_NAME} PRIVATE GGML_IQK_FLASH_ATTENTION)
|
|
endif()
|
|
```
|
|
|
|
### 2d. Top-level CMakeLists.txt
|
|
|
|
In the root `CMakeLists.txt`, add the options and pass them through:
|
|
|
|
```cmake
|
|
option(GGML_IQK_MUL_MAT "Enable optimized IQK matrix multiplications" OFF)
|
|
option(GGML_IQK_FLASH_ATTENTION "Enable IQK Flash Attention kernels" OFF)
|
|
```
|
|
|
|
And ensure they're passed to the ggml subdirectory via `add_subdirectory(ggml)` or `set()`.
|
|
|
|
## Step 3: Adapt the Forward Dispatch (THE HARD PART)
|
|
|
|
This is where the structural mismatch matters. In ik_llama, IQK hooks into `ggml_compute_forward_mul_mat()` in `ggml.c`. In TurboQuant, that function lives in `ggml-cpu/ggml-cpu.c` with a different signature.
|
|
|
|
### 3a. Locate the dispatch function
|
|
|
|
In TurboQuant's `ggml/src/ggml-cpu/ggml-cpu.c`, find `ggml_compute_forward_mul_mat()`. This is where MUL_MAT operations are dispatched.
|
|
|
|
### 3b. Port the IQK hooks
|
|
|
|
From ik_llama's `ggml.c` lines 18016-18111, port these three code blocks into TurboQuant's `ggml_compute_forward_mul_mat()`:
|
|
|
|
**Block 1: Early exit fast path (ik_llama lines 18016-18024)**
|
|
```c
|
|
#ifdef GGML_USE_IQK_MULMAT
|
|
// When src1 is already in vec_dot_type and dst is F32, use IQK directly
|
|
if (src1->type == vec_dot_type && dst->type == GGML_TYPE_F32) {
|
|
iqk_mul_mat_4d(ne01, ne11, ne00, ne02, ne03, ne12, ne13,
|
|
src0->data, src1->data, dst->data,
|
|
src0->type, vec_dot_type,
|
|
params->ith, params->nth);
|
|
return;
|
|
}
|
|
#endif
|
|
```
|
|
|
|
**Block 2: Custom quantization (ik_llama lines 18039-18047)**
|
|
```c
|
|
#ifdef GGML_USE_IQK_MULMAT
|
|
// Replace standard from_float with IQK's faster quantization
|
|
if (src1->type != vec_dot_type) {
|
|
iqk_quantize_any(src1->type, vec_dot_type, ne10, ne11, ne12, ne13,
|
|
src1->data, wdata, params->ith, params->nth);
|
|
}
|
|
#else
|
|
// Original llama.cpp quantization path
|
|
ggml_quantize_mat(src1->type, vec_dot_type, ne10, ne11, ne12, ne13,
|
|
src1->data, wdata, params->ith, params->nth);
|
|
#endif
|
|
```
|
|
|
|
**Block 3: Main matmul + node fusion (ik_llama lines 18083-18111)**
|
|
```c
|
|
#ifdef GGML_USE_IQK_MULMAT
|
|
iqk_mul_mat_4d(ne01, ne11, ne00, ne02, ne03, ne12, ne13,
|
|
src0->data, wdata, dst->data,
|
|
vec_dot_type, GGML_TYPE_F32,
|
|
params->ith, params->nth);
|
|
#else
|
|
// Original ggml_compute_forward_mul_mat_batched path
|
|
...
|
|
#endif
|
|
```
|
|
|
|
### 3c. Adapt params->shared to params->threadpool
|
|
|
|
ik_llama's IQK uses `params->shared` for barrier and abort callbacks. TurboQuant uses `params->threadpool`. You need to:
|
|
|
|
1. Check if `iqk_mul_mat_4d()` and related functions actually use `params->shared` (they likely don't for the basic path — the shared struct is mainly for synchronization in multi-threaded dispatch)
|
|
2. If they do, create an adapter that wraps `params->threadpool` into the interface IQK expects
|
|
3. If they don't, this is a non-issue
|
|
|
|
### 3d. Adapt type_traits
|
|
|
|
ik_llama modifies `type_traits[]` entries to use `GGML_TYPE_Q8_2_X4` as `vec_dot_type` for q4_0, q4_1, q5_0, q5_1, q6_0, q8_0 (ik_llama ggml.c lines 714-856).
|
|
|
|
In TurboQuant, the equivalent is `type_traits_cpu[]` in `ggml-cpu/ggml-cpu.c` or `ggml-cpu/traits.cpp`. Add the same modifications:
|
|
|
|
```c
|
|
#ifdef GGML_USE_IQK_MULMAT
|
|
// Override vec_dot_type for legacy quants to use IQK's repacked format
|
|
type_traits_cpu[GGML_TYPE_Q4_0].vec_dot_type = GGML_TYPE_Q8_2_X4;
|
|
type_traits_cpu[GGML_TYPE_Q4_1].vec_dot_type = GGML_TYPE_Q8_2_X4;
|
|
type_traits_cpu[GGML_TYPE_Q5_0].vec_dot_type = GGML_TYPE_Q8_2_X4;
|
|
type_traits_cpu[GGML_TYPE_Q5_1].vec_dot_type = GGML_TYPE_Q8_2_X4;
|
|
type_traits_cpu[GGML_TYPE_Q6_0].vec_dot_type = GGML_TYPE_Q8_2_X4;
|
|
type_traits_cpu[GGML_TYPE_Q8_0].vec_dot_type = GGML_TYPE_Q8_2_X4;
|
|
#endif
|
|
```
|
|
|
|
## Step 4: Add IQK Includes
|
|
|
|
In `ggml-cpu/ggml-cpu.c`, add at the top:
|
|
|
|
```c
|
|
#ifdef GGML_USE_IQK_MULMAT
|
|
#include "iqk/iqk_quantize.h"
|
|
#include "iqk/iqk_cpu_ops.h"
|
|
#include "iqk/iqk_mul_mat.h"
|
|
#include "iqk/iqk_config.h"
|
|
#endif
|
|
```
|
|
|
|
## Step 5: Node Fusion (Optional, Advanced)
|
|
|
|
ik_llama fuses consecutive MUL_MAT nodes sharing the same src1 (lines 18083-18111). This requires the `int ggml_compute_forward()` signature that returns an updated node index.
|
|
|
|
**Option A (recommended): Skip node fusion initially.** Get the basic IQK fast path working first. The 5x prefill speedup comes primarily from the optimized GEMM kernels, not from node fusion. Node fusion adds maybe 10-20% on top.
|
|
|
|
**Option B (full port):** Modify TurboQuant's forward dispatch signature to match ik_llama's. This requires:
|
|
1. Changing `void ggml_compute_forward(...)` to `int ggml_compute_forward(...)` in ggml-cpu.c
|
|
2. Updating the dispatch loop in ggml-cpu.c to track node_n
|
|
3. Updating all callers of ggml_compute_forward
|
|
|
|
This is invasive and risky. Do it only after the basic path works.
|
|
|
|
**Option C (hybrid):** Use TurboQuant's existing `ggml_cpu_try_fuse_ops()` mechanism (line 3162 of ggml-cpu.c) to implement IQK's MUL_MAT fusion as a separate fusion pass. This avoids changing the forward dispatch signature.
|
|
|
|
## Step 6: Build and Test
|
|
|
|
```bash
|
|
cd atomic-llama-cpp-turboquant
|
|
cmake -B build -DGGML_CUDA=ON -DGGML_IQK_MUL_MAT=ON -DCMAKE_BUILD_TYPE=Release
|
|
cmake --build build --config Release -j$(nproc)
|
|
```
|
|
|
|
### Test 1: Verify IQK is active
|
|
|
|
```bash
|
|
./build/bin/llama-cli -m <model.gguf> -p "Hello" -n 1 -lv 3 2>&1 | grep -i iqk
|
|
```
|
|
|
|
Should show IQK initialization messages.
|
|
|
|
### Test 2: CPU-only prefill benchmark
|
|
|
|
```bash
|
|
./build/bin/llama-bench -m <model.gguf> -ngl 0 -t 16
|
|
```
|
|
|
|
Compare pp512 before and after. Expected: 60 t/s → ~300 t/s.
|
|
|
|
### Test 3: GPU + CPU hybrid
|
|
|
|
```bash
|
|
./build/bin/llama-bench -m <model.gguf> -ngl auto -t 16
|
|
```
|
|
|
|
### Test 4: Full server test with Ling 3.0 Flash
|
|
|
|
```bash
|
|
./build/bin/llama-server \
|
|
-m Ling-3.0-flash-Q4_K_S.gguf \
|
|
--jinja -ngl 99 -c 32768 \
|
|
--temp 0.6 --top-p 0.95 --top-k 20 \
|
|
--host 127.0.0.1 --port 8080
|
|
```
|
|
|
|
Verify:
|
|
- Model loads without errors (especially `blk.0.ssm_f.weight`)
|
|
- Tool calls work correctly (TurboQuant's bugfix)
|
|
- KV cache compression works (TurboQuant's turbo3/turbo4 types)
|
|
- Prefill speed is ~300 t/s
|
|
- Decode speed is ~9-11 t/s (DDR4 bandwidth limit)
|
|
|
|
## Known Risks
|
|
|
|
1. **GATED_DELTA_NET**: Both forks have this op but implementations may differ. ik_llama has `iqk_fused_delta_net()` which is called when IQK is available. Verify this path is correctly guarded by `#ifdef GGML_USE_IQK_MULMAT` and doesn't conflict with TurboQuant's implementation.
|
|
|
|
2. **SIMD detection**: IQK uses compile-time SIMD detection (`__AVX2__`, `__AVX512__`, `__aarch64__`). TurboQuant's build system may set these differently. Verify the right code paths are selected.
|
|
|
|
3. **Repacked quantization types**: IQK introduces `Q8_K_R8`, `Q8_K_R16` as first-class GGUF types. These need to be registered in TurboQuant's type system. Check if `ggml-common.h` in TurboQuant already has these or needs them added.
|
|
|
|
4. **Thread count**: IQK's tiling strategy assumes certain thread counts. The `func16` path (16-row interleaved) is only used on AVX512. On AVX2 (your Zen 3), the `funcs[]` array is used instead. Verify the dispatch selects the right path.
|
|
|
|
## File Inventory
|
|
|
|
### Files to CREATE (new):
|
|
- `ggml/src/iqk/` — entire directory (36 files, ~2.1 MB)
|
|
- `ggml/src/ggml-cpu/CMakeLists.txt` — modified (add IQK sources)
|
|
|
|
### Files to MODIFY:
|
|
- `ggml/src/CMakeLists.txt` — add GGML_IQK_MUL_MAT option
|
|
- `CMakeLists.txt` (root) — add option passthrough
|
|
- `ggml/src/ggml-cpu/ggml-cpu.c` — add IQK hooks in ggml_compute_forward_mul_mat()
|
|
|
|
### Files to NOT TOUCH:
|
|
- `ggml/src/ggml-turbo-quant.c` — KV cache compression (orthogonal)
|
|
- `ggml/src/ggml-cpu/ggml-cpu-impl.h` — params struct (no change needed)
|
|
- Any MTP/NextN files — speculative decoding (orthogonal)
|
|
- Any bailingmoe3 architecture files — model-specific (orthogonal)
|
|
|
|
## Time Estimate
|
|
|
|
- Step 1-2 (copy + cmake): 30 minutes
|
|
- Step 3 (forward dispatch adaptation): 2-4 hours (the hard part)
|
|
- Step 4-5 (includes + optional fusion): 30 minutes
|
|
- Step 6 (build + test): 1-2 hours
|
|
- **Total: 4-7 hours for a competent C/C++ developer**
|
|
|
|
## Verification Checklist
|
|
|
|
- [ ] `cmake -DGGML_IQK_MUL_MAT=ON` configures without errors
|
|
- [ ] Build completes without errors
|
|
- [ ] IQK initialization message appears in verbose logs
|
|
- [ ] Ling 3.0 Flash loads without `blk.0.ssm_f.weight` errors
|
|
- [ ] Prefill speed improves from ~60 to ~300 t/s
|
|
- [ ] Decode speed is unchanged (~9-11 t/s)
|
|
- [ ] Tool calls work correctly
|
|
- [ ] KV cache compression (turbo3/turbo4) works
|
|
- [ ] `/no_think` parameter works
|
|
- [ ] Multi-slot (-np > 1) works if previously working
|