| .. |
|
fa
|
Support GigaChat3 (#995)
|
2025-11-24 06:55:14 +01:00 |
|
iqk_common.h
|
Unroll for loop for repacked BF16 MATMUL (#1047)
|
2025-12-08 06:09:45 +01:00 |
|
iqk_config.h
|
Fix termux/android build (#336)
|
2025-04-21 09:13:46 +02:00 |
|
iqk_cpu_ops.cpp
|
Hadamard transforms for K-cache - CPU only (#1033)
|
2025-12-04 06:51:11 +01:00 |
|
iqk_cpu_ops.h
|
Hadamard transforms for K-cache - CPU only (#1033)
|
2025-12-04 06:51:11 +01:00 |
|
iqk_flash_attn.cpp
|
Better CPU SWA (#757)
|
2025-09-04 11:58:16 +02:00 |
|
iqk_flash_impl.h
|
Enable CUDA graphs for MoE models + GPT-OSS support (#689)
|
2025-08-15 09:18:07 +03:00 |
|
iqk_gemm_1bit.cpp
|
AVX512+AVXVNNI GEMM implementation for quants using Q8_K for activations (#710)
|
2025-08-22 06:27:07 +03:00 |
|
iqk_gemm_1bit.h
|
Faster iq1_s GEMM via repacking to Q8_0_R8 (#517)
|
2025-06-11 15:01:34 +03:00 |
|
iqk_gemm_floats.cpp
|
Unroll for loop for repacked BF16 MATMUL (#1047)
|
2025-12-08 06:09:45 +01:00 |
|
iqk_gemm_floats.h
|
Refactor iqk_mul_mat.cpp (#435)
|
2025-05-22 10:05:51 +03:00 |
|
iqk_gemm_iqk_quants.cpp
|
AVX512+AVXVNNI GEMM implementation for quants using Q8_K for activations (#710)
|
2025-08-22 06:27:07 +03:00 |
|
iqk_gemm_iqk_quants.h
|
Much faster CPU prompt processing (part 2) (#533)
|
2025-06-18 07:29:33 +03:00 |
|
iqk_gemm_iquants.cpp
|
AVX512+AVXVNNI GEMM implementation for quants using Q8_K for activations (#710)
|
2025-08-22 06:27:07 +03:00 |
|
iqk_gemm_iquants.h
|
IQ2_XXS: much faster CPU prompt processing (#515)
|
2025-06-11 11:12:30 +03:00 |
|
iqk_gemm_kquants.cpp
|
AVX512+AVXVNNI GEMM implementation for quants using Q8_K for activations (#710)
|
2025-08-22 06:27:07 +03:00 |
|
iqk_gemm_kquants.h
|
Faster CPU prompt processing for Q4_K and Q5_K (#525)
|
2025-06-13 07:58:15 +03:00 |
|
iqk_gemm_ktquants.cpp
|
Adding IQ1_KT - 1.75 bpw SOTA quants (#616)
|
2025-07-20 10:05:23 +02:00 |
|
iqk_gemm_ktquants.h
|
Trellis quants: faster CPU prompt processing (#482)
|
2025-06-01 15:24:05 +03:00 |
|
iqk_gemm_legacy_quants.cpp
|
Add missing AVX512 operators for MSVC (#948)
|
2025-11-14 06:58:51 +02:00 |
|
iqk_gemm_legacy_quants.h
|
Much faster CPU prompt processing (part 3) (#534)
|
2025-06-18 15:30:56 +03:00 |
|
iqk_mul_mat.cpp
|
Support GigaChat3 (#995)
|
2025-11-24 06:55:14 +01:00 |
|
iqk_mul_mat.h
|
cpu: fused softmax+topk (#794)
|
2025-09-24 09:02:21 +02:00 |
|
iqk_quantize.cpp
|
Hadamard transforms for K-cache - CPU only (#1033)
|
2025-12-04 06:51:11 +01:00 |
|
iqk_quantize.h
|
Check for NaNs while loading the model. (#727)
|
2025-08-27 19:00:17 +03:00 |
|
iqk_utils.h
|
Add missing AVX512 operators for MSVC (#948)
|
2025-11-14 06:58:51 +02:00 |