SamuelOliveirads
|
0d75eee35a
|
remove duplicated code and unnecesary refactor
|
2026-06-14 16:02:02 -03:00 |
SamuelOliveirads
|
3a1d46c4d1
|
Merge remote-tracking branch 'origin/main' into feat/dflash-implementation
# Conflicts:
# common/common.cpp
# common/speculative.cpp
# convert_hf_to_gguf.py
# examples/server/server-context.cpp
# examples/server/server-context.h
# src/llama-arch.cpp
# src/llama-arch.h
# src/llama-model.cpp
# src/llama.cpp
|
2026-06-13 17:27:52 -03:00 |
Kawrakow
|
11c3546235
|
Support for alternative Gemma4 assistant (#1937)
|
2026-06-09 09:30:12 +02:00 |
SamuelOliveirads
|
dc43cdf06b
|
move dflash for it own file
|
2026-06-02 10:22:13 -03:00 |
SamuelOliveirads
|
3d73312d9d
|
apply workspace support for KV cache
|
2026-06-01 09:55:34 -03:00 |
SamuelOliveirads
|
ed403dca27
|
Use windows update in kv cache
|
2026-05-31 14:51:21 -03:00 |
SamuelOliveirads
|
1369e68471
|
fix graph mask, swa layers and tokens positions
|
2026-05-31 11:12:03 -03:00 |
SamuelOliveirads
|
532499836e
|
improve DFlash caching and profiling capabilities
|
2026-05-30 21:36:10 -03:00 |
SamuelOliveirads
|
9f5f70cf7e
|
implement target position tracking and context management
|
2026-05-29 23:11:38 -03:00 |
SamuelOliveirads
|
82cff238fe
|
Initial dflash implementation
|
2026-05-28 18:57:58 -03:00 |
Samuel Oliveira Alves
|
11a1fea9e2
|
Move embedding management to speculative (#1825)
* refactor speculative decoding with companion context and draft result structures
* feat: add common speculative feature handling in server context
* refactor: move embedings outside server
* feat: harden draft input hidden state in llama context
* remove unused functions
* refactor: streamline speculative feature handling and remove unused code
* remove redundant code
* remove more unused variables
* refactor: implement speculative feature handling
|
2026-05-20 17:42:48 +03:00 |