* Gemma4: WIP * Gemma4: WIP - runs with totally wrong results * Gemma4: WIP - add CPU 512, 512 FA * Gemma4: WIP It gives a meaningful response in llama-cli, but PPL is still much too high. Is this due to tokenizer issues? * Gemma4: this works I had forgotten the softcap on the final output. * Remove log * Gemma4: WIP E4B/E2B * Gemma4: Q4B/E2B appear to work now * gemma4: tokenizer fixes |
||
|---|---|---|
| .. | ||
| cmake | ||
| include | ||
| src | ||
| .gitignore | ||
| CMakeLists.txt | ||