atomic-llama-cpp-turboquant
C++
★ 1
updated 1mo ago
⑂ fork
llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% throughput).
No plain-English explanation yet — one is being written right now. Check back in a minute.