gitmyhub

atomic-llama-cpp-turboquant

C++ ★ 1 updated 1mo ago ⑂ fork

llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% throughput).

No plain-English explanation yet — one is being written right now. Check back in a minute.