gitmyhub

tq

Python ★ 1 updated 4mo ago

Run local LLMs with maximum context on minimum hardware. One command: hardware detection + TurboQuant KV cache compression + OpenAI-compatible server.

No plain-English explanation yet — one is being written right now. Check back in a minute.