tq
Python
★ 1
updated 4mo ago
Run local LLMs with maximum context on minimum hardware. One command: hardware detection + TurboQuant KV cache compression + OpenAI-compatible server.
No plain-English explanation yet — one is being written right now. Check back in a minute.