gitmyhub

tinycompress

Python ★ 0 updated 3mo ago

Implemented and benchmarked LLM inference compression: int4/int8 quantization, GPTQ-like calibration, int8 KV cache, pruning, distillation, speculative decoding, torch.compile, and ONNX. Every number from a logged, self-audited run.

No plain-English explanation yet — one is being written right now. Check back in a minute.