tinycompress
Python
★ 0
updated 3mo ago
Implemented and benchmarked LLM inference compression: int4/int8 quantization, GPTQ-like calibration, int8 KV cache, pruning, distillation, speculative decoding, torch.compile, and ONNX. Every number from a logged, self-audited run.
No plain-English explanation yet — one is being written right now. Check back in a minute.