gitmyhub

Differential-KV

★ 0 updated 15d ago ⑂ fork

Differential-KV (DKV) is a sparse KV-cache inference runtime designed for high-efficiency, memory-bounded long-context Large Language Model (LLM) inference across Apple Silicon (MLX) and CUDA GPUs.

No plain-English explanation yet — one is being written right now. Check back in a minute.