3-day longest streak
-
vLLM-Moet
A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official (NV)FP4 checkpoint's quality on consumer Blackwell cards
Sass ★ 444 1h agoExplain → -
qwentin
27B at 256k context on one RTX 5090 — FP6 + 4-bit KV + MTP spec-decode, hand-written SM120 tensor-core kernels, OpenAI API.
Cuda ★ 11 17d agoExplain → -
cubit
Open-source SM120 (Blackwell) SASS assembler — encode, schedule & assemble GPU machine code, including all tensor-core instructions.
Rust ★ 6 1mo agoExplain → -
LLobotomy
Surgical uncensoring of large language models via Optimal Transport activation hooks
Python ★ 6 4mo agoExplain → -
rtx-p2p
PCIe P2P on consumer RTX 5090 via one GSP firmware registry key — no driver patches.
Shell ★ 3 3mo agoExplain → -
blackwell-isa
Reverse-engineered instruction set architecture database for NVIDIA SM120 (Blackwell) GPUs.
HTML ★ 2 1mo agoExplain → -
mercury
Mercury — LD_PRELOAD that unlocks ptxas's 949 internal scheduling knobs.
C ★ 2 5mo agoExplain → -
vllm ⑂
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 1 1h agoExplain → -
msa-120
FlashAttention BF16/FP8 for NVIDIA SM120 — per-warp HMMA + QMMA.SF block-scaled, no TMA/TMEM required. Port of MiniMax MSA.
Python ★ 1 1mo agoExplain → -
sasskit
SM120 SASS assembler toolkit
Python ★ 1 1mo agoExplain → -
aztec-packages ⑂
Fork of aztec-packages containing bb_rs to create bindings to Barretenberg
★ 0 6mo agoExplain →
No repos match these filters.