46-day current streak·46-day longest streak
AI Practitioner & Data-Driven Growth Specialist building local LLM infrastructure, benchmarking models, publishing results --- What I Do I build local LLM inference stacks from source on consumer hardware, benchmark…
AI Practitioner & Data-Driven Growth Specialist
*building local LLM infrastructure, benchmarking models, publishing results*


---
What I Do
I build local LLM inference stacks from source on consumer hardware, benchmark models systematically, and publish datasets on HuggingFace. I also build analytics dashboards and have scaled a tech community to 20,000+ members.
Current Focus:
- local inference optimisation (llama.cpp, CUDA..)
- systematic benchmarks across dense, MoE, and hybrid architectures
- quantisation testing (GGUF Q4_K_M, IQ4_XS, turboquant turbo2/turbo3)
- context window scaling analysis and VRAM profiling
- publishing benchmark datasets on HuggingFace
---
🛠️ Tech Stack
AI / ML
!CUDA !llama.cpp !HuggingFace !Pythondata & analytics
!Dune Analytics !SQLfrontend
!React !Next.js !TypeScriptinfra & tools
!Linux !Bash !Git---
📊 Background
- Community Lead @ Nous Research - Hermes community, local inference guides
- AI / ML Practitioner - local LLM inference, model evaluation, HuggingFace contributor
- Growth Lead @ Yari Finance - DeFi protocol growth, partnerships, on-chain analytics
- Founder @ BeraLand - built a 20K+ member blockchain community from zero
- 15+ Dune dashboards tracking $1B+ in trading volume
- Master's in Corporate & Market Finance - KPMG background
🎓 Learning Journey
---
*I write about AI infrastructure, local inference, and model evaluation on 𝕏*
-
llm-bench-rig ★ PINNED
Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publishable cards.
Python ★ 24 3d agoExplain → -
hermes-recipes ★ PINNED
tested hermes agent recipes: configs, deploys, mcp, automations. copy, run, build.
Python ★ 101 13d agoExplain → -
sm120-field-guide ★ PINNED
AI on consumer Blackwell — a field guide for the RTX 50-series (sm_120). Fixes, footguns, and honest measurement from one RTX 5090.
★ 25 13d agoExplain → -
notwitcheer
My GitHub profile — AI practitioner building local LLM inference stacks and publishing benchmark datasets.
★ 1 16d agoExplain → -
DreamServer ⑂
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
★ 1 1mo agoExplain → -
hermes-agent-fork ⑂
The agent that grows with you
Python ★ 0 1d agoExplain → -
vllm ⑂
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 0 11d agoExplain → -
sglang ⑂
SGLang is a high-performance serving framework for large language models and multimodal models.
★ 0 11d agoExplain → -
SpecForge ⑂
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
★ 0 11d agoExplain →
No repos match these filters.