A journal of building, experimenting, and running frontier AI models
on consumer hardware. Thoughts on coding agents, self-hosting, and the
pursuit of AI sovereignty.
2026-07-11
You Can Just Download More Tokens/Sec
Deep dive into the RTX PRO 6K community's vLLM fork — DSpark. Custom
CUDA kernels, multi-step speculation, FP8 attention, and what it takes to
squeeze 1,200+ tok/s from DeepSeek V4 Flash on 4 consumer GPUs.
AI
Self-hosting
vLLM
CUDA
2026-07-02
In Search of the Frontier at Home
From a Reddit-trained seq2seq chatbot to running frontier-scale models
on consumer hardware. A deep dive into GLM 5.2, DeepSeek V4 Flash,
MiniMax M3, and all the tradeoffs involved in running large models locally.
AI
Self-hosting
Benchmarking