Theme
Reflections

A journal of building, experimenting, and running frontier AI models on consumer hardware. Thoughts on coding agents, self-hosting, and the pursuit of AI sovereignty.

2026-07-11
You Can Just Download More Tokens/Sec
Deep dive into the RTX PRO 6K community's vLLM fork — DSpark. Custom CUDA kernels, multi-step speculation, FP8 attention, and what it takes to squeeze 1,200+ tok/s from DeepSeek V4 Flash on 4 consumer GPUs.
AI Self-hosting vLLM CUDA
2026-07-02
In Search of the Frontier at Home
From a Reddit-trained seq2seq chatbot to running frontier-scale models on consumer hardware. A deep dive into GLM 5.2, DeepSeek V4 Flash, MiniMax M3, and all the tradeoffs involved in running large models locally.
AI Self-hosting Benchmarking
~/rss

Subscribe via RSS