Theme
Reflections

A journal of building, experimenting, and running frontier AI models on consumer hardware. Thoughts on coding agents, self-hosting, and the pursuit of AI sovereignty.

2026-08-05
The Right Harness Is All You Need
A 13-model Terminal-Bench shootout with the PCIe Gen 5 switch topology. DeepSeek V4 Flash 0731, Poolside Laguna S 2.1, Qwen 3.6, and a community EXL3 quant of GLM 5.2 that punches far above its weight class.
AI Self-hosting Benchmarking PCIe
2026-07-11
You Can Just Download More Tokens/Sec
Deep dive into the RTX PRO 6K community's vLLM fork — DSpark. Custom CUDA kernels, multi-step speculation, FP8 attention, and what it takes to squeeze 1,200+ tok/s from DeepSeek V4 Flash on 4 consumer GPUs.
AI Self-hosting vLLM CUDA
2026-07-02
In Search of the Frontier at Home
From a Reddit-trained seq2seq chatbot to running frontier-scale models on consumer hardware. A deep dive into GLM 5.2, DeepSeek V4 Flash, MiniMax M3, and all the tradeoffs involved in running large models locally.
AI Self-hosting Benchmarking
~/rss

Subscribe via RSS