A journal of building, experimenting, and running frontier AI models
on consumer hardware. Thoughts on coding agents, self-hosting, and the
pursuit of AI sovereignty.
2026-09-19
GLM Is Still King
GLM-5.3 Flash NVFP4 vs DeepSeek-V4.1 Flash head to head on Terminal-Bench
v2.1 with OMP — a dead heat at 73/89, but two very different ways to
get there.
AI
Self-hosting
Benchmarking
2026-09-03
All Roads Lead Back to GLM
Every round of model shopping, I end up back at GLM-5.2. Why the contenders
keep falling short at benchmark time, the EXL3 quant that keeps winning, and
the speed tax I keep paying.
AI
Self-hosting
GLM
2026-08-05
The Right Harness Is All You Need
A 13-model Terminal-Bench shootout with the PCIe Gen 5 switch topology.
DeepSeek V4 Flash 0731, Poolside Laguna S 2.1, Qwen 3.6, and a community
EXL3 quant of GLM 5.2 that punches far above its weight class.
AI
Self-hosting
Benchmarking
PCIe
2026-07-11
You Can Just Download More Tokens/Sec
Deep dive into the RTX PRO 6K community's vLLM fork — DSpark. Custom
CUDA kernels, multi-step speculation, FP8 attention, and what it takes to
squeeze 1,200+ tok/s from DeepSeek V4 Flash on 4 consumer GPUs.
AI
Self-hosting
vLLM
CUDA
2026-07-02
In Search of the Frontier at Home
From a Reddit-trained seq2seq chatbot to running frontier-scale models
on consumer hardware. A deep dive into GLM 5.2, DeepSeek V4 Flash,
MiniMax M3, and all the tradeoffs involved in running large models locally.
AI
Self-hosting
Benchmarking