Blog

๐Ÿ“„ Quantization (Or: Why Your 7B Model Fits in 4GB)

14GB of weights on an 8GB GPU. Quantization is the polite way of telling the model to fit.

Written on: Mon Sep 21 2026
๐Ÿ“„ Make Illegal States Unrepresentable (Or: Stop Letting Your Types Lie To You)

Your isLoading/isError/data triangle of lies, and the enum that ends it.

Written on: Mon Sep 21 2026
๐Ÿ“„ Domain Context to your LLM? RAG!

Let's load some context into a tiny LLM and make the answers actually good.

Written on: Tue Jul 07 2026
๐Ÿ“„ ONNX: The Model Exchange Standard That Actually Stuck

What ONNX actually is under the hood, how the graph format works, and why it became the lingua franca for ML deployment.

Written on: Mon Jun 01 2026
๐Ÿ“„ Running Your Own AI Is Actually Good Now

Open WebUI and Qwen changed what local AI actually feels like to use.

Written on: Sun May 10 2026
Page 1 of 4Next ยป