Blog
๐ Quantization (Or: Why Your 7B Model Fits in 4GB)
14GB of weights on an 8GB GPU. Quantization is the polite way of telling the model to fit.
Written on: Mon Sep 21 2026๐ Make Illegal States Unrepresentable (Or: Stop Letting Your Types Lie To You)
Your isLoading/isError/data triangle of lies, and the enum that ends it.
Written on: Mon Sep 21 2026๐ Domain Context to your LLM? RAG!
Let's load some context into a tiny LLM and make the answers actually good.
Written on: Tue Jul 07 2026๐ ONNX: The Model Exchange Standard That Actually Stuck
What ONNX actually is under the hood, how the graph format works, and why it became the lingua franca for ML deployment.
Written on: Mon Jun 01 2026๐ Running Your Own AI Is Actually Good Now
Open WebUI and Qwen changed what local AI actually feels like to use.
Written on: Sun May 10 2026Page 1 of 4Next ยป