Notes on self-hosting, AI stacks and web engineering
Published on Nov 2025
Explore how to build a customizable local AI stack using Llama.cpp, with tools like Open Web UI for advanced interfaces, Tabby for code completion, and Comfy UI for workflow design. Learn to optimize GPU performance, achieve 128k context scaling through ROPE, and integrate embedding models for enhanced code understanding. Discover streamlined setups with Docker Compose and draft models to boost speed, all tailored for developers seeking self-hosted AI capabilities.
Published on Mar 2025