Topic
AI & LLMs
Models you can run on your own hardware, prompt patterns that ship, agent frameworks that don't catch fire, and the awkward questions nobody answers in the breathless launch posts. Ollama, vLLM, llama.cpp, LocalAI, plus the quieter stuff — embeddings, RAG, evals, and figuring out when the cloud API is actually the right answer. If you'd rather understand the trade-offs than chase benchmarks, you'll feel at home here.
87 articles in this topic.
Featured posts
-
Multiple Claude Code Accounts, One Config
Run several Claude Code accounts on one Linux machine. CLAUDE_CONFIG_DIR plus symlinks share your skills, hooks and history while each login stays isolated.
15 min read -
The Free Tier Rug Pull
Free AI tiers are loans against a future price, paid in your data, your architecture, or your time. Here's the collateral to check before you build on one.
7 min read -
Real-ESRGAN & Upscaling Tools
Image upscaling with Real-ESRGAN, super-resolution AI models, ComfyUI workflows, and practical upscaler comparison for self-hosted setups.
9 min read -
ControlNet & LoRA: Advanced Image Control
Master ControlNet and LoRA for precise image generation control in Stable Diffusion. Learn Canny edges, depth, pose, scribbles, tiling, and style stacking in ComfyUI.
13 min read -
API vs Self-Hosted LLMs: The Real Cost
API vs self-hosted LLM cost reality, GPU TCO, privacy, latency, break-even math. When paying OpenAI/Anthropic wins. When local wins.
10 min read -
Free AI Image Gen vs Your Own GPU
Free AI image generators win on convenience but lose fast on volume, style consistency, and control. Here's the real break-even against running your own GPU.
12 min read
All AI & LLMs articles
- Multiple Claude Code Accounts, One Config
- The Free Tier Rug Pull
- Real-ESRGAN & Upscaling Tools
- ControlNet & LoRA: Advanced Image Control
- API vs Self-Hosted LLMs: The Real Cost
- Free AI Image Gen vs Your Own GPU
- RAG Beyond Vector Search: BM25, Hybrid, Re-ranking
- How to Stretch a Free LLM Tier
- The Free AI Stack: What $0 Gets You
- Agentic Browsers: Browser-Use & Skyvern Reviewed
- Local Voice Assistant: Whisper + Piper + Home Assistant
- Your Agent Doesn't Need a Shell
- Llamafile: Single-Binary LLMs That Actually Just Run
- DeepSeek V4 Flash: 28 Cents a Million
- Where Should Your Coding Agent Run?
- DIY Perplexity: SearXNG + Local LLM = Private Web Search
- Local Coding Agents Need Less Context
- LM Studio vs Jan vs GPT4All: Desktop LLM Clients
- RAGAS: Evaluating RAG Without Vibes
- KV Cache Quantization: Free LLM Context, Almost
- Aider & Cline: Terminal AI Coding That Actually Ships
- Mixture of Experts (MoE) for Self-Hosters, Demystified
- Python Libraries Worth Your Time in 2026
- Speculative Decoding: Faster LLMs With a Tiny Sidekick
- Karakeep: Self-Hosted Bookmarks With AI Tagging
- Stop Feeding the AI Your Whole Repo
- RTK vs snip vs lean-ctx: Token Killers
- Hailo-8 vs Coral: AI Accelerators for the Edge
- AI Swarm Audited My 840-Post Blog
- Used GPU Buying Guide for Home Lab LLMs
- Claude Code in a Homelab Workflow
- Self-Host a Local AI Coding Workhorse
- Give Your AI Agent a Cheap Intern
- Claude Code + SearXNG: Private Web Search
- Dify: Visual Agent Workflows
- OpenRouter vs LiteLLM
- Function Calling in Local LLMs
- Gemma 4 vs Qwen3.6
- AnythingLLM as Knowledge Base
- Local Vision LLMs Worth Running in 2026
- MCP Servers: Tools for LLMs
- RAG Evaluation with Ragas
- LLM Distillation Explained
- Open WebUI Tools, Functions and Pipelines Explained
- Self-Supervised Learning Explained
- Continue.dev vs Cody vs Tabby: AI Code Help Without the Cloud
- Qdrant vs Weaviate vs Chroma: Vector DB Showdown
- LangChain vs LlamaIndex: RAG Framework Showdown
- Beyond RAG: When a Virtual Filesystem Works Better
- Running Gemma 4 Locally with Ollama
- 1-Bit LLMs: The Quantization Endgame
- AMD Lemonade: Local LLM Serving for AMD GPUs
- When to Use Structured Output (JSON Mode) in LLMs
- Using AI to Find Security Bugs in Your Code
- LLM Temperature and top_p Explained Without the Math
- GPU Memory Math: Will This Model Actually Fit?
- LLM Backends: vLLM vs llama.cpp vs Ollama
- RAG Chunking: Why Chunk Size Is Everything
- LiteLLM & vLLM: One API to Rule All Your Models
- System Prompts: The LLM Feature Most People Ignore
- LLM Quantization: Q4_K_M Isn't Always the Best Choice
- Running Multiple Ollama Models Without Running Out of RAM
- Piper vs Coqui: Text-to-Speech on Your Own Hardware (Because AWS Polly Charges Per Character Like It's 1999 SMS)
- Context Window vs Token Limit: Not the Same Thing
- The Embedding Model Choice Nobody Explains
- Ollama Keep Alive: Unload Models from VRAM
- RAG on a Budget: Building a Knowledge Base with Ollama & ChromaDB
- Stable Diffusion vs ComfyUI vs Fooocus: AI Image Generation at Home
- n8n + LLM: Building Automations That Actually Think
- Text Generation Web UI vs KoboldCpp: Power User LLM Interfaces
- LangGraph vs CrewAI vs AutoGen: AI Agent Frameworks for Mere Mortals
- Open WebUI vs LibreChat: Self-Hosted ChatGPT Alternatives Compared
- LLM Fine-Tuning for Mortals: LoRA, QLoRA, and Your Gaming GPU
- Ollama Beyond the Basics: Model Management, Custom Models, and Optimization
- Whisper & Faster-Whisper: Self-Hosted Speech-to-Text That Actually Works
- Continue.dev vs Cody vs Tabby: AI Code Assists That Live on Your Machine
- CUDA vs ROCm vs CPU: Running AI on Whatever GPU You've Got
- Flowise vs Langflow: Build AI Pipelines Without Writing a Novel
- n8n vs Node-RED: Automate Everything Without Learning to Code (Much)
- Key Parameters of Large Language Models
- Prompts for Image Generation in Stable Diffusion
- Prompt Engineering for Generative AI 101
- Large Language Model Formats and Quantization
- Exploring the Diverse World of LLM Models
- Ollama: Powerful Language Models on Your Own Machine
- Unleash the Power of LLMs with LocalAI
- Machine Learning models (AI)