Blog
Technical deep-dives, case studies, and field notes from working with customers on cloud and database engineering.
-
Instant at 10 Million Rows, Dead at a Billion: A Decision Framework for Billion-Row SQL
A query that's instant at 10M rows can time out at a billion — same SQL, same engine. It didn't regress; it crossed two invisible boundaries. An engine-neutral decision framework: pick by engine class (row-store / columnar / distributed) and memory budget (fits RAM vs spills to disk), with an 8-signal crossover checklist and cross-engine benchmarks (one desktop beat a 21-node cluster on a 1.1B-row scan).
-
How to Identify the Best RAG Model for Your Use Case
Teams building RAG almost always ask "which model?" — and mean the generation LLM. But a RAG system is not one model; it is at least three stacked decisions: the embedding model that decides what gets found, the reranker that decides what survives, and the generator that decides how the answer reads. A vendor-neutral framework for choosing each layer by the failure you can't tolerate — grounded in Anthropic's published numbers (retrieval failure 5.7% → 1.9%, generator unchanged).
-
Find the Root Blocker: Build an Evidence-First PostgreSQL Lock Agent (Part 1)
Part 1 of a 2-part series. The query your dashboard shows waiting is almost never the one holding the lock. Trace the real root blocker with pg_blocking_pids, tell deadlocks from lock timeouts, normalize five failure modes into one machine-checkable report, and keep the whole agent read-only so it physically cannot amplify an outage — all backed by a runnable public lab.
-
Top AI Agent Evaluation Frameworks to Know in 2026 — Pick by the Layer You Need to Test, Not by Star Count
You don't pick an agent-eval framework by GitHub stars — you pick by the evaluation layer you need to test. A five-bucket taxonomy (trajectory, CI, observability, benchmarks, provider-native) with one representative tool per layer, real versions and star counts, pass^k reliability numbers, and a shift-left/monitor-right lifecycle loop.
-
Read-Only Is a Lie: The Postgres MCP Server Mistakes to Avoid When You Wire an Agent to Your Database
A read-only MCP tool is a lie until the database enforces it. The exploratory-to-operational spectrum for wiring an agent to Postgres — and the four independent layers (AST parse, read-only transaction, least-privilege role, statement timeout) that earn "read-only" for real.
-
When Not to Use Postgres: A Decision Framework for the Four Walls Where One Engine Isn't Enough
One ACID engine now absorbs vectors, time-series, queues, search, and documents. The consolidation case for defaulting to Postgres — and the four specific walls where a specialist still wins, each named with the number that justifies it.
-
From Prompts to Loop Engineering: The Workflow Shift in AI-Native Development
You're not getting better at prompting — the skill moved. Prompt, context, harness, loop engineering are one staircase, and in 2026 the unit of work you own has climbed all the way up to the iteration loop. A first-person walk through the four eras of AI-native development.
-
Spend Fewer Tokens, Get Better Code: A Context Engineering Guide for AI Code Assistants (Part 1 of 3)
Anthropic cut tool context by 85%. Accuracy improved from 49% to 74%. Five context engineering practices that make your AI code assistant produce better output — while spending fewer tokens.
-
Invisible Compound Savings: Caching, Workflow Discipline, and the Habits That Add Up (Part 2 of 3)
90% of your AI prompt context repeats across every request. Prompt caching gives you 90% off. The retry tax costs you 1.4x. Here is how structural habits compound into invisible savings.
-
From PRUs to AI Credits: The Token-Based Bill Is Already Here (Part 3 of 3)
On June 1, 2026 GitHub Copilot retires Premium Request Units and replaces them with token-metered GitHub AI Credits. A task taxonomy keyed to GitHub's official Lightweight / Versatile / Powerful categories, the new cost math in actual dollars (~$4 vs ~$30 per developer day), and the complete three-layer optimization playbook.
-
PostgreSQL EXPLAIN BUFFERS: How We Cut Checkout Latency 96%
A real-world e-commerce case study: one word added to EXPLAIN ANALYZE diagnosed a checkout regression from 50ms to 1.2s that three days of network debugging missed. The three tuning levers that fixed it.