blog

Specs, drift, and code that stays true to intent.

One post a month on spec-driven development and verifying AI-generated code. Every claim links to its source.

· 4 min read

Coding Agent Evaluation Is Splintering Into Smaller Checks

Recent benchmark churn, open eval releases, and vendor guardrails show coding-agent trust is moving from one score to many small, evidence-anchored checks.

· 4 min read

Spec Conformance Is the Metric Benchmarks Skip

AI coding benchmarks hit new highs in 2026, but review debt and requirement drift show why passing tests isn't the same as meeting a spec.