blog
Specs, drift, and code that stays true to intent.
One post a month on spec-driven development and verifying AI-generated code. Every claim links to its source.
· 4 min read
Coding Agent Evaluation Is Splintering Into Smaller Checks
Recent benchmark churn, open eval releases, and vendor guardrails show coding-agent trust is moving from one score to many small, evidence-anchored checks.
· 4 min readSpec Conformance Is the Metric Benchmarks Skip
AI coding benchmarks hit new highs in 2026, but review debt and requirement drift show why passing tests isn't the same as meeting a spec.