Evaluation-driven development for agents: a regress-gate that can't fail your build
9 min
Agents need tests the way code does, but the assertion is fuzzy. Here's the EDD loop I built into ReplayGate: deterministic offline replay that gates the build with real exit codes, an LLM judge that only ever advises, and the honest cost of pinning replay to a hash.