A recent Google paper showed spec-driven test generation raised bug detection by 9.8 percentage points on a sample from their codebase. But their spec is read out of the code, which leaves the half I care about unresolved. I have been arguing for months that a test suite written by the same model that wrote the code cannot really disagree with it and to support my claim I've developed a new open source Python library.
Software testing has been moving in one direction for twenty years, and Specification-Driven Development (SDD) is where that movement has recently arrived. This article argues for one more step: Independent SDD, or ISDD. TDD said the tests are the specification. BDD was the answer to that. SDD is the version that arrived with the agents.
Currently, Spec-Driven Development divides the work. It does not divide who has the knowledge. The same specification goes to the planner, the test generator and the coding agent. My argument is to cut along that line: give the coding agent the decisions and withhold the acceptance criteria, so the test suite can tell the code it is wrong.
A team at Google measured the step before withholding. "Grounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test Generation" does something narrower than my argument. They first ask it to reason about the code and write down its contract. That document becomes a cognitive scaffold and the tests are generated from it. The results, on production bugs from Google's own codebase, showed that more than half the time, a test suite generated...
Source: Towards Data Science · Summarized by HeadlinesBriefing