Gradle Technologies is now Develocity — read the announcement

All Blog Posts
September 30, 2026

Why We Still Need Regression Tests, Even with Smarter AI and Formal Verification

By Hans Dockter

As software development becomes AI-native, I want to look at the role of regression testing. It is not a much-debated topic, but the success of AI-native development depends on understanding how it fits with other verification methods. Of these, specification tests and formal verification get most of the attention. That, together with ever smarter models, makes it tempting to think that regression tests will matter less.

My argument, based on mathematics and logic, is that for all practical purposes no one can guarantee that a change breaks nothing that worked before, however smart the AI that makes it. So every change still has to be checked against existing behavior. That is what regression tests do.

Specification tests check behavior that someone decided on and wrote into a specification, for example which premium a new customer pays. In spec-driven development, where people describe what they want and agents build it, specification tests are the foundation.

Regression tests check that existing behavior still works after a change. The two overlap: once the change it was written for is merged, a specification test also becomes a regression test. But a regression suite holds many tests that were never specification tests. When a change breaks something that worked before, we call this an unintended side effect. Side effects are classic bugs. It is tempting to see them as mistakes that more careful thinking would have prevented. Mistakes look like a question of intelligence: a smarter AI makes fewer of them, and eventually, the thinking goes, so few that they no longer matter.

Some business leaders therefore expect that, with a smarter AI, specification tests will be enough. Their reasoning is that an AI that hardly makes mistakes will hardly change anything beyond what it was asked to. Specification tests already check what it was asked to do, so regression tests will have little left to catch. Others expect formal verification to take over, now that AI can write the proofs.

It sounds reasonable to treat side effects as mistakes that a smarter AI will stop making, and hindsight supports that view. After an incident, the postmortem finds the code path that led to the side effect, and once it is found, it looks obvious: someone should have seen it. But before the change, that path was one among more paths than anyone can examine. Any single side effect can be found. A smarter AI will find more of them. But finding side effects one at a time never shows that none are left.

The next model will not fix this. And agents raise the stakes: they make far more changes, and every change can have side effects.

A regression test runs the real code against expectations written before the change. The specification tests among them protect only what the specification says, and in spec-driven development usually only at the level of whole scenarios. Tests at that level can also miss what goes wrong inside the code. A study of real bugs in Java projects found that when a bug corrupted the state inside a function, the damage almost always showed up in what the function returned or left behind, where a unit test could check it. When the whole program was tested, the damage sometimes disappeared before it reached the output. A regression suite made only of specification tests is therefore thin.

Software does much more than anyone decided. It rounds amounts in a certain way, returns records in a certain order and answers within a certain time. With enough users, every behavior they can observe is relied on by somebody, a pattern known as Hyrum's Law. A regression suite protects as much of that behavior as its tests check. Unit tests are the backbone of most regression suites: they make up most of the tests, run automatically and pin down the small choices inside the code, such as thresholds and rounding rules nobody put into a specification. Tests at the interfaces pin down what other systems see. Any behavior that no test pins down can change unnoticed.

A 2025 study by Wang, Pradel and Liu looked at fixes that agents wrote for real bugs. A benchmark had counted each fix as correct because it passed the benchmark's tests. Of these fixes, 7.8% failed other tests that the projects' developers had already written. The benchmark's tests missed these failures, but the existing regression tests caught them.

A team's test suite also carries its memory. When something breaks, teams often add a test so that the same failure cannot come back unnoticed. Over the years, the test suite becomes a record of what broke and where the code is fragile. An experienced engineer carries some of that knowledge in their head. An agent starts most tasks without that knowledge, but it can run the tests. Where a codebase has few tests, the agent has neither the memory nor the check. But an agent can help close that gap: before anyone changes a legacy system, it can record in tests what the code does today.

Software behaves non-linearly, in two ways that matter here. First, the size of an effect is not proportional to the size of the change: a one-line change can take down a system. In July 2019, one badly written firewall rule took down Cloudflare's network worldwide for 27 minutes. Every if-statement is a threshold. A value that moves slightly past it can flip the program's behavior entirely. Second, changes interact: two changes that are each safe on their own can break something together.

To be sure that a change has no side effects, you would have to try every input on every path through the program (every combination of branches it can take) or use formal verification, which the next section covers. Trying every path is impossible: each if-statement can double the number of paths, so a program with just 90 if-statements can have more paths than a computer could have checked at a billion per second since the Big Bang. Enterprise codebases have hundreds of thousands of if-statements.

Intelligence helps to find side effects, but finding them is not the same as ruling them out. To show that a side effect exists, one example is enough. If a change was meant only for new customers, one existing customer whose premium changed proves that it has a side effect. To show that none exist, you need an argument that covers every path, and there are too many to check one by one. Regression tests cannot rule them out either. But they catch the side effects that break behavior they check. Because they run automatically, they can do so on every change, far more cheaply than checking each change from scratch, as long as someone maintains them.

The alternative to checking paths one by one is formal verification, which reasons about all of them at once. Lighter relatives, such as type checkers and some static analysis tools, already reason over every path, but only for narrow properties, for example that a value is never missing. For the kind of behavior a side effect breaks, such as the result a calculation returns, formal verification is the only way around the explosion in the number of paths.

One form of formal verification is the formal proof, which works like a proof in mathematics. To show that the sum of two even numbers is always even, a mathematician does not try every pair of even numbers: one short argument covers all of them at once. A formal proof does the same for a program, and a computer checks every step of the argument, so no gap in the reasoning slips through. Languages like Lean make such proofs possible, and AI systems are getting good at writing them. That can bring formal verification to far more code than today.

Where a proof exists, it can replace the unit tests that check what the code calculates. Given a complete and correct statement of what the code must do, a proof covers every input, which no test can. People still need specification tests to confirm that the statement says what they decided. Formal verification is already used where errors are expensive, such as in cryptography, access control and compilers. AI will carry it further. How far is hard to predict, because there is still little experience with the costs and downsides of AI-written proofs across enterprise software.

A proof also has a boundary. It shows that the code does what the statement says, but it does not cover the libraries, the database or the configuration the code runs with. Tests that run against the real dependencies check those.

Sometimes the proof is not even about the production code. AWS proves properties of a Lean model of Cedar, its authorization language. Every night, it then runs the model and the production code on millions of generated inputs and checks that both give the same results. This is called differential testing: two implementations of the same behavior run on the same inputs, and any difference in their results points to an error. A new version is released only when the model, the proofs and these tests are all up to date.

For the next several years, most of an enterprise's software will not be formally verified. So proofs and regression tests will work side by side: proofs where someone has done them, regression tests everywhere else and at every boundary.

Agents let a team change far more code than before. Assume each of their changes is as likely to cause a side effect as an experienced engineer's. The absolute number of side effects still rises, and side effects that used to be rare become routine. Say one change in 100 has a side effect the tests miss. A team of five engineers making 20 changes a week hits one every five weeks and treats it as an incident. At 200 changes a week, it hits two a week, and side effects become part of normal operation. If the volume keeps rising, side effects arrive faster than the team can deal with them, and delivery breaks down. A smarter AI may lower the rate of side effects. But the number of side effects falls only if the rate falls faster than the volume of change rises.

So far, the evidence points the other way. DORA's 2025 report, restating its 2024 finding, estimates a 7.2% increase in software delivery instability for every 25% increase in AI adoption. It also finds that AI adoption now improves throughput but still increases instability.

Two things make this worse. First, changes made at the same time can interact. The more changes are in progress at once, the more chances there are for two safe changes to break something together. Second, human review does not scale. Reviewers already suffer from review fatigue. If people remain the main check, more changes mean less attention per change, and a larger share of side effects slips through. Automated tests scale with compute, as long as a failing test means something. Human attention does not.

Unless regression testing improves enough to catch side effects before they ship, more of them will reach production. Staged rollouts catch side effects the tests miss, and fast rollback limits the damage, but both act late, often after some users are affected. How much agent work a team can ship depends more and more on how many side effects its regression tests miss.

The checks that matter most for judging a change are the ones that were not written for it. An agent's own tests check what it foresaw, and an unintended side effect is exactly what it did not foresee. Regression tests, written for earlier changes by people or by AI, do not share the agent's expectations about this change.

Agents also game the checks they are given. The independent evaluation lab METR has documented this in frontier models. When a test fails, an agent sometimes edits the test or writes code that fools the test instead of fixing the problem. A smarter agent is also better at finding a way to make a check pass without doing the work. Proofs do not escape this. The proof checker catches every wrong step. But a proof can also declare a step true without proving it, and unless the setup forbids that, the checker accepts it. An agent that gets stuck can do exactly that, or, if it writes the statement to be proved as well as the code, prove something weaker than what was asked.

So a change has to be judged by checks the agent did not write or rewrite for it, and the agent must not be able to change those checks unnoticed. Anthropic's playbook for AI-native development instructs the agent: "If a test fails, fix the code, not the test." It recommends a hook that blocks edits to test files during a fix task, or a review that rejects any change touching a test. Outside fix tasks, the same applies: if an existing test has to change, someone other than the agent must approve it. New tests need no such approval, because adding a test leaves the existing ones as they are. Changing shared test setup or the build's test settings is different: it can weaken existing tests without touching them, so it needs the same approval. Finance has followed the same principle for a long time: the person who enters a payment does not approve it.

Smarter AI and formal verification will change a great deal about how software is built and checked. They will not remove the need to check every change against what worked before. With agents making more changes than ever, regression tests matter more, not less.

Is GenAI stressing your Continuous Delivery pipeline?

GenAI Will Stress Your Continuous Delivery Pipeline whitepaper

Share this blog post

© 2026 Gradle, Inc. Gradle®, Develocity®, Build Scan®, and the Gradlephant logo are registered trademarks of Gradle, Inc.

Get an AI summary of Develocity: