A program can pass every unit test and the product can still be broken. Maybe the database schema is wrong, or two services disagree on an API. Or everything works individually, but the checkout flow still fails for the user.
That’s because unit tests, integration tests, and end-to-end tests each catch a different kind of bug.
Take an online store. A unit test checks one piece of logic on its own, such as whether the order-total function applies a 10% discount correctly, or gives free shipping on orders over $50. There’s no database and no other services involved. Which is why unit tests are fast enough to run every time the code changes.

An integration test checks whether the parts work together: whether the checkout API writes the right data to the database, or sends the payment service the fields it expects. Checkout logic can pass every unit test and still fail here, because the SQL is wrong or the two services disagree on an API format.

An end-to-end test does what a customer does. It puts something in the cart and checks out, then makes sure the confirmation page appears. That covers everything from the front end down to the services behind it.

The broader the test, the more of the real system it exercises. But it takes more work to set up and longer to run, and when it fails it’s harder to debug. A failing unit test points to a small amount of code. A failing end-to-end test could be failing almost anywhere along the path.
And much of that path isn’t our code. A browser might render a button late, or a payment sandbox might time out. Or the test environment might have drifted out of sync with production. These failures come and go, so the same test can pass on one run and fail on the next, and a failure doesn’t say whether the code or the environment is at fault. Teams often end up rerunning flaky tests until they pass, which teaches them to ignore failures.
Hence the testing pyramid: many small, fast tests at the bottom, fewer integration tests above them, and a small set of end-to-end tests for the most important workflows. A common rule of thumb, popularized by Google’s testing guidance, is roughly 70% unit, 20% integration, and 10% end-to-end. Treat those numbers as a starting point. A product that’s mostly glue between services may need more integration tests. Some teams argue for a “testing trophy” that puts most of the weight there, since integration tests catch the boundary bugs that cause real failures. Either way, the aim is to test each thing at the lowest layer that gives us the confidence we need.

AI writing test? Can we use agents?
So what changes when AI writes the code? Mostly the speed and the volume of change. An agent can make a large change, write tests for it, and keep fixing and rerunning until they pass.
The danger is an agent writing the implementation and then writing tests from that same implementation. It can carry one wrong assumption into both. The code is wrong and the tests agree with it. So everything passes.
Suppose the requirement says orders of $50 or more get free shipping. The agent writes if total > 50, then generates tests checking that a $60 order ships free and a $40 order doesn’t. Both pass. A customer whose order comes to exactly $50.00 gets charged for shipping, and no test notices, because the tests were derived from the code’s own reading of the rule.
So the expected behavior has to come from outside the code, from something like a specification or an API contract. Acceptance criteria work too. For important changes, we can add a separate review step, or a separate agent to verify the result. Don’t let the agent that built it grade its own work.
A second agent has a weakness of its own, though. It’s often the same kind of model, trained on similar data, so it can share the same blind spots and approve the same wrong assumption. It only helps if it checks against the specification instead of just reading the code. And it isn’t free: every change pays for it in cost and latency.
AI makes these feedback loops much cheaper to build and run. But it can’t tell us that a $50 order was supposed to ship free. Deciding what “correct” means is still our job.