Mutation Testing with mutmut: Which Bugs Your Tests Ignore
Measure whether your AI-written tests actually verify behaviour using mutmut
Coverage says a line ran. Mutation testing says whether anything would notice if that line were wrong. The idea takes a sentence; making it finish before you lose interest takes a plan.

Coverage tells you a line ran. Mutation testing tells you whether anything would have noticed if that line were wrong — which is the question you actually had, and the one that matters most when the tests were written by a model rather than by someone who understood the requirement.
The idea is simple enough to explain in a sentence. Getting it to finish before you lose interest is the part that needs a plan, and it is the reason most teams try it once and abandon it.
How it works, in one function
A mutation tool makes small, deliberate edits to your source — flipping a comparison, changing a constant, removing a statement — and runs your test suite against each edited copy. Each copy is a mutant.

If your suite fails, the mutant is killed: your tests noticed the bug. If your suite still passes, the mutant survived, and you have found a change to your code that no test objects to.
Survivors are the output. Not the percentage — the list. Each one is a specific sentence of the form "you could change this and nothing would complain", and you either write a test or decide you do not care.
Running it
pip install mutmutConfigure what to mutate, and be specific. The single most common first mistake is pointing it at the whole package:
# pyproject.toml
[tool.mutmut]
paths_to_mutate = [ "src/billing/" ]
tests_dir = [ "tests/billing/" ]mutmut run
mutmut results # the survivors
mutmut show 47 # the diff for one of them
mutmut browse # interactive, and the nicest way to work through a listFlags and configuration keys have moved between mutmut's major versions — the 3.x line reworked both substantially — so check mutmut --help against the version you installed rather than trusting a snippet, this one included.
The arithmetic that decides whether this is usable
This is the part that gets skipped, and it is why mutation testing has a reputation for being impractical.
The cost is number of mutants × time for one test run. Both numbers are larger than people expect. A modest module produces hundreds of mutants; a mid-sized codebase produces thousands. If your suite takes thirty seconds and the tool generates two thousand mutants, that is roughly seventeen hours.
Three things bring it back into the realm of the possible, and you generally want all three:
- Mutate one module, not the codebase. Pick the code where a silent bug would actually hurt — pricing, permissions, anything that touches money or access — and leave the rest alone. Mutation testing your serialisation helpers is a poor use of an afternoon.
- Point it at the matching tests. If
tests_diris your whole suite, every mutant runs every unrelated test. Narrowing this is usually the single biggest speed-up available. - Run it on the diff, not the repository. In CI, mutate only files changed in the pull request. That turns an overnight job into a few minutes and puts the result in front of the person who can act on it.
Run the full thing once as a baseline, on a weekend or overnight, and then never again — the useful ongoing signal is per-change.
Reading a survivor properly
Say mutmut reports this:
--- src/billing/discount.py
+++ mutant 47
- if pct > 50:
+ if pct >= 50:
raise ValueError("discount too large")Your suite passes either way. That tells you something precise: no test exercises a discount of exactly 50. The boundary is untested, and the boundary is where this kind of bug lives. Whether 50 should be allowed or rejected is a product question — the point is that your tests currently have no opinion, so either answer would ship.
The fix is one test, and it is a better test than the three that were already there:
def test_fifty_percent_is_allowed():
assert apply_discount(100, 50) == 50.00
def test_above_fifty_is_rejected():
with pytest.raises(ValueError):
apply_discount(100, 50.01)Do that half a dozen times and you notice the pattern: survivors cluster on boundaries, on error branches, and on conditions with more than one term. Those are exactly the places example-based tests — and models writing example-based tests — tend not to go.
Equivalent mutants, and why 100% is the wrong goal
Some survivors cannot be killed, because the mutation did not change behaviour at all.
# Original # Mutant — identical behaviour
for i in range(len(items)): for i in range(0, len(items)):
... ...No test can distinguish those, because there is nothing to distinguish. These are equivalent mutants, and detecting them automatically is undecidable in the general case — it reduces to program equivalence. No tool will ever solve this for you; every mutation testing run produces some proportion of survivors that are simply noise.
The practical consequence is that chasing a 100% mutation score is worse than useless. It sends you writing contorted tests to kill mutants that represent no real defect. Treat the score as a trend on a module you care about, and treat the survivor list as the deliverable.
When a survivor is genuinely equivalent, or genuinely not worth a test, exclude it and move on — with a comment saying why, so the next person does not re-litigate it.
Why this pairs specifically well with generated tests
A model writing tests optimises, in effect, for tests that pass. It has seen the implementation, it writes assertions consistent with the implementation, and it produces a suite that is green on day one and silent on the day something breaks.
Mutation testing measures precisely that gap, because it does not care whether tests pass on correct code — only whether they fail on incorrect code. It is the one tool that answers the question a careful review of generated tests is trying to answer, without relying on a reviewer's patience holding out to the fortieth test.
A workable routine: generate the tests, run coverage to find what was missed entirely, then run mutation testing on the two or three modules that matter to find what was covered but not checked. The second step routinely finds things the first declared fine.
Where it fits, and where it does not
- Worth it: business logic with rules and boundaries — pricing, entitlements, validation, permissions, anything with money or access in it.
- Not worth it: glue code, thin wrappers, serialisation, anything where the tests are effectively checking that a library still works.
- Not a substitute for: integration tests, or for having thought about what the code is meant to do. A suite can have a high mutation score and still test the wrong requirement thoroughly.
Versions and scope
Commands and configuration above target recent mutmut; the 3.x releases changed configuration handling and the results interface substantially, so verify against your installed version. Equivalent alternatives exist elsewhere — Stryker for JavaScript, TypeScript and .NET, PIT for the JVM — and the reasoning transfers directly.
No runtime figures here are measurements of ours. The seventeen-hour example is arithmetic from the stated assumptions rather than an observed run, and your numbers will depend entirely on suite speed and mutant count.


