Flaky Tests, Found and Fixed
An agent finds unstable tests and fixes them, or reports why they fail, so a red build means something again.
Flaky Tests Teach the Team to Ignore Red Builds
A test that fails one run in five gets retried, then ignored, then skipped. Every flaky test costs a rerun and a little trust in the suite, and fixing one means reproducing a failure that doesn't happen on demand. It's rarely anyone's priority.
An Agent Chases the Failures Nobody Has Time For
On a schedule, an agent looks at unstable tests, runs them in an isolated environment, and classifies each failure as a product, test, or environment issue. When the cause is in the test, it opens a pull request with the fix. When it isn't, it reports why the test fails.
auth › refreshes session
Cause: race condition in test setup
Common Questions
How does it find flaky tests?
It runs the tests you point it to, repeatedly if needed, and classifies each failure as a product, test, or environment issue.
Will it just disable a flaky test?
It fixes the test when the cause is clear, or reports why it fails. Any change arrives as a pull request for a person to review.
Can it run on a schedule?
Yes. Run it nightly or weekly with a cron schedule, or start it by hand.
Which models can it use?
Any major provider, plus local and private models. Bring your own keys or subscription, and there's no markup on tokens.
Related: Failed build triage, Quality assurance docs.
Keep Your Attention for the Hard Problems
Start with one repo and one recurring task.