Flaky tests: why end-to-end suites break every sprint, and how to fix them
· 4 min read
A flaky test is a test that passes and fails on the same code without any change. Most flaky end-to-end tests are caused by timing (the test acts before the page is ready), brittle selectors (the test depends on markup that changes), shared or leftover test data, and unstable environments. Each cause has a specific fix, and the triage table below helps you find which one you are dealing with.
Why flaky tests are expensive
Flaky tests cost more than the minutes spent re-running them:
- Trust erodes. Once a suite "sometimes fails", people re-run until green and stop reading failures. Real bugs then pass through the same door.
- Releases slow down. Every red build becomes an investigation.
- Maintenance eats the sprint. Engineers spend their time repairing tests instead of shipping features.
Flaky test triage table
| Symptom | Likely cause | First fix |
|---|---|---|
| Fails on CI, passes locally | Timing, resources, environment differences | Replace fixed sleeps with waits for a visible state; check CI resources |
| "Element not found" after a UI change | Brittle selector | Target roles, labels or test IDs, not CSS paths or text that changes |
| Fails only when run with other tests | Shared state or test order | Give each test its own data; reset state between tests |
| Fails at certain times of day | Dates, time zones, scheduled jobs | Freeze or control time in the test; avoid "today" in assertions |
| Fails randomly on async actions | Race condition | Wait for the network call or UI state the step depends on |
| Fails when a third-party service is slow | External dependency | Stub it for functional tests; test the real integration separately |
| Fails after data cleanup | Leftover or missing test data | Create the data each test needs inside the test |
The five root causes, in detail
1. Timing and waits
Fixed sleeps (wait 2 seconds) are too long on a fast run and too short on a slow one. Wait for a condition, not a
duration: the button is enabled, the list has rows, the request finished.
2. Brittle selectors
A test that clicks div:nth-child(3) > button breaks when a designer adds a wrapper. Prefer accessible roles and
labels ("the Apply button") or dedicated test IDs. The deeper problem is that scripts describe how the page is built
instead of what the user is trying to do.
3. Test data
Tests that share accounts, carts or records interfere with each other. Each test should create what it needs and not depend on what the previous test left behind.
4. Environment instability
Shared staging environments get deployed to in the middle of a run, run out of resources or point at a service that is down. Pin the build under test and record which environment a failure happened in.
5. Genuine product race conditions
Sometimes the flake is real: the app has a race condition that users hit too. Before you "fix" a flaky test, check whether it found a bug. Suppressing it may hide the most valuable failure your suite ever produced.
A process that keeps flakiness under control
- Record every failure with a category: product bug, test problem or environment problem.
- Quarantine, don't delete: move a known-flaky test to a separate job so it doesn't block releases, and give it an owner and a deadline.
- Track the flake rate per test and fix the worst ones first.
- Never merge a new test that hasn't passed several consecutive runs.
Why scripted UI suites keep breaking
Even with perfect practices, a scripted suite is code that mirrors your UI. Every UI change becomes a test change. That is why teams with large suites often spend more time maintaining tests than writing new ones.
Two approaches reduce that burden:
- Self-healing locators try alternative selectors when one breaks. They cut maintenance but you still own the scripts.
- Intent-based execution describes steps in plain language ("apply the coupon WINTER25") and lets an agent find the right element in a real browser at run time. There is no selector for the team to maintain.
The second approach is part of what we call AI-native QA.
How testdart avoids script maintenance
testdart turns requirements and test scenarios directly into executable browser tests, so there is no script or framework for your team to write or maintain. Tests run in a real, headed Chrome browser, and each run records its steps and screenshots so you can see exactly what the agent did. When a test fails, the report explains why it failed and which requirement it affects, which makes "is this a real bug?" a quick question. See why teams switch.
FAQ
What is an acceptable flaky test rate? As close to zero as you can get for tests that block releases. Any known-flaky test should be quarantined with an owner.
Should I automatically retry failed tests? Retries hide flakiness. If you use them, record every retry and review tests that needed one.
Are flaky tests a sign of bad engineers? No. They are a sign that the suite depends on things that change: timing, markup, data and environments.
See it on your own flow.
testdart reads your requirements, writes the test cases, runs them in a real browser and shows you what broke.