Flaky tests are a data problem
before they are a code problem.
When an automated test passes on Monday and fails on Tuesday with no code change, teams usually blame the script. More often, the data or environment underneath it changed. Fix the data first and many flaky tests stabilise on their own.
Summarise this blog post with:
Every automation team knows the moment. A test that passed yesterday fails today. Nobody changed the code or the script. Someone reruns it and it passes. The defect is closed as “cannot reproduce”, the test is tagged as flaky, and the team quietly stops trusting it. Multiply that by dozens of tests and the automation suite becomes something people run and then check by hand.
The usual response is to look at the script: add waits, retries, better locators. Sometimes that is the answer. But in enterprise systems, the cause is frequently somewhere else entirely. The data the test depended on changed.
Why flaky tests matter more than they seem
Flaky tests do more damage than their numbers suggest. Each false failure costs investigation time. More importantly, each one teaches the team that a red result might not mean anything. Once that belief takes hold, genuine failures are dismissed as noise, and defects slip through an automation suite that technically caught them.
Flakiness also hides the return on automation. When teams rerun suites manually to be sure, the same tests are executed twice and the savings disappear. Restoring trust is therefore not a nice-to-have. It is what makes automation worth having.
There is a cultural cost too. Engineers who spend their days investigating false failures lose confidence in automation as a practice, and new automation initiatives face scepticism before they start. Fixing flakiness is often the fastest way to rebuild support for automation across a delivery organisation.
How data causes flakiness
Data-driven flakiness takes a few recurring forms:
- Shared data: several tests, or several testers, use the same records. One test changes a customer’s status, and another test that expects the original status fails.
- Consumed data: a test uses up the data it needs, such as closing an account or redeeming a voucher, so the next run finds nothing to work with.
- Time-sensitive data: records age past a threshold, dates cross a billing cycle or a policy expires, and expected results no longer apply.
- Refreshed environments: a scheduled data refresh replaces the records tests were built around, without anyone updating the tests.
- Order dependence: a test only passes when another runs first to create its data, so running tests in parallel or a different order causes failures.
None of these show up in the script. All of them produce exactly the pattern teams call flaky.
A test that fails without a code change is telling you something changed. Often it is the data, not the software.
Give every test the data it needs
The most effective fix is to make each test responsible for its own data. Rather than relying on records that happen to exist in an environment, a test should create or reserve the data it needs, use it and leave the environment in a known state. Where creation through the interface is slow, data can be prepared through APIs or directly in the database before tests run.
For large regression packs, this is where automated test data generation pays back. Generating full test condition variations from a protected source means each run starts with the data it expects, in the volumes it needs. PinnacleQM’s HealthTest does this for healthcare, creating scripts, expected results and condition variations from obfuscated records, and running millions of checks a day with consistent results.
Make data repeatable and protected
Repeatable data is data that can be restored to the same state on demand. That means versioning test data sets alongside the tests that use them, and being able to refresh an environment to a known baseline before a run. When a refresh is required for other reasons, tests and data should be updated together.
Protection matters too. Many organisations reach for production copies to get realistic data, which introduces both privacy risk and instability, because production data keeps changing. Obfuscated data with referential integrity preserved gives realism without exposure, and because it is prepared deliberately, it can be kept stable between runs.
Control environment state
Data is the most common cause of flakiness, but environment state is close behind. Shared environments where other teams deploy, configure or test at the same time introduce changes nobody planned for. Integrations to external services may be slow, unavailable or return different results. Scheduled jobs may run midway through a test.
Mature teams manage these deliberately: environment booking and readiness checks before runs, service virtualisation or stubs for unreliable external dependencies, and awareness of batch schedules. When environment instability cannot be removed, it should at least be visible, so a failure caused by an unavailable service is reported as such rather than as a defect.
Recording environment state alongside test results helps with diagnosis. When a failure occurs, knowing which build was deployed, which data baseline was loaded and which dependencies were available turns a vague intermittent failure into a specific, explainable event.
Measure flakiness, then remove it
What gets measured gets fixed. Track the rate of false failures across the suite, the tests responsible for most of them and the root cause of each: data, environment, timing or script. Most suites follow a familiar pattern where a small number of tests cause most of the noise. Fixing those first restores trust quickly.
Script improvements still have their place. Automation that separates business intent from technical detail, and adapts when applications change, removes a class of failures caused by interface changes. When one bank’s core platform upgrade changed 1,383 screens, PinnacleQM’s Enginuity relearned them in two hours and the regression pack ran without script changes. Combined with disciplined data, that is what reliable automation looks like.
Trust is the goal
The aim is not zero failures. It is failures that mean something. When data and environments are controlled, a red result points to a real change in the software, and teams act on it with confidence. That is when automation stops being a second job and starts delivering the speed it promised.
Assurance practice lead
Works with programme sponsors on go-live decisions and independent assurance across banking, government and utilities.
Flaky tests,
answered plainly.
Common questions about making automated results reliable.
What is a flaky test?
A flaky test is an automated test that sometimes passes and sometimes fails without any change to the software under test. Because its results are unreliable, teams stop trusting it, investigate false failures by hand and may dismiss genuine failures as noise, which undermines the value of the whole automation suite.
Are flaky tests usually caused by bad scripts?
Sometimes, but not usually in enterprise systems. Shared, consumed or time-sensitive test data, environment refreshes, order dependence between tests and unstable shared environments are common causes. These produce the same intermittent pattern as script problems, so it is worth checking data and environment state before rewriting tests.
How do we stop tests interfering with each other's data?
Make each test responsible for its own data. Tests should create or reserve the records they need, use them and leave the environment in a known state, rather than relying on data that happens to exist. Preparing data through APIs or the database keeps this fast, and it allows tests to run in parallel safely.
Should we use production data to make tests more realistic?
Production copies add realism but also privacy risk and instability, because production data keeps changing. Obfuscated data with referential integrity preserved gives realistic behaviour without exposing personal information, and because it is prepared deliberately, it can be kept stable and restored to a known baseline between test runs.
How should we track and reduce flakiness?
Measure the false failure rate across the suite, identify which tests cause most of the noise and record the root cause of each: data, environment, timing or script. Fixing the small number of worst offenders first restores trust quickly. Then address systemic causes, such as data management and environment control, to prevent new flakiness.
More on reliable automation.
7 min read
A test data strategy that survives privacy review
How to give teams realistic test data while meeting privacy obligations, including obfuscation, synthetic data and residency.
7 min read
How much regression testing can you safely automate?
A practical way to decide which regression tests to automate, which to keep manual, and how to measure the return.
7 min read
Why DevOps maturity, not tooling, decides your release cadence
Tools do not make releases faster by themselves. The practices, ownership and feedback loops that actually change cadence.
Make your
automation results trustworthy.
If your team reruns automated tests by hand to be sure, we can find the causes of flakiness and help restore trust in your suite.
- Share your automation suite and recent failure patterns.
- We analyse root causes across data, environments and scripts.
- You receive a prioritised plan to stabilise your suite.