Contact Us
Performance

Why your performance test
environment lies to you.

A performance test is only as honest as the environment it runs in. Smaller servers, thinner data, stubbed integrations and quiet networks all flatter results. Here is how to spot the distortions and still get numbers you can trust.

7 min read

Summarise this blog post with:

The performance tests passed comfortably. Response times were well within target at the expected peak. Two weeks after go-live, the first month-end run takes three times longer than planned and online users see timeouts. The performance team re-runs the same tests in the same environment and gets the same comfortable results.

Nothing was wrong with the tests. The environment they ran in was not the one the business uses. It flattered the system, and nobody quantified by how much.

Why environments differ

Full production replicas are expensive, and most organisations make reasonable compromises. The trouble is that each compromise changes performance, and the effects are rarely documented. A test environment that is “production-like” in architecture can still differ in almost every factor that determines how fast the system responds.

The compromises also drift. An environment that closely matched production when it was built may fall behind as production grows, infrastructure is upgraded and configuration changes are applied in one place but not the other. Without regular comparison, a once reliable environment quietly becomes misleading.

The usual distortions

The differences that most often mislead performance results are consistent across organisations:

  • Infrastructure: fewer or smaller servers, different storage, shared virtual hosts or different cloud instance types.
  • Data volume: databases a fraction of production size, so queries, indexes and caches behave very differently.
  • Integrations: external services replaced by stubs that respond instantly and never fail.
  • Network: test traffic that never crosses the firewalls, load balancers, gateways and distances production traffic does.
  • Configuration: different connection pools, thread limits, logging levels, caching settings or security controls.
  • Background load: no batch jobs, reporting, backups or other applications competing for the same resources.

Most of these make test results look better than production will. Some, such as smaller servers, can make them look worse. Either way, the numbers are not the numbers the business will experience.

Nothing was wrong with the tests. The environment they ran in was not the one the business uses.

Data volume: the most underestimated gap

Of all the differences, data volume is the one most often underestimated. A query that returns instantly against ten thousand rows can take seconds against ten million. Indexes that are never needed on small tables become critical on large ones. Caches that hold the entire test data set cannot hold production’s.

Scaling test data to production volumes closes much of this gap. Because copying production introduces privacy risk, the practical approach is to scale obfuscated or synthetic data to match production size and distribution. PinnacleQM’s Enigma platform scales data sets this way for performance and scalability testing, keeping volumes realistic without exposing personal information.

Integrations: the invisible bottleneck

Stubs are necessary when partner systems are unavailable, but a stub that responds in a millisecond hides one of the most common causes of production slowness. Real integrations add latency, occasionally time out and have their own capacity limits.

Better stubs help. Service virtualisation can simulate realistic response times, variability and failure rates based on measured behaviour of the real service. Where possible, a subset of tests should run against real integrations, even at reduced volume, to calibrate the simulations.

Integration limits matter as well as integration speed. A partner service may accept only a certain number of concurrent requests or throttle traffic above an agreed rate. If the stub has no such limit, a test can generate traffic the real partner would reject, and the resulting failures will only appear in production.

Making results trustworthy anyway

Few organisations can afford a perfect replica, and they do not need one. They need to know how their environment differs and what that means for results. A practical approach has four parts.

First, document the differences between test and production, factor by factor, and estimate the likely direction and scale of each effect. Second, close the gaps that matter most, usually data volume, integration behaviour and key configuration settings, which are often cheaper to fix than infrastructure. Third, interpret results in light of the remaining gaps rather than reporting them as production predictions; state clearly what the environment can and cannot tell you. Fourth, add background load where production has it, such as batch processing at month-end, so tests reflect the conditions that matter to the business.

Reporting should make these adjustments visible. A result stated as “meets target in the test environment; production risk moderate because integration latency was simulated” is far more useful than a green tick. It tells decision-makers what the test proved, what it assumed and where to watch closely after release.

Confirm in production, carefully

The final check is measurement where users actually are. End-user monitoring in production, and synthetic journeys run on a schedule, show whether real performance matches what testing predicted. PinnacleQM’s UXperform verifies, tracks and reports end-user performance across systems and devices, which gives teams a feedback loop between test results and real experience.

That feedback loop is what improves the environment over time. When a production measurement differs from a test result, the reason usually points to an environment gap that can be documented and, where worthwhile, closed before the next release.

Honest numbers beat comfortable ones

A performance environment does not need to be perfect to be useful. It needs to be understood. Teams that know how their environment differs from production, close the gaps that matter and interpret results honestly will make far better capacity decisions than teams that trust comfortable numbers from an environment nobody has examined.

Assurance practice lead

Works with programme sponsors on go-live decisions and independent assurance across banking, government and utilities.

Straight answers

Performance environments,
answered plainly.

Common questions about trustworthy performance results.

Does a performance test environment need to match production exactly?

Not exactly, but its differences from production need to be known and understood. Document how infrastructure, data volume, integrations, network, configuration and background load differ, close the gaps that matter most, and interpret results with the remaining differences in mind. A well-understood compromise is more useful than an unexamined replica.

Why do performance tests pass but production is slow?

Usually because the test environment flatters the system. Smaller data volumes make queries faster, stubbed integrations respond instantly, test traffic avoids production network paths, and there is no competing batch or reporting load. Each difference improves test results, so the combined effect can hide problems that only appear in production.

How much test data do we need for performance testing?

Enough to match production size and distribution for the tables and data sets that drive performance. Small databases make queries, indexes and caches behave differently. Scaling obfuscated or synthetic data to production volumes gives realistic behaviour without the privacy risk of copying production data into a test environment.

How should integrations be handled in performance tests?

Use service virtualisation that simulates realistic response times, variability and occasional failures based on measured behaviour of the real services, rather than stubs that respond instantly. Where possible, run a subset of tests against real integrations, even at reduced volume, to calibrate the simulations and confirm their accuracy.

How can we check whether our performance testing was accurate?

Measure end-user performance in production after release, using monitoring and scheduled synthetic journeys, and compare it with test predictions. Differences usually point to specific environment gaps, such as data volume or integration latency, that can be documented and, where worthwhile, closed before the next release.

Get performance numbers
you can trust.

If your performance tests pass but production tells a different story, we can help you understand your environment and make results reliable.

  1. Tell us about your environments and recent results.
  2. We assess the gaps between test and production.
  3. You receive a plan to make results trustworthy.