Contact Us
Test data and privacy

A test data strategy
that survives privacy review.

Testers want production data because it finds real defects. Privacy teams refuse it because it exposes real people. A good test data strategy gives both sides what they need, and it starts with a few clear principles.

7 min read

Summarise this blog post with:

Most organisations have had the same argument. A delivery team asks for a copy of production data because their tests keep missing defects that only appear with real records. The privacy officer says no, because the copy would expose customers, patients or staff to people who have no need to see their information. Both are right. Without a strategy that reconciles them, the argument repeats on every project, and one side usually loses in ways that cost more later.

Why copying production fails

Copying production data into test environments is the fastest path to realistic testing, and the fastest path to a privacy incident. Test environments typically have wider access than production, weaker monitoring and longer retention. They are used by contractors, partners and sometimes offshore teams. Privacy obligations do not relax because data is being used for testing, and an exposure from a test system is still an exposure.

There is also a practical problem. Once a production copy exists in a test environment, it tends to be copied again, refreshed irregularly and forgotten. Nobody can say with confidence where personal information now lives.

Copies also create obligations that outlive the project. Data retained in test environments must still be protected, reported in data inventories and included in breach assessments. When an organisation is asked where a person’s information is held, forgotten test copies make that question hard to answer honestly.

Why simple masking and invented data fail too

The usual alternatives have their own weaknesses. Simple masking replaces names and identifiers, but it often breaks the relationships between records. A customer’s accounts no longer link to the customer, or a patient’s appointments no longer link to the patient, so tests fail for reasons that have nothing to do with the software. Masking also misses personal information in free-text fields, comments and notes.

Hand-made test data avoids exposure but rarely reflects reality. It lacks the volumes, edge cases and messy combinations that production contains, so defects that only appear with real-world data slip through to go-live.

Principles of a strategy that holds up

A test data strategy that survives privacy review rests on a small number of principles:

  • Synthetic or obfuscated by default: production data is never used raw in test environments unless there is a documented, approved exception.
  • Referential integrity preserved: obfuscation keeps relationships between records intact, so data behaves as it would in production.
  • Free text included: personal information in comments, notes and nicknames is found and replaced, not just structured fields.
  • Volume on demand: data can be scaled to production volumes for performance and migration testing without copying more real records.
  • Residency respected: data stays in its country of origin where obligations require it, including data used by offshore teams.
  • Governed and evidenced: access, handling and retention are controlled, and the controls can be demonstrated to auditors.

Privacy obligations do not relax because data is being used for testing. An exposure from a test system is still an exposure.

Obfuscation that keeps data useful

Smart obfuscation is the technique that reconciles realism and privacy. It replaces identifying values consistently, so the same customer is replaced the same way everywhere they appear, and relationships across tables and systems survive. It detects personal information in free text and replaces it. And it keeps values plausible, so validation rules, calculations and reports behave normally.

PinnacleQM’s Enigma platform works this way, and its healthcare counterpart, HealthTest, applies the same principles to patient records and HL7 messages. One senior integration project manager used Enigma to obfuscate millions of patient records daily and to demonstrate that a $200 million programme complied with data privacy regulations. In another engagement, obfuscated production records allowed a state healthcare provider to test more than five million transactions a day.

Synthetic data and scaling

Obfuscation is not the only tool. Synthetic data, generated to match production patterns rather than derived from real records, removes privacy risk almost entirely and is ideal for training environments, early development and new systems with no production history. Scaling techniques then grow either obfuscated or synthetic data to the volumes performance and migration tests need, so capacity limits are found without copying more real records.

Most mature strategies use both: obfuscated data where realism depends on real-world patterns, and synthetic data where it does not.

The choice between them should be made per environment, not per project. A development sandbox may only ever need synthetic data, while an integration test environment may need obfuscated data that reflects real-world combinations. Recording that decision, and the reason for it, is part of the evidence privacy teams will want to see.

Governance a privacy officer can check

Technology is only half the strategy. The other half is governance that privacy and security teams can verify. Document which data sources feed which environments and how each is protected. Control who can access test data and review that access regularly. Set retention periods and enforce them automatically. Keep records showing that obfuscation was applied and verified before data was released to a test environment.

Where offshore or partner teams are involved, introduce them only after security, privacy, data-access and residency controls are agreed, and give them obfuscated data by default. Management systems aligned with standards such as ISO/IEC 27001 and ISO/IEC 27701 give this governance a recognised structure.

Ending the argument

When a strategy like this is in place, the argument between testers and privacy officers largely disappears. Testers get data that behaves like production. Privacy officers get evidence that no real person is exposed. Projects stop negotiating data access from scratch, and test cycles start on time because data is prepared in advance rather than requested at the last minute.

The strategy does not need to be perfect on day one. Start with the most sensitive systems, apply the principles, prove them, and extend. Each project that follows inherits a process that already passes review.

It also makes change easier. When a new system, vendor or offshore team joins, the rules for what data they receive and how it is protected are already written. Onboarding becomes a matter of applying an agreed standard rather than reopening a debate, which shortens lead times and removes a recurring source of friction between delivery and privacy teams.

Assurance practice lead

Works with programme sponsors on go-live decisions and independent assurance across banking, government and utilities.

Straight answers

Test data privacy,
answered plainly.

Common questions about realistic, privacy-safe test data.

Can we use production data for testing if access is restricted?

Restricting access reduces risk but does not remove it. Test environments usually have wider access, weaker monitoring and longer retention than production, and privacy obligations still apply. The safer default is obfuscated or synthetic data, with any use of raw production data treated as a documented, approved exception with specific controls.

What is the difference between data masking and smart obfuscation?

Simple masking replaces sensitive values, often inconsistently, and can break the relationships between records. Smart obfuscation replaces identifying values consistently across tables and systems, preserves referential integrity, keeps values plausible and finds personal information in free-text fields, so tests behave as they would with real data.

When should we use synthetic data instead of obfuscated data?

Synthetic data suits training environments, early development and new systems without production history, because it contains no real personal information at all. Obfuscated data suits testing where realism depends on real-world patterns, edge cases and volumes. Most mature test data strategies use both, chosen according to the purpose of each environment.

Can offshore teams use test data safely?

Yes, with the right controls. Offshore teams should be introduced only after security, privacy, data-access and residency controls are agreed, and should work with obfuscated or synthetic data by default. Where residency rules require it, the data itself stays in its country of origin while offshore teams access it under controlled conditions.

How do we prove our test data is compliant?

Keep records of data sources, obfuscation rules and verification results, and show that personal information was removed before data entered each test environment. Document access controls, regular access reviews and enforced retention periods. Management systems aligned with ISO/IEC 27001 and ISO/IEC 27701 give auditors a recognised framework to assess.

Give teams data
privacy can approve.

If test data is slowing your projects or worrying your privacy team, we can help you design a strategy that satisfies both.

  1. Share your systems, data sensitivity and current approach.
  2. We assess obfuscation, synthetic data and governance options.
  3. You receive a practical test data strategy and plan.