Six steps to achieving trustworthy synthetic data in healthcare

Høvik, Norway, XX June 2026 – A new position paper from DNV outlines a structured approach to building trust in synthetic data used in healthcare, identifying six practical steps to support evidence-based evaluation and responsible adoption. 

Synthetic data, artificially generated data designed to reflect real-world patterns, is increasingly being explored to support artificial intelligence (AI) innovation in healthcare, particularly where access to sensitive patient data is limited. However, ensuring that such data are reliable, safe, and fit for purpose remains a key challenge. 

The paper called How can I trust my fake data? sets out a six-step approach to evaluating synthetic data quality across its lifecycle. These steps include defining the intended use, assessing training data quality, evaluating generation performance, validating synthetic data against real-world benchmarks, carrying out analytical and clinical validation, and documenting outcomes through structured “synthetic data cards.” 

“Synthetic data holds significant potential, but its safe use depends on structured and transparent evaluation,” said Serena Elizabeth Marshall, Principal Researcher at DNV. “By grounding assessments in evidence and aligning them with regulatory expectations, organisations can better understand whether data is fit for purpose.” 

The recommended approach begins with clearly defining the purpose of the synthetic dataset, which determines relevant quality criteria and metrics. Quality must then be assessed at multiple stages – from the underlying training data to the performance of downstream applications – ensuring that risks relating to bias, privacy, and representativity are actively managed.  

A key element of the framework is the evaluation of synthetic data against five dimensions of quality: similarity, utility, privacy, fairness, and environmental impact. These dimensions provide a comprehensive basis for assessing whether a dataset can be safely and effectively used in healthcare contexts. 

The paper also emphasises the importance of transparency. It introduces “synthetic data cards” as a standardised way to document how data were generated, the assumptions made, evaluation results, and known limitations. This documentation supports informed decision-making by developers, regulators, and users. 

DNV’s position is that building trust in synthetic data requires robust, evidence-based evaluation combined with transparent documentation and multidisciplinary collaboration. This approach can help ensure that synthetic datasets are suitable for their intended use and compliant with existing regulations governing healthcare, AI, and data protection. 

“Synthetic data opens up new possibilities across medicine – from improving the development of diagnostic tools to enabling research into rare diseases where real-world data are limited,” said Carla Ferreira, Team Leader, AI Risk and Assurance Science at DNV. “With the right validation and safeguards in place, it can support more robust testing, faster innovation, and ultimately better-informed clinical decisions.”