Machine Learning
Synthetic data for open medical research: current challenges and perspectives
Published on
Open science practices have many advocated benefits. At their core is the sharing of open data, especially open research data. Although their value is widely recognized, such data are rarely available in the healthcare domain. One of the main barriers to their dissemination is the fear of statistical disclosure by data custodians. Synthetic data represent a promising way to address this concern by altering the dependency of the released data on the original ones. Despite their intuitive appeal, synthetic data themselves come with many difficulties that have so far prevented their widespread adoption. In this article, we review the main ethical, technical, and logistical challenges associated with data synthesis for open medical research, and discuss some promising directions to overcome them. These current challenges and lines of research are placed in historical perspective and illustrated with several examples formalized within a unified statistical framework.