Machine Learning

Synthetic data for open medical research: what practitioners should know

Publié le

Auteurs : Rémy Chapelle, Bruno Falissard, Mohammed Sedki

Publicly releasing research data is at the core of the open science principles. In medical research, however, this practice conflicts with confidentiality requirements for the collected data. One appealing solution to this dilemma is to release altered versions of these data instead of the original ones. Synthetic data are gaining increasing attention for such uses by clinicians, statisticians, and ethicists in the healthcare domain; yet, this interest is slow to materialize. One reason for that is a lack of familiarity among practitioners with synthetic data, in part explained by their abstract and technical nature. In this article, we address this need by reviewing, in a rigorous yet accessible way, the main ethical, technical, and logistical aspects of data synthesis that may be relevant to non-specialist practitioners who wish to engage with synthetic data in support of open research. In addition to the general principles for generating and evaluating synthetic data, this includes the current challenges that still hinder their democratization, as well as some research directions to overcome them.