More data, more images, more simulated cases: in the field of artificial intelligence, abundance is often presented as a condition for performance. The idea seems logical. A model learns from examples; the more it encounters, the better it should understand the world. But a medical image is not an ordinary example. It carries an anatomy, a possible pathology, an acquisition protocol and a clinical context. It may contribute to a decision about surveillance, treatment or intervention. Its value therefore depends not only on its appearance, but on the quality of the information it contains. Generative models nevertheless open up considerable possibilities. They can produce synthetic images to enrich databases, represent rare diseases, compensate for the lack of certain patient profiles or test algorithms in situations that are difficult to assemble in reality. The trap would be to believe that each new image automatically improves learning.
Quantity can give a false impression of diversity
A dataset can contain thousands of images while remaining narrow. The examinations may come from the same machines, the same protocols, the same institutions or from populations with little diversity. In this case, generating more images from this base does not necessarily correct its limitations. The generative model may simply produce new variations of an already incomplete world. The difference between visual variety and clinical diversity then becomes essential. Slightly modifying a texture, a contrast or the shape of a lesion creates a new image. This does not necessarily mean that the model will encounter a new and clinically relevant situation. Useful diversity concerns ages, body types, disease stages, scanners, protocols, artifacts and the atypical forms of a given pathology. A generated image is only of interest if it introduces a variation that truly matters for the task under study.Synthetic data do not replace real data
Image generation can be particularly valuable when data are scarce or difficult to annotate. It can enrich a training set or help better represent certain cases. But it does not allow one to dispense with the real. A generative model learns from the data it is given. If certain populations or certain forms of disease are absent, it cannot faithfully recreate them through imagination alone. It is more likely to produce an approximation based on what it already knows. Synthetic data can therefore complement real data, but they do not replace the collection of examinations that are diverse, well annotated and representative of clinical practice.Producing in order to answer a question
Before generating thousands of images, one must define their function. Is the goal to train a model to detect a lesion? To test its robustness to noise or to patient motion? To simulate a rare pathology? To better represent an underrepresented population? Each objective imposes different requirements. To train a model for pulmonary nodule detection, for example, it is not enough to create round shapes within a lung. The nodules must present plausible sizes, densities, locations and anatomical relationships. A few precisely controlled images may then hold more value than thousands of generic outputs.From accumulation to selection
True maturity will probably not consist in generating the largest possible number of medical images. It will consist in selecting those that meet a specific need, verifying their consistency and measuring their effect on the trained models. Synthetic images can enrich medical artificial intelligence. But abundance, on its own, is not proof of quality. In imaging, seeing more does not necessarily mean seeing better. Everything depends on what the new images actually make it possible to learn.