Using images from open-access medical literature to train AI models raises a fair question: is this actually ethical, given that many of these images ultimately originate from real patients?
The short answer is that published, open-access medical literature has already been through an ethics and consent process before it reaches a public archive. Case studies and research papers that include patient images are governed by journal ethics policies, typically requiring documented patient consent for publication, or de-identification sufficient to remove identifying features. This isn't a step we add, it's a prerequisite that already exists in academic medical publishing, and it's part of why open-access literature is a meaningfully different category from, say, scraping medical images from the open web.
That said, "the journal handled consent" isn't a reason to stop thinking about ethics at the dataset-building stage. A few things matter beyond the initial publication approval. First, aggregation changes context: an image that was ethically appropriate to publish as one case study in one journal takes on a different character when it becomes one of a million pairs in a training corpus used to build commercial or research AI systems. We think this shift deserves acknowledgment, not just a pass-through assumption that original publication consent covers every downstream use indefinitely.
Second, de-identification in published literature isn't always perfect. Manual review, the same process used to assess medical relevance, is also an opportunity to flag content where identifying information appears to have persisted despite a journal's de-identification efforts, and to exclude it rather than propagate the gap forward.
Third, transparency about sourcing, discussed elsewhere in this series, is itself an ethical practice, not just a scientific one. Being clear about where data comes from allows the broader community, including patient advocacy groups, ethicists, and other researchers, to scrutinize and challenge sourcing practices that don't hold up, which is a meaningful check that disappears when provenance isn't disclosed.
None of this makes ethical questions in medical AI data sourcing fully resolved, it's an area the field continues to work through as AI applications expand faster than governance frameworks always keep pace with. But starting from genuinely open-access, already-published, already-consented literature is a materially different and more defensible starting point than sourcing that skips those steps. It's also why we consider ethical sourcing and dataset transparency to be connected practices, not separate ones.

Comments
Post a Comment