WSEAS Transactions on Computer Research
Print ISSN: 1991-8755, E-ISSN: 2415-1521
Volume 13, 2025
From Ultrasound Image Collection to De-identification and Re-identification: A Practical Pipeline
Authors: , ,
Search Articles
Abstract: Many AI research initiatives consider medical images a crucial resource to improve or enhance healthcare outcomes. The lack of high-resolution real-world image datasets, detailed annotations, and clinical relevance forces researchers to use public datasets as an alternative. The latter often impacts the accuracy of results and impedes further advancements of AI in this field. Meanwhile, in limited scenarios where researchers can collect real-world data, ensuring patient privacy becomes their primary concern. To minimize the risk of private information disclosure, images must be de-identified in a way that preserves their research value. Numerous studies focusing on de-identification approaches are available in the literature. However, there are often gaps or missing points in creating a real valuable dataset because simply de-identifying images is not sufficient. Creating medical image datasets for AI research projects involves many steps beyond just protecting patient identity. This study contributes to the existing research by presenting a comprehensive process for creating a clean and safe ultrasound images dataset, using real data as a basis. The authors introduce a real-world pipeline named UltraSafe, which serves as a semi-automated or automated tool that considers all the necessary steps, such as on-site ultrasound data collection from a private clinic, data cleaning, annotation, de-identification, and re-identification.
Keywords:
Ultrasound images, Data collection, De-identification, Re-identification, Artificial intelligence, Protected health information
Pages: 644-652
DOI: 10.37394/232018.2025.13.57