Design of an Iterative Explainable AI-Guided Watermarking for Transparent and Clinically Safe Medical Image Security
Keywords:
Watermarking, Explainable AI, Medical Imaging, Interpretability, Image Security, Counterfactual ValidationAbstract
The authentication of medical images in a system that spans hospitals, teleradiology, and AI diagnostic pipelines must be secure, auditable, and safe. Existing watermarking methods attach ownership signatures that do not consider the radiological salience of parts of the image; embedding location decisions are neither transparent nor is their diagnostic safety considered. In this study, we propose and outline a 5-stage explainable watermarking pipeline for medical images, which, to our knowledge, has not been defined in the literature in this composite form. It uses the Saliency Synthesis Watermark Prior (SSWP) to combine Grad-CAM, Guided Backpropagation, DeepSHAP, and clinician-relevant maps to find radiologically safe regions into which to embed the watermarked signal. It uses an Interpretable Wavelet-Attention Watermark Encoder (IWA-WE), which embeds the watermarked signal into diagnostically neutral wavelet coefficients. The Counterfactual Embedding Validation Network (CEV-Net) quantifies the prediction changes due to watermarking. The Transparent Latent Robustness Optimizer (TLRO) fine-tunes the latent representation, subject to interpretability constraints. The Explainable Security Integrity Fusion Benchmark (XSIF-Bench) measures various attributes in combination. Using tests on NIH ChestX-ray14, BraTS, Messidor-2, and another fundus set, we show that the method produces results with a PSNR between 43.7-47.1 dB and SSIM values of 0.982-0.993, and watermark recovery from four attacks at levels of 96.9-98.5%. These results support a diagnostic-safe, auditable medical image authentication protocol.

