Similar Accuracy, Different Explanations: A Multi-Metric XAI Assessment of Masonry Brick Segmentation Models


Öztürk O., Şeker D. Z.

BUILDINGS (BASEL), cilt.16, sa.19, ss.3873-3892, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 16 Sayı: 19
  • Basım Tarihi: 2026
  • Doi Numarası: 10.3390/buildings16193873
  • Dergi Adı: BUILDINGS (BASEL)
  • Derginin Tarandığı İndeksler: Applied Science & Technology Source, Natural Science Collection (ProQuest), Scopus, Materials Science & Engineering Collection (ProQuest), Technology Collection (ProQuest), Science Citation Index Expanded (SCI-EXPANDED), Avery, Compendex, INSPEC, Directory of Open Access Journals
  • Sayfa Sayıları: ss.3873-3892
  • Recep Tayyip Erdoğan Üniversitesi Adresli: Evet

Özet

Deep learning-based semantic segmentation is increasingly used for automated structural health monitoring (SHM) of masonry infrastructure, yet model evaluation is still primarily based on segmentation metrics. Such metrics capture predictive performance, while the underlying decision-making mechanisms of the models remain unclear. In this study, the relationship between segmentation performance and explainability was investigated for Attention U-Net, U-Net++, and SegFormer-B2 in masonry wall segmentation. All models achieved successful brick segmentation, while SegFormer-B2 showed a modest advantage across all segmentation metrics, with a Dice score of 0.9665 and an IoU of 0.9462. Explanation characteristics were examined using Seg-Grad-CAM, Integrated Gradients, and SLIC-LIME. Attention-gate coefficient maps were additionally analyzed for Attention U-Net. Despite their comparable segmentation performance, the models exhibited distinct spatial explanation patterns, and the quantitative XAI results varied across evaluation metrics. Seg-Grad-CAM achieved the highest Pointing Game and Mask IoU values, while SLIC-LIME produced a low Deletion AUC despite minimal overlap with ground-truth masks. High spatial alignment does not guarantee explanation quality, so method selection depends directly on the evaluated criterion. The findings indicate that segmentation accuracy and explanation behavior provide complementary information for model evaluation.