Search for collections on FTS Digilib

Cross-Domain Faithfulness Evaluation of SHAP and Attention-Based Explanations in Transformer NLP Models

Firmawan, Dony Bahtera and Darnoto, Brian Rizqi Paradisiaca (2026) Cross-Domain Faithfulness Evaluation of SHAP and Attention-Based Explanations in Transformer NLP Models. Journal of Computing Theories and Applications, 4 (1). pp. 146-163. ISSN 3024-9104

[thumbnail of 16258-Article Text-60025-1-10-20260721.pdf] Text
16258-Article Text-60025-1-10-20260721.pdf - Published Version
Available under License Creative Commons Attribution.

Download (623kB)

Abstract

Transformer-based models such as BERT, RoBERTa, DistilBERT, and DeBERTa have achieved remarkable performance across a wide range of natural language processing (NLP) tasks. However, their decision-making processes remain difficult to interpret, particularly in high-risk applications such as hate speech detection, where unreliable explanations may undermine model transparency, trust, and accountability. This study investigates whether explainability methods remain faithful and stable under domain shift in transformer-based text classification. Four transformer architectures were fine-tuned and evaluated on two linguistically distinct datasets: IMDb Movie Reviews and Hate Speech Offensive. Model performance and explanation quality were assessed using classification accuracy, macro F1-score, top-k token-removal faithfulness analysis, and cross-domain Spearman rank correlation. Experimental results show that DeBERTa achieved the highest classification performance, reaching accuracies of 95.6% on IMDb and 91.3% on Hate Speech. Across all evaluated models and datasets, SHAP consistently produced higher faithfulness scores than attention-based explanations. Cross-domain analysis further revealed reduced agreement between SHAP and attention-based explanations under domain shift, indicating lower explanation consistency across linguistically distinct domains. Qualitative error analysis further showed that implicit sentiment, sarcasm, and domain-specific slang remain major sources of prediction errors. Overall, the results demonstrate that superior predictive performance does not necessarily correspond to higher explanation faithfulness or stronger cross-domain stability. These findings highlight the importance of jointly evaluating predictive performance, explanation faithfulness, and explanation robustness when developing trustworthy transformer-based NLP systems.

Item Type: Article
Subjects: Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Depositing User: dl fts
Date Deposited: 21 Jul 2026 03:26
Last Modified: 21 Jul 2026 03:26
URI: https://dl.futuretechsci.org/id/eprint/197

Actions (login required)

View Item
View Item