Dual-penalty evasion framework to fool white-box explainable AI auditors; arXiv:2608.00566v1 reports attacks on post-hoc explainers like LIME, SHAP, and Integrated Gradients
Read the original at arxiv.org→arXiv:2608.00566v1 Announce Type: new Abstract: Post-hoc model explainers such as LIME, SHAP, and Integrated Gradients are widely deployed to audit models in high-stakes sensitive domains, including finance,...
Original headline: "Crushing the Evidence: A Dual-Penalty Evasion Framework for Fooling White-Box Explainable AI Auditors"
Coverage timeline
- Aug 4, 04:00 UTC arXiv cs.LG lead source Crushing the Evidence: A Dual-Penalty Evasion Framework for Fooling White-Box Explainable AI Auditors