Robust Explanations using Diverse Adversarially Trained Ensembles, Multi-Modal Contrastive Learning, and Attribution-based Confidence Metrics
Public Project Summary
Rickard Ewetz, University of Central Florida (Lead Principal Investigator)
Sumit Kumar Jha, University of Texas at San Antonio (Co-Principal Investigator)
Modern neural networks have surpassed human-level performance for many challenging tasks and are rapidly being deployed within real-world applications. The phenomenal performance of these networks comes from a large number of parameters and the non-linear interactions among them. This complex and high-dimensional dynamics makes it difficult for a human to understand and visualize why an AI system makes a particular decision. Robust methods for explaining AI decisions are needed to enhance the broad adoption of AI in scientific applications.
While there has been a lot of work on explainable AI over the last decade, the emphasis often has been on the use of model gradients. We are not proposing a refinement to one of these gradient-based methods; instead, we are leveraging several new ideas, including symbolic reasoning, adversarially robust networks, synthesizing counterfactual inputs, multi-modal perceivers, self-supervised learning, and ensembles of diverse models. Our goal is to create robust multi-modal explanations using unsupervised learning for data including images, time series, text in the form of visualizations including super-pixels/text for images, and rules such as natural laws for time series. Robustness will be achieved by a combination of robust diverse ensembles of neural networks and faster adversarial training using stochastic generators. Attribution-based confidence metrics will be used to identify the envelope of the correct operation of neural networks. All work will be performed on state-of-the-art networks such as Transfomers and Perceivers.
Our proposed effort will initiate a new wave of explainable AI. This will have a strong impact on many scientific disciplines where AI is being used for pursuing challenging tasks that ultimately require human trust, such as life sciences, deep reinforcement learning, physics, and cyber-physical systems. The project is also expected to broadly impact the scientific community by enabling the analysis of large datasets and models for scientific discovery.