{"doi":"10.52294/001c.85074","title":"Improving the Interpretability of fMRI Decoding Using Deep Neural Networks and Adversarial Robustness","abstract":"While deep neural networks (DNNs) are being increasingly used to make predictions from high-dimensional, complex data, they are widely seen as uninterpretable \"black boxes\", as it can be difficult to discover what input information is used to make predictions. This ability is particularly important for applications in cognitive neuroscience and neuroinformatics. A saliency map is a common approach for producing interpretable visualizations of the relative importance of input features in prediction. However, many methods for creating these maps fail due to focusing too much on the input or being extremely sensitive to small input noise. It is also challenging to evaluate how well saliency maps correspond to the truly relevant input information, given that ground truth is not always available. In this paper, we briefly review a variety of methods for producing gradient-based saliency maps, and present a new adversarial training method we developed to make DNNs robust to input noise with the goal of improving interpretability. We introduce two quantitative evaluation procedures for saliency map methods in functional magnetic resonance imaging (fMRI), applicable whenever a model is being trained to decode some information from imaging data. We describe the rationale for the procedures using a synthetic dataset where the complex activation structure is known. We then use them to evaluate saliency maps produced for linear and DNN models performing task decoding in the Human Connectome Project (HCP) dataset. Our key finding is that saliency maps produced with different methods vary widely in interpretability, as measured by our evaluation procedures in synthetic and HCP data. Strikingly, even when linear and DNN models decode at comparable levels of performance, gradient-based saliency maps from the DNN score higher on interpretability than maps derived from the linear model (via weights or gradient). Finally, saliency maps produced with our adversarial training method outperform those produced with alternative methods.","journal":"Aperture Neuro","year":2023,"id":371974,"datarank":0.0,"base_score":0.0,"endowment":0.0,"self_citation_contribution":0.0,"citation_network_contribution":0.0,"self_endowment_contribution":0.0,"citer_contribution":0.0,"corpus_percentile":null,"corpus_rank":null,"citation_count":5,"citer_count":0,"citers_with_citation_signal":0,"citers_with_endowment":0,"datacite_reuse_total":0,"is_dataset":false,"is_dataset_confidence":0.9536,"is_data_producer":false,"deposit_databanks":null,"is_oa":true,"file_count":0,"downloads":0,"has_version_chain":false,"published_date":"2023-01-01","fair_score":null,"fair_percentile":null,"algorithm_id":"datarank_citation_only_1hop_v6","ranking_scope":"data_only","authors":[{"id":659934,"name":"Dustin Moraczewski","orcid":"0000-0002-0422-3135","position":1,"is_corresponding":false},{"id":1086736,"name":"Ka Chun Lam","orcid":"0000-0003-2131-4386","position":2,"is_corresponding":false},{"id":343600,"name":"Adam G. Thomas","orcid":null,"position":3,"is_corresponding":false},{"id":1132033,"name":"Francisco Pereira","orcid":"0009-0008-9119-9487","position":4,"is_corresponding":false},{"id":772407,"name":"Patrick McClure","orcid":"0000-0002-7187-5731","position":0,"is_corresponding":true}],"reference_count":46,"raw_metadata":null,"created_at":"2026-07-19T01:15:49.453761Z","pmid":"40895392","pmcid":null,"fwci":null,"citation_percentile":null,"influential_citations":0,"oa_status":null,"license":null,"views":0,"total_file_size_bytes":0,"version_count":0,"fair_f":null,"fair_a":null,"fair_i":null,"fair_r":null,"fair_zscore":null,"fair_rationale":null,"fair_model":null,"fair_agent_version":null,"fair_fulltext_source":null,"fair_has_llm":null,"fair_computed_at":null,"clinical_trials":[],"software_tools":[],"db_accessions":[],"linked_datasets":[],"topics":[]}