{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/benchmarking-deep-learning-interpretability","title":"Benchmarking Deep Learning Interpretability in Time Series Predictions","arxiv_id":"2010.13924","date":"2020-10-26","proceeding":"NeurIPS 2020 12","authors":["Aya Abdelsalam Ismail","Mohamed Gunady","Héctor Corrada Bravo","Soheil Feizi"],"abstract":"Saliency methods are used extensively to highlight the importance of input features in model predictions. These methods are mostly used in vision and language tasks, and their applications to time series data is relatively unexplored. In this paper, we set out to extensively compare the performance of various saliency-based interpretability methods across diverse neural architectures, including Recurrent Neural Network, Temporal Convolutional Networks, and Transformers in a new benchmark of synthetic time series data. We propose and report multiple metrics to empirically evaluate the performance of saliency methods for detecting feature importance over time using both precision (i.e., whether identified features contain meaningful signals) and recall (i.e., the number of features with signal identified as important). Through several experiments, we show that (i) in general, network architectures and saliency methods fail to reliably and accurately identify feature importance over time in time series data, (ii) this failure is mainly due to the conflation of time and feature domains, and (iii) the quality of saliency maps can be improved substantially by using our proposed two-step temporal saliency rescaling (TSR) approach that first calculates the importance of each time step before calculating the importance of each feature at a time step.","url_abs":"https://arxiv.org/abs/2010.13924v1","url_pdf":"https://arxiv.org/pdf/2010.13924v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"benchmarking-deep-learning-interpretability","repo_url":"https://github.com/ayaabdelsalam91/TS-Interpretability-Benchmark","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"feature-importance","task_name":"Feature Importance"},{"task_slug":"time-series-1","task_name":"Time Series"},{"task_slug":"time-series","task_name":"Time Series Analysis"}],"methods":[{"method_slug":"interpretability","method_name":"Interpretability"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2010.13924","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2010.13924"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ayaabdelsalam91/TS-Interpretability-Benchmark","reach":null}],"summary":{"ran_draft_wrong":2,"ran_honours":1,"unverified":1},"by_repo_kind":{"official":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"94be2b4c3aa07cf9","entry":"attention","repo":"ayaabdelsalam91/TS-Interpretability-Benchmark","repo_kind":"official","path":"Scripts/Models/Transformer.py","file_url":"https://github.com/ayaabdelsalam91/TS-Interpretability-Benchmark/blob/HEAD/Scripts/Models/Transformer.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"94be2b4c3aa07cf9"}},{"code_sha256_prefix":"a3368005f95cadc1","entry":"creatMask","repo":"ayaabdelsalam91/TS-Interpretability-Benchmark","repo_kind":"official","path":"Scripts/Models/Transformer.py","file_url":"https://github.com/ayaabdelsalam91/TS-Interpretability-Benchmark/blob/HEAD/Scripts/Models/Transformer.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a3368005f95cadc1"}},{"code_sha256_prefix":"891b8ebab395921f","entry":"get_clones","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"891b8ebab395921f"}},{"code_sha256_prefix":"0c32c2b12d8473af","entry":"parse_arguments","repo":"ayaabdelsalam91/TS-Interpretability-Benchmark","repo_kind":"official","path":"Scripts/run_benchmark.py","file_url":"https://github.com/ayaabdelsalam91/TS-Interpretability-Benchmark/blob/HEAD/Scripts/run_benchmark.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0c32c2b12d8473af"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}