{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/distributed-and-parallel-time-series-feature","title":"Distributed and parallel time series feature extraction for industrial big data applications","arxiv_id":"1610.07717","date":"2016-10-25","proceeding":null,"authors":["Maximilian Christ","Andreas W. Kempa-Liehr","Michael Feindt"],"abstract":"The all-relevant problem of feature selection is the identification of all\nstrongly and weakly relevant attributes. This problem is especially hard to\nsolve for time series classification and regression in industrial applications\nsuch as predictive maintenance or production line optimization, for which each\nlabel or regression target is associated with several time series and\nmeta-information simultaneously. Here, we are proposing an efficient, scalable\nfeature extraction algorithm for time series, which filters the available\nfeatures in an early stage of the machine learning pipeline with respect to\ntheir significance for the classification or regression task, while controlling\nthe expected percentage of selected but irrelevant features. The proposed\nalgorithm combines established feature extraction methods with a feature\nimportance filter. It has a low computational complexity, allows to start on a\nproblem with only limited domain knowledge available, can be trivially\nparallelized, is highly scalable and based on well studied non-parametric\nhypothesis tests. We benchmark our proposed algorithm on all binary\nclassification problems of the UCR time series classification archive as well\nas time series from a production line optimization project and simulated\nstochastic processes with underlying qualitative change of dynamics.","url_abs":"http://arxiv.org/abs/1610.07717v3","url_pdf":"http://arxiv.org/pdf/1610.07717v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"distributed-and-parallel-time-series-feature","repo_url":"https://github.com/blue-yonder/tsfresh","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"distributed-and-parallel-time-series-feature","repo_url":"https://github.com/Jllamoza/bbva_attrition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"distributed-and-parallel-time-series-feature","repo_url":"https://github.com/Jllamoza/bbva_attrition_FRESH","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"binary-classification","task_name":"Binary Classification"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"feature-importance","task_name":"Feature Importance"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"time-series-1","task_name":"Time Series"},{"task_slug":"time-series","task_name":"Time Series Analysis"},{"task_slug":"time-series-classification","task_name":"Time Series Classification"},{"task_slug":"feature-selection","task_name":"feature selection"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1610.07717","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}