{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/audiovisual-speaker-tracking-using-nonlinear","title":"Audiovisual Speaker Tracking using Nonlinear Dynamical Systems with Dynamic Stream Weights","arxiv_id":"1903.06031","date":"2019-03-14","proceeding":null,"authors":["Christopher Schymura","Dorothea Kolossa"],"abstract":"Data fusion plays an important role in many technical applications that\nrequire efficient processing of multimodal sensory observations. A prominent\nexample is audiovisual signal processing, which has gained increasing attention\nin automatic speech recognition, speaker localization and related tasks. If\nappropriately combined with acoustic information, additional visual cues can\nhelp to improve the performance in these applications, especially under adverse\nacoustic conditions. A dynamic weighting of acoustic and visual streams based\non instantaneous sensor reliability measures is an efficient approach to data\nfusion in this context. This paper presents a framework that extends the\nwell-established theory of nonlinear dynamical systems with the notion of\ndynamic stream weights for an arbitrary number of sensory observations. It\ncomprises a recursive state estimator based on the Gaussian filtering paradigm,\nwhich incorporates dynamic stream weights into a framework closely related to\nthe extended Kalman filter. Additionally, a convex optimization approach to\nestimate oracle dynamic stream weights in fully observed dynamical systems\nutilizing a Dirichlet prior is presented. This serves as a basis for a generic\nparameter learning framework of dynamic stream weight estimators. The proposed\nsystem is application-independent and can be easily adapted to specific tasks\nand requirements. A study using audiovisual speaker tracking tasks is\nconsidered as an exemplary application in this work. An improved tracking\nperformance of the dynamic stream weight-based estimation framework over\nstate-of-the-art methods is demonstrated in the experiments.","url_abs":"http://arxiv.org/abs/1903.06031v1","url_pdf":"http://arxiv.org/pdf/1903.06031v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"audiovisual-speaker-tracking-using-nonlinear","repo_url":"https://github.com/rub-ksv/avtrack","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}