{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/direction-of-arrival-with-one-microphone-a","title":"Direction of Arrival with One Microphone, a few LEGOs, and Non-Negative Matrix Factorization","arxiv_id":"1801.03740","date":"2018-08-28","proceeding":null,"authors":[],"abstract":"Conventional approaches to sound source localization require at least two\nmicrophones. It is known, however, that people with unilateral hearing loss can\nalso localize sounds. Monaural localization is possible thanks to the\nscattering by the head, though it hinges on learning the spectra of the various\nsources. We take inspiration from this human ability to propose algorithms for\naccurate sound source localization using a single microphone embedded in an\narbitrary scattering structure. The structure modifies the frequency response\nof the microphone in a direction-dependent way giving each direction a\nsignature. While knowing those signatures is sufficient to localize sources of\nwhite noise, localizing speech is much more challenging: it is an ill-posed\ninverse problem which we regularize by prior knowledge in the form of learned\nnon-negative dictionaries. We demonstrate a monaural speech localization\nalgorithm based on non-negative matrix factorization that does not depend on\nsophisticated, designed scatterers. In fact, we show experimental results with\nad hoc scatterers made of LEGO bricks. Even with these rudimentary structures\nwe can accurately localize arbitrary speakers; that is, we do not need to learn\nthe dictionary for the particular speaker to be localized. Finally, we discuss\nmulti-source localization and the related limitations of our approach.","url_abs":"http://arxiv.org/abs/1801.03740v3","url_pdf":"http://arxiv.org/pdf/1801.03740v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"direction-of-arrival-with-one-microphone-a","repo_url":"https://github.com/swing-research/scatsense","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"sound-source-localization","task_name":"Sound Source Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}