{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sacrificing-information-for-the-greater-good","title":"Sacrificing information for the greater good: how to select photometric bands for optimal accuracy","arxiv_id":"1511.05424","date":"2015-11-17","proceeding":null,"authors":["Kristoffer Stensbo-Smidt","Fabian Gieseke","Christian Igel","Andrew Zirm","Kim Steenstrup Pedersen"],"abstract":"Large-scale surveys make huge amounts of photometric data available. Because\nof the sheer amount of objects, spectral data cannot be obtained for all of\nthem. Therefore it is important to devise techniques for reliably estimating\nphysical properties of objects from photometric information alone. These\nestimates are needed to automatically identify interesting objects worth a\nfollow-up investigation as well as to produce the required data for a\nstatistical analysis of the space covered by a survey. We argue that machine\nlearning techniques are suitable to compute these estimates accurately and\nefficiently. This study promotes a feature selection algorithm, which selects\nthe most informative magnitudes and colours for a given task of estimating\nphysical quantities from photometric data alone. Using k nearest neighbours\nregression, a well-known non-parametric machine learning method, we show that\nusing the found features significantly increases the accuracy of the\nestimations compared to using standard features and standard methods. We\nillustrate the usefulness of the approach by estimating specific star formation\nrates (sSFRs) and redshifts (photo-z's) using only the broad-band photometry\nfrom the Sloan Digital Sky Survey (SDSS). For estimating sSFRs, we demonstrate\nthat our method produces better estimates than traditional spectral energy\ndistribution (SED) fitting. For estimating photo-z's, we show that our method\nproduces more accurate photo-z's than the method employed by SDSS. The study\nhighlights the general importance of performing proper model selection to\nimprove the results of machine learning systems and how feature selection can\nprovide insights into the predictive relevance of particular input features.","url_abs":"http://arxiv.org/abs/1511.05424v2","url_pdf":"http://arxiv.org/pdf/1511.05424v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sacrificing-information-for-the-greater-good","repo_url":"https://github.com/gieseke/speedynn","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"model-selection","task_name":"Model Selection"},{"task_slug":"feature-selection","task_name":"feature selection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}