{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-intelligent-are-convolutional-neural","title":"How intelligent are convolutional neural networks?","arxiv_id":"1709.06126","date":"2017-09-18","proceeding":null,"authors":["Zhennan Yan","Xiang Sean Zhou"],"abstract":"Motivated by the Gestalt pattern theory, and the Winograd Challenge for\nlanguage understanding, we design synthetic experiments to investigate a deep\nlearning algorithm's ability to infer simple (at least for human) visual\nconcepts, such as symmetry, from examples. A visual concept is represented by\nrandomly generated, positive as well as negative, example images. We then test\nthe ability and speed of algorithms (and humans) to learn the concept from\nthese images. The training and testing are performed progressively in multiple\nrounds, with each subsequent round deliberately designed to be more complex and\nconfusing than the previous round(s), especially if the concept was not grasped\nby the learner. However, if the concept was understood, all the deliberate\ntests would become trivially easy. Our experiments show that humans can often\ninfer a semantic concept quickly after looking at only a very small number of\nexamples (this is often referred to as an \"aha moment\": a moment of sudden\nrealization), and performs perfectly during all testing rounds (except for\ncareless mistakes). On the contrary, deep convolutional neural networks (DCNN)\ncould approximate some concepts statistically, but only after seeing many\n(x10^4) more examples. And it will still make obvious mistakes, especially\nduring deliberate testing rounds or on samples outside the training\ndistributions. This signals a lack of true \"understanding\", or a failure to\nreach the right \"formula\" for the semantics. We did find that some concepts are\neasier for DCNN than others. For example, simple \"counting\" is more learnable\nthan \"symmetry\", while \"uniformity\" or \"conformance\" are much more difficult\nfor DCNN to learn. To conclude, we propose an \"Aha Challenge\" for visual\nperception, calling for focused and quantitative research on Gestalt-style\nmachine intelligence using limited training examples.","url_abs":"http://arxiv.org/abs/1709.06126v2","url_pdf":"http://arxiv.org/pdf/1709.06126v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-intelligent-are-convolutional-neural","repo_url":"https://github.com/zhennany/synthetic","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[{"method_slug":"dcnn","method_name":"DCNN"},{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}