{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/im-promptu-in-context-composition-from-image-1","title":"Im-Promptu: In-Context Composition from Image Prompts","arxiv_id":"2305.17262","date":"2023-05-26","proceeding":"NeurIPS 2023 11","authors":["Bhishma Dedhia","Michael Chang","Jake C. Snell","Thomas L. Griffiths","Niraj K. Jha"],"abstract":"Large language models are few-shot learners that can solve diverse tasks from a handful of demonstrations. This implicit understanding of tasks suggests that the attention mechanisms over word tokens may play a role in analogical reasoning. In this work, we investigate whether analogical reasoning can enable in-context composition over composable elements of visual stimuli. First, we introduce a suite of three benchmarks to test the generalization properties of a visual in-context learner. We formalize the notion of an analogy-based in-context learner and use it to design a meta-learning framework called Im-Promptu. Whereas the requisite token granularity for language is well established, the appropriate compositional granularity for enabling in-context generalization in visual stimuli is usually unspecified. To this end, we use Im-Promptu to train multiple agents with different levels of compositionality, including vector representations, patch representations, and object slots. Our experiments reveal tradeoffs between extrapolation abilities and the degree of compositionality, with non-compositional representations extending learned composition rules to unseen domains but performing poorly on combinatorial tasks. Patch-based representations require patches to contain entire objects for robust extrapolation. At the same time, object-centric tokenizers coupled with a cross-attention module generate consistent and high-fidelity solutions, with these inductive biases being particularly crucial for compositional generalization. Lastly, we demonstrate a use case of Im-Promptu as an intuitive programming interface for image generation.","url_abs":"https://arxiv.org/abs/2305.17262v3","url_pdf":"https://arxiv.org/pdf/2305.17262v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"im-promptu-in-context-composition-from-image-1","repo_url":"https://github.com/JHA-Lab/impromptu","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"meta-learning","task_name":"Meta-Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2305.17262","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2305.17262"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/JHA-Lab/impromptu","reach":null}],"summary":{"ran_fixture":1,"ran_honours":1,"unverified":1},"by_repo_kind":{"listed":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"8ad9768f4ab9a2c4","entry":"augment_analogy","repo":"JHA-Lab/impromptu","repo_kind":"listed","path":"train_scripts/train_monolithic.py","file_url":"https://github.com/JHA-Lab/impromptu/blob/HEAD/train_scripts/train_monolithic.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"BSD-3-Clause-Clear","inline_ok":false,"mcp_get_code":{"code_sha256":"8ad9768f4ab9a2c4"}},{"code_sha256_prefix":"ea06ca059adeb8e4","entry":"linear_warmup","repo":"JHA-Lab/impromptu","repo_kind":"listed","path":"train_scripts/train_monolithic.py","file_url":"https://github.com/JHA-Lab/impromptu/blob/HEAD/train_scripts/train_monolithic.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"BSD-3-Clause-Clear","inline_ok":false,"mcp_get_code":{"code_sha256":"ea06ca059adeb8e4"}},{"code_sha256_prefix":"a27c4262fce22d72","entry":"visualize_generation","repo":"JHA-Lab/impromptu","repo_kind":"listed","path":"train_scripts/train_monolithic.py","file_url":"https://github.com/JHA-Lab/impromptu/blob/HEAD/train_scripts/train_monolithic.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause-Clear","inline_ok":false,"mcp_get_code":{"code_sha256":"a27c4262fce22d72"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}