{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mctbench-multimodal-cognition-towards-text","title":"MCTBench: Multimodal Cognition towards Text-Rich Visual Scenes Benchmark","arxiv_id":"2410.11538","date":"2024-10-15","proceeding":null,"authors":["Bin Shan","Xiang Fei","Wei Shi","An-Lan Wang","Guozhi Tang","Lei Liao","Jingqun Tang","Xiang Bai","Can Huang"],"abstract":"The comprehension of text-rich visual scenes has become a focal point for evaluating Multi-modal Large Language Models (MLLMs) due to their widespread applications. Current benchmarks tailored to the scenario emphasize perceptual capabilities, while overlooking the assessment of cognitive abilities. To address this limitation, we introduce a Multimodal benchmark towards Text-rich visual scenes, to evaluate the Cognitive capabilities of MLLMs through visual reasoning and content-creation tasks (MCTBench). To mitigate potential evaluation bias from the varying distributions of datasets, MCTBench incorporates several perception tasks (e.g., scene text recognition) to ensure a consistent comparison of both the cognitive and perceptual capabilities of MLLMs. To improve the efficiency and fairness of content-creation evaluation, we conduct an automatic evaluation pipeline. Evaluations of various MLLMs on MCTBench reveal that, despite their impressive perceptual capabilities, their cognition abilities require enhancement. We hope MCTBench will offer the community an efficient resource to explore and enhance cognitive capabilities towards text-rich visual scenes.","url_abs":"https://arxiv.org/abs/2410.11538v1","url_pdf":"https://arxiv.org/pdf/2410.11538v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mctbench-multimodal-cognition-towards-text","repo_url":"https://github.com/xfey/mctbench","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"fairness","task_name":"Fairness"},{"task_slug":"scene-text-recognition","task_name":"Scene Text Recognition"},{"task_slug":"visual-reasoning","task_name":"Visual Reasoning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.11538","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.11538"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/xfey/mctbench","reach":null}],"summary":{"ran_draft_wrong":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"9eb0b14fc790de41","entry":"load_jsonl","repo":"xfey/mctbench","repo_kind":"official","path":"eval/choice_stat.py","file_url":"https://github.com/xfey/mctbench/blob/HEAD/eval/choice_stat.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"9eb0b14fc790de41"}},{"code_sha256_prefix":"7e7617ca02a9a5dc","entry":"parse_response_to_choice","repo":"xfey/mctbench","repo_kind":"official","path":"eval/choice_stat.py","file_url":"https://github.com/xfey/mctbench/blob/HEAD/eval/choice_stat.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"7e7617ca02a9a5dc"}},{"code_sha256_prefix":"1ee95b2457644335","entry":"preproc_answer_from_json","repo":"xfey/mctbench","repo_kind":"official","path":"eval/choice_stat.py","file_url":"https://github.com/xfey/mctbench/blob/HEAD/eval/choice_stat.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"1ee95b2457644335"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}