{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mmlongbench-doc-benchmarking-long-context","title":"MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations","arxiv_id":"2407.01523","date":"2024-07-01","proceeding":null,"authors":["Yubo Ma","Yuhang Zang","Liangyu Chen","Meiqi Chen","Yizhu Jiao","Xinze Li","Xinyuan Lu","Ziyu Liu","Yan Ma","Xiaoyi Dong","Pan Zhang","Liangming Pan","Yu-Gang Jiang","Jiaqi Wang","Yixin Cao","Aixin Sun"],"abstract":"Understanding documents with rich layouts and multi-modal components is a long-standing and practical task. Recent Large Vision-Language Models (LVLMs) have made remarkable strides in various tasks, particularly in single-page document understanding (DU). However, their abilities on long-context DU remain an open problem. This work presents MMLongBench-Doc, a long-context, multi-modal benchmark comprising 1,062 expert-annotated questions. Distinct from previous datasets, it is constructed upon 130 lengthy PDF-formatted documents with an average of 49.4 pages and 20,971 textual tokens. Towards comprehensive evaluation, answers to these questions rely on pieces of evidence from (1) different sources (text, image, chart, table, and layout structure) and (2) various locations (i.e. page number). Moreover, 33.2% of the questions are cross-page questions requiring evidence across multiple pages. 22.8% of the questions are designed to be unanswerable for detecting potential hallucinations. Experiments on 14 LVLMs demonstrate that long-context DU greatly challenges current models. Notably, the best-performing model, GPT-4o, achieves an F1 score of only 42.7%, while the second-best, GPT-4V, scores 31.4%. Furthermore, 12 LVLMs (all except GPT-4o and GPT-4V) even present worse performance than their LLM counterparts which are fed with lossy-parsed OCR documents. These results validate the necessity of future research toward more capable long-context LVLMs. Project Page: https://mayubo2333.github.io/MMLongBench-Doc","url_abs":"https://arxiv.org/abs/2407.01523v3","url_pdf":"https://arxiv.org/pdf/2407.01523v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mmlongbench-doc-benchmarking-long-context","repo_url":"https://github.com/mayubo2333/mmlongbench-doc","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition (OCR)"},{"task_slug":"document-understanding","task_name":"document understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2407.01523","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2407.01523"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mayubo2333/mmlongbench-doc","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":2,"ran_draft_wrong":1,"ran_fixture":1,"unverified":7},"by_repo_kind":{"listed":{"samples":11,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"57e95166ade5f25c","entry":"build_transform","repo":"mayubo2333/mmlongbench-doc","repo_kind":"listed","path":"models/internvl_chat.py","file_url":"https://github.com/mayubo2333/mmlongbench-doc/blob/HEAD/models/internvl_chat.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"57e95166ade5f25c"}},{"code_sha256_prefix":"d610b5eabe0c9db1","entry":"dynamic_preprocess","repo":"mayubo2333/mmlongbench-doc","repo_kind":"listed","path":"models/internvl_chat.py","file_url":"https://github.com/mayubo2333/mmlongbench-doc/blob/HEAD/models/internvl_chat.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d610b5eabe0c9db1"}},{"code_sha256_prefix":"bd77f5f8067f18e9","entry":"find_closest_aspect_ratio","repo":"mayubo2333/mmlongbench-doc","repo_kind":"listed","path":"models/internvl_chat.py","file_url":"https://github.com/mayubo2333/mmlongbench-doc/blob/HEAD/models/internvl_chat.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"bd77f5f8067f18e9"}},{"code_sha256_prefix":"90a8e73fa6de72a4","entry":"levenshtein_distance","repo":"mayubo2333/mmlongbench-doc","repo_kind":"listed","path":"eval/eval_score.py","file_url":"https://github.com/mayubo2333/mmlongbench-doc/blob/HEAD/eval/eval_score.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"90a8e73fa6de72a4"}},{"code_sha256_prefix":"b9646cb1dd61841b","entry":"anls_compute","repo":"mayubo2333/mmlongbench-doc","repo_kind":"listed","path":"eval/eval_score.py","file_url":"https://github.com/mayubo2333/mmlongbench-doc/blob/HEAD/eval/eval_score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b9646cb1dd61841b"}},{"code_sha256_prefix":"18507a5fed3525e6","entry":"concat_images","repo":"mayubo2333/mmlongbench-doc","repo_kind":"listed","path":"run_lvlm.py","file_url":"https://github.com/mayubo2333/mmlongbench-doc/blob/HEAD/run_lvlm.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"18507a5fed3525e6"}},{"code_sha256_prefix":"f9d80f7332b0b8a1","entry":"encode_image_to_base64","repo":"mayubo2333/mmlongbench-doc","repo_kind":"listed","path":"run_api.py","file_url":"https://github.com/mayubo2333/mmlongbench-doc/blob/HEAD/run_api.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f9d80f7332b0b8a1"}},{"code_sha256_prefix":"9677a8f9d8728dbb","entry":"get_response_concat","repo":"mayubo2333/mmlongbench-doc","repo_kind":"listed","path":"models/minicpm_llama3.py","file_url":"https://github.com/mayubo2333/mmlongbench-doc/blob/HEAD/models/minicpm_llama3.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9677a8f9d8728dbb"}},{"code_sha256_prefix":"2e9e62e3fa0f5b69","entry":"init_model","repo":"mayubo2333/mmlongbench-doc","repo_kind":"listed","path":"models/internlm_xc2_4khd.py","file_url":"https://github.com/mayubo2333/mmlongbench-doc/blob/HEAD/models/internlm_xc2_4khd.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"2e9e62e3fa0f5b69"}},{"code_sha256_prefix":"5e1f94bb8dee0afa","entry":"is_float_equal","repo":"mayubo2333/mmlongbench-doc","repo_kind":"listed","path":"eval/eval_score.py","file_url":"https://github.com/mayubo2333/mmlongbench-doc/blob/HEAD/eval/eval_score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"5e1f94bb8dee0afa"}},{"code_sha256_prefix":"3c76aa11708f6da6","entry":"load_model","repo":"mayubo2333/mmlongbench-doc","repo_kind":"listed","path":"run_lvlm.py","file_url":"https://github.com/mayubo2333/mmlongbench-doc/blob/HEAD/run_lvlm.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3c76aa11708f6da6"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}