{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/surgical-vqla-transformer-with-gated-vision","title":"Surgical-VQLA: Transformer with Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery","arxiv_id":"2305.11692","date":"2023-05-19","proceeding":null,"authors":["Long Bai","Mobarakol Islam","Lalithkumar Seenivasan","Hongliang Ren"],"abstract":"Despite the availability of computer-aided simulators and recorded videos of surgical procedures, junior residents still heavily rely on experts to answer their queries. However, expert surgeons are often overloaded with clinical and academic workloads and limit their time in answering. For this purpose, we develop a surgical question-answering system to facilitate robot-assisted surgical scene and activity understanding from recorded videos. Most of the existing VQA methods require an object detector and regions based feature extractor to extract visual features and fuse them with the embedded text of the question for answer generation. However, (1) surgical object detection model is scarce due to smaller datasets and lack of bounding box annotation; (2) current fusion strategy of heterogeneous modalities like text and image is naive; (3) the localized answering is missing, which is crucial in complex surgical scenarios. In this paper, we propose Visual Question Localized-Answering in Robotic Surgery (Surgical-VQLA) to localize the specific surgical area during the answer prediction. To deal with the fusion of the heterogeneous modalities, we design gated vision-language embedding (GVLE) to build input patches for the Language Vision Transformer (LViT) to predict the answer. To get localization, we add the detection head in parallel with the prediction head of the LViT. We also integrate GIoU loss to boost localization performance by preserving the accuracy of the question-answering model. We annotate two datasets of VQLA by utilizing publicly available surgical videos from MICCAI challenges EndoVis-17 and 18. Our validation results suggest that Surgical-VQLA can better understand the surgical scene and localize the specific area related to the question-answering. GVLE presents an efficient language-vision embedding technique by showing superior performance over the existing benchmarks.","url_abs":"https://arxiv.org/abs/2305.11692v1","url_pdf":"https://arxiv.org/pdf/2305.11692v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"surgical-vqla-transformer-with-gated-vision","repo_url":"https://github.com/longbai1006/surgical-vqla","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"surgical-vqla-transformer-with-gated-vision","repo_url":"https://github.com/yuyangdu01/llm-cl-vqa","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"answer-generation","task_name":"Answer Generation"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"},{"method_slug":"vision-transformer","method_name":"Vision Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2305.11692","atlas_url":"https://app.syntology.ai/?focus=2305.11692","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2305.11692"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/yuyangdu01/llm-cl-vqa","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/longbai1006/surgical-vqla","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":7},"by_repo_kind":{"official":{"samples":7,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"53daad1d7b99fbdf","entry":"accuracy","repo":"longbai1006/surgical-vqla","repo_kind":"official","path":"utils.py","file_url":"https://github.com/longbai1006/surgical-vqla/blob/HEAD/utils.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"53daad1d7b99fbdf"}},{"code_sha256_prefix":"19988d5d6c4a4d7c","entry":"calc_acc","repo":"longbai1006/surgical-vqla","repo_kind":"official","path":"utils.py","file_url":"https://github.com/longbai1006/surgical-vqla/blob/HEAD/utils.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"19988d5d6c4a4d7c"}},{"code_sha256_prefix":"9fd07c1d7cd49c46","entry":"calc_classwise_acc","repo":"longbai1006/surgical-vqla","repo_kind":"official","path":"utils.py","file_url":"https://github.com/longbai1006/surgical-vqla/blob/HEAD/utils.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9fd07c1d7cd49c46"}},{"code_sha256_prefix":"e21bb85b8e1b4836","entry":"do_nms","repo":"longbai1006/surgical-vqla","repo_kind":"official","path":"dataset/feature_extraction/visual_bert/modeling_frcnn.py","file_url":"https://github.com/longbai1006/surgical-vqla/blob/HEAD/dataset/feature_extraction/visual_bert/modeling_frcnn.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e21bb85b8e1b4836"}},{"code_sha256_prefix":"cfa4887da544c990","entry":"norm_box","repo":"longbai1006/surgical-vqla","repo_kind":"official","path":"dataset/feature_extraction/visual_bert/modeling_frcnn.py","file_url":"https://github.com/longbai1006/surgical-vqla/blob/HEAD/dataset/feature_extraction/visual_bert/modeling_frcnn.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cfa4887da544c990"}},{"code_sha256_prefix":"54a2fff601881b32","entry":"pad_list_tensors","repo":"longbai1006/surgical-vqla","repo_kind":"official","path":"dataset/feature_extraction/visual_bert/modeling_frcnn.py","file_url":"https://github.com/longbai1006/surgical-vqla/blob/HEAD/dataset/feature_extraction/visual_bert/modeling_frcnn.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"54a2fff601881b32"}},{"code_sha256_prefix":"77586e7080d66922","entry":"tryload","repo":"longbai1006/surgical-vqla","repo_kind":"official","path":"dataset/feature_extraction/visual_bert/extracting_data.py","file_url":"https://github.com/longbai1006/surgical-vqla/blob/HEAD/dataset/feature_extraction/visual_bert/extracting_data.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"77586e7080d66922"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}