{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bridge-the-gap-between-vqa-and-human-behavior","title":"Bridge the Gap Between VQA and Human Behavior on Omnidirectional Video: A Large-Scale Dataset and a Deep Learning Model","arxiv_id":"1807.10990","date":"2018-07-29","proceeding":null,"authors":["Chen Li","Mai Xu","Xinzhe Du","Zulin Wang"],"abstract":"Omnidirectional video enables spherical stimuli with the $360 \\times 180^\n\\circ$ viewing range. Meanwhile, only the viewport region of omnidirectional\nvideo can be seen by the observer through head movement (HM), and an even\nsmaller region within the viewport can be clearly perceived through eye\nmovement (EM). Thus, the subjective quality of omnidirectional video may be\ncorrelated with HM and EM of human behavior. To fill in the gap between\nsubjective quality and human behavior, this paper proposes a large-scale visual\nquality assessment (VQA) dataset of omnidirectional video, called VQA-OV, which\ncollects 60 reference sequences and 540 impaired sequences. Our VQA-OV dataset\nprovides not only the subjective quality scores of sequences but also the HM\nand EM data of subjects. By mining our dataset, we find that the subjective\nquality of omnidirectional video is indeed related to HM and EM. Hence, we\ndevelop a deep learning model, which embeds HM and EM, for objective VQA on\nomnidirectional video. Experimental results show that our model significantly\nimproves the state-of-the-art performance of VQA on omnidirectional video.","url_abs":"http://arxiv.org/abs/1807.10990v1","url_pdf":"http://arxiv.org/pdf/1807.10990v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bridge-the-gap-between-vqa-and-human-behavior","repo_url":"https://github.com/Archer-Tatsu/VQA-ODV","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[{"slug":"vqa-ov","name":"VQA-OV","full_name":"Visual Quality Assessment of Omnidirectional Video"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}