Papers › GPT-4o: Visual perception performance of multimodal large language models in piglet...

GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding

14 Jun 2024arXiv:2406.09781archive 2025-07-28

Yiqi Wu, Xiaodan Hu, Ziming Fu, Siling Zhou, Jiangong Li

Animal ethology is an crucial aspect of animal research, and animal behavior labeling is the foundation for studying animal behavior. This process typically involves labeling video clips with behavioral semantic tags, a task that is complex, subjective, and multimodal. With the rapid development of multimodal large language models(LLMs), new application have emerged for animal behavior understanding tasks in livestock scenarios. This study evaluates the visual perception capabilities of multimodal LLMs in animal activity recognition. To achieve this, we created piglet test data comprising close-up video clips of individual piglets and annotated full-shot video clips. These data were used to assess the performance of four multimodal LLMs-Video-LLaMA, MiniGPT4-Video, Video-Chat2, and GPT-4 omni (GPT-4o)-in piglet activity understanding. Through comprehensive evaluation across five dimensions, including counting, actor referring, semantic correspondence, time perception, and robustness, we found that while current multimodal LLMs require improvement in semantic correspondence and time perception, they have initially demonstrated visual perception capabilities for animal activity recognition. Notably, GPT-4o showed outstanding performance, with Video-Chat2 and GPT-4o exhibiting significantly better semantic correspondence and time perception in close-up video clips compared to full-shot clips. The initial evaluation experiments in this study validate the potential of multimodal large language models in livestock scene video understanding and provide new directions and references for future research on animal behavior video understanding. Furthermore, by deeply exploring the influence of visual prompts on multimodal large language models, we expect to enhance the accuracy and efficiency of animal behavior recognition in livestock scenarios through human visual processing methods.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Activity RecognitionMMR totalSemantic correspondenceVideo UnderstandingZero-Shot Video Question Answer

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
MMR total MRR-Benchmark GPT-4o Total Column Score 457 #2 of 14 Archive leaderboard report
Zero-Shot Video Question Answer Video-MME GPT-4o Accuracy (%) 77.2 #3 of 11 Archive leaderboard report
Zero-Shot Video Question Answer Video-MME GPT-4o mini Accuracy (%) 68.9 #5 of 11 Archive leaderboard report
Zero-Shot Video Question Answer Video-MME (w/o subs) GPT-4o Accuracy (%) 70.3 #3 of 9 Archive leaderboard report
Zero-Shot Video Question Answer Video-MME (w/o subs) GPT-4o mini Accuracy (%) 62.3 #6 of 9 Archive leaderboard report
Zero-Shot Video Question Answer Zero-shot Video Question Answering on LongVideoBench GPT-4o Accuracy (% ) 64.0 #3 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutGPT-4Label SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections