Opening book details…
Can I read Visual Perception in Multimodal LLMs on EtoBox?
Visual Perception in Multimodal LLMs by thaotran is a document available to read on EtoBox.
What is Visual Perception in Multimodal LLMs about?
This study evaluates the visual perception capabilities of multimodal large language models (LLMs) in recognizing piglet activities through annotated video clips. It assesses the performance of four models—Video-LLaMA, MiniGPT4-Video, Video-Chat2, and GPT-4 omni (GPT-4o)—across five dimensions, revealing that while improvements are needed in semantic correspondence and time perception, GPT-4o demonstrated superior performance. The findings highlight the potential of multimodal LLMs in livestock video unders
- Author
- thaotran
- Language
- EN