Opening book details…
Can I read LLaVA: Large Language & Vision Assistant on EtoBox?
LLaVA: Large Language & Vision Assistant by classaen9 is a document available to read on EtoBox.
What is LLaVA: Large Language & Vision Assistant about?
The document discusses various Large Multi-Modal Models, including ViT, CLIP, and LLaVA, highlighting their development and methodologies. ViT introduced the use of Transformer networks for image classification, while CLIP utilized natural language supervision for image representation learning. LLaVA focuses on creating multimodal instruction-following data and proposes methods for generating visual instruction data using large models like GPT-4.
- Author
- classaen9
- Language
- EN