Ollama model
llava
π LLaVA is a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding. Updated to version 1.6.
- Pulls
- 14.7M
- Tags
- 98
- Updated
- Feb 1, 2024
- Listed sizes
- 7b Β· 13b Β· 34b
Run locally
Copy one command. You stay in control.
The command downloads the default tag if needed, then starts it. Confirm the exact file size, license, context, and hardware requirement on the official page first.
ollama run llava- 1
Install Ollama on your Mac, Windows, or Linux computer.
- 2
Open Terminal and paste the copied command.
- 3
Press Enter to download the model if needed and start it.
Capabilities
What its library labels mean
- Vision
- Understands images and other supported visual input.