[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
-
Updated
Aug 12, 2024 - Python
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
An open-source implementation for training LLaVA-NeXT.
[CVPR'24] RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
(AAAI 2024) BLIVA: A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions
🧘🏻♂️KarmaVLM (相生):A family of high efficiency and powerful visual language model.
Multimodal Instruction Tuning for Llama 3
Build a simple basic multimodal large model from scratch. 从零搭建一个简单的基础多模态大模型🤖
[ACM MMGR '24] 🔍 Shotluck Holmes: A family of small-scale LLVMs for shot-level video understanding
Docker image for LLaVA: Large Language and Vision Assistant
PyTorch implementation of OpenAI's CLIP model for image classification, visual search, and visual question answering (VQA).
Efficient Video Question Answering
To associate your repository with the visual-language-learning topic, visit your repo's landing page and select "manage topics."