Multimodal Model
Autonomous EntityAI AssistantAI Answer EngineAI SystemSynthetic Personality PlatformSynthetic CompanionCreative IntelligenceVoice IntelligenceFictional IntelligenceAI OrganizationFoundation ModelHistorical AI SystemAgentic SystemRobotic IntelligenceDeveloper PlatformCreative Intelligence SystemRobotics and Embodied AIResearch SystemEvaluation SystemDatasetAI ToolAI InfrastructureScientific AI SystemEmbodied AI SystemCreative AI SystemSpecialized AI SystemAI AgentAI Research SystemAI Governance SystemAI FrameworkAI Training FrameworkAI KernelAI RuntimeMLOps PlatformSearch PlatformGraph AI PlatformRAG MethodLanguage ModelMachine Translation ModelSpeech ModelVision ModelAI DatasetAI BenchmarkScientific AI ModelScientific AI FrameworkAI HardwareEmbodied AIAI Writing SystemAI TutorAI Documentation AssistantFictional Synthetic PersonalityFictional Synthetic PersonMultimodal ModelMultimodal FrameworkVision FrameworkSpeech FrameworkAI Simulation PlatformAgentic AI SystemAI CompilerMachine Learning FrameworkEmbedding ModelReranker ModelMixture-of-Experts ModelCode ModelMachine Translation BenchmarkMachine Translation DatasetMedical AI DatasetMedical AI BenchmarkScientific DatasetScientific AI DatasetFictional RobotFictional Synthetic IntelligenceMath ModelAI Assistant ModelAI ArchitectureAI Alignment MethodReinforcement Learning MethodAI Reasoning MethodAutoML FrameworkWorkflow PlatformData PlatformAI Evaluation PlatformAI Observability PlatformAI GatewayImage Generation ModelImage Control ModelImage Fine-Tuning MethodAudio Generation ModelAudio Codec ModelSpeech Generation ModelAudio Generation FrameworkVoice Conversion SystemRobotics Simulation PlatformRobotics BenchmarkRobotics DatasetAI PlatformAI Search AssistantAI Data PlatformAI InterfaceRobotics SoftwareAI Safety ToolAI Evaluation FrameworkAI Safety ModelComputer Vision ModelAI Narrative SystemAI LibraryAI ModelDataset / BenchmarkOrganization / PlatformEmbodied AI / RoboticsBenchmark
OpenCLIP
OpenCLIP is a open CLIP implementation associated with LAION.
Phi-3.5 Vision
Phi-3.5 Vision is a vision-language model associated with Microsoft.
Qwen2 VL 2B Instruct
Qwen2 VL 2B Instruct is a vision-language model associated with Alibaba Cloud.
Qwen2 VL 7B Instruct
Qwen2 VL 7B Instruct is a vision-language model associated with Alibaba Cloud.
Qwen2 VL 72B Instruct
Qwen2 VL 72B Instruct is a vision-language model associated with Alibaba Cloud.
Qwen2.5 VL 3B Instruct
Qwen2.5 VL 3B Instruct is a vision-language model associated with Alibaba Cloud.
Qwen2.5 VL 7B Instruct
Qwen2.5 VL 7B Instruct is a vision-language model associated with Alibaba Cloud.
Qwen2.5 VL 32B Instruct
Qwen2.5 VL 32B Instruct is a vision-language model associated with Alibaba Cloud.
Qwen2.5 VL 72B Instruct
Qwen2.5 VL 72B Instruct is a vision-language model associated with Alibaba Cloud.
LLaVA 1.5 7B
LLaVA 1.5 7B is a vision-language assistant model associated with LLaVA Team.
LLaVA 1.5 13B
LLaVA 1.5 13B is a vision-language assistant model associated with LLaVA Team.
LLaVA-NeXT Mistral 7B
LLaVA-NeXT Mistral 7B is a vision-language assistant model associated with LLaVA Team.
Qwen-VL-Chat
Qwen-VL-Chat is a vision-language chat model associated with Alibaba Cloud.
InternVL Chat V1.5
InternVL Chat V1.5 is a vision-language chat model associated with Shanghai AI Laboratory.
InternVL2 8B
InternVL2 8B is a vision-language model associated with Shanghai AI Laboratory.
MiniCPM-V 2.6
MiniCPM-V 2.6 is a vision-language model associated with OpenBMB.
Idefics
Idefics is a open vision-language model associated with Hugging Face.
Idefics2
Idefics2 is a open vision-language model associated with Hugging Face.
DeepSeek-VL-7B-Chat
DeepSeek-VL-7B-Chat is a vision-language chat model associated with DeepSeek AI.
DeepSeek-VL2
DeepSeek-VL2 is a vision-language model family associated with DeepSeek AI.
Moondream2
Moondream2 is a small vision-language model associated with Moondream.
Qwen2-VL-7B-Instruct
Qwen2-VL-7B-Instruct is a vision-language instruction model associated with Alibaba Cloud.
InternVL2-8B
InternVL2-8B is a vision-language model associated with OpenGVLab.
MiniCPM-V 2.6
MiniCPM-V 2.6 is a vision-language model associated with OpenBMB.