- 788kFlorence-2vision-multimodal
image-text-to-text
788kMIT - 247kPhi-3-Vision-128K-Instructvision-multimodal
text-generation
247kMIT - 25kLLaVAvision-multimodal
[NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
25kApache-2.0 - 23k3D Gaussian Splattingvision-multimodal
Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
23kNot specified - 20kQwen3-VLvision-multimodal
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
20kApache-2.0 - 20kSAM 2vision-multimodal
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
20kApache-2.0 - 14kOpen3Dvision-multimodal
Open3D: A Modern Library for 3D Data Processing
14kNot specified - 13kMeshroomvision-multimodal
Node-based Visual Programming Toolbox
13kNot specified - 12knerfstudiovision-multimodal
A collaboration friendly studio for NeRFs
12kApache-2.0 - 10kInternVLvision-multimodal
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
10kMIT - SponsorReach 50,000+ buyers
Enterprise buyers looking for private AI solutions see your brand here.
- 9.9kmoondreamvision-multimodal
tiny vision language model
9.9kApache-2.0 - 6.0kPixtral-12B-2409vision-multimodal
Natively multimodal model with a 12B parameter decoder and 400M parameter vision encoder, supporting variable image sizes and 128k context length.
6.0kApache-2.0 - 5.3kDeepSeek-VL2vision-multimodal
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
5.3kMIT - 84Falcon2-11B-VLMvision-multimodal
image-text-to-text
84unknown
Stop paying for AI APIs. Everything here runs on your hardware.
Enterprise buyers looking for private AI solutions see your brand here.
No fluff. No spam. Join 12,000+ builders.