VLM Awesome
VLM / MLLM / Vision Awesome
视觉语言模型、多模态大语言模型和视觉任务资源索引。
VLM / MLLM
- VLM - Vision Language Model。
- MLLM - Multimodal Large Language Model。
- 常见结构:视觉编码器 + Projector / Vision-Language Adapter + Language Model。
- Projector 用于将视觉特征映射到语言模型表示空间,常见组件包括 Cross-Attention Module。
- Bounding box 坐标格式必须按模型确认:
- Qwen:
(xmin, ymin, xmax, ymax),通常为 0-1 浮点数。 - Gemini:
(ymin, xmin, ymax, xmax),通常为 0-1000 整数。
- Qwen:
视觉任务
- Document OCR / Handwriting OCR。
- Visual QA / Image QA。
- Visual Reasoning。
- Image Classification。
- Document Understanding。
- Video Understanding。
- Object Detection / Object Counting。
- Object Grounding:返回目标 Bounding Box 坐标。
- Computer Agent:屏幕理解与交互操作。
Models 与资源
- Qwen2-VL / Qwen3-VL。
- SmolVLM 256M
- 512px 图像约使用 64 image tokens。
- 浏览器实时 DemoHuggingFaceHF Space:webml-community/smolvlm-realtime-webgpuhuggingface.co · huggingface.co/spaces/webml-community/smolvlm-realtime-webgpu
- Google VideoPrismHuggingFaceHF:google/videoprismhuggingface.co · huggingface.co/google/videoprism
- Video Understanding。
- OpenGVLab/InternVLGitHubOpenGVLab/InternVLhttps://github.com/OpenGVLab/InternVL · InternVL 3.5 · MPO - Mixed Performance Optimization · feat: support internvl ggml-org/llama.cpp#9403笔记:InternVL
- haotian-liu/LLaVAGitHubhaotian-liu/LLaVA
- Vicuna + CLIP 的视觉语言助手。
- google-deepmind/gemmaGitHubgoogle-deepmind/gemma
- Gemma 3 的 4B、12B、27B 版本支持 Vision + Text。
- microsoft/OmniParserGitHubmicrosoft/OmniParsermicrosoft/OmniParser · 识别 UI 交互元素, 屏幕分析 · Pure Vision Based GUI Agent · hf microsoft/OmniParser-v2.0 · demo microsoft/OmniParser-v2笔记:OmniParser
- UI 元素识别和屏幕分析,用于视觉 GUI Agent。
- microsoft/MagmaGitHubmicrosoft/Magma
- 视觉规划、机器人操作、环境交互和视觉导航。
- OpenGVLab/VisualPRM-8B-v1_1HuggingFaceHF:OpenGVLab/VisualPRM-8B-v1_1huggingface.co · huggingface.co/OpenGVLab/VisualPRM-8B-v1_1
- Process Reward Model。
参考
关联信息
反向链接、本文链接的其他页面和外部资料。
反向链接
- AI Model笔记 · VLM / MLLM / Vision
AI Model Awesome 总索引 · LLM 与开放模型 · 图像生成 · 视频生成 · OCR · ASR / STT · TTS · VAD · 音乐生成 · VLM / MLLM / Vision · Agent 与 Coding
- AI Model Awesome笔记 · VLM / MLLM / Vision
按模型能力和使用场景组织的 AI 模型资源索引,保留模型架构、厂商、版本、推理资源、量化、微调和扩散模型历史资料。新增内容优先写入对应专题页面。 · LLM 与开放模型 · 图像生成 · 视频生成 · OCR 与文档理解 · ASR / STT
- LLM Awesome笔记 · VLM / MLLM / Vision
| Date | Model Series | Size | Context Window | Creator | Notes | · | 2026-04-07 | GLM 5.1 | 754B MoE | 200K | Zhipu AI | 大规模稀疏 MoE |
References
GitHub5 条
- google-deepmind/gemmagithub.com/google-deepmind/gemma
- haotian-liu/LLaVAgithub.com/haotian-liu/LLaVA
- microsoft/Magmagithub.com/microsoft/Magma
- microsoft/OmniParsergithub.com/microsoft/OmniParser
- OpenGVLab/InternVLgithub.com/OpenGVLab/InternVL
HuggingFace3 条
- HF Space:webml-community/smolvlm-realtime-webgpuhuggingface.co/spaces/webml-community/smolvlm-realtime-webgpu
- HF:google/videoprismhuggingface.co/google/videoprism
- HF:OpenGVLab/VisualPRM-8B-v1_1huggingface.co/OpenGVLab/VisualPRM-8B-v1_1
其他外链2 条
- blog.roboflow.com/multimodal-vision-modelsblog.roboflow.com/multimodal-vision-models
- simedw.com/2025/07/10/gemini-bounding-boxessimedw.com/2025/07/10/gemini-bounding-boxes