{"repo":"zli12321/Vision-Language-Models-Overview","free":true,"listed":false,"github":"https://github.com/zli12321/Vision-Language-Models-Overview","clone":"git clone https://github.com/zli12321/Vision-Language-Models-Overview.git","description":"A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.","language":"HTML","stars":698,"topics":["blip2","claude","clip","deepseek","gemini-pro","gpt-4v","llama-vision-model","llava","multimodal-models","qwen-vl"],"license":null,"category":"ai-agents","readme_excerpt":"Benchmark and Evaluations, RL Alignment, Applications, and Challenges of Large Vision Language Models 🌐 Language : English · 简体中文 A most Frontend Collection and survey of vision-language model papers, and models GitHub repository --- 🧭 The Evolution of VLM Architectures VLM design has gone through four distinct architectural eras in just six years — and Era 3 has split into two parallel branches. Early models kept frozen vision and language towers, aligned contrastively (CLIP) or bridged by a learnable connector into a frozen LM (BLIP-2, Flamingo). The 2023–2025 generation made a pretrained LLM the trunk and treated vision as a bolt-on adapter (LLaVA, Qwen2.5-VL, GPT-4V). The 2025–2026 generation drops the bridge entirely and early-fuses all modalities into a single transformer — forking along the output axis — and in 2026 the trunk is becoming a world model that predicts and acts: - Era 3a — Native Multimodal Input → Text Out. Image, video, and (sometimes) audio enter a single early-fused token stream, but generation is still autoregressive text. This is the design used by today's general-purpose flagships: Qwen3.5 / Qwen3.6, Gemma 4, Gemini 3, GPT-5.4, Phi-4-Reasoning-Vision, Claude Opus 4.6, Nemotron 3 Nano Omni . - Era 3b — Omni-Modal Unified I/O. The same fused trunk plus dedicated image / video decoder (VAE / DiT / flow-matching) and/or audio codec decoder heads, so the model can also generate images, video, and speech — via autoregression or, increasingly, discrete d","default_branch":null,"files":null,"tree":[],"storefront":"/r/zli12321","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/zli12321/Vision-Language-Models-Overview/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}