Vision Language Model
Breaking News
Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver...
WeatherNext 3 ingests live geostationary satellite mosaics, refreshes hourly, and outputs 5 km forecasts across Search, Gemini, Maps.
Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That...
Cohere has released Parse (parse-v5.0), a 2.3B-parameter vision language model that converts PDFs, slides and images into Markdown with HTML tables, bounding boxes and image descriptions. It runs at $1.50 per 1,000 pages through the API, or on dedicated Model Vault instances from $2,500 a month. Cohere reports a ParseBench score of 79.2, ahead of Mistral OCR 4, Azure Document Intelligence and Databricks AI Parse — but that figure averages three of the benchmark's five dimensions and drops charts and visual grounding entirely.
Create a Reasoning-Focused LLM: A Practical Guide to Streaming, Curating, and...
This tutorial provides a complete workflow for building a compact, reasoning-focused language model. By streaming the SupraLabs reasoning corpus from Hugging Face, we apply quality filters and curate data for Supervised Fine-Tuning (SFT). Using SmolLM2-135M-Instruct and LoRA, we demonstrate an end-to-end pipeline—from dataset analysis and heuristic cleaning to efficient training and inference—enabling the development of specialized small models without excessive resource requirements
Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens,...
Liquid AI released LFM2.5-VL-3B, a 3.1B-parameter vision-language model built for on-device deployment. It averages 80.7 on ScreenSpot-v2 and lifts RefCOCO grounding from 57.1 to 87.9. Function calling is new to the VL line, with ToolSandbox moving from 26.4 to 59.5. The model fits in roughly 3 GB and decodes 228 tokens/s on an Apple M5 Max.
Dyna Robotics Introduces Dyna-2: A World-Action Model Pre-Trained on 1 Million...
Dyna Robotics has released Dyna-2, a world-action model pre-trained on more than one million hours of egocentric human video. The technical report establishes three results: a scaling law on human data to 1M hours, the first transfer of that law to unseen robot data, and evidence that video co-training drives cross-embodiment generalization.
Alibaba Qwen Releases Qwen3.8-Max: A 2.4 Trillion Parameter MoE Model and...
Alibaba's Qwen team moved Qwen3.8-Max from preview to general availability, with published per-token pricing and open weights due next week. The 2.4T parameter MoE model accepts text, image and video input across a 1M-token context. No benchmark table has been published.
Onton Releases Ontology 1: A Neurosymbolic Search Model That is 2.7x...
Onton, a San Francisco-based search and discovery company, has released Ontology 1, a neurosymbolic model for complex, conversational, multimodal product search. On a 90-query...
Google DeepMind Ships Three Physical AI Models For Whole Body Control,...
Google DeepMind has released Gemini Robotics 2, the intelligence layer for its next generation of robots. The release ships three models: a vision-language-action model for whole body humanoid control, Gemini Robotics ER 2 for embodied reasoning and task orchestration, and an on-device VLA that adapts to new robot bodies in hours. One checkpoint drives Apptronik Apollo 2 and a Franka Duo. Only ER 2 is publicly available.
Meet Token Saver: An Open-Source MCP Extension Using Local Hybrid RAG...
Marktechpost AI has released Token Saver, an open-source MCP extension for Claude Desktop that uses local Hybrid RAG to slash PDF token consumption by up to 99% while ensuring absolute document privacy.
Datalab Marker v2 vs MinerU, Docling, and Liteparse: Benchmark Breakdown
Datalab rewrote Marker as a three-mode pipeline. Version 2 hits 76.0 on olmOCR-bench and sustains 2.9 pages per second on one B200 — over 5× MinerU's pipeline backend, while beating Docling on both accuracy and speed. Here's how it compares against MinerU, Docling and LiteParse, and which one fits your use case.
Best Local LLMs You Can Run on a Single 24GB GPU...
A single 24GB GPU is the practical floor for serious local inference. This guide compares six open-weight models that fit one card at Q4_K_M. It covers Qwen3.6, Gemma 4, Mistral Small, gpt-oss-20b, and DeepSeek-R1-Distill. Each entry lists VRAM fit, licensing, and the job it does best.
Meet LingBot-World-Infinity: An Open Causal World Model With An Agentic Harness
Robbyant, Ant Group's embodied-intelligence unit, has released LingBot-World-Infinity (LingBot-World 2.0). It is a 14B causal video generation model that behaves as an interactive world simulator. The core idea is the Mixture of Bidirectional and Autoregressive (MoBA) attention mask, paired with distribution matching distillation applied over long self-rollout trajectories. Together they target long-horizon drift, the failure mode that smears textures and warps geometry in most interactive world models. A Director-Pilot agentic harness wraps the generator, where a VLM proposes events and the Diffusion Transformer renders them. The report shows a single 60-minute uninterrupted session covering 20 scenarios. But the release is thinner than the paper: one checkpoint, a 480P reference script, no deployment code, no quantitative benchmarks, and a non-commercial CC BY-NC-SA 4.0 license.
Ant Group’s Robbyant Open-Sources LingBot-Vision: A 1B Boundary-Centric Vision Foundation Model...
Ant Group's Robbyant open-sourced LingBot-Vision, a self-supervised ViT family for dense spatial perception. Masked boundary modeling makes image boundaries a native training signal. The 1B backbone matches or surpasses larger models, and initializes LingBot-Depth 2.0.
Baidu Releases Unlimited OCR, a 3B Model That Keeps the KV...
Baidu open-sourced Unlimited OCR, a 3B-parameter MoE model that parses dozens of document pages in a single forward pass. Its Reference Sliding Window Attention (R-SWA) holds the KV cache constant, so memory and latency stay flat as output grows. It scores 93.23 on OmniDocBench v1.5, beating the DeepSeek OCR baseline by 6.22 points, under an MIT license.
Zyphra Release Zamba2-VL: Hybrid Mamba2–Transformer Vision-Language Models That Cut Time-to-First-Token by...
Zyphra has released Zamba2-VL, a family of open vision-language models at 1.2B, 2.7B, and 7B parameters. The models use a hybrid Mamba2 state-space and Transformer backbone, shipping under Apache 2.0. They stay competitive with comparable Transformer VLMs while cutting time-to-first-token by about an order of magnitude.
StepFun Releases Step 3.7 Flash: A 198B MoE Vision-Language Model for...
StepFun releases Step 3.7 Flash, a 198B MoE model with native vision, 256k context, and Advisor Mode.
NVIDIA Introduces SANA-WM: A 2.6B-Parameter Open-Source World Model That Generates Minute-Scale...
Researchers from NVIDIA introduce SANA-WM, an open-source camera-controlled world model that generates 60-second, 720p videos with precise 6-DoF camera control — trained on 64 H100 GPUs and deployable on a single RTX 5090.
Microsoft Research’s World-R1 Uses Flow-GRPO and 3D-Aware Rewards to Inject Geometric...
Microsoft Research's World-R1 Uses Reinforcement Learning to Force 3D Consistency Into Text-to-Video Models
Meta AI Releases Sapiens2: A High-Resolution Human-Centric Vision Model for Pose,...
Meta Reality Labs releases a new foundation model family for human-centric vision that pushes pose estimation, segmentation, and 3D geometry to new state-of-the-art levels — all from a single backbone.
Qwen Team Open-Sources Qwen3.6-35B-A3B: A Sparse MoE Vision-Language Model with 3B...
Qwen Team Open-Sources Qwen3.6-35B-A3B: A Sparse MoE Vision-Language Model with 3B Active Parameters and Agentic Coding Capabilities
Liquid AI Releases LFM2.5-VL-450M: a 450M-Parameter Vision-Language Model with Bounding Box...
Liquid AI just released LFM2.5-VL-450M, an updated version of its earlier LFM2-VL-450M vision-language model. The new release introduces bounding box prediction, improved instruction following,...
Meta AI Releases EUPE: A Compact Vision Encoder Family Under 100M...
Running powerful AI on your smartphone isn't just a hardware problem — it's a model architecture problem. Most state-of-the-art vision encoders are enormous, and...
TII Releases Falcon Perception: A 0.6B-Parameter Early-Fusion Transformer for Open-Vocabulary Grounding and Segmentation from...
In the current landscape of computer vision, the standard operating procedure involves a modular 'Lego-brick' approach: a pre-trained vision encoder for feature extraction paired...
IBM Releases Granite 4.0 3B Vision: A New Vision Language Model...
IBM has announced the release of Granite 4.0 3B Vision, a vision-language model (VLM) engineered specifically for enterprise-grade document data extraction. Departing from the...
Microsoft Releases Phi-4-Reasoning-Vision-15B: A Compact Multimodal Model for Math, Science, and...
Microsoft has released Phi-4-reasoning-vision-15B, a 15 billion parameter open-weight multimodal reasoning model designed for image and text tasks that require both perception and selective...
Physical Intelligence Team Unveils MEM for Robots: A Multi-Scale Memory System...
Current end-to-end robotic policies, specifically Vision-Language-Action (VLA) models, typically operate on a single observation or a very short history. This 'lack of memory' makes...
Google AI Just Released Nano-Banana 2: The New AI Model Featuring...
In the escalating 'race of "smaller, faster, cheaper' AI, Google just dropped a heavy-hitting payload. The tech giant officially unveiled Nano-Banana 2 (technically designated...
NVIDIA Releases DreamDojo: An Open-Source Robot World Model Trained on 44,711...
Building simulators for robots has been a long term challenge. Traditional engines require manual coding of physics and perfect 3D models. NVIDIA is changing...
Google Introduces Agentic Vision in Gemini 3 Flash for Active Image...
Frontier multimodal models usually process an image in a single pass. If they miss a serial number on a chip or a small symbol...
Ant Group Releases LingBot-VLA, A Vision Language Action Foundation Model For...
How do you build a single vision language action model that can control many different dual arm robots in the real world? LingBot-VLA is Ant...
Liquid AI Releases LFM2.5: A Compact AI Model Family For Real...
Liquid AI has introduced LFM2.5, a new generation of small foundation models built on the LFM2 architecture and focused at on device and edge...
Zhipu AI Releases GLM-4.6V: A 128K Context Vision Language Model with...
Zhipu AI has open sourced the GLM-4.6V series as a pair of vision language models that treat images, video and tools as first class...
Jina AI Releases Jina-VLM: A 2.4B Multilingual Vision Language Model Focused...
Jina AI has released Jina-VLM, a 2.4B parameter vision language model that targets multilingual visual question answering and document understanding on constrained hardware. The...
Tencent Hunyuan Releases HunyuanOCR: a 1B Parameter End to End OCR...
Tencent Hunyuan has released HunyuanOCR, a 1B parameter vision language model that is specialized for OCR and document understanding. The model is built on...
Uni-MoE-2.0-Omni: An Open Qwen2.5-7B Based Omnimodal MoE for Text, Image, Audio...
How do you build one open model that can reliably understand text, images, audio and video while still running efficiently? A team of researchers...
Baidu Releases ERNIE-4.5-VL-28B-A3B-Thinking: An Open-Source and Compact Multimodal Reasoning Model Under...
How can we get large model level multimodal reasoning for documents, charts and videos while running only a 3B class model in production? Baidu...
Liquid AI’s LFM2-VL-3B Brings a 3B Parameter Vision Language Model (VLM)...
Liquid AI released LFM2-VL-3B, a 3B parameter vision language model for image text to text tasks. It extends the LFM2-VL family beyond the 450M...
Baidu’s PaddlePaddle Team Releases PaddleOCR-VL (0.9B): a NaViT-style + ERNIE-4.5-0.3B VLM...
How do you convert complex, multilingual documents—dense layouts, small scripts, formulas, charts, and handwriting—into faithful structured Markdown/JSON with state-of-the-art accuracy while keeping inference latency...
Alibaba’s Qwen AI Releases Compact Dense Qwen3-VL 4B/8B (Instruct & Thinking)...
Do you actually need a giant VLM when dense Qwen3-VL 4B/8B (Instruct/Thinking) with FP8 runs in low VRAM yet retains 256K→1M context and the...
IBM AI Releases Granite-Docling-258M: An Open-Source, Enterprise-Ready Document AI Model
IBM has released Granite-Docling-258M, an open-source (Apache-2.0) vision-language model designed specifically for end-to-end document conversion. The model targets layout-faithful extraction—tables, code, equations, lists, captions,...
Qwen Team Introduces Qwen-Image-Edit: The Image Editing Version of Qwen-Image with...
In the domain of multimodal AI, instruction-based image editing models are transforming how users interact with visual content. Just released in August 2025 by...
Zhipu AI Releases GLM-4.5V: Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Zhipu AI has officially released and open-sourced GLM-4.5V, a next-generation vision-language model (VLM) that significantly advances the state of open multimodal AI. Based on...
Tencent Open Sources Hunyuan-A13B: A 13B Active Parameter MoE Model with...
Tencent's Hunyuan team has introduced Hunyuan-A13B, a new open-source large language model built on a sparse Mixture-of-Experts (MoE) architecture. While the model consists of...
Alibaba Qwen Team Releases Qwen-VLo: A Unified Multimodal Understanding and Generation...
The Alibaba Qwen team has introduced Qwen-VLo, a new addition to its Qwen model family, designed to unify multimodal understanding and generation within a...
DeepSeek Researchers Open-Sourced a Personal Project named ‘nano-vLLM’: A Lightweight vLLM...
The DeepSeek Researchers just released a super cool personal project named 'nano-vLLM', a minimalistic and efficient implementation of the vLLM (virtual Large Language Model)...
Meta AI Releases Web-SSL: A Scalable and Language-Free Approach to Visual...
In recent years, contrastive language-image models such as CLIP have established themselves as a default choice for learning vision representations, particularly in multimodal applications...
Long-Context Multimodal Understanding No Longer Requires Massive Models: NVIDIA AI Introduces...
In recent years, vision-language models (VLMs) have advanced significantly in bridging image, video, and textual modalities. Yet, a persistent limitation remains: the inability to...
Transformer Meets Diffusion: How the Transfusion Architecture Empowers GPT-4o’s Creativity
OpenAI’s GPT-4o represents a new milestone in multimodal AI: a single model capable of generating fluent text and high-quality images in the same output...
Google DeepMind Research Releases SigLIP2: A Family of New Multilingual Vision-Language...
Modern vision-language models have transformed how we process visual data, yet they often fall short when it comes to fine-grained localization and dense feature...
Google DeepMind Releases PaliGemma 2 Mix: New Instruction Vision Language Models...
Vision‐language models (VLMs) have long promised to bridge the gap between image understanding and natural language processing. Yet, practical challenges persist. Traditional VLMs often...
IBM AI Releases Granite-Vision-3.1-2B: A Small Vision Language Model with Super...
The integration of visual and textual data in artificial intelligence presents a complex challenge. Traditional models often struggle to interpret structured visual documents such...
NVIDIA AI Releases Eagle2 Series Vision-Language Model: Achieving SOTA Results Across...
Vision-Language Models (VLMs) have significantly expanded AI’s ability to process multimodal information, yet they face persistent challenges. Proprietary models such as GPT-4V and Gemini-1.5-Pro...
Qwen AI Releases Qwen2.5-VL: A Powerful Vision-Language Model for Seamless Computer...
In the evolving landscape of artificial intelligence, integrating vision and language capabilities remains a complex challenge. Traditional models often struggle with tasks requiring a...
Qwen Team Releases QvQ: An Open-Weight Model for Multimodal Reasoning
Multimodal reasoning—the ability to process and integrate information from diverse data sources such as text, images, and video—remains a demanding area of research in...
Meta AI Releases Apollo: A New Family of Video-LMMs Large Multimodal...
While multimodal models (LMMs) have advanced significantly for text and image tasks, video-based models remain underdeveloped. Videos are inherently complex, combining spatial and temporal...
Tsinghua University Researchers Released the GLM-Edge Series: A Family of AI...
The rapid development of artificial intelligence (AI) has produced models with powerful capabilities, such as language understanding and vision processing. However, deploying these models...
Hugging Face Releases SmolVLM: A 2B Parameter Vision-Language Model for On-Device...
In recent years, there has been a growing demand for machine learning models capable of handling visual and language tasks effectively, without relying on...
Nexa AI Releases OmniVision-968M: World’s Smallest Vision Language Model with 9x...
Edge AI has long faced the challenge of balancing efficiency and effectiveness. Deploying Vision Language Models (VLMs) on edge devices is difficult due to...
Are Small Language Models Really the Future of Language Models? Allen...
Multimodal models represent a significant advancement in artificial intelligence by enabling systems to process and understand data from multiple sources, like text and images....
Pixtral 12B Released by Mistral AI: A Revolutionary Multimodal AI Model...
The release of Pixtral 12B by Mistral AI represents a groundbreaking leap in the multimodal large language model powered by an impressive 12 billion...
CogVLM2: Advancing Multimodal Visual Language Models for Enhanced Image, Video Understanding,...
Large Language Models (LLMs), initially limited to text-based processing, faced significant challenges in comprehending visual data. This limitation led to the development of Visual...
Qwen2-VL Released: The Latest Version of the Vision Language Models based...
Researchers at Alibaba have announced the release of Qwen2-VL, the latest iteration of vision language models based on Qwen2 within the Qwen model family....
NVEagle Released by NVIDIA: A Super Impressive Vision Language Model that...
Multimodal large language models (MLLMs) represent a significant leap in artificial intelligence by combining visual and linguistic information to understand better and interpret complex...
































































