Molmo2

VLM
A fully open Vision-Language Model with video grounding capabilities
Author
Published

February 3, 2026

Last Updated

August 8, 2026

Molmo2 (Multimodal Open Language Model 2) is a fully open Vision-Language Model (VLM) family developed by the Allen Institute for AI and the University of Washington. Its defining capability is video grounding: identifying when and where an event or object appears within a video.

This book examines the open data pipeline, model family, and grounding evaluations behind that capability, including the reported comparisons with proprietary systems.

Paper · Code · Demo