Molmo2
VLM
A fully open Vision-Language Model with video grounding capabilities
Molmo2 (Multimodal Open Language Model 2) is a fully open Vision-Language Model (VLM) family developed by the Allen Institute for AI and the University of Washington. Its defining capability is video grounding: identifying when and where an event or object appears within a video.
This book examines the open data pipeline, model family, and grounding evaluations behind that capability, including the reported comparisons with proprietary systems.