SAM 2 is a segmentation model that enables fast, precise selection of any object in any video or image.
A click, box, or mask on any image or frame of video.
SAM 2 is the first unified model for segmenting objects across images and videos. Prompt it with a click, box, or mask to select any object on any image or frame of video.
A single architecture segments still images and video frames with consistent, high-quality results.
Select any object precisely using any prompt type, then use additional prompts to refine predictions.
A data engine that extends annotation from images to video powers SAM 2 training at scale.
Checkpoints, inference code, and the web demo are publicly available for research and applications.
State-of-the-art segmentation that is faster, more robust, and easier to use than ever before.
Segment any object, now in any video or image.
Select one or multiple objects in a video frame and adjust predictions across frames.
Strong performance on objects and videos never seen during training, enabling real-world applications.
Streaming inference designed for efficient, interactive video processing.
Top results for object segmentation in both videos and images.
A unified model with a straightforward design and fast inference speed.
The next generation of Meta Segment Anything, bringing image and video together.
Questions about SAM 2, segmentation, and how to get started.
Track an object across any video interactively with as little as a single click, and create fun effects.