Comparison between the ground-truth annotation, the initial SAM mask, and our refined mask on the Figurines scene. In this cluttered roomtop scene with many small objects, the initial SAM mask often produces fragmented or overly fine-grained predictions, whereas our method yields cleaner object-level masks with more consistent identities.