why isn't this being compared to siglip2 (also from google)? because that one isn't fully multimodal? or because it's a different org/team?