That's wonderful. I was going off an older version of the Artificial Analysis page for GLM-5.3-Flash https://artificialanalysis.ai/models/glm-5-3-flash. The page is updated now to show that it does support multi-modal image inputs.