I tried out the multimodality and this is the biggest win, imo.
You can give it a CCTV image (like of a train platform) and ask it to quickly decide actions such as triggering an automated auditory alert, deferring to a larger model for more detailed analysis, etc.