You don't think it's useful to learn whether a model's "intelligence" generalizes beyond the tasks and modalities it is usually optimized for?
Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.
Absolutely not. Makes about as much sense as judging a car based on how good an airplane it makes.