I'm a little bit disappointed that vision seems to fall before language at scale.
It seems pretty counter intuitive that we can't do vision significantly better with specialized techniques.