the ceiling of abilities seem to be growing steadily but the floor of errors seems to not change. New models can do more and more but still fail at seemingly (to human) simple tasks