That’s surprising even gpt-image-1 was able to handle the equestrian astronaut prompt pretty trivially, albeit yellow-tinged as all hell.
For the heck of it, I ratcheted up the difficulty of the prompt (two-headed horse, astronaut with the visor up, etc.), and gpt-image-2 got it right every time.