A big wtf at the LM Arena scores: https://arena.ai/leaderboard/text-to-image
gpt-image-2.5-sunburst: 1421
gpt-image-2.5-flare: 1399
gpt-image-2 (medium): 1381
mai-image-2.6: 1331
Even with LM Arena being flawed, this is significant. I was planning to do a writeup on the original gpt-image-2 as it crushed every complex image comprehension benchmark I had...I'm glad I procrastinated since ChatGPT Images 2.5 seems like an even better starting point to test out what these models can actually do nowadays.
Many people still think AI images output the wrong number of fingers on a regular basis. (EDIT: this was an ironic comment to make in hindsight and I own it)
> Many people still think AI images output the wrong number of fingers on a regular basis.
Last time I used a frontier image model it made me a seal with three hands so…
I mean there's literally one of their example images with hands with the wrong number of fingers.
> Many people still think AI images output the wrong number of fingers on a regular basis.
There's literally an image of a dude with 3 fingers in the Composite Party Photo.