Heh, there's one of mine: https://stoppels.ch/goalposts/?c=39727943
"GPT-4 looks at original ASCII art of a foot, not copied from the web, and says it is a foot."
The vote is currently 64% yes, 18% no.
Just now I asked Opus 5.5 to generate an ASCII art foot, and it did a passable job. It's not great, but it's a foot. Then I pasted it into ChatGPT (whatever they're serving to the free tier by default, which seems to be 5.6 Luna), and it said it was a "train/locomotive": https://chatgpt.com/share/6abeaa39-cc80-83ed-851f-29370db089...
Maybe it's Opus's fault for drawing a bad foot but I think it's fair to say LLMs are still pretty bad at ASCII art (without additional tool calling etc).
Mm, I kina agree with the AI on this one:
(_)(_)(_) represents the wheels
They do look rather wheel-like; I have to assume you see them as toes though?It's like the duck-bunny picture to me. If I focus on the "wheels", I see a steam train locomotive (but perhaps I'm only seeing that because I read your comment?); if I look at the ankle I see a foot.
Readers: before you vote or comment, look at that foot.
I think I would have failed this test!
I think the problem is that you're using basing your conclusion from the cheap/dumb models available on the free tier of services. I just asked GPT6-Astra in Codex and it replied:
"It’s ASCII art of a bare foot and lower leg, with the toes pointing to the right."
No tool calling, just an immediate reply with the correct answer.
A better test would be using an image to ascii converter tool to rule out bad ascii drawing from Opus.
I'm a human, and that's not a foot, it's a smokestack.
But as you pointed out, while that absolves ChatGPT, it makes Opus look worse.
Opus 5.5 was able to parse and understand an ASCII art foot when I pasted one in.
This is a Rorschach test, not a foot. If you'd shown this to me without telling me what it was meant to be first, I'd have guessed a crematorium.
To be fair if a human was given a linear sequence representing ascii art you couldn't tell either
For me, 6.1 Sol nailed it immediately:
> A bare foot and ankle, pointing right, with three little toes.
I wonder how much of the wide variation in perceptions of LLM capabilities is driven by the gulf between free models and frontier models. Luna getting something wrong is not always great evidence for LLMs be unable to do that thing.
Edit: for curious skeptics without access to 6.1 Sol, I tried 3 times and it got it all 3 times. Convo share link: https://chatgpt.com/share/e/6abeb955-7614-832e-a5e1-b1bd134f...