logoalt Hacker News

My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”

122 pointsby thebigshipyesterday at 7:42 PM54 commentsview on HN

Comments

zirkuswurstikustoday at 4:04 AM

https://imgur.com/a/2DFUpGZ

ChatGPT MMD

show 2 replies
jnwatsontoday at 2:02 AM

Fable 5 on Max knocks it out of the park:

https://imgur.com/a/usR8K7G

Definitely has some creative flourishes.

(I made no extra prompting. Just the above text. Single shot.)

show 1 reply
hn_throwaway_99yesterday at 8:32 PM

I thought this was great, and hilarious. Kudos to Opus 5, I thought it was the only one that came close to passing. Interestingly, I thought many of the failures drew the frog face OK, and they had some type of big blob for the jaw, so they knew "Hapsburg jaw" meant a protruding jaw, but it wasn't really connected to the frog face in any way that made sense.

Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.

show 1 reply
krisoftyesterday at 11:54 PM

Curious that none of the attempts draw the frog from side profile. If i have to draw this i would immediately know that drawing a recognisable frog is the easy part of job. Expressing a particular jaw shape and melding it on the frog is the hard part. And jaw shapes are more prominent from the side.

Even absence of thinking this through you would think that some frogs will be from the front, some from the side. Just by chance. And yet all appears to go for the harder pose.

show 1 reply
aboundtoday at 1:32 AM

My personal benchmark is a directory containing a bunch of research papers on the physics of popping popcorn kernels, and a prompt about creating high fidelity, photo realistic 3D models of all the different kinds of popped kernels. Fable (surprisingly? unsurprisingly?) refused to do it last time I tried, and the results from other models are, well, fine, but there's still plenty of headroom on this particular one.

show 1 reply
thebigshipyesterday at 8:47 PM

Hi all, the site is getting hugged to death, thank you, was not expecting this kind of warm response. I will be working to make this more reliable, in the meantime, sign up for my newsletter: https://www.jaymollica.com/blog/

also my favorite SVG was def the google/gemini-3.6-flash

edit: ok better now I think

show 1 reply
riazrizviyesterday at 10:41 PM

The secret to great interview questions and challenge tests is keeping them secret. Posting them on HN and getting them onto the front page puts them in jeopardy.

wren6991yesterday at 9:04 PM

Opus 5 clearly frogmaxxed.

gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.

show 2 replies
evan_yesterday at 9:28 PM

Hopsburg Jaw

getnormalityyesterday at 8:28 PM

This is a strong benchmark! None of these could be remotely mistaken for human art. Opus 5 comes closest.

rush86999yesterday at 9:20 PM

It's opus 5 > Kimi K3 > grok 4.5

That's a pretty good benchmark

ianberdinyesterday at 10:05 PM

Check out my MacBook SVG benchmark. From my experience, it demonstrates the Real model’s behavior. However, I notice the errors it makes, which are similar to the mistakes made by the mistake model in code.

https://playcode.io/blog/macbook-svg-benchmark

buffer_overlordtoday at 2:23 AM

They all suck

ricardobeatyesterday at 10:06 PM

Gemini 2.5 Pro fails, but has a distinctive art style that is quite nice. It seems to understand shading to a much higher level than all other models.

dehrmannyesterday at 8:50 PM

How do models approach SVG generation? In one version, I imagine them actually trying to reason about them as an LLM. In another, I imagine something closer to a GAN.

linksnapzzyesterday at 8:51 PM

A friend’s favorite prompt is “Batman & Julia Child; in the kitchen laughing at a ham”. Sounds simple, but has been surprisingly tough.

show 1 reply
k1eyesterday at 10:38 PM

No gpt-5.6 sol and no fable?

leumonyesterday at 8:39 PM

Can you also try the new deepseek v4 flash?

show 1 reply
gerdesjyesterday at 9:56 PM

Please could we have a human generated image to compare the AI generated tosh with?

MiroslavPokornyyesterday at 9:46 PM

My test is to ask AI to pick up all the rubbish at the beach.

show 1 reply
csomaryesterday at 10:39 PM

Here is GLM 5.2 (https://codeinput.com/s/HAO0qTxw2ia) which is still inferior to Opus. I can't find Qwen 3.8 which now is my daily driver replacing GLM. This SVG test matches my experience when working with the different models. The other models can get the details right but their output is structured in a way that makes little or no sense.

I also did a timeline from 4.7 to 5.2: https://codeinput.com/s/7oK2IIA7qRO The improvements in models looks much less impressive with this test.

throwuxiytayqyesterday at 10:26 PM

My personal human benchmark: "Jump on one leg, while reciting the national anthem of Latvia, translated to Spanish, backwards, while drawing a frog with a brush held by toes of the other leg, on the ceiling". So far they're not doing very good but I'm sure they'll improve over time.

show 1 reply
epolanskiyesterday at 10:20 PM

I don't get the point of these benchmarks, what are they supposed to represent practically?

show 2 replies
epolanskiyesterday at 9:37 PM

Gemini 3.6 flash is crazy good.

Would've wanted to see also DS4 flash.

show 1 reply
troupoyesterday at 8:50 PM

Mine is any variations on mammoths in various situations, or anthropomorphic. Since mammoths are invariably majestically going from one place to another in any of the books, models have hard time imagining anything but that.

Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)

AlienRobotyesterday at 11:05 PM

Am I the only one who thinks it's incredible that an LLM can do this, and at the same time it's ridiculous to expect it to be capable of doing it, even thought it clearly can do it?

thebigshipyesterday at 7:42 PM

I think this one has advantages over the “pelican riding a bicycle” one because it hinges on an anatomical feature that many models associate with royalty, “habsburg” being a lineage and “habsburg jaw” being an anatomical feature.

Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.

Mistral returned byte-identical output across separate calls.

Gemini narrates its work in 65 comments; Llama says nothing.

If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.

show 2 replies
sixtyjyesterday at 8:23 PM

For those who don’t know a Habsburg jaw also known as mandibular prognathism, it is a genetic condition characterized by a protruding lower jaw, which was notably prevalent among members of the Habsburg royal family due to their history of inbreeding. This condition often resulted in significant facial deformities and difficulties with eating and speaking.

slipperybelugatoday at 2:33 AM

[dead]

kindawindayesterday at 9:49 PM

[flagged]

NemoNobodyyesterday at 10:54 PM

[flagged]