logoalt Hacker News

simonwtoday at 7:42 PM13 repliesview on HN

  llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle"
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

4.2266 cents, 38 seconds.

For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.

UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

The most expensive was reasoning level xhigh - 7.5 cents, 1m34s.

And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...


Replies

ipsum2today at 9:55 PM

All of the links show "Error: Gist API returned 403".

drusepthtoday at 7:45 PM

Is there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality.

show 9 replies
hollowturtletoday at 9:03 PM

Is there any point anymore regarding this svg test? I would not be surprised if in the training they're fine tuned for this task too

show 2 replies
tintortoday at 9:14 PM

Did any LLM so far draw pelican knees correctly and have them bend in opposite direction from human knees? Knees of many animals bend opposite to humans.

Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.

show 1 reply
drob518today at 8:52 PM

Simon, at this point I really wonder if teams aren’t gaming this. You should pick a random animal doing a random thing every time.

jonahxtoday at 7:57 PM

If you have a grading rubric, huge points off for adding arms instead of using the wings as arms!

show 1 reply
jonplacketttoday at 8:22 PM

Has any ab tried to game this yet and just made the most amazing pelican by hand and always reply with that?

jttnrtoday at 8:33 PM

I wonder, given Simons reputation in AI benchmarking, whether model providers try to train or tweak their models to perform better at drawing bicycles and pelicans?

EugeneOZtoday at 8:37 PM

Absolutely BRUTAL! :)

Thank you for doing this, I love your benchmark the most!

0xbadcafebeetoday at 9:29 PM

For all the comments of "I'm sure they're fine-tuning for pelicans": https://dylancastillo.co/posts/pelicanmaxxing.html

  "Sorry, HN haters, but there’s little evidence that AI labs are pelicanmaxxing.
   Or at least they’re not doing it in a plainly obvious manner."
tomrodtoday at 7:55 PM

What does the mean pelican look like at this point?

Also 3X token use vs. 1.2

show 1 reply
jmknitoday at 7:45 PM

lol

Definitely an upgrade over 1.2

NoOneCares44today at 8:23 PM

[flagged]

show 2 replies