logoalt Hacker News

steve-atx-7600today at 1:57 AM5 repliesview on HN

You would not expect the developers of the model to optimize for a well known benchmark?


Replies

simonwtoday at 4:17 AM

Here's 'Generate an SVG of a ring-tailed lemur riding an electric scooter' at reasoning level max: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Quote from the thinking trace:

> I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.

It's pretty solid - face is a little wonky but excellent tail and scooter.

show 2 replies
y1n0today at 2:14 AM

I would expect it happens more organically where people’s discussion of the benchmark and posted results find their way into the training data, like anything else on the internet.

kibaetoday at 2:55 AM

There was a HN post that tested this hypothesis a few weeks ago: https://news.ycombinator.com/item?id=49010129

Kranartoday at 2:20 AM

What is this supposed to mean exactly? Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles? Like how exactly do you expect them to be doing that anyways? Hiring graphic artists to create SVGs of bike riding pelicans and feeding thousands of them into the model's training set?

show 2 replies
mi_lktoday at 2:24 AM

It stopped being a valuable benchmark proxy quite a few model versions ago. Simon knows it, so is everyone who’s serious about it

Treat it like a bit as is

show 1 reply