logoalt Hacker News

jmugantoday at 5:56 PM4 repliesview on HN

A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)


Replies

IsTomtoday at 7:22 PM

Aren't they still bad at understanding how bicycle frame works? Especially the steering part?

show 4 replies
maxutilitytoday at 6:05 PM

Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting.

A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.

show 2 replies
dofmtoday at 6:34 PM

But it's another benchmark on how good models are at generating intensely average, unwanted things with unthinking design. Just scaled up.

charcircuittoday at 7:33 PM

Bad? It has a charming style. I would watch the whole book if it was made like this.

show 2 replies