logoalt Hacker News

Karpathy’s Pelican

311 pointsby delichontoday at 4:05 AM247 commentsview on HN

https://xcancel.com/karpathy/status/2083749667410727319


Comments

stackedinsertertoday at 6:26 PM

It would be better to ask model to render segmented 3d, with placeholders, like magenta is water, blue is sky, green is grass, purple is Frodo's face, etc, then pass the result through img2img model to properly "render" it.

forrestthewoodstoday at 5:19 PM

As a former gamedev watching non-gamedev AI talk about games is so amusing. They really truly do not understand anything about games or consumer entertainment.

There’s a reason AI slop games have literally zero engagement. Last summer that stupid flying game blew up. Maybe a million people “played” the game. Where play means they clicked a link and checked it out not because of what the game was but solely because of how it was made.

In terms of concurrent players that game wouldn’t have cracked the Top 5,000 on Steam.

My metric for AI games is “number of players who spent more than 15 minutes playing”. I’m not aware of any vibeslop that has achieved 1 such player.

Now obviously LLMs are transformative for game dev. But “hyper custom worlds you can drop into” shows an extreme ignorance of what players want imho.

epolanskitoday at 4:45 PM

I wish there was a timeline where I never ever had to see the pelican SVG test ever again.

show 1 reply
mvdtnztoday at 8:49 PM

> I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free".

Except it's not ~free, it cost ~$10. And no one in their right mind would ever exchange $10 for that crap output except in this brief moment that we're in because it's fun and surprising to see what will happen. The actual result is as close to useless as it's possible to be - it's not interesting in its own right, it's not aesthetically pleasing, nor funny, nor informative. It's just slop.

nozzlegeartoday at 5:30 PM

Don't miss Elon Musk's reply:

> @elonmusk 13h

> Yah

> 158 replies, 74 reposts, 1400 likes

Thanks Elon, you goofy fuck

https://xcancel.com/elonmusk/status/2083761408932458568

show 2 replies
angoragoatstoday at 8:52 PM

Can we please link directly to XCancel and leave the fascist garbage site behind?

miltonlosttoday at 5:35 PM

Tech bros continue wasting money to make the absolute worst art

show 2 replies
shapefrogtoday at 5:13 PM

"Check the current situation and make a new Iran Lego (tm) truth bomb video."

hansmayertoday at 8:18 PM

[dead]

trlhaqtoday at 5:00 PM

[flagged]

show 2 replies
theproblemisyoutoday at 5:31 PM

[flagged]

yourewrongsorrytoday at 5:15 PM

[flagged]

andrewstuarttoday at 7:03 PM

This is equally bad as a pelican test.

LLMs should be tested in the same way people should be tested for a job interview (but often aren’t) - with tasks RELEVANT to usage.

So you don’t just randomly pick some random thing to make the LLM randomly do (like many job interviewers do).

You start with clear statements about real world usage scenarios. THEN you come up with tests that give insight to how well the LLM/hob seeker gets the job done.

Please, stop coming up with random tests like it’s Microsoft in 1990 and you’re asking job seekers how the would move Mount Fuji, as a way of assessing their programming skills.

No stupid irrelevant pelicans on bicycles and no stupid renderings of Lord Of The Rings. Unless those are relevant use cases.

Any test that anyone comes up with must clearly state the context and how the outcome is measured.

show 1 reply
hn22fazjsvtoday at 6:08 PM

Screenshotting for later

matchagauchotoday at 4:51 PM

It's difficult to think in exponentials.

But this demonstrates we're a couple orders of magnitude away from generating 1:1 hyper-personalized entertainment and media for individuals, rather than the masses.

show 6 replies