I keep seeing Astra make beautiful 3d stuff online, yet when I feed it some old school RuneScape assets (even tried with some very detailed guidelines) and asked it to generate some new plausible assets it failed horribly.
I think there’s still something really off with current (frontier) models when it comes to creating “novel” stuff? Even 2004 style graphics..
Or am promoting it wrong?
It’s a tiring YouTube formula:
“Model X CHANGES EVERYTHING!” (surprised Oculus Rift face without the headset)
Some one-shot prompts for quick impact.
I get it with the incentive structure — honestly not trying to shit on any creators (hate the game, not the player) but seeing it every day is just getting tiring.
This sounds unconvincing, because a) pelican test is subjective, there's simply nothing to leak as it has no available direct answers and maybe an extremely faint preference signal, and b) the same small models actually do perform well when you change the subject. Some models are genuinely trained to be better at some domain, such as 2D layouts or vector graphics in this case. It all depends on particular recipes and datasets. Which is the actual reason these tests are poor as vibe checks: they don't do anything to disentangle generalization, memorization, and training preference. One-shotting popular software in particular is definitely not a good test of anything as memorization is going to dominate it.
AAII is also not very useful, neither is any generic score/benchmark. If you want a weather forecast you aren't looking at the average temperature of Earth.
(actually when did the term "one-shot" get hijacked to mean something other than "one example"?..)
I used to think that way about SVGBench, after all labs can just train on the test set, right? It turns out the task was highly generalizable. Try designing a logo and you quickly see the gap between models visually. Even though there is still a gap between "shiny demo SVG" and actual real-world use.
Same thing happened with MineBench, basically SVGBench+3D, until that got "saturated".
Remember spinning hexagon bench[1]? Or the AI World Clocks[2]? Yeah, that used to be hard for frontier models.
Creating games is the next iteration that still has some signal left. Assets + Game logic + UI + Sound, it let's you assess a model's "taste" very quickly.
What else is left, once all these benchmarks get saturated?
> That’s the problem, these tests can’t tell you how good a model is anymore because it’s trivial for labs to optimise for exactly these tests by the next release.
But they don't prove the claim. Are the models amazing at recreating Minecraft, but the second you swap the word Minecraft out with another game or a custom game, it shits the bed?
That's not what I see. My feed is full of people using Astra to recreate all sorts of games from Diablo to some random idea they came up with, in ridiculously polished detail like animations that would have taken me weeks of iteration in gpt-5.6-sol but it was a single shot by Astra.
"When a measure becomes a target, it ceases to be a good measure."
Goodhart's law strikes again.
https://en.wikipedia.org/wiki/Goodhart%27s_law
However, what is meaningful is whether something is able to create usefully adjacent output, like "let's make Minecraft, but with marching cubes, subdivision surfaces, and global illumination... and behaviorally accurate pandas..." (or something like that).
I have an 11 year old, and most of his game ideas are adjacent to other games he's played. He can make those now, or, at least, enough that he can see where it works and where it doesn't.
Compared to a few years ago, that's pretty cool.
I use Minecraft (and Warcraft and an arcade flying simulator) as a silly but directionally correct indication of the models' capabilitites.
Compare Astra[0] with GPT 5.4[1] which was OpenAI's state of the art just six months ago.
(all tests on more models with code and prompts available here: https://senko.net/vibecode-bench )
Yes, it's not a scientific benchmark but it's a good heuristic.
For a better eval, create a one-page prompt / mini spec related to whatever you're using the LLMs for, and see how well a particular one works for what's important to you.
0: https://senko.net/vibecode-bench/2026/rts-gpt-6-astra.html
The benchmark I want to see people adopt is:
Build a flowsheet based steady state chemical process simulator, then use it to simulate and optimize a full scale oil refinery.
1) Building a solver engine that works at this scale is not a trivial problem, and the successful ones rely more on heuristics than some categorically different solution approach.
2) Defining the engineering equations relevant to this task is relies on understanding what level of fidelity is required to answer the questions people ask of steady state process models.
3) Knowing the thermophysical properties of chemicals and crude oils is possible from the open literature, but the information is diffuse and different correlations are applicable in different situations.
4) Creating a GUI which converts a flowsheet into matrix math is non-trivial, although a sequential modular approach is a bit easier.
5) Defining large scale models in such a way that they solve robustly is as much art as science. For example, completely closed recycle loops like refrigeration systems are a nightmare for solvers, so it is often better to define them in an open-loop way.
6) Optimization involves knowing the relevant commodity prices, but more importantly how to define the constraints on the model so it doesn't just say to produce infinite gasoline.
7) Troubleshooting the inevitable convergence failures is also as much art as science. There are a large number of diagnostic techniques, but fundamentally you need to be able to relate what is happening during the solver iterations with the intent of your model because more often than not the problem is that you've asserted something impossible, redundant, or irrelevant.
But they're not recreating Minecraft. They're recreating one of the most common tutorials on the internet that shows a simple voxel overworld and nothing else. That's not Minecraft.
I don't remember seeing playstation controller SVGs before, I doubt they benchmaxxed that
Definitely not.
It has to be Doom or Crysis, aren't they the ones people usually ask if it can run?
Though I somewhat think some benchmarks are silly like the article says. I saw someone on YouTube recently take a picture of a building across the street from them (seemed like it was in NYC), and asked GPT 6 Astra to make it in blender. It did a surprisingly good job in 30 minutes. So though these benchmarks don't seem to mean much you could always add a touch of randomness to them like the person in the YouTube video did, but the problem with that is how would you compare the benchmarks in any clear way if they aren't even consistent? Regardless it appears LLMs are getting this good at the general task and not just at the particular instances of said task.
I partially owe my career to benchmarking AI with Minecraft, so, I'm going to disagree with the author on this one. Games are a great way to test a model, it's not that deep.
I cant select text nor click links on this page with firefox.
I’ve used Astra for the past day and a half. My layperson’s review is that it is impressive at computer use and 3D reasoning, and fails in similar ways to 5.6 Sol at similar rates when it comes to coding. I have no idea how it scored so high on SWE benchmarks because so far it has been very “mid” as the kids say.
Most of my and my peers PortCos run their own eval and benchmark sets, simply because they know what they need best.
The reality is, capabilities have largely converged across foundation models over the last 18 months, and much of the value add is coming from the harness layer itself now.
This has been the operating assumption for me and my peers, and has largely played out that way.
That said, this has always been an issue with benchmarking since the very beginning. DB Benchmarks, compute benchmarks, and others that were external facing were always inherently a content and product marketing tool. The actual internal benchmarking used to model, understand, and enhance your product was always a closely held secret.
Most of these conversations are happening, but largely in person and not on HN.
It'd be more plausible if the current SOTA models weren't so good at generating SVG with all subjects instead of just pelican.
We're witnessing the mission getting fucking accomplished [0].
[dead]
@notch Hey Markus, can I get a refund on my alpha distro of minecraft? I think the currency is worth more than it used to be considering how many versions there are now :)
Right answer wrong thesis.
The example of 'labs are optimizing for this' is the wrong thesis.
AI arbitrarily generating something of 'apparent sophistication' is not that hard - being able to produce it to spec that has invariable vague elements - and then being able to rationally modify it is the problem.
Look at image gen: you press the magic button and get a 'Pelican on a Bike' - but you can't just change the 'hat' of the Pelican. You have to press the magic button again, and you get a whole different Pelican on a Bike.
This is the fundamental conceit.
It's akin to the conceit that 'writing the code is the work' - but it's not - it's the research, the design, docs, integration and all that 'know-how' that has to 'sit somewhere' so the shape of the thing output can be adapted and moulded.
Arbitrary code output is not quite worthless but almost, it's the the '80% that means another 80% and then anther 80% to go'. It's like a nicer stating point.
The research capabilities of the AI, which don't make for nice demos, are arguably more powerful.