The doom demo would have to be reproduced to confirm what their model is capable of. Oh but it's all closed source, so who knows.
It reminds. Me of Devin. Took a while to debunk. Not saying Jev is a fraud , but the gap between structuring typed output and playing a game involving logical interpretation of frames made of pixels, screams unstructured interpretation they made and forgot to mention.
They mention in their blog post that the model is working on text rather than pixels in the Doom demo.