They mention in their blog post that the model is working on text rather than pixels in the Doom demo.