Gonna be rude and say I don't think this is an interesting or useful benchmark tbh.
Clearly some new RLAAS/dataset/env is being used for this now (it doesn't even seem that complicated, you have one LLM judge whether gameplay is recognizable as the original game or not and another trying to implement a logically/semantically identical version of the game). It's why the performance improvement on this workload has been so dramatic.
Everything is going to go from 0->1 on this benchmark in short order because of that.