From my skim of their paper, they have gotten robodojo 14% versus previous sota 9%. So its an improvement, but doesnt seem like a reliable model yet, which is fine. I also notice the model becomes worse with more data on the printer refilling task, so thats either a bad run or its a data point against the robustness. The transfer learning results are meaningful. 75% on new complex tasks from <10 hours of demo is really good. I believe this is a 2x improvement over the SOTA. 10 trials per task is standard for these papers but is pretty weak statistically. Memory tasks its weaker at but they have acknowledged it, and I cant see why this arch would have problems w that in the future.