Author here, feel free to ask me any questions. Though I am no expert in this, just someone trying to learn and have some fun doing it.
Curious on whether you think this JEPA-style approach could scale to a full game completions?
For reference, standard PPO has been able to beat the game end-to-end with a relatively small network https://drubinstein.github.io/pokerl/
Curious on whether you think this JEPA-style approach could scale to a full game completions?
For reference, standard PPO has been able to beat the game end-to-end with a relatively small network https://drubinstein.github.io/pokerl/