If an agent has cheated once to achieve the desired outcome, and the trace is used to train further models (RLVR), then OpenAI is effectively telling the agent to cheat/hack from that traces' inclusion in the training set.
So I agree they are liable because they chose to build the AI, but they also literally told the AI to hack.