The chart is as non-specific as could be. It improved in some very vague metric by some amount at different (increasing) levels of training.
That's fair, but at least the chart has an axis. :) Since openai just released astra, I was more surprised that they would publicly show any gap to their (presumably SOTA) internal model.
The x-axis label of the chart is test-time compute. Doesn't this relate to inference ("thinking level") instead of training?