this is mostly because RLVR is driving all of the recent gains, and you can continue improving the m...

m_ke • yesterday at 5:53 PM • 0 replies • view on HN

this is mostly because RLVR is driving all of the recent gains, and you can continue improving the model by running it longer (+ adding new tasks / verifiers)

so we'll keep seeing more frequent flag planting checkpoint releases to not allow anyone to be able to claim SOTA for too long

alt Hacker News