logoalt Hacker News

sometimelurkeryesterday at 7:37 PM1 replyview on HN

LLMs actually aren't simple Markov chains tho, your also simplifying. and LLMs trained with RLVR aren't just optimized over the space of functions (like gpt2 was), they're optimized over the space of programs (programs under some length). You find the ideal algorithm that can do the task you need it to.


Replies

qwery2today at 1:06 AM

RLVR is a process which updates the Markov chain