LLMs actually aren't simple Markov chains tho, your also simplifying. and LLMs trained with RLV...

sometimelurker • yesterday at 7:37 PM • 1 reply • view on HN

LLMs actually aren't simple Markov chains tho, your also simplifying. and LLMs trained with RLVR aren't just optimized over the space of functions (like gpt2 was), they're optimized over the space of programs (programs under some length). You find the ideal algorithm that can do the task you need it to.

Replies

qwery2 • today at 1:06 AM

RLVR is a process which updates the Markov chain

alt Hacker News

Replies