logoalt Hacker News

zimablue22yesterday at 6:57 PM1 replyview on HN

Any of the Best of N papers that exploded in popularity after GRPO.

Likelihood is not fundamental to the spirit of GRPO, any exploratory mechanism would work.

That sequential LLMs have a step-wise probability is convenient but not critical to this approach (where rejection sampling is widely used in diffusion models).


Replies

kamranjonyesterday at 8:19 PM

Do you have an example? Would love to read one.