logoalt Hacker News

yorwbatoday at 6:25 PM1 replyview on HN

GRPO increases the likelihood of samples that are better than average, not just the single best, and decreases that of samples that are worse than average. This method doesn't even involve an explicit likelihood, so it's a completely different mechanism.

A comparison with minibatch optimal transport is in appendix A.2 of the paper.


Replies

zimablue22today at 6:34 PM

You're right about the original GRPO proposal, but there are simplified variants that do just use best of K sampling.

GRPO (or GRPO like approaches) for diffusion/flow matching similarly can be likelihood free.

show 1 reply