logoalt Hacker News

extryesterday at 4:01 PM2 repliesview on HN

It's because Fable is just synthetic RL tasks + scale. The secret has been out for awhile now.


Replies

causalyesterday at 4:05 PM

Does not explain timing

show 3 replies
lossoloyesterday at 5:20 PM

This is basically the answer, they generate A LOT of synthetic task rollouts in parallel, then use RL on the resulting reward signals to improve the model. Add scale to this and you have a Fable class model.