alt
Hacker News
dwaltrip
•
today at 12:21 AM
•
0 replies
•
view on HN
RL is part of the model’s training. It changes the weights.
What distinction are you drawing?