logoalt Hacker News

esskaytoday at 10:49 AM3 repliesview on HN

I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.


Replies

utilize1808today at 11:23 AM

It's logical to serve the best version (quant) of the model at the beginning so that users keep testing it. It is also reasonable to think that the developer of the model tried to test various quant levels by gradually degrading the model's capabilities.

show 1 reply
daveyoungtoday at 10:54 AM

Two potentials from my pov:

1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus.

2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe.

I am leaning towards 1.

show 3 replies
rfootoday at 10:53 AM

lol don't shout out the obvious