logoalt Hacker News

dsiegel2275yesterday at 2:41 PM5 repliesview on HN

The focus of the study was the different harness approaches and how they scale across model sizes. The fact that they used any particular set of models is irrelevant.


Replies

klooneyyesterday at 3:01 PM

I'm not totally convinced that models are fungible, the claudes/gpts/Gemini all have pretty individual feels when you're working with them. I wouldn't be surprised if the approaches don't scale or even work the same in a poly model setup

show 1 reply
Tychoyesterday at 5:24 PM

But the behaviour of the system can totally change under different scales.

comparing x10 to x100 doesn’t necessary inform you about x100_000 to x1_000_000

siva7yesterday at 5:36 PM

That can only be assumed by someone with zero clue about harness engineering.

So i can skip this study.

hiddencostyesterday at 2:47 PM

Nope. Sorry. Not how this works.

show 1 reply
shermantanktopyesterday at 2:46 PM

Agree. Harnesses are effective because they interact with the underlying model effectively. If the latest models were fundamentally different, excluding them would be a miss. But I don’t think they are, at least not in ways that would affect these observations.

show 1 reply