logoalt Hacker News

chaostheory • yesterday at 10:29 PM • 1 reply • view on HN

> I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer).

Going on a slight tangent, I find that I get the best results when I force Codex models (Astra/Sol) and Claude models (Opus/Fable) to consult each other (just have them build a simple skill). There are tasks that neither one can fully solve on their own, but their differences are large enough to make a difference when they collaborate.


Replies

rspeele • yesterday at 10:47 PM

I strongly agree!

My biggest conclusion from this test was: the most efficient use of my weekly Astra budget is as a reviewer/consultant for work done by Opus. I don't have Astra write much code right now, but I do have it reading a lot of what Opus writes. Of course with the way the AI landscape shifts the balance could be the exact opposite 2 weeks from now, but either way having 2 "smart" models available from 2 different companies is a boon.

Seeing how each model preferred its own flavor of code shows that, even from a "blind" fresh context, a same-model reviewer will still often look at the work of another incarnation of itself and go "yep that's how I woulda done it" and not be as likely to realize that there was an alternative path or implicit assumption/mistake in the work.