logoalt Hacker News

cheikhcheikhtoday at 5:37 PM1 replyview on HN

did you actually verify that it's output in those scenarios is good ? in my experience opus has been a disappointment and constantly trailing behind actually solving hard problems versus the OpenAI models. I'll say that both have terrible writing style though.


Replies

drnick1today at 6:10 PM

> did you actually verify that it's output in those scenarios is good ?

Yes, in the sense that it reproduced results in the paper or known solutions obtained by other methods. In fact, Opus is very good at checking it's own work in my experience.