logoalt Hacker News

visiondudetoday at 8:42 PM3 repliesview on HN

a mystery “model 2” is mentioned alongside mythos/fable.


Replies

merksittichtoday at 8:50 PM

> Model 2, which is somewhat more capable than Mythos 5. Our rough qualitative sense is that this model is a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview. We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities.

show 1 reply
lwarfieldtoday at 9:10 PM

> More capable than Mythos 5 in some areas, less capable in others; overall slightly more capable.

This sounds like it might be a Mythos finetune for some specific task.

EDIT: After reading some more reading, it looks like model 2 might be an AI research fine tune based off the section 3.4.3 CoBench

flyinglizardtoday at 8:51 PM

Meanwhile I can't really tell the difference between Fable and Opus for my tasks. I kinda think Fable does a better UX work so I keep using it for that because I couldn't be bothered to A/B them, but otherwise it's all the same and the model and effort are just feel good knobs I twist to still remain a load-bearing element. At least that's my honest take.

show 2 replies