logoalt Hacker News

vidarhyesterday at 12:35 AM0 repliesview on HN

If doing a lot of heavy lifting there. Not only is it not a given that they'll get the correct answer for a lot of simpler tasks in fewer tokens, but smaller models are often available at far higher tokens/second inference.

There are certainly tasks where fable will be faster and/or cheaper, but there are plenty of tasks where even Haiku is as fast or faster and cheaper, or where you can e.g. get away with models like gpt-oss that you can get from inference providers providing 10x+ the token/second speed.

If you don't use enough tokens that relying only on Fable becomes a problem, then keep using just Fable. Personally, for my $200/week Max subscription I'd run out of the weekly quota for Fable in a day. At API pricing I'd go bankrupt if I tried doing the things I do with cheaper models using Fable.