noticeably smarter remains to be seen in practice. For now, Haiku is a bit more expensive than Luna on < 100k token, but I just don't have any agentic work below 100k, so this is going to be 5x more expensive than shown on these charts. It's hardly competitive ...
No? Don't use these lower end models to work on small well defined tasks with a frontier model orchestrating? This approach works very well for me and don't have any issue staying under 100k. I have no idea if it's cheaper but it does seem to be much faster for tasks like QA.