logoalt Hacker News

mark_l_watsontoday at 11:35 AM2 repliesview on HN

This seems like a smart move, given their ability to host efficiently. I approve of efforts to make the cost of inference for smaller useful models slowly approach 'close to zero' and there are many good paths for getting there. It is useful for companies to get fast hosting for the class of smaller models they may end up hosting in house.


Replies

jonas_scholztoday at 11:40 AM

I really hope they dont stop at the small models though! The bigger ones that dont fit on a single GPU are more interesting I think

mips_avatartoday at 2:04 PM

Problem is right now the biggest GPU boxes they have is single rtx pro 6000s.