Sadly until china scales production of hardware it really isn’t economical to run this stuff yourself. It is good it exists though to put pressure against the labs.
Honestly imo this is just proof apple will win in the end. Eventually a phone will be able to run a model good enough to do most things and it then is game over.
> it really isn’t economical to run this stuff yourself
Quantised models running overnight go most of the way for non-coding tasks.
>> do most things and it then is game over.
For the Hyperscalers...and Oracle...cant wait for the day...
Model-on-Chip is coming. GPU are for general computing but have a huge bottle neck for doing model inference.
Even not being able to significantly update a model that is burned on a chip the performance gains are immense. You also don't need the latest chip fabs to make them drastically reducing the cost.